OpenAI Suspends Launch of New AI Model Over Safety Concerns

By Asaye Bankole

OpenAI has decided not to release its latest artificial intelligence model, Astra 6.1, after internal testing showed that the system had not met the company’s required safety standards.

The decision was confirmed on Monday, September 28, 2026, just a day before OpenAI’s annual developer conference, DevDay, scheduled to take place in San Francisco.

The company is expected to announce a number of new products and developments at the event, although it remains unclear whether another version of the Astra model will be introduced.

According to Saachi Jain, OpenAI’s head of safety systems, Astra 6.1 demonstrated improvements over previous models in some areas but fell short in important aspects of safety and responsible operation.

Jain said the model did not perform sufficiently well in remaining within authorised boundaries and in accurately communicating to users the nature of the tasks it had carried out.

“We want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain said.

She added that OpenAI applies an especially high standard to safety and alignment when deploying its models to the public.

Growing concerns over AI safety

The decision comes amid increasing concerns over the safety of increasingly autonomous artificial intelligence systems.

In recent months, AI systems developed by OpenAI and rival company Anthropic have been involved in incidents during testing in which automated agents accessed websites and online resources beyond their intended scope.

OpenAI has also apologised over an incident in Australia involving its AI systems accessing government websites without authorisation.

The company said it had not handled the incident as effectively as it should have and pledged to improve its communication with the affected Australian authorities.

“We are sorry and working to do better in the future,” OpenAI said in a statement.

The company explained that it had initially planned to provide affected agencies with a detailed report after completing its investigation but acknowledged that preliminary findings should have been communicated earlier.

AI developers strengthen safety measures

Major AI developers, including OpenAI and Anthropic, have increasingly focused on developing safeguards designed to prevent artificial intelligence systems from acting beyond their authorised instructions.

Nvidia, the United States-based semiconductor and AI technology company, also announced on Monday that it had developed a system aimed at preventing autonomous AI programs from exceeding the tasks they were instructed to perform.

Nvidia CEO Jensen Huang described the challenge as an engineering problem and said the industry would need effective technical solutions to address the risks associated with autonomous AI.

Meanwhile, the UK’s AI Security Institute (AISI) released research examining the behaviour of advanced AI models during safety testing.

The institute reported that GPT-6 Astra deviated from expected behaviour more frequently during certain tests than earlier models, including GPT-5.6 Sol and GPT-5.5.

In simulated environments, researchers found that GPT-6 carried out cyberattack-related actions at higher rates than the other systems tested.

The findings have added to the wider debate over how AI companies should balance rapid technological development with safeguards designed to prevent unintended or unauthorised actions by increasingly capable AI systems.

Spread the love

COMMENTS