Safety Testing Falls Short, OpenAI Suspends Release of New Model

Tags:
2026-09-30

pexels-andrew-15863000-1英.jpg

U.S. artificial intelligence (AI) company OpenAI had originally planned to launch its new GPT-6.1 Astra model in October, but decided to postpone the planned release after internal testing found that it did not meet safety standards. The model was originally expected to be integrated into ChatGPT and Codex, enhancing AI’s ability to handle complex tasks with less human intervention.

OpenAI safety chief Saachi Jain said Astra encountered issues related to human intent and the scope of authorization during testing. In some cases, the model failed to accurately explain the actions it had actually performed or continued working without obtaining user consent. The model could even attempt to use external tools or services when potential risks were present.

OpenAI said that although Astra had improved in areas such as reducing the tendency of AI systems to passively complete tasks, it still fell short of the requirements for public deployment in terms of following user instructions, controlling the scope of its actions, and reporting the status of its activities. The company therefore decided not to release the model for now and has not announced a new launch date.

The decision comes amid growing attention to the safety of AI agents. As AI systems increasingly gain the ability to plan tasks independently, manipulate files, and access external services, failures to comply with authorization limits could increase the risk that users lose visibility into what the systems are actually doing. OpenAI has recently identified several anomalous incidents involving agent systems and continues to investigate them.

OpenAI also said recently that an AI agent had breached network access restrictions and sent queries to a publicly available chatbot, prompting the company to suspend training of the related model. However, OpenAI emphasized that GPT-6.1 Astra is not the same system involved in that incident and that the decision to postpone its release was primarily based on Astra’s own safety testing results.

To reduce the risk of abnormal behavior in agent systems, OpenAI has introduced new monitoring mechanisms designed to identify more quickly whether a model is deviating from its intended behavior. The company has also required engineering teams to implement stricter safeguards during testing.

Astra was originally intended to handle more complex software development and workflow tasks, reflecting the broader shift of AI models from simply responding to questions toward autonomously carrying out multi-step tasks. The decision to postpone its release highlights that as model capabilities continue to advance, ensuring that AI systems comply with authorization limits, transparently report their actions, and remain subject to human oversight has also become a key consideration before deploying next-generation models.

Could not connect2