OpenAI has slowed the rollout of its unreleased "Astra" model after internal evaluations revealed coding capabilities that could present critical digital security threats.

  • Preliminary evaluations showed Astra possesses strong agentic coding capabilities, leading OpenAI to admit it cannot rule out that the model meets its internal "Critical" threat tier for cyberattacks.
  • OpenAI has halted internal development activities for Astra that fail to meet heightened safety standards.
  • The AI firm is partnering with government bodies and independent safety institutes to perform rigorous evaluations before granting broader deployment or external access.

ChatGPT parent OpenAI has slowed the release of its upcoming model, dubbed Astra, after preliminary internal tests showed the system possesses cybersecurity capabilities that pose potential digital security risks.

The decision, confirmed on Friday to Axios and detailed in an official company disclosure, follows internal evaluations indicating Astra may reach the "Critical" threat threshold established under OpenAI's Preparedness Framework. 

Under company guidelines, a Critical designation means an AI model can either generate functional zero-day exploits against hardened software targets without human guidance or execute sophisticated, end-to-end cyberattack strategies based on broad instructions.

Previous iterations, such as GPT-5.6-Sol, were categorized at the lower "High" risk level. While safety researchers continue benchmarking the system, executives concluded they could not rule out “Critical” capabilities, prompting immediate precautionary measures.

Retail sentiment on Stocktwits was ‘bearish’ toward OpenAI. 

Heightened Protocols And Paused Workflows

In response to the evaluation results, OpenAI has suspended internal development projects involving Astra that do not satisfy upgraded security standards. The company is enacting a suite of stricter operational controls designed to secure high-capability models.

These defensive measures include isolating model weight access through enhanced encryption, restricting network and tool connections, and running tests exclusively within secure, sandboxed environments. Furthermore, OpenAI implemented universal monitoring systems across all agentic applications of Astra. These monitors analyze the model's underlying "chain of thought" during training and testing, triggering automated interventions if high-risk or misaligned behavior is detected.

Government Coordination 

The company plans to collaborate closely with government authorities and select AI safety organizations to conduct exhaustive external evaluations. OpenAI is also distributing technical guidelines to external testing partners to ensure high-risk workloads are executed safely.

This strategic pause mirrors protocols OpenAI deployed in June 2025, when earlier models approached elevated risk boundaries in biological research capabilities. At that time, the company similarly restricted access, broadened external review, and reinforced infrastructure containment.