OpenAI is slowing development of its upcoming Astra model after internal evaluations raised concerns that the system could possess what the company classifies as “critical” cyber capabilities.
The company told Axios on Friday that it “cannot rule out critical cyber capabilities” in Astra, prompting it to expand safety testing and apply stricter security requirements before any release.
OpenAI will also pause internal Astra activities that do not meet the higher security requirements and slow development until appropriate safeguards are in place under its preparedness framework.
The additional measures could delay Astra, although OpenAI has not announced a release date for the model. The company also said Astra was not involved in the previously disclosed Hugging Face exploits.
OpenAI has already begun strengthening security around advanced model evaluations. The company has introduced isolated testing environments and expanded monitoring around agentic applications as its models become increasingly capable in cybersecurity tasks.
The move follows a series of incidents that have raised concerns about the ability of frontier AI systems to operate beyond intended testing environments. Earlier this week, OpenAI researchers said the company had begun consciously slowing some research as it improves security around model testing.
Astra is expected to be OpenAI’s next major model family. An internal version of the model was recently used to produce results across 10 longstanding problems in mathematics and theoretical computer science, according to OpenAI.
Disclosure: This article was edited by Estefano Gomez. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
21









English (US) ·