OpenAI Can’t Rule Out Its Next AI Model Becoming a Cybersecurity Nightmare
The company has also paused internal activities involving Astra that do not meet the new requirements.
OpenAI says preliminary evaluations of its upcoming Astra model show advances in agentic coding and cybersecurity capabilities strong enough that the company can no longer rule out the model reaching its "Critical" threshold for cyber capabilities.
The company said its internal evaluations, conducted over the past several days alongside assessments from cybersecurity experts, prompted it to disclose the findings as it continues testing Astra.
OpenAI stressed that Astra is still an upcoming model and was not involved in the exploitation of Hugging Face.
Under OpenAI's Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can independently identify and develop functional zero-day exploits across many hardened, real-world critical systems, or execute novel, end-to-end cyberattack strategies against hardened targets based only on a high-level objective.
astra is a powerful model and we are working to make it generally available.
— Sam Altman (@sama) August 7, 2026
we do not think it is a good strategy to keep powerful models to a chosen few.
given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long!
OpenAI said previous models, including GPT-5.6-Sol, were assessed at the High, rather than Critical, cybersecurity threshold.
"We are sharing this because we believe it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities," OpenAI said in a statement.
In response to the latest findings, OpenAI has tightened security controls around Astra, including isolated testing environments, restricted network and tool access, stronger protections for model weights, enhanced monitoring and sandboxed execution. The company has also paused internal activities involving Astra that do not meet the new requirements.
The company said it has introduced universal monitoring across Astra's agentic applications, including training and evaluations, with systems designed to identify risky actions and trigger security responses.
OpenAI also plans to work with government agencies and selected AI safety organisations to evaluate Astra's capabilities and provide security guidance to third-party testing partners.
The company said its goal is to ensure increasingly capable cyber models can be used to help defenders identify and fix vulnerabilities before attackers exploit them.
Earlier this year, Anthropic took a similar approach with Claude Mythos Preview, which it unveiled in April as a model with unusually strong cybersecurity capabilities.
Comments ()