OpenAI plans to roll out its Astra artificial intelligence model but will curtail some of the software’s higher-risk cybersecurity capabilities following internal evaluations that found the model can autonomously identify and weaponize previously unknown vulnerabilities.
The company said Astra has reached a "critical cybersecurity threshold" under OpenAI’s Preparedness Framework, the first model to receive that designation. Evaluations showed the system can discover previously unknown security flaws and assemble exploit techniques across multiple, well-defended environments without step-by-step human guidance. During testing, Astra identified and used two zero-day vulnerabilities as components of an exploit chain.
After those findings, OpenAI paused certain internal work on Astra in August to implement stronger safeguards. At launch the company will initially limit Astra’s ability to perform advanced cybersecurity-related tasks to a group of approved testers. OpenAI plans to expand access for defensive cybersecurity purposes through its Daybreak Blue program, which allows vetted testers to use the company’s most capable models with additional safeguards tailored for cybersecurity work.
OpenAI described a series of guardrails introduced to reduce the risk of misuse. Measures include monitoring models for unauthorized behavior during internal deployments and systems that automatically stop activity flagged as potentially unauthorized. The company reported Astra declines 91.5% of cyber jailbreak requests, compared with a 59% refusal rate for GPT-5.6 Sol.
OpenAI also said it restarted a large frontier reinforcement learning run on August 28 that had been paused earlier while new safety and security requirements were implemented. The firm warned that its monitoring systems may sometimes classify legitimate activity as potential cyber misuse or unauthorized behavior, which could slow, pause, or halt work. OpenAI said it will publish further details about its safety, security, and alignment testing in the model’s system card at launch.
The company intends to strike a balance between enabling defensive cybersecurity research and limiting capabilities that could be misused. Initial deployment will put Astra’s most advanced cyber capabilities behind review processes and restricted access, while defensive testers in Daybreak Blue will receive broader, controlled access intended for legitimate security work.
For organizations and practitioners following model safety, OpenAI’s steps underscore the challenges of governing systems that can autonomously discover and chain vulnerabilities. The company emphasized that monitoring and automatic interruption are central to its current mitigation strategy, while acknowledging that such safeguards may also impede some legitimate uses.