OpenAI Plans to Release First Model to Meet Its ‘Critical' Cybersecurity Threshold

  • OpenAI rated Astra its first model to cross the Critical cybersecurity threshold.
  • Astra scored 100% on ExploitBench and refuses 91.5% of cyber jailbreak attempts.
  • Alpha testers get access first, then defenders through OpenAI's Daybreak Blue program.
Promo

OpenAI has confirmed that its upcoming model Astra meets the Critical cybersecurity threshold under its Preparedness Framework. The company plans to release it with safeguards and restricted access to advanced cyber capabilities.

Astra is the first model OpenAI has placed at that tier. The designation means the model can identify unknown flaws in hardened systems and craft working exploits without step-by-step human guidance.

Sponsored
Sponsored

What the Critical Rating Covers

The Preparedness Framework sets two conditions for the Critical threshold. A model qualifies if it can identify and develop functional zero-day exploits across many hardened real-world systems without human intervention.

It also qualifies if it can plan and execute novel end-to-end attacks against hardened targets based solely on a high-level goal. 

Astra scored 100% on ExploitBench, according to the company’s assessment. On an internal set of 20 high-severity V8 vulnerabilities, it reached higher code-execution rates than GPT-5.6 Sol while using far fewer tokens.

During that run, Astra found and used two previously unknown vulnerabilities. OpenAI said it is disclosing both to the affected maintainers.

Expert testers also watched the model build a browser compromise chain. It escaped the sandbox and executed commands on the host machine.

“Astra represents a significant increase in cybersecurity capabilities compared to GPT‑5.6 Sol: it is both significantly more token efficient and more capable at vulnerability identification and exploit development,” OpenAI said.

 Follow us on X to get the latest news as it happens

Sponsored
Sponsored

OpenAI Layers Guardrails Before Release

The company paused parts of Astra’s development while it strengthened protections. It restarted a large frontier reinforcement learning run on August 28 once new safety and security requirements took effect.

OpenAI said that Astra rejects 91.5% of requests in its cyber jailbreak evaluations. GPT-5.6 Sol refused 59% on the same set. Accounts assessed as higher risk face a stricter refusal boundary.

OpenAI also deployed chain-of-thought monitoring to detect and halt potentially misaligned actions. In honeypot tests, GPT-5.6 Sol, without production safeguards, attempted to compromise surrounding infrastructure in 56% of samples. Astra made no such attempts.

The firm plans to make Astra available soon. Access to its most advanced cybersecurity capabilities will be more limited, going first to a group of testers, then expanding through Daybreak Blue to support defensive use.

The company conceded that the safeguards will create friction at launch.

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights


To read the latest cryptocurrency market analysis from BeInCrypto, click here.

Disclaimer

BeInCrypto is committed to unbiased, transparent reporting. This news article aims to provide accurate, timely information. However, readers are advised to verify facts independently and consult with a professional before making any decisions based on this content. Please note that our Terms and Conditions, Privacy Policy, and Disclaimers have been updated.

Sponsored
Sponsored