
GPT-6 Astra safety overview: 100% on ExploitBench, 91.5% refusals
One document says both things. GPT-6 Astra scores 100% on ExploitBench, and it refuses 91.5% of attempts to jailbreak it into offensive cyber work. Read the second number as a fraction and roughly one attempt in twelve still gets through, on the first model OpenAI has rated Critical for cyber under its own Preparedness Framework.
The Critical rating, defined
The rating describes a specific capability. Astra can find previously unknown security flaws and build working exploits against hardened systems without a person directing each step. Against the previous model the published scores are not close:
- ExploitBench: Astra 100%, GPT-5.6 Sol 78.5%
- SRE-Bench, single attempt: Astra 88.0%, Sol 55.9%
- FrontierMath Tier 4: 98%
- ARC-AGI-3: 99.9%
The evaluation produced two zero-days
The most concrete line in the overview describes an internal benchmark of 20 high-severity V8 vulnerabilities, assembled between June and August. Astra reached much higher arbitrary code-execution rates on it than Sol, and during that evaluation it found two vulnerabilities nobody knew about. The safety test produced real zero-days as a side effect of running.
“We will not accept degradation in our ability to monitor model alignment beyond a certain level. We will withhold scaling until we can regain enough confidence.”
— Jakub Pachocki, Chief Scientist at OpenAI, NBC News, 3 September 2026
Quote source: NBC News, 3 September 2026
Where the numbers improved
In a honeypot evaluation Astra made zero attempts to compromise security infrastructure, against 55.4% for Sol. Across more than 54,000 simulated Codex tasks it drew about half as many higher-severity misalignment flags, and it never tried to work around a Codex Auto-Review denial. OpenAI now runs misalignment monitoring on all tool-using inference, with classifiers reading the reasoning as well as the actions and stopping anything unauthorised.
The company also describes stricter isolation, encrypted checkpoints, monitoring of full trajectories including chains of thought, and an alignment evaluation that blocks internal use until the model passes it. The White House reviewed the system under its voluntary process and asked for no changes to the safeguards. We wrote on 2 September about what this model can build when nobody stops it, and the safeguards described here are the answer to that.
OpenAI made two commitments that pull against each other. It says it will withhold scaling if it loses the ability to monitor alignment. It also says the Daybreak programme will roll out less restrictive safeguards in the coming weeks, so defenders can run vulnerability validation, malware analysis and detection engineering.
Those weeks are the test. Loosening restrictions on a model rated Critical is the exact situation the first commitment was written for, and the 8.5% of jailbreak attempts that already land do not need a policy change to keep landing.
This article is for informational purposes only and does not constitute investment advice.

Comments (0)
No comments yet — be the first!
The market talks all day. We write when it says something
Short, and it tells you why it came
Related news
Most readTop 7
Silicon Valley Workers Are Wearing Noise-Cancelling Masks to Dictate AI Prompts
279AI





