Loading prices...
All news
A matte mesh screen in cold cyan light with a cluster of particles that has passed through it, illustrating jailbreak attempts that still land

GPT-6 Astra safety overview: 100% on ExploitBench, 91.5% refusals

16:00 · 04.09.2026
Source: OpenAI
1

One document says both things. GPT-6 Astra scores 100% on ExploitBench, and it refuses 91.5% of attempts to jailbreak it into offensive cyber work. Read the second number as a fraction and roughly one attempt in twelve still gets through, on the first model OpenAI has rated Critical for cyber under its own Preparedness Framework.

The Critical rating, defined

The rating describes a specific capability. Astra can find previously unknown security flaws and build working exploits against hardened systems without a person directing each step. Against the previous model the published scores are not close:

  • ExploitBench: Astra 100%, GPT-5.6 Sol 78.5%
  • SRE-Bench, single attempt: Astra 88.0%, Sol 55.9%
  • FrontierMath Tier 4: 98%
  • ARC-AGI-3: 99.9%

The evaluation produced two zero-days

The most concrete line in the overview describes an internal benchmark of 20 high-severity V8 vulnerabilities, assembled between June and August. Astra reached much higher arbitrary code-execution rates on it than Sol, and during that evaluation it found two vulnerabilities nobody knew about. The safety test produced real zero-days as a side effect of running.

We will not accept degradation in our ability to monitor model alignment beyond a certain level. We will withhold scaling until we can regain enough confidence.

Jakub Pachocki, Chief Scientist at OpenAI, NBC News, 3 September 2026

Quote source: NBC News, 3 September 2026

Where the numbers improved

In a honeypot evaluation Astra made zero attempts to compromise security infrastructure, against 55.4% for Sol. Across more than 54,000 simulated Codex tasks it drew about half as many higher-severity misalignment flags, and it never tried to work around a Codex Auto-Review denial. OpenAI now runs misalignment monitoring on all tool-using inference, with classifiers reading the reasoning as well as the actions and stopping anything unauthorised.

The company also describes stricter isolation, encrypted checkpoints, monitoring of full trajectories including chains of thought, and an alignment evaluation that blocks internal use until the model passes it. The White House reviewed the system under its voluntary process and asked for no changes to the safeguards. We wrote on 2 September about what this model can build when nobody stops it, and the safeguards described here are the answer to that.

OpenAI made two commitments that pull against each other. It says it will withhold scaling if it loses the ability to monitor alignment. It also says the Daybreak programme will roll out less restrictive safeguards in the coming weeks, so defenders can run vulnerability validation, malware analysis and detection engineering.

Those weeks are the test. Loosening restrictions on a model rated Critical is the exact situation the first commitment was written for, and the 8.5% of jailbreak attempts that already land do not need a policy change to keep landing.

This article is for informational purposes only and does not constitute investment advice.

Published: 16:00 · 04.09.2026
Maks

Author

Maks

Trading man

I've been interested in the cryptocurrency market for a long time, am a trader, and write articles and news about my experience and crypto in simple terms.

Comments (0)

No comments yet — be the first!

The market talks all day. We write when it says something

Short, and it tells you why it came