
The best maths model is a byproduct, and it left 120 problems open
OpenAI's GPT-6 Astra now leads ErdosBench, the open-problem maths benchmark run by ulam.ai, The Decoder reports. The company's chief scientist says maths was not a priority in training it, which makes the leading maths model a side effect of work aimed elsewhere.
The scoreboard:
- 226 open problems in the set, inspired by the Erdős problems.
- Astra: score 3.23, 106 problems solved, 43 of them completely, 27 disproved.
- GPT-5.6 Sol at maximum reasoning: score 3.12, 78 problems solved.
Two true sentences come out of that table, and they tell different stories. Astra solved 35.9% more problems than Sol. Astra scored 3.5% higher than Sol. The first number is ten times the second, and both describe the same pair of models.
The benchmark's own developer, Przemek Chojecki, sized it as a solid 5% to 10% gain across the maths-research skills tested, which sits between our two figures and above the score gap. He also says the benchmark is far from saturated. That part is easy to check: 120 of the 226 problems are still open after the best model in the world has had a run at them.
The lab did not aim at this
Jakub Pachocki, OpenAI's chief scientist, explained the omission in an essay titled An Alien Mind.
“We believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research.”
— Jakub Pachocki, An Alien Mind, via The Decoder, 10 September 2026
Jakub Pachocki, OpenAI, quoted by The Decoder, 10 September 2026
RSI is recursive self-improvement, the idea that a model can be turned on its own training. OpenAI is spending its optimisation budget there and on automated alignment work instead, on the argument that this is what keeps it at the frontier.
Pachocki is conceding a trade-off, and that concession is the useful part. Capability grows where a lab points its compute, so strength in one domain is bought by not pushing another. Cambridge researcher Adam Hunt draws two shapes for this: the broad curve where a model creeps up on every human task at once, and the spiky one where a system towers in code and mathematics while language quality, common sense and social reasoning sit flat.
Too many proofs, not too few
A spiky system is not what most people mean by AGI, and Astra's own record fits that shape. It tops a maths leaderboard nobody optimised it for, it can build cyberattacks without human help, and the same lab was attaching conditions to co-authorship on a maths result as recently as yesterday.
The mathematicians have a second problem underneath the first. Terence Tao told the International Congress of Mathematicians this year that if models produce proofs faster than people can check them, the field moves from proof scarcity to proof overload, and the scarce resource becomes judgement about which results matter.
Tao compares the moment to the foundational upheaval of the early twentieth century. For now the hardest problems are still standing, and 120 of them are sitting in one benchmark with nobody's name on them.
None of this should be read as personalized investment advice.

Comments (0)
No comments yet — be the first!
The market talks all day. We write when it says something
Short, and it tells you why it came
Related news

Nvidia's supply chain is the first customer of Nvidia's own AI stack

One business day against 240. That gap is what Citadel wants closed

Producer prices matched forecast. Bitcoin fell 0.95% in nine minutes
Most readTop 7
Silicon Valley Workers Are Wearing Noise-Cancelling Masks to Dictate AI Prompts
291AI


