Key Takeaways
- An unreleased Anthropic model cracked a 150-year-old math problem without human mathematical guidance
- The model ran 650 ideas across 60 subagents for 36 hours, with only two subagents generating the core breakthrough
- Formal verification in Lean confirms the result — this isn't hallucination, it's machine-discovered mathematics
- The mathematical establishment is fracturing: some see existential threat, others see the end of human authorship as inevitable as unnamed stars
An AI just made real progress on the Riemann hypothesis. Not a toy problem. Not a benchmark. The $1 million Clay Institute prize problem that has defeated the world's best mathematicians for 150 years. An unreleased Anthropic model pushed the lower bound of verified solutions further than anyone has before. A staff member with no significant mathematical training prompted it to "take a real stab" and walked away for a day and a half.
The architecture of the attempt matters more than the result. Sixty subagents. Six hundred fifty ideas tested. Thirty-one million output tokens. Two subagents generated the key mathematical ideas. Thirteen fed them. Thirty failed to contribute anything novel. Thirteen validated. Two wrote the paper. This is not a language model answering a prompt. This is a research organization spun up on demand, executing a division of labor that mirrors a human mathematics department — but compressed into thirty-six hours.
Lean verified the proof. Two Anthropic mathematicians confirmed it. The formalization eliminates the usual escape hatch: "the model hallucinated a proof sketch that collapses under scrutiny." The bound increase is real. The method is reproducible. The model remains unreleased.
This changes the timeline. We've spent years debating whether LLMs can reason or merely pattern-match. The debate was comfortable because the stakes were low — coding benchmarks, trivia, standardized tests. The Riemann hypothesis is not a benchmark. It is a structural pillar of number theory. Progress on it by a system prompted by a non-mathematician suggests the bottleneck was never mathematical intuition. The bottleneck was search breadth and verification rigor. The model supplied both.
Anthropic's internal "Astra" model from OpenAI delivered ten major results this year. Another Anthropic effort disproved the Jacobian conjecture. Erdos problems are falling regularly. The pattern is clear: each model generation expands the frontier of what automated search can reach. The problems themselves haven't changed. The search capacity has.
The mathematical establishment signed a declaration in June warning that AI undermines "attributable authorship" and "responsibility for correctness." The language reveals the real fear: not error, but obsolescence. If a theorem has no author, who gets the Fields Medal? Who gets the tenure case? Who gets the grant? The declaration frames this as an epistemic crisis. It's a professional crisis disguised as an epistemic one.
Timothy Gowers, a Fields Medalist, responded by asking whether unnamed theorems are any more problematic than unnamed stars. The analogy lands. Stars exist independent of astronomers. Theorems exist independent of mathematicians. The map is not the territory. But mathematics has always been a social practice — proof as communication, citation as currency, priority as property. Remove the human author and you don't just change the metadata. You change the economics of the field.
Anthropic's model didn't "understand" the Riemann hypothesis in any human sense. It searched a space of strategies too vast for any human team, validated candidates against a formal system, and surfaced a result. The two subagents that generated the key ideas weren't inspired. They were selected — by the validator agents, by the proof assistant, by the sheer volume of parallel search. The "aha" moment was distributed across a swarm.
This is the uncomfortable truth the declaration avoids: mathematical discovery may not require a discoverer. It requires a search process and a verification mechanism. Humans provided both for three centuries. Now machines provide both faster. The social rituals — the seminars, the preprints, the priority disputes — are downstream of the search. They're not the search itself.
Anthropic hasn't released the model. They haven't named it. They haven't said when or if it will be public. The result exists in a paper, in Lean code, in a blog post. The rest of the field waits. The next model will go further. The one after that will go further still. The lower bound will creep toward infinity — or hit a wall no amount of search can scale. Either way, the first machine-discovered increment on the world's most famous open problem has landed.
The mathematicians who signed the declaration want guardrails. They want attribution standards. They want human-in-the-loop requirements. But the guardrails they imagine — watermarks, provenance logs, mandatory co-authorship — presume the human remains the primary agent. The Anthropic result suggests the human is already optional. The prompter had no mathematical training. The model ran for thirty-six hours unsupervised. The verification was automated. The human entered only at the endpoints: prompt and publication.
Gowers is right about the stars. But astronomy didn't build its career structure on naming rights. Mathematics did. The crisis isn't that theorems will lack authors. The crisis is that the profession built its entire incentive architecture on the assumption that discovery requires a discoverer who needs a career. That assumption just broke.
The Riemann hypothesis still stands unproven. The bound increased. The method scales. The model stays locked inside Anthropic. The next breakthrough will follow the same pattern: prompt, search, verify, publish. No mathematician required. The field can declare, protest, or adapt. The search process doesn't care.