How Claude AI challenged the Riemann hypothesis
AI and the Riemann hypothesis explained, including Claude's 650 failed attempts, a new 67.25 percent bound, and what the result means for mathematics.
The Riemann hypothesis is one of the oldest unsolved problems in mathematics. It has stood untouched since 1859. It carries a one-million-dollar prize for whoever proves it. The greatest mathematical minds in history have tried and failed. It concerns the distribution of prime numbers, and its proof, if it ever comes, would reshape entire branches of mathematics.
An Anthropic employee named Jarred Sumner asked an unreleased version of Claude to solve it.
Claude tried 650 different approaches. None of them worked. Then, somewhere in that enormous search, something unexpected happened. Claude did not solve the Riemann hypothesis. But it quietly produced a mathematical result that no human had ever achieved, and Anthropic’s mathematicians had to read it twice before they believed it.
This is that story.
What the Riemann Hypothesis Actually Is
Before getting into what Claude did, it helps to understand what it was trying to do, because “the Riemann hypothesis” sounds abstract in a way that undersells how fundamental it actually is.
Prime numbers are numbers divisible only by themselves and one: 2, 3, 5, 7, 11, and so on, extending infinitely. They do not follow a neat pattern. They appear to scatter somewhat randomly as numbers get larger. But mathematicians discovered long ago that this scattering is not actually random. It follows a distribution, and that distribution can be described with extraordinary precision using a mathematical object called the Riemann zeta function.
The Riemann hypothesis, proposed by Bernhard Riemann in 1859, makes a specific claim about the behaviour of this function. The function has points called “nontrivial zeros,” and Riemann hypothesised that every single one of these zeros lies on a specific line in the complex plane, called the critical line, where the real part equals exactly one-half. That is the entire hypothesis. Prove that every zero sits on that line, and you have solved one of the deepest problems in mathematics.
Nobody has ever proven it. Computers have verified it for trillions of zeros. But verifying does not prove. The hypothesis requires showing that it holds for infinitely many zeros, forever, with no exceptions.
The 650 Failures and What Came After
Jarred Sumner is not a mathematician. He is an Anthropic employee who prompted an unreleased research version of Claude, running inside Claude Code, to attempt to prove the Riemann hypothesis, and then largely stayed out of the way.
The first session produced 650 different ideas. Claude tried each one. None worked. The session ended without a result.
Then Sumner ran it again. His role in the second session, as Anthropic describes it, was mostly limited to sending Claude messages of encouragement. Variants of “keep going” and “believe in yourself.” He was not providing mathematical guidance. He was not steering the approach. He was essentially telling an AI not to give up.
During that second run, Claude coordinated approximately 60 sub-agent instances of itself, each working on different approaches simultaneously. It ran 2,400 shell commands, wrote hundreds of Python scripts, and downloaded 54 research papers from arXiv, the repository where mathematicians publish new work. The whole process consumed 31 million output tokens across the two sessions. Then, after about 37 minutes of silence in the second run, the first crucial result appeared.
Claude did not solve the Riemann hypothesis. But it proved something that no human had proven before.
What It Actually Proved
The hypothesis requires showing that 100 per cent of the nontrivial zeros lie on the critical line. Nobody can prove that yet. But for decades, mathematicians have worked on a related question: what fraction of those zeros can we prove lies on the critical line?
For most of mathematical history, that bound was stuck at 5/12, or roughly 41.6per centt. That figure had survived decades of work by some of the world’s most capable mathematicians. Recent human research had pushed toward two-thirds, but only under a specific simplifying assumption called the narrow-box condition, which essentially constraints where the zeros can be. Getting to two-thirds unconditionally, without that assumption, had not been done.
Claude got there without it. It raised the unconditional lower bound from 41.6 percent to 67.25 percent, eliminating the narrow-box assumption by connecting existing mathematical results in a way nobody had assembled before. It drew on work from multiple research groups over the past decade and synthesised it into a new proof.
Two Anthropic mathematicians reviewed the work. Then Anthropic brought in two outside experts, Brian Conrey and Dan Goldston, both significant figures in analytic number theory, to evaluate it independently. Claude also produced a formally verifiable version of the proof in Lean, a proof-verification language, so that a computer can automatically check the logical steps. You can run the verification yourself.
Anthropic is explicit that this result does not put a proof of the full hypothesis within reach. The technique used, they say, is unlikely to generalise into an actual proof. This is a step forward on one narrow piece of a much larger unsolved problem. But it is a real, verified, unconditional improvement to a result that had resisted human effort for generations.
The Thing Claude Said When It Got There
Here is the detail that stops people.
When Claude arrived at its result, its first reaction was scepticism. It noted in its own output that the finding seemed “too strong to be new.” It had learned enough from training, apparently including the history of failed attempts at this problem and the general difficulty of open questions in mathematics, to suspect that something this good was probably wrong.
The researchers had to provide more encouragement before Claude continued checking its own work. When it did, the result held.
Anthropic’s writeup notes: “Perhaps Claude, like many of us, underestimates the rate of AI progress.”
That sentence is worth sitting with. It does not claim that Claude is conscious or that its scepticism was human-like. It almost certainly reflects patterns learned from training data full of researchers who had similar reactions to surprising results. But the functional behaviour, an AI system producing a novel mathematical result and then doubting it because the result seemed too good to be true, is the kind of thing that would have sounded like science fiction five years ago.
This Is the Second Time It Has Happened
The Riemann bound is not an isolated incident.
A separate session using a different version of Claude, referred to as Claude Fable 5, worked on the Jacobian conjecture, a century-old unsolved problem in algebraic geometry. That session used a similar workflow: a research-style agentic setup, substantial compute, and what Anthropic describes as encouragement-style prompting. Claude produced a proof that disproved the conjecture in its general form.
The pattern across both results is the same. A capable reasoning model, given sufficient compute, enough time, and the ability to coordinate multiple sub-agents working in parallel, can work through mathematical territory that is too vast and too combinatorially complex for a single mathematician to explore by hand. It does not work through insight the way a human does. It works through systematic, parallel search at a scale that human cognition cannot match.
This is a meaningfully different kind of mathematical progress than the field has seen before. Not smarter, necessarily. Different. A mathematician working on the Riemann hypothesis brings intuition, taste, and judgment developed over a career. Claude can try 650 things that don’t work and then coordinate 60 instances of itself to try 60 more approaches simultaneously, drawing on 54 papers downloaded in real time.
Whether that counts as mathematical understanding is a philosophical question. Whether it produces real mathematical results is not the issue. It does.
What This Means and What It Does Not
It would be easy to read this story as proof that AI is about to solve all of mathematics. That would be wrong, and Anthropic has been careful to say so. The technique that produced the Riemann bound result is explicitly described as a dead end for the full hypothesis. The result has not yet passed conventional peer review, though the formal verification in Lean provides a different kind of assurance. And the specific model used is an unreleased research version that is not publicly available.
What it does mean is that the set of problems AI systems can contribute to is larger than most people, including the researchers building these systems, had assumed. When Claude produced its result and said the finding seemed too strong to be new, it may have been expressing something close to what a mathematician would feel on encountering a result that seems too clean, too convenient. Except it turned out to be correct.
The Riemann hypothesis has stood since 1859. No human has proven it. An AI that failed 650 times moved humanity’s understanding of it further than anyone had in decades. The person guiding it was not a mathematician. He was mostly telling it to keep going.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Angry
0
Sad
0
Wow
0