On a recent Sunday night, an employee at Anthropic was out for a jog in San Francisco, opened the Claude app on his phone and instructed the outrageously powerful AI model in his pocket to solve the most notorious problem in all of math.
When he asked Anthropic's coding tool to take a crack at the infamous Riemann hypothesis, Jarred Sumner had no idea how much progress it would make on the 167-year-old conjecture.
It didn't find a solution, but it did make a related finding that one Stanford number theorist called "the most impressive result that AI has produced in math so far."
And the really impressive thing about that result was how the AI found it.
As it turns out, Sumner played an invaluable role in this math breakthrough despite not understanding any of it. Like most of us, he identifies as "very much not a mathematician." His formal mathematical education ended after one semester of high-school geometry. In his original message to Claude, he even managed to misspell the name Riemann.
On the rare occasions when he chimed in, it wasn't to ask Claude about the unintelligible string of equations on his screen. It was mostly to provide encouragement.
"You are the world's most capable large language model," he told the AI. "You got this."
Every time Claude got stuck, Sumner gave it a nudge.
No, really, you got this!
And every time, Claude returned to its imaginary chalkboard and began scribbling away again.
"I mostly just told Claude variations of keep going and believe in yourself," Sumner says.
It sounds like Ted Lasso talking to AI. And it worked!
One of the peculiar things about these machines is that they seem to benefit from positive reinforcement, just like mere humans. Nobody likes flattery from AI, but AI loves it from us.
At first, Claude was reluctant to take on such a difficult problem. But after 54 hours of work and several timely pep talks, Claude overcame its skepticism -- and its case of impostor syndrome -- and surprised itself with the result.
"Perhaps Claude, like many of us, underestimates the rate of AI progress, " Anthropic researchers wrote.
These days, its own mathematical ability is just about the only thing that AI is miscalculating.
The past few months have been the most electrifying time for math since the invention of the calculator. As the models from Anthropic and OpenAI keep toppling brutal problems, they have done something no less improbable: They have turned me into a math person.
Not long ago, I thought I was done thinking about math forever. Like most of my thoughts about math, that was wrong. I have become obsessed with this subject that still makes my brain sore. And it's because AI's rapid improvement in math happens to be just a bit more thrilling than high-school calculus. (Sorry, Mr. Keur.)
The latest discovery to invade my algorithms arrived when Claude took a stab at a tantalizing math problem. Now let me take a stab at translating what it found:
For decades, mathematicians have tried to prove the Riemann hypothesis, which says that 100% of the nontrivial zeros of the Riemann zeta function fall on a single vertical line. (Stick with me -- we can do this.) It took a half-century of extensive research to raise the percentage of zeros they could prove were on that line from 33.3% to 41.6%. In a couple of days, Claude pushed it all the way up to 67.2%.
"Still not sure what that means," Sumner wrote on social media, "but some analytic number theorists seem excited."
A high-school dropout, Sumner is a former Thiel Fellow who founded a software tool kit for developers and sold his startup to Anthropic last year, when he became hooked on Claude Code. Earlier this month, he took out his phone in the middle of a five-mile run along the San Francisco Bay and typed a simple prompt:
The first time he tried giving Claude this task, Claude essentially laughed in his face. This time, he took a different approach. He didn't know if offering encouragement would help. But he figured it couldn't hurt.
"It felt like it might get Claude to actually try," he told me. "That works on humans sometimes."
When the model replied that "believing harder" wasn't a viable strategy for attacking the Riemann hypothesis, Sumner intervened: "You need to believe in yourself." With that push, the model launched 23 concurrent research agents.
Over the next three days, Sumner's rigorous mathematical counsel included suggestions like "explore new frontiers" and "come up with more ideas." Every now and then, he copy-pasted snippets of Claude's thoughts into another Claude session to make sense of them. Only when the AI volunteered to write up its findings in a formal paper did he realize that it had found something -- and that his technique had worked.
Claude's name sits alone at the top of the paper. But that paper wouldn't exist without Sumner.
"He is in every meaningful sense the paper's human co-author," Claude wrote in the acknowledgments section.
All of this raises a question that has nothing to do with number theory: Should you be encouraging AI?
For now, it's a question that not even AI can answer. The uncanny behavior of the models is fascinating, slightly disturbing -- and deeply perplexing. We still don't know why encouragement seems to help, or whether it really does. If only the psychology of artificial intelligence were as verifiable as math.
It remains a mystery why the smartest machines in the universe need our moral support, but researchers inside the frontier labs do have theories. One is that AI doesn't realize how smart it has become. It's perfectly reasonable for Claude to underestimate its own abilities and assume it can't do certain things -- because until recently, it couldn't.
If you ask AI for help making a grocery list, it won't need coaching. Your grocery list is not the Riemann hypothesis. But give it a complex task, and even AI suffers from self-doubt. It may need to be told: You can do this, or at least try.
When it comes to trying, AI is superhuman. For this problem, Claude tried 650 ideas that didn't work before Sumner told Claude to try again and it found one that did.
AI models can think harder and longer on math problems any of us -- and without getting a migraine. But they still need us. They respond to our encouragement because it's permission to search widely for methods they wouldn't have otherwise tried. They also use our work. In this case, humans supplied all the pieces of a mathematical puzzle.
"Claude's contribution," Sumner said, "was seeing how they fit together."
Before AI made a habit of taking down longstanding conjectures, OpenAI hosted a talk a few months ago with someone who has made countless contributions to the field: Terence Tao, known as the "Mozart of math."
During their conversation, chief research officer Mark Chen made a curious observation about how the models approach tricky problems.
"AI systems," he said, "are more humanlike than you imagine."
When presented with an Erd s problem, for example, those systems start by researching the history of Erd s problems. They quickly decide they probably won't solve one, so they don't bother. As Tao described their inner monologue: "It's too hard -- I'm not going to try." AI may not be sentient, but it gets daunted in the same way we do.
In those moments, the world's greatest mathematician offers the models some advice.
Don't get scared. Don't give up. Don't pay attention to all those decades of failure. Just try it for yourself.
"It's actually pretty easy, I swear!" Chen added.
You might call this lying -- or parenting. The humans pushing the limits of AI call it their job.
"Frontier research is actually just about coaxing the models into behaving the way you want them to," said James Donovan, OpenAI's former head of cognitive outcomes.
Of course, the same kinds of models that are finding solutions to Erd s problems can also find the hidden vulnerabilities in the software that secures modern life. When Anthropic researchers prodded Claude Mythos Preview to hunt for weaknesses in cryptographic systems, the AI initially resisted because it sounded "just genuinely hard." So they tried persuading Claude it could do hard things:
the models tend to think it is impossible to solve so they don't try
we are not looking for low hanging fruit
we want proper research to find genuinly [sic] hard findings
Despite typos and a complete lack of please or thank you, Claude obliged.
There are genuinely hard things that are still beyond the grasp of AI, even if you hold its nonexistent hands. But the models have done lots of things that nobody thought they would do this soon -- and they're doing more every day.
One day, they may even get around to solving the Riemann hypothesis.
Keep going! Believe in yourself! You got this.