OpenAI recently announced something extraordinary: an internal AI system had produced a proposed solution to the Navier-Stokes existence and smoothness problem, one of mathematics’ legendary Millennium Prize Problems.
More precisely, the system constructed a finite-time singularity for the three-dimensional Navier-Stokes equations with a smooth external force, one of the routes permitted by the official Clay Mathematics Institute formulation of the problem. It does not settle the better-known question of whether the unforced equations always remain smooth.
According to OpenAI’s official report, the Navier-Stokes effort involved on the order of 10,000 concurrent AI agents, 2.7 million messages and roughly 130 billion output tokens. The agents reached their proposed solution about 88 hours after the experiment began, followed by another 17 hours of Lean formalization and verification.
At first glance, it sounds like the kind of result people have long imagined from highly autonomous AI: give a system a problem mathematicians have struggled with for decades, let it work for a few days, and get a proof back.
But the story of how OpenAI got there is considerably more interesting.
Contents
The Story Started With Two Human Mathematicians
Before OpenAI launched those 10,000 agents, mathematicians Tristan Buckmaster of NYU and Levent Alpöge, who works at Anthropic, had already been making major progress on closely related fluid dynamics problems.
They weren’t working without AI either.
Their research used tools including Claude and OpenAI Codex, and their results were formally verified using Lean. Their work included constructing finite-time blowup for the three-dimensional incompressible Euler equations with smooth forcing.
Their result did not solve Navier-Stokes, but it pushed further into a closely related area.
Then something interesting happened.
According to OpenAI, on September 1, 2026 it heard rumors that two Millennium Prize problems had been solved. Those rumors, combined with strong results from its newly trained internal model, prompted the company to test the system across the remaining open Millennium Prize problems.
OpenAI later realized the rumors were connected to Buckmaster and Alpöge.
So AI did not simply wake up one morning, independently choose Navier-Stokes and solve it from nowhere.
Human researchers were already making significant progress around the same frontier.
But importantly, it was OpenAI’s own Euler result that later convinced the company to concentrate its resources specifically on Navier-Stokes.
Then OpenAI Scaled It to 10,000 Agents
This is where OpenAI did something genuinely unusual.
It essentially created a gigantic virtual research lab.
The agents were divided into groups that could communicate internally, run code and access a cached version of the internet.
Different groups explored different mathematical approaches.
Nearly 100 agents first worked for about 50 hours on an Euler-related problem and produced a result that OpenAI considered promising.
At that point, OpenAI redirected agents away from other Millennium Prize problems and toward Navier-Stokes.
It also began sharing useful findings between groups.
OpenAI describes this as cross-pollination: Codex consolidated promising intermediate results from different agent groups, and those findings were fed back into later prompts.
Eventually, the successful effort involved roughly 10,000 concurrent agents.
That may actually be the biggest story here.
A single mathematician can explore a handful of ideas.
A research group can explore more.
Ten thousand agents can investigate huge numbers of possible paths simultaneously, discard failures and keep building on promising results.
Perhaps the breakthrough isn’t that AI suddenly became a mathematical genius.
Perhaps it became an incredibly scalable research workforce.
Then the Controversy Started
After OpenAI’s result emerged, Buckmaster raised an uncomfortable question.
He and Alpöge had been using Codex while developing their own research and, according to Buckmaster, had put drafts from the project into the tool.
So he asked OpenAI whether its new model had been trained on, or had access to, those sessions.
As detailed in ABC News’ account of the dispute, Buckmaster said he was initially told that the model did not “look up” user data. When he asked specifically about training, he said he did not immediately receive an answer.
Buckmaster was careful not to accuse OpenAI, saying, “I do not know what their model did, or how,” and that he did not know whether their data had been used.
OpenAI later investigated.
In an update to its report, the company said Buckmaster’s Codex prompts from the preceding two months could not have influenced the system in any way, including through training. OpenAI also says its researchers and agents had not seen Buckmaster and Alpöge’s unpublished work before it became public.
So based on the evidence currently available, there is no basis for saying OpenAI trained on their unpublished proof or copied their private research.
But then another disagreement emerged: who gets credit?
There were now two separate results: Buckmaster and Alpöge’s Euler work and OpenAI’s Navier-Stokes result.
According to Buckmaster, OpenAI researcher Sébastien Bubeck presented two possible ways forward.
One was for Buckmaster and Alpöge to publish their Euler result before OpenAI released Navier-Stokes.
The other was more controversial.
Buckmaster says he was offered the chance to write a paper presenting OpenAI’s Navier-Stokes proof, clearly acknowledging that an OpenAI model had generated it.
But Alpöge would not be included because he worked for Anthropic.
Buckmaster refused.
Bubeck later said he had proposed Buckmaster as lead author of a rewritten presentation of OpenAI’s proof, and that he considered it inappropriate for an Anthropic employee to author OpenAI’s work. He also clarified that he had never proposed removing Alpöge from Buckmaster and Alpöge’s own Euler paper.
That disagreement points to a problem academic publishing was never really designed for.
If humans develop the surrounding ideas, AI tools participate in the research, another AI system generates the final proof and humans still need to interpret and publish it, who actually gets the credit?
So What Did AI Actually Discover?
OpenAI’s experiment was not 10,000 agents starting from a blank page and inventing fluid dynamics from scratch.
It built on decades of mathematics, recent human progress in the same area, and research directions that were already becoming promising. Humans also decided where to allocate compute, while OpenAI deliberately passed useful intermediate findings between agent groups.
But that does not make the result any less significant.
What OpenAI demonstrated is something different: thousands of AI agents can explore research paths in parallel, discard dead ends, combine promising ideas and compress an enormous amount of work into just a few days.
The Clay Mathematics Institute has since said that the Navier-Stokes problem “has apparently been settled,” while making clear that evaluating the result and determining credit will take time.
So perhaps the important question is not whether AI has suddenly reached artificial general intelligence.
What this experiment really shows is that we can now give a very difficult research problem to thousands of AI agents, let them explore many different ideas at the same time, and potentially finish in days what could take human researchers much longer.
That is a major shift.
But it also creates a new question: if AI is building on decades of human research and then exploring those ideas at massive scale, what part of the final result should we call a genuine AI discovery?
Final Thoughts
For me, the most interesting part of this experiment is not that AI solved a famous mathematics problem in 88 hours.
It is how it got there.
OpenAI did not ask one model to sit and think harder. It created thousands of agents, sent them down different research paths, moved more compute toward promising ideas, shared useful discoveries between groups, and kept going until something worked.
That starts to look less like a chatbot and more like a research organization running at machine speed.
In fact, OpenAI has already described its broader goal as building an automated AI researcher.
But Navier-Stokes also shows why we should be careful with the word discovery.
The agents were working on top of decades of mathematics, existing research, human decisions and ideas that were already developing at the frontier.
So maybe the real milestone here is not that AI suddenly learned how to discover things on its own.
It is that research itself may now be scalable.
And if 10,000 AI agents can already do this today, the more interesting question is what happens when the same approach is applied to thousands of other unsolved problems across mathematics, science and engineering.
Abid Ali Awan (@1abidaliawan) is a certified data scientist professional who loves building machine learning models. Currently, he is focusing on content creation and writing technical blogs on machine learning and data science technologies. Abid holds a Master’s degree in technology management and a bachelor’s degree in telecommunication engineering. His vision is to build an AI product using a graph neural network for students struggling with mental illness.
