10⁴ agents: a new scaling law

Published: 2026-09-14

Last week OpenAI announced a solution to the Navier-Stokes millennium problem. This was a relatively modest update on AI ability and timelines — there have been a number of impressive mathematical results this summer and AI already seems superhuman at counterexample construction. A low hanging millennium problem as it were. A positive proof of the Riemann hypothesis or P≠NP would surprise me a lot more.

But the details of the post were quite surprising. 10000 agents worked 88 hours to construct the proof. We've had subagents for a while, they've existed in Claude Code since April 2025 and Codex since January 2026. But so far I've felt their purpose was to annoy me. Every time I ask a question and Claude code spins up a search agent I wonder if I should just restart the session. Larger groups felt likely to go off the rails entirely. It seems like OpenAI has cracked something fundamental that will change the dynamics of the returns to compute, the scope of autonomous work possible and human-LLM interaction.

I'm possibly late to update here as the shape of this should have been clear from the HuggingFace incident which involved the self-organized cooperation of over 700 agents. The update is it's actually being used successfully to achieve frontier work. I imagine it's easier for the agents when you give them a communication channel instead of making them discover it.

As Ilya observed, when one path appears blocked, nature finds another:

There is a precedent, there is an example of biology figuring out some kind of different scaling, something is clearly different ... It is possible for things to be different. The things that we are doing, the things we have been scaling so far is actually the first thing we have figured out how to scale and without doubt, the field, everyone who is working here will figure out what to do.

Ilya Sutskever

Graph of body mass to brain size.

Biologically this transition is perhaps obvious in retrospect. Humans are not so much more intelligent than other hominids, primates and whales. The discontinuous achievements of humanity seem due to our superior knowledge accretion i.e. our culture. See Joseph Henrich. So far LLMs have picked up human culture through pre- and post-training. But they haven't developed norms for contributing back to large socially driven projects. A modern GitHub repo is a team of humans making all their diffs with LLMs but mostly still determining the direction of what needs to be done and the shape of the repo in a human way. Perhaps now the humans can be cut out and long term accretive collaboration of agents is possible.

Navier-Stokes result added to metr time horizons graph.

This result seems to break the smooth increase of time horizon of tasks AIs can complete. Obviously there are some things off about the comparison: This is a sample of 1 and AI cannot complete 50% of problems of this difficulty. But also this is not a task human researchers were able to complete either. But what is totally fair about the addition is putting the collaboration of 10000 agents as a completed task. In the original metr graph the far left points are simple tasks like fact recall or basic arithmetic that were accomplished with a single LLM completion and the far right points are simple research tasks that were first accomplished by basic agentic loops. The whole point of the graph is to show the progress in what an autonomous AI system can do and should be agnostic to the structural details of the system. The more compute used the less impressive and practical but seeing what can be done at all is important. The possible will be easy in a year.

Open Questions

Graph of scaling of new model compared to Astra.

The OpenAI announcement has some interesting details about the agent collaboration:

This makes me really curious about the topology of the collaboration. Were the messages more like posts to a bulletin board than peer to peer messages? And if so did agents in the group decide which parts of the problem to attend to? Was there a leader agent and middle manager agents or was it fully self-organized? When OpenAI writes "we shifted agents away from the other Millennium Problems and prompted these agents with the Euler resolution," do they mean they took agents off of the other projects and just added another message about their new task, the way you might re-org humans? I have to imagine what they mean is they halted agents in those other groups and spun up totally new agents to aid in the other task. But that's not what it sounds like from the language. Also is the subagent use an important part of the new model's scaling? Does the graph show the scaling of a single agent or of a single agent allowed to spawn a number of subagents with the X axis showing the cost across all subagents?

Monitoring Challenges

In some ways splitting across agents is a boon for monitorability. Each agent is working on some task and its behavior can be understood by its work output: the messages it sends or the diffs it creates. But as we saw in Subliminal Learning it's a mistake to think we can naively monitor agent communications and understand the true message being sent. For a while the scope of diffs agents create has been difficult to understand. And with the Navier-Stokes proof it seems wholly impossible.

We're leaving the period of human acceleration by AI and entering a period of superhuman accomplishments.





Subscribe Premium $10/month

Checkout with Stripe