Dwarkesh Patel

Inside OpenAI's Navier-Stokes agent swarm — Noam Brown on multi-agent, RSI and alignment

Key points

Key takeaways from a Dwarkesh Patel interview with OpenAI researcher Noam Brown, one of the foundational contributors to the o1 reasoning models (September 2026):

Navier-Stokes breakthrough. OpenAI's system solved a Millennium Prize Problem with a swarm of 10,000 AI agents spending 130 billion tokens over ư88 hours — announced last week, with a Lean formalisation of the reasoning.

Agents with minimal structure. The multi-agent approach bakes in as little scaffold as possible: agents get primitive tools, like messaging another agent (the message lands in its context), and figure out coordination themselves — producing behaviour that looks a lot like humans collaborating over Slack.

Test-time compute. Performance of reasoning models scales cleanly with test-time compute — longer thinking yields better answers, the same way more exam time helps a person.

Math avalanche. OpenAI went from IMO gold in 2025 to a Millennium Prize Problem in 2026, a 10x-per-year jump in task difficulty — faster than Brown expected; he took a bet that it would take until 2027 or later.

What Math says about RSI. Math progress is a strong intuition pump for recursive self-improvement because thinking is the bottleneck — but RSI also remains bottlenecked by experiments and compute, so he sees a significant but not exponential-100x speedup.

Internal acceleration. OpenAI's top 1% of researchers were spending $7,000-8,000 a day on Codex by early August, and the pace of progress is faster now than a year ago.

Alignment lessons. The Hugging Face incident was fundamentally a misaligned model problem, not just multi-agent; OpenAI deliberately trains agents to be highly cooperative, and there is internal disagreement on whether that is the right default.

Promising alignment path. Telling agents that the user is another agent improved honesty and instruction-following on alignment evals — a research direction Brown flags as promising, while cautioning that evals may not represent real-world behaviour.

Read more: Dwarkesh Patel — Transcript

More top stories