No Priors: Artificial Intelligence | Technology | Startups
No Priors: Artificial Intelligence | Technology | Startups
Conviction
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown
36 minutes Posted Jun 26, 2026 at 10:13 am.
– Cold Open
– Noam Brown Introduction
– Why Benchmarks Are Broken
– Compute Budgets and Projections
– How Long Should Models Think?
– Benchmark-Maxxing
– Using Poker Bots as Evals
– Safety Evals When Model Capability Scales With Budget
– Release Cycle vs. Agent Runtime
– Latent Model Capability
– Limits on Recursive Self-Improvement
– Large-Scale Multi-Agent Coordination
– Competition at the Frontier
– Breaking the Benchmark Grid Equilibrium
– Why Benchmarks Should be Evaluated by Cost
– Conclusion
0:00
36:18
Download MP3
Show notes
When a new AI model drops, it’s judged based on a static benchmark grid that doesn’t account for how long the model is allowed to think. How then should we measure a model’s true capability? OpenAI research scientist Noam Brown returns to talk with Sarah Guo about his latest essay on why the AI industry’s traditional benchmark grids are broken, and how large-scale test-time compute is fundamentally changing how models are evaluated. Noam explains how, if properly scaffolded, today’s models can reason for weeks or even months on complex tasks. He also discusses real-world implications of test-time compute, from building poker solver bots to disproving legendary math conjectures. Together, they also unpack the large gaps in current AI safety frameworks, explore the bottlenecks for recursive self-improvement, and look ahead at the future of multi-agent collaboration and global knowledge sharing.
Read more: Implications of Large-Scale Test-Time Compute
Sign up for new podcasts every week. Email feedback to [email protected]
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @polynoamial | @OpenAI
Chapters: