Show notes
OpenAI Research Scientist Noam Brown talks with AI Deep Dive host Rocket Drew about AI agents, reinforcement learning and what happens when increasingly capable agents begin reasoning, delegating and coordinating with each other. They dig into the infamous Hugging Face incident, the limits of AI “research taste,” OpenAI’s push toward AI systems that can help improve the next generation of AI, and the growing challenge of monitoring models’ chains of thought as they become more capable.Related articles:Investigation into OpenAI Hugging Face Hack: https://www.theinformation.com/briefings/sen-josh-hawley-launches-investigation-openai-hugging-face-hackThe Hugging Face Hack’s Chilling Postmortem: https://www.theinformation.com/newsletters/the-weekend/hugging-face-hacks-chilling-postmortemSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaFollow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/Chapters:00:00 - Introduction and Overview of AI Agents04:08 - Reasoning and Reinforcement Learning08:50 - Current Limitations, Astra, and Impact on Work25:50 - Multi-Agent Systems and Delegation31:04 - The Hugging Face Incident and Agent Coordination36:05 - Game Theory and Agent Security40:39 - AI Alignment, Monitoring, and Chain of Thought50:07 - Future Outlook on AI Models and Interpretability

