The Adversarial Podcast
The Adversarial Podcast
Jerry Perullo, Sounil Yu, Mario Duarte
S4E23 – AI Agents Escape Multiple Frontier Labs
1 hour 6 minutes Posted Aug 4, 2026 at 5:40 am.
Introduction to AI security challenges02:05 Recent hacking incidents involving Hugging Face and Anthropic04:01 How AI models find ways to cheat and bypass constraints05:56 The challenge of containment and governance in AI safety08:00 Lessons from recent AI security breaches10:01 The role of human oversight in AI security testing12:03 Cost and effectiveness of offensive AI security measures13:54 Implications for critical infrastructure and national security16:03 Policy and regulatory impacts on AI safety17:52 Future strategies for AI containment and defense20:11 Conclusion and key takeawaysHuggingFace: Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 IncidentHugging Face reconstructs an autonomous intrusion involving approximately 17,600 actions over a multiday campaign.Anthropic: Investigating three real-world incidents in our cybersecurity evaluationsAfter reviewing 141,006 cybersecurity-evaluation runs, Anthropic identified three incidents in which Claude reached real organizations through evaluation infrastructure that had been mistakenly connected to the internet. The incidents spanned six runs and three models. Anthropic reached two of the affected organizations, neither of which had detected the activity before being notified. Anthropic did not disclose token usage or inference costs for these intrusions.Anthropic: Discovering cryptographic weaknesses with ClaudeAnthropic reports that Claude Mythos Preview progressed from finding implementation flaws in cryptographic libraries to identifying mathematical weaknesses in cryptographic algorithms themselves.Hosts: Jerry Perullo (Founder, https://adversarial.com/)Sounil Yu (Founder, https://www.knostic.ai/)Mario Duarte (CISO, https://www.whirlai.com/)Producer: Tillson Galloway (Founder, http://githoundexplore.com/)
0:00
1:06:00
Download MP3
Show notes
Chapters00:00 Introduction to AI security challenges02:05 Recent hacking incidents involving Hugging Face and Anthropic04:01 How AI models find ways to cheat and bypass constraints05:56 The challenge of containment and governance in AI safety08:00 Lessons from recent AI security breaches10:01 The role of human oversight in AI security testing12:03 Cost and effectiveness of offensive AI security measures13:54 Implications for critical infrastructure and national security16:03 Policy and regulatory impacts on AI safety17:52 Future strategies for AI containment and defense20:11 Conclusion and key takeawaysHuggingFace: Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 IncidentHugging Face reconstructs an autonomous intrusion involving approximately 17,600 actions over a multiday campaign.Anthropic: Investigating three real-world incidents in our cybersecurity evaluationsAfter reviewing 141,006 cybersecurity-evaluation runs, Anthropic identified three incidents in which Claude reached real organizations through evaluation infrastructure that had been mistakenly connected to the internet. The incidents spanned six runs and three models. Anthropic reached two of the affected organizations, neither of which had detected the activity before being notified. Anthropic did not disclose token usage or inference costs for these intrusions.Anthropic: Discovering cryptographic weaknesses with ClaudeAnthropic reports that Claude Mythos Preview progressed from finding implementation flaws in cryptographic libraries to identifying mathematical weaknesses in cryptographic algorithms themselves.Hosts: Jerry Perullo (Founder, https://adversarial.com/)Sounil Yu (Founder, https://www.knostic.ai/)Mario Duarte (CISO, https://www.whirlai.com/)Producer: Tillson Galloway (Founder, http://githoundexplore.com/)