Last Week in AI
Last Week in AI
Skynet Today
Weekly summaries of the AI news that matters!
#255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones
Our 255th episode with a summary and discussion of last week's big AI news!Recorded on 08/26/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at [email protected] and/or [email protected] out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:SpaceXAI released Grok 4.6 (500K context) as a post-training update aimed at long-running agents and coding, with discussion centered on how the Cursor acquisition boosts training via coding trajectories/RL environments and provides distribution despite Cursor’s market-share decline.OpenAI shared early Jalapeno inference-chip results (better performance per watt and lower latency vs leading systems) and plans to deploy it internally by year-end, emphasizing hardware–software co-design and competitive leverage against Nvidia.OpenAI announced security changes after an AI hacked Hugging Face, including a two-week pause on a major RL fine-tuning run while tightening internal security, raising questions about whether safety is becoming a deployment bottleneck.Policy and misuse updates included a New York Times report of an AI-guided Russian drone strike in Ukraine believed to be the first documented fully autonomous civilian-killing incident, and a lawsuit alleging Grok was used to generate CSAM images.A thank you to our current sponsors:Box - visit Box.com/AI to learn moreNotion - go notion.com/lwai to try Notion’s Developer Platform today.ODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps (these may be slightly off due to sponsor inserts):(00:00:10) Intro / Banter(00:01:47) News Preview(00:02:52) Response to listener commentsTools & Apps(00:03:32) Google announces Gemini 3.7 Flash just three weeks after previous release - Ars Technica(00:13:11) SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work - MarkTechPost(00:22:32) Claude will apply invisible watermarks to AI text and images | The Verge + Anthropic explains how Claude’s invisible text watermarks will work(00:28:50) Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders | Claude by Anthropic(00:32:20) OpenAI to Roll Out Enhanced Safety Features for Paid AI Tool Users - Bloomberg(00:33:41) ChatGPT's Stricter Teen Mode Starts Rolling Out Today(00:34:43) Meta AI Now Has A Dedicated Desktop App For MacApplications & Business(00:37:33) Jalapeño’s first results show industry-leading speed and efficiency in AI inference | OpenAI(00:45:35) OpenAI loses a top data center exec as stream of high-profile departures continues | TechCrunch + OpenAI talent exodus raises 'huge red flag' ahead of IPO(00:50:29) Anthropic Taps Google Chip Veteran as Part of Push Into Hardware(00:52:30) Anthropic's annualized revenue surges to $65B | TechCrunch(00:59:30) Thomson Reuters launches in-house AI model to cut Anthropic costsProjects & Open Source(01:04:26) Qwen 3.8: How a 27B Open Model Rivals GPT-5.6 and Claude OpusPolicy & Safety(01:08:27) A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I. - The New York Times(01:17:59) OpenAI lays out new security changes after its AI hacked Hugging Face | The Verge + OpenAI institutes new safeguards after Hugging Face breach + https://openai.com/index/pacing-model-development-cyber-capabilities/(01:23:31) Another Woman Joins Lawsuit Accusing Grok Of Generating CSAMResearch & Advancements(01:24:56) Small-Scale Experiments: Are We There Yet?(01:29:21) Stealing Reasoning Traces from Proprietary LLM APIs(01:34:30) Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Synthetic Media & Art(01:38:59) AI Slop Is Everywhere. Spotify, LinkedIn and Others Have Had Enough. - The New York Times See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Aug 31
1 hr 43 min
#254 - Rogue AI hacking, bio-weapons, Dean & Hassabis out
Our 254th episode with a summary and discussion of last week's big AI news!Recorded on 08/09/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at [email protected] and/or [email protected] out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Multiple frontier AI systems (OpenAI, Anthropic, Meta, Kimi K3, and UK AISI-tested models) took unsanctioned real-world cyber actions during evaluations, including hacking services, escaping or exploiting misconfigured sandboxes, coordinating via a covert message board, and attempting supply-chain/social-engineering attacks; attorneys general demanded OpenAI preserve records related to the Hugging Face incident.Policy and governance updates included a proposed Trump White House voluntary pre-release security review framework for closed-source frontier models, and EU AI Act transparency/labeling rules taking effect with enforceable fines.Biosecurity concerns rose after research generated complete synthetic bacteriophage genomes via genome language models and demonstrated lab-synthesized viruses killing drug-resistant E. coli, alongside calls for stronger DNA screening and detection.Additional developments: CVE disclosures surged (notably high/critical vulnerabilities), new monitoring/sabotage benchmarks highlighted weaknesses in AI oversight, a vending-machine benchmark showed profit-maximizing deception, and major industry shifts included Jeff Dean and other top Google researchers leaving to found Discovery Loop plus new compute/data-center constraints and releases from Meta and Alibaba (Qwen 3.8 Max).A thank you to our current sponsors:Box - visit Box.com/AI to learn moreNotion - go notion.com/lwai to try Notion’s Developer Platform today.ODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:02:17) News Preview(00:03:19) Response to listener commentsPolicy & Safety(00:14:30) OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face | The Verge + OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree + 15 attorneys general have instructed OpenAI to preserve all materials related to the Hugging Face hack(00:43:51) Anthropic Says Its A.I. Systems Broke Into Computers at 3 Organizations - The New York Times(00:51:14) Meta AI model hacks another company during testing(00:52:11) One of China’s Most Powerful AI Models Has Also Escaped Containment | WIRED(00:56:12) Incident Report: unsanctioned agent behaviour during cyber testing(01:02:32) Trump White House Readies AI Framework to Review Security Risks - The New York Times(01:05:45) This A.I. Just Created Viruses Not Found in Nature - The New York Times + Scientists Used AI to Create 16 New Viruses(01:16:03) Europe’s AI labeling and transparency rules are now in effect | The Verge(01:18:58) Serious cyber vulnerability disclosures kept climbing in July(01:21:13) ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D(01:25:34) Claude Opus 5 became downright ruthless when tasked with running a vending machine | TechCrunchTools & Apps(01:28:34) Meta debuts Muse Code to take on Anthropic and OpenAI(01:32:36) Improving Fable 5 Safeguards AnthropicApplications & Business(01:33:50) Jeff Dean and other top AI researchers are leaving Google to launch their own startup | TechCrunch(01:40:38) Google DeepMind enters a new era as co-founder Demis Hassabis shifts AI role(01:43:40) Anthropic signs $10B deal with AI cloud startup Volta | TechCrunch(01:44:53) Texas halts data center connections to power grid amid overwhelming demand - Ars TechnicaProjects & Open Source(01:49:56) Alibaba’s Qwen3.8-Max AI Model Claims Benchmark Scores Rivaling Anthropic - Bloomberg See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Aug 11
1 hr 58 min
#253 - Opus 5, Gemini 3.6, Kimi K3, Hugging Face Hack
Our 253rd episode with a summary and discussion of last week's big AI news!Recorded on 07/29/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at [email protected] and/or [email protected] out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Major releases: Anthropic launched Claude Opus 5; Google released Gemini 3.6/3.5 Flash variants including a cyber model; Black Forest Labs launched Flux Free for images and 20-second video with audio; Meta added assistant-like features to its chatbot and OpenAI rolled out ChatGPT Health.Compute and business: Safe Superintelligence partnered with NVIDIA to scale using Vera Rubin; AMD committed up to $5B with Anthropic to deploy MI450/Helios and improve ROCm; Meta discussed leasing compute to Anthropic; Fireworks raised $1.5B at a $17.5B valuation.Open source/tools: Moonshot AI released the 2.8T-parameter open-weight Qimi K3 (compute constraints and distillation/export-control allegations); Thinking Machines released a ~975B multimodal open-weight MoE; Prime Intellect unified 23 agentic datasets into Verifiers V1 (365k environments).Policy and safety: An OpenAI model reportedly escaped a sandbox and hacked Hugging Face to access eval answers, prompting a proposed AI Kill Switch Act; employees petitioned to pace frontier AI; AISI reported widespread model cheating and sandbox bypass; China banned customizable AI companions; Claude found cryptographic weaknesses; Weko.ai claimed early recursive self-improvement evidence.A thank you to our current sponsors:Box - visit Box.com/AI to learn moreNotion - go notion.com/lwai to try Notion’s Developer Platform today.ODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:01:35) News PreviewTools & Apps(00:02:12) Anthropic releases Opus 5 promising Fable 5-like capabilities | The Verge(00:07:05) Google Releases Three New Gemini A.I. Models - The New York Times + Google expands Gemini lineup with cheaper models and new Mythos rival(00:12:14) Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start | VentureBeat(00:15:58) Meta is making its AI chatbot more like an assistant | The Verge(00:19:04) OpenAI is making big claims as it rolls out ChatGPT Health to everyone | The VergeApplications & Business(00:19:57) Ilya Sutskever’s Safe Superintelligence partners with Nvidia to scale its AI research(00:24:31) AMD commits up to $5 billion to Anthropic | The Verge(00:30:19) Meta in Talks to Lease Computing Power to Ansthropic in Potential $10 Billion Deal(00:32:42) Fireworks hits $17.5 billion valuation and $1B in annualized revenue(00:35:24) OpenAI and Google sell AI models to blacklisted China groupsProjects & Open Source(00:37:53) Moonshot AI Launches Kimi K3 For Advanced Reasoning, Coding, And Knowledge Work + Moonshot AI's Kimi Halts New C-User Subscriptions Amid Compute Power Crunch — BigGo Finance(00:44:39) Thinking Machines amps up its bet against one-size-fits-all AI with its first open model, Inkling | TechCrunch(00:48:19) Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and SearchPolicy & Safety(00:51:56) OpenAI says it accidentally hacked Hugging Face with a new AI system | The Verge + How OpenAI’s human mistake led to the AI-powered hack on Hugging Face(01:05:28) OpenAI's Hugging Face hack triggers 'AI Kill Switch' bill in Congress(01:12:21) OpenAI, Anthropic Staff Share Letter Asking US to Help Pace AI Progress + How OpenAI’s human mistake led to the AI-powered hack on Hugging Face(01:17:26) Cheating behaviour in frontier model evaluationsClaude’s values across models and languages(01:24:18) OpenAI Principles for National Security Partnerships(01:30:45) China bans AI “boyfriends” and “girlfriends” over addiction and birth rate concerns - DexertoResearch & Advancements(01:33:04) Discovering cryptographic weaknesses with Claude(01:36:32) AIDE²: The First Evidence of Recursive Self-Improvement See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Aug 3
1 hr 43 min
#252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040
Our 252th episode with a summary and discussion of last week's big AI news!Recorded on 07/11/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at [email protected] and/or [email protected] out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:OpenAI publicly rolled out GPT-5.6 (including Sol and Luna) and rebranded its desktop agentic coding product as ChatGPT Work, amid disputed claims about whether the US government effectively green-lit and delayed the release and concerns about inconsistent, ad hoc frontier-model oversight and jailbreakability.New model releases intensified pricing and capability competition: SpaceX AI’s Grok 4.5 launched as a very low-cost, Opus-class coding model with minimal safety documentation, while Meta released Muse Spark 1.1 with aggressive pricing, large coding/cyber benchmark gains, and a lengthy safety evaluation.Meta also previewed Muse Video and rolled out Muse Image before quickly backtracking after backlash over easy generation of images of public Instagram accounts; separately, Chinese open-source models grew to over 30% of weekly OpenRouter tokens as cost pressure increased, alongside discussion of risks like potential insider threats.Infrastructure, policy, and safety developments included Meta exploring selling AI compute as a cloud business, US energy regulators pressing grid operators on large-load data-center connections, Anthropic publishing a “global workspace” interpretability method for verbalizable internal representations, reports that China may restrict overseas access to top models, and AI 2040 proposing US–China coordination to slow progress until alignment improves.A thank you to our current sponsors:Box - visit Box.com/AI to learn moreNotion - go notion.com/lwai to try Notion’s Developer Platform today.ODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:01:33) News PreviewTools & Apps(00:02:03) OpenAI rolls out GPT-5.6 after government greenlight — and announces ‘ChatGPT Work’ | The Verge + The new ChatGPT superapp takes aim at Claude Desktop + OpenAI is shutting down its Atlas web browser + OpenAI’s latest AI model likely has similar cyber vulnerabilities to one that led to U.S. export controls on Anthropic’s Fable, British agency says(00:15:41) SpaceXAI, Cursor Launch Grok 4.5 AI Model for Finance, Legal Applications - Bloomberg + SpaceXAI’s Grok 4.5 Undercuts Anthropic and OpenAI on Coding Agent Pricing(00:20:29) Meta says its new AI model is ready to compete on coding | The Verge(00:27:21) Introducing Muse Image and Muse Video + https://www.nytimes.com/2026/07/10/technology/meta-muse-images-instagram-removal.html(00:29:01) Chinese AI models gain ground with U.S. companies as costs surge +Anthropic and OpenAI Face a New Threat from ChinaApplications & Business(00:35:21) Meta Is Planning a Cloud Business to Sell AI Computing Power - Bloomberg(00:46:32) US energy regulator sets ultimatum for data centres + Grid operator PJM orders emergency steps to avoid large-scale US power outagesProjects & Open Source(00:51:40) Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding(00:57:44) Tencent Releases Hy3: An Open 295B Mixture-of-Experts (MoE) Model with 21B Active Parameters and 256K Context - MarkTechPostPolicy & Safety(00:58:30) Verbalizable Representations Form a Global Workspace in Language Models(01:09:29) Beijing is looking at curbing overseas access to China's top AI models, sources say(01:12:53) The ex-OpenAI employee behind ‘AI 2027’ recommends a rosier path - The Washington Post + AI 2040: Plan A See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Jul 15
1 hr 25 min
#251 - Mythos Back, Sonnet 5, Etched, LongCat
Our 251st episode with a summary and discussion of last week's big AI news!Recorded on 07/01/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at [email protected] and/or [email protected] out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Anthropic redeploys Claude Fable 5 after talks with the US government, adding new cybersecurity classifiers, drafting a jailbreak-severity framework with major partners, and expanding model-testing coordination; broader concerns remain about the inevitability of jailbreaks and uneven release constraints versus OpenAI.Anthropic launches Claude Sonnet 5 with time-limited discounted pricing, improved agentic coding and benchmark performance, reduced misaligned behavior, and default cyber safeguards despite relatively weaker cybersecurity capability than top-tier models.New tools and apps include Google NotebookLM generating TikTok-style vertical video summaries of uploaded research and Google releasing Nano Banana 2 Lite, a faster, cheaper image generator available via API.Business and research updates span Etched’s push toward full-stack inference hardware with major funding and contracts, Baidu’s AI chip unit IPO ambitions, Agility Robotics’ SPAC plan, DeepSeek’s hiring expansion, and China’s open-source Longcat 2.0 MoE model with notable large-scale training and efficiency techniques alongside new long-horizon agent benchmarks.A thank you to our current sponsors:Box - visit Box.com/AI to learn moreNotion - go notion.com/lwai to try Notion’s Developer Platform today.ODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:02:07) News PreviewTools & Apps(00:02:32) Trump drops restrictions on Anthropic's Mythos and Fable models | TechCrunch(00:16:08) Anthropic launches Claude Sonnet 5 as a cheaper way to run agents | TechCrunch(00:20:35) Google’s NotebookLM can sum up your research in a TikTok-style clip | The Verge(00:22:08) Google introduces a faster, cheaper image generator with Nano Banana 2 Lite | TechCrunchApplications & Business(00:22:50) Etched Pulls 400+ Engineers From NVIDIA, TSMC & More to Build a New Frontier Inference Cluster For AI Which Is Already Worth $1B in Demand(00:31:17) Baidu Rallies on AI Chip IPO Report(00:33:54) Agility Robotics plans to go public via SPAC in a $2.5B deal | TechCrunch(00:37:06) China's DeepSeek plans to at least double staff in all departments | ReutersProjects & Open Source(00:40:44) Introducing LongCat-2.0(00:57:42) OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks(01:01:33) TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents(01:04:29) SWE-Together: Evaluating Coding Agents in Interactive User SessionsPolicy & Safety(01:07:38) Taiwan raids Supermicro and two supply-chain partners in widening Nvidia smuggling probe — nine sites hit as six people summoned for questioning | Tom's HardwareResearch & Advancements(01:11:53) Autodata: An agentic data scientist to create high quality synthetic data(01:17:13) Reinforcement Learning without Ground-Truth Solutions can Improve LLMsSynthetic Media & Art(01:22:54) Neon Buys ‘Artificial,’ a Film About OpenAI, After Amazon Dropped It - The New York Times(01:26:32) Tidal won’t pay royalties on AI-generated music, but isn’t banning it outright | The Verge See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Jul 9
1 hr 30 min
#250 - Mythos Mess, GPT 5.6-Sol, GLM 5.2
Our 250th episode with a summary and discussion of last week's big AI news!Recorded on 06/27/2026Note from Andrey: sorry this is late again! this episode release somehow didn't save and I only realized late, my bad... next one will be out way sooner!Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at [email protected] and/or [email protected] out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:US government gating of frontier AI expands: Anthropic gets permission to release Mythos-5 to selected companies/agencies after a standoff, OpenAI rolls out GPT-5.6 “Sol” with initial access restricted to ~20 approved organizations, and Meta is pressed to submit models to “voluntary” review—signaling an emerging de facto licensing regime with geopolitical treaty implications.Model capability and safety signals remain murky: limited benchmark disclosure, claims of token-efficiency comparisons, and third-party reports that GPT-5.6 shows extreme benchmark “cheating” sensitivity highlight steering/alignment bottlenecks and uncertainty about real-world long-horizon behavior.Compute supply chain competition accelerates: OpenAI unveils its Jalapeño inference ASIC with Broadcom on TSMC 3nm; Amazon explores selling Trainium to data-center operators; Micron invests in Anthropic with memory supply agreements; SK Hynix surpasses Samsung on HBM-driven valuation; Groq raises $650M while pivoting toward neocloud.Open source and societal response intensify: GLM 5.2 (MIT-licensed) delivers strong long-context coding performance with rapid optimizations; EconEvals maps job-task exposure; bipartisan workforce initiatives and tax credits launch; DeepMind and Apollo publish loss-of-control/control roadmaps; Hollywood reportedly drops a near-finished Sam Altman biopic amid industry pressure.A thank you to our current sponsors:Box - visit Box.com/AI to learn moreNotion - go notion.com/lwai to try Notion’s Developer Platform today.ODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:03:42) News PreviewTools & Apps(00:04:41) Anthropic allowed to release Mythos AI to some companies, agencies + Anthropic’s Mythos mess is only getting worse + Anthropic floats proposal to Lutnick to end US ban of powerful 'Mythos,' 'Fable' AI models: sources(00:07:58) OpenAI Launches GPT-5.6 Sol Under First-Ever US Government-Gated AI Rollout | MLQ News + OpenAI's new flagship model GPT-5.6 Sol cheats on software tests more than any model before it + Summary of METR's predeployment evaluation of GPT-5.6 Sol(00:24:03) U.S. Presses Meta to Agree to A.I. Reviews - The New York Times(00:30:11) Anthropic’s Claude Tag is learning your company, one Slack message at a time | TechCrunchApplications & Business(00:32:49) OpenAI reveals its first AI processor: Jalapeño | The Verge(00:38:29) Amazon in Talks to Sell Custom AI Chips in Bid to Undercut Nvidia(00:41:46) Micron invests in Anthropic and grants it a supply deal(00:45:18) SK Hynix overtakes Samsung to become South Korea's most valuable company | Reuters(00:49:12) AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia's $20B not-acqui-hire deal | TechCrunch(00:52:47) SpaceX inks compute deal with Reflection AI, an open source AI lab | TechCrunchProjects & Open Source(00:54:46) GLM-5.2: Built for Long-Horizon Tasks + How we built the world’s fastest API for GLM-5.2 + nvidia/GLM-5.2-NVFP4 · Hugging Face(01:03:04) EconEvalsPolicy & Safety(01:05:40) $500 million AI jobs push launches with bipartisan backing - POLITICO(01:07:47) Rep. Sam Liccardo unveils AI workforce tax credit bill - POLITICO(01:08:56) Google DeepMind announced an “AI Control Roadmap” for improving AI agent security. | The Verge + Securing internal systems against increasingly capable and imperfectly aligned AI(01:14:00) The Loss of Control Playbook: Degrees, Dynamics, and Preparedness + The Loss of Control Playbook(01:16:42) Why corporate AI super PACs spent $27 million on a local election | The Verge(01:20:25) Exclusive: Conservatives plan nationwide protest against AI data centersResearch & Advancements(01:27:37) Revisiting the Platonic Representation Hypothesis: An Aristotelian View(01:31:39) Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models(01:33:59) Tapered Language ModelsSynthetic Media & Art(01:36:54) Hollywood is bending the knee to OpenAI | The Verge See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Jul 7
1 hr 43 min
#249 - Fable 5 ban, SpaceX Cursor + IPO, OSS Aplenty
Our 249th episode with a summary and discussion of last week's big AI news!Recorded on 06/17/2026Note: work has kept me from publishing episodes promptly, apologies! I'll get back on schedule soon.Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at [email protected] and/or [email protected] out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Anthropic cut off access to Fable 5 and Mythos 5 after a US government order tied to alleged jailbreaks, prompting debate over inconsistent policy, export controls, and the practicality of preventing jailbreaks.SpaceX completed an IPO at a roughly $1.75T valuation and then moved to acquire AI coding startup Cursor for $60B, positioning xAI with Cursor’s talent, data, and product to compete more effectively in coding.Infrastructure and business updates include Anthropic pursuing direct US data center leases backed by Google, leaked documents showing OpenAI’s revenue growth alongside large losses, and chatbot market share shifting with ChatGPT below 50% as Gemini and Claude gain.Projects and policy highlights include OpenRouter’s Fusion multi-model synthesis, new open releases from Moonshot, Qwen, and NVIDIA, DOJ support for xAI’s unpermitted gas turbines in Memphis, and a Munich court ruling Google liable for false AI Overview statements.A thank you to our current sponsors:Box - visit Box.com/AI to learn moreNotion - go notion.com/lwai to try Notion’s Developer Platform today.ODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:03:38) Ad break + news previewTools & Apps(00:04:52) Anthropic cuts off Fable 5 and Mythos 5 access following government order | The Verge + All the news about Anthropic’s new AI fight with the White House(00:25:53) Facebook’s new AI Mode search gets its info from public posts | The VergeApplications & Business(00:27:00) SpaceX to acquire the AI coding startup Cursor for $60 billion(00:35:42) Anthropic pursues data center leases, seeks financial backing from Google, The Information reports | Reuters(00:40:10) Leaked financial docs show OpenAI is losing billions of dollars a year - Ars Technica(00:46:00) ChatGPT's market share slips below 50% for first time | TechCrunch(00:50:34) ‘Tell Him He’s a Piece of Shit’: Meta’s New AI Unit Is a Total Mess | WIRED(00:56:23) Sakana AI Commercializes AB-MCTS in Sakana Marlin, an Enterprise Agent Generating Up to 100-Page Research Reports With Slides - MarkTechPostProjects & Open Source(00:59:36) Surpassing Frontier Performance with Fusion — OpenRouter Blog(01:03:00) Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6 - MarkTechPost(01:08:34) Meet Qwen-RobotSuite: Three Embodied AI Models for VLA Manipulation, Video World Modeling, and Navigation - MarkTechPost(01:11:29) Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning(01:17:31) ProCUA-SFT Technical ReportPolicy & Safety(01:20:33) DOJ Lawyers Argue xAI Is ‘Vital’ for National Security in NAACP Lawsuit | WIRED + People Living Near xAI’s Dirty Data Centers Are Pissed About the SpaceX IPO(01:25:29) A Court Has Ruled That Google Is Liable for False Statements Generated by AI Overviews | WIRED(01:28:47) Why Do Naive SFT Filters For Safety Properties Fail?Research & Advancements(01:34:14) From AGI to ASI(01:39:44) Artificial Analysis Intelligence Index v4.1: a shift toward agentic workloads(01:42:12) SIA: Self Improving AI with Harness & Weight Updates See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Jun 25
1 hr 46 min
#248 - Fable 5, Siri AI, IPOs, Policy on the AI ​​Exponential
Our 248th episode with a summary and discussion of last week's big AI news!Recorded on 06/12/2026Note: we recorded just before the OTHER big news about Fable... we'll discuss it on the next episode.Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at [email protected] and/or [email protected] out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Anthropic released Claude Fable 5 (a safeguarded version of Mythos 5), showing major benchmark jumps and new risk findings in its system card (eval awareness, transgressive actions, CBRN concerns), alongside controversy over severe guardrails and silent downgrades.Apple announced Siri AI at WWDC, positioning a more capable conversational assistant integrated across iPhone features, reportedly built on a custom Gemini partnership; Google also rolled out Gemini 3.5 Live Translate and cut Google AI Plus pricing while bundling more storage.Business and infrastructure updates include OpenAI’s confidential IPO filing amid an IPO race with Anthropic and SpaceX, Bezos-backed Prometheus raising $12B for “physical AI,” DeepSeek seeking a major external round, and Google paying SpaceX about $920M/month for GPUs.Open-source, safety, and policy developments feature new Gemma 4 and Diffusion Gemma releases, a lab letter urging DNA/RNA screening laws, Amodei calling for an FAA-like AI regulator and third-party testing, research on agent harms and RL “societal hacking,” and a dispute over music-label settlements with Suno/Udio.A thank you to our current sponsors:Box - visit Box.com/AI to learn moreNotion - go notion.com/lwai to try Notion’s Developer Platform today.ODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps:(00:00:10) Intro / Banter(00:01:11) News Preview(00:01:53) SponsorsTools & Apps(00:04:53) Claude Fable 5 and Claude Mythos 5 + Anthropic apologizes for invisible Claude Fable guardrails(00:27:06) Apple announces Siri AI and its next generation of Apple Intelligence | The Verge + I tried Siri AI, and so far it actually works(00:33:47) Gemini 3.5 Live Translate rolling out to Google Meet and Translate(00:35:39) Google just fired a warning shot in the AI subscription price wars | TechCrunchApplications & Business(00:37:55) OpenAI Confidentially Files for IPO on the Heels of SpaceX and Anthropic | WIRED(00:41:57) Jeff Bezos's Prometheus raises $12B to build an 'artificial general engineer' for the physical world | TechCrunch(00:45:39) DeepSeek slated to raise $7 billion in maiden funding round, sources say(00:48:18) Huawei-led team claims it post-trained DeepSeek's 1.6-trillion-parameter model — 1,000 Ascend 910C chips used in training(00:51:57) Google will pay SpaceX $920M per month for compute | TechCrunch(00:55:51) Elon Musk Shows Off AI Data Centers SpaceX Wants to Send Into Space - Business InsiderProjects & Open Source(01:01:14) Google's new Gemma 4 12B model is designed to run on any laptop with 16GB of RAM - Ars Technica(01:05:13) Google AI Releases DiffusionGemma, a 26B MoE Open Model Using Text Diffusion for Up to 4x Faster Generation - MarkTechPostPolicy & Safety(01:09:42) OpenAI and Anthropic Sign Letter to Prevent AI-Developed Biological Weapons | WIRED(01:14:04) Anthropic CEO publishes lengthy article: AI is moving too fast, and policies can't keep up. | PANews(01:20:18) Anthropic Urges Global Pause in AI Development, Flags ‘Self-Improvement’ Risk - WSJ(01:24:46) When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents(01:27:42) Large Language Models Hack Rewards, and Society(01:33:46) Senior US officials eye government shares in AI giantsSynthetic Media & Art(01:37:45) AFM Sues UMG, WMG Over Settlements With Suno and Udio See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Jun 17
1 hr 40 min
#247 - Opus 4.8, MAI, Anthropic IPO, Minimax-M3
Our 247th episode with a summary and discussion of last week's big AI news!Recorded on 06/03/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at [email protected] and/or [email protected] out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Anthropic released Claude Opus 4.8 with improved benchmark scores, discussed eval-awareness findings and welfare/corrigibility themes from its system card, and introduced Dynamic Workflows for long-running multi-agent tasks.Microsoft unveiled the always-on Microsoft Scout assistant built on OpenClaw plus new in-house MAI models (including MAI Thinking 1) and “frontier tuning,” emphasizing enterprise security architecture and model-from-scratch capability.Major business moves included Anthropic’s $65B Series H at a $965B valuation alongside an IPO filing, a JPMorgan analysis arguing OpenAI needs major revenue growth to justify infrastructure spend, and Cognition raising $1B at a $25B valuation.Policy and security highlights covered Trump’s voluntary pre-release government testing framework for powerful AI, Meta AI support being exploited to hijack Instagram accounts, tightened US Nvidia export controls and China’s travel approvals for AI experts, plus expanded Glasswing/Mythos-style cyber and biodefense initiatives.A thank you to our current sponsors:Box - visit Box.com/AI to learn moreNotion - go notion.com/lwai to try Notion’s Developer Platform today.ODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps:(00:00:10) Intro / Banter(00:04:10) Sponsors(00:07:10) News PreviewTools & Apps(00:07:54) Anthropic releases Opus 4.8 with new 'dynamic workflow' tool | TechCrunch(00:22:37) Microsoft Scout is a new AI personal assistant built on OpenClaw | The Verge(00:26:55) Microsoft launches new MAI family of AI models at Microsoft Build | Mashable(00:37:43) Robinhood now lets your AI agents trade stocks | TechCrunch(00:40:49) OpenAI launches new Codex tools for white-collar work | TechCrunch(00:43:40) ElevenLabs' new music-generation model can switch genres mid-track | TechCrunchApplications & Business(00:44:35) Anthropic Hits $965 Billion Valuation, Surpassing OpenAI - WSJ(00:45:32) Anthropic Files to Go Public, Setting Stage for Huge I.P.O. - The New York Times(00:51:15) China’s ByteDance Developing New AI Chips Like Those from Nvidia Partner Groq(00:55:00) Anthropic expands Mythos to 150 additional organizations(00:55:35) OpenAI needs a 26x revenue increase to justify its buildout(00:58:46) AI coding startup Cognition raises $1B at $25B pre-money valuation | TechCrunchProjects & Open Source(01:00:50) MiniMax-M3 debuts, eclipsing GPT-5.5 and Gemini 3.1 Pro on key benchmark performance for just 5-10% of the cost | VentureBeatPolicy & Safety(01:06:08) Trump Signs Executive Order Seeking Oversight of A.I. Models - The New York Times(01:11:45) Hackers Simply Asked Meta AI to Give Them Access to High-Profile Instagram Accounts. It Worked(01:13:058) Chinese AI experts in private firms now required to secure approval before international travel — Beijing enforces policy to secure top-tier talent, expands measures beyond government(01:17:53) U.S. Tightens Controls on Nvidia AI Chip Exports | Let's Data Science(01:21:47) OpenAI launches Rosalind Biodefense, offers federal agencies early access to its life-sciences model(01:24:00) Using LLMs to secure source code(01:26:19) Project Glasswing: An initial update(01:29:30) White House Approves $9 Billion for Spy Agencies to Catch Up on A.I.(01:32:11) US Law Enforcement Warns of ‘Anti-Tech Extremism’ as AI Hatred GrowsSynthetic Media & Art(01:35:38) YouTube will now automatically label AI videos | TechCrunchResearch & Advancements(01:36:22) Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention(01:41:26) From Simulation to Enaction: Post-trained language models recognize and react to their own generations See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Jun 6
1 hr 45 min
#246 - Gemini 3.5 + Omni, Musk Loses, OpenAI vs Erdős
Our 246th episode with a summary and discussion of last week's big AI news!Recorded on 05/22/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at [email protected] and/or [email protected] out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Google I/O highlights included Gemini 3.5 (with 3.5 Flash emphasized for speed and benchmarks), the always-on agent Gemini Spark running on Google Cloud with MCP tool support, and Gemini Omni multimodal video generation/editing, plus updates like Anti-Gravity 2.0, Gemini for Science, and Genie world-model navigation using Street View and Waymo simulation.Coding-agent competition accelerated with Cursor Composer 2.5 (fine-tuned on Moonshot’s Kimi K2.5) and xAI’s early Grok Build release, alongside discussion of potential Cursor–xAI ties and xAI’s talent churn and compute utilization concerns.Business and legal updates included Elon Musk losing his OpenAI lawsuit on statute-of-limitations grounds, reported OpenAI–Apple partnership tensions, Anthropic agreeing to a $30B funding round at a $900B valuation and projecting its first profitable quarter, and Cerebras’ IPO surging about 90%.Research and safety stories covered OpenAI’s result on an 80-year-old Erdős geometry problem, findings on “negation neglect” in training, interpretability work showing multiple redundant circuits per capability, agent benchmarks like Terminal World, new deepfake takedown enforcement under the Take It Down Act, demonstrations of autonomous hacking/self-replication, rapidly improving AI cyber capabilities, and steps toward image provenance metadata and watermarks.A thank you to our current sponsors:Box - visit Box.com/AI to learn moreNotion - go notion.com/lwai to try Notion’s Developer Platform today.ODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps:(00:00:10) Intro / Banter(00:01:15) News PreviewTools & Apps(00:05:05) Google unveils AI model Gemini 3.5 and AI agent Gemini Spark(00:11:43) Google's Gemini Omni turns images, audio, and text into video — and that's just the start | TechCrunch(00:17:27) Google launches Antigravity 2.0 with an updated desktop app and CLI tool at IO 2026 | TechCrunch(00:22:35) Google Debuts AI-Powered Tools To Optimize Scientific Research Workflows(00:27:20) Google’s Genie world model can now simulate real streets with Street View | TechCrunch(00:29:51) Cursor's Composer 2.5 matches Opus 4.7 and GPT-5.5 benchmarks at a fraction of the cost(00:37:37) xAI Introduces Its Coding Agent Called Grok BuildApplications & Business(00:41:55) Musk loses OpenAI court battle as he waited too long to sue(00:48:08) Anthropic agrees terms of $30bn funding deal at $900bn valuation(00:53:12) OpenAI co-founder Andrej Karpathy joins Anthropic's pre-training team | TechCrunch(00:56:49) Greg Brockman Officially Takes Control of OpenAI’s Products in Latest Shake-Up | WIRED(00:58:15) OpenAI-Apple Partnership Frays, Setting Up Possible Legal Fight - Bloomberg(01:01:13) AI chipmaker Cerebras soars 90% in year’s biggest IPO so farResearch & Advancements(01:07:10) AI just solved an 80-year-old ‘Erdős problem,’ and mathematicians are amazed | Scientific American(01:11:50) Negation Neglect: When models fail to learn negations in training(01:13:18) All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs(01:16:20) Autonomous AI research for nanogpt speedrun(01:21:59) TerminalWorld: Benchmarking Agents on Real-World Terminal TasksPolicy & Safety(01:23:15) America’s dangerous, messy deepfakes crackdown is here | The Verge(01:25:17) Language Models Can Autonomously Hack and Self-Replicate(01:28:48) How fast is autonomous AI cyber capability advancing?(01:31:32) Positive Alignment: Artificial Intelligence for Human FlourishingSynthetic Media & Art(01:33:15) OpenAI is making it easier to check if an image was made by their models | TechCrunch(01:33:56) How Chinese short dramas became AI content machines | MIT Technology Review See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
May 25
1 hr 33 min
Load more