Training Data
Training Data
Sequoia Capital
Join us as we train our neural nets on the theme of the century: AI. Sonya Huang, Pat Grady and more Sequoia Capital partners host conversations with leading AI builders and researchers to ask critical questions and develop a deeper understanding of the evolving technologies—and their implications for technology, business and society. The content of this podcast does not constitute investment advice, an offer to provide investment advisory services, or an offer to sell or solicitation of an offer to buy an interest in any investment fund.
Making Cities Awesome: Peregrine’s Nick Noone & Ben Rudolph
Most public safety technology companies grow by collecting more data. Peregrine inverted the model: no sensors, no new data, a business built on connecting the data and information cities already own. Co-founders Nick Noone and Ben Rudolph received more than two dozen no's before San Pablo PD let them in the door in February 2018. Today, Peregrine powers law enforcement, emergency medical services, fire and rescue, and other services in more than 400 cities and communities globally. Nick and Ben explain their north star for data sovereignty, and discuss how Peregrine's philosophy and privacy-first approach to data access and ownership preserves individual privacy and cities' sovereignty. They walk through how AI and long-horizon agents are being deployed: a cold case agent that reproduced an exoneration detectives had reached by hand, a Wisconsin county that placed a suspect using cell records buried in 300GB of evidence, identifying threats to a synagogue, root-causing an escalation in weather-related incidents, and more. 00:00 Introduction 02:07 What Forward Deployed Engineering Means 03:58 What Silicon Valley Gets Wrong 05:23 UNHCR, Dimagi And Downstream Data Problems 08:25 Why Cities, Why Safety 10:45 Two Dozen Nos And San Pablo PD 14:19 Building Through Defund The Police 18:16 The Inversion Of The Collection Model 21:20 Data Ownership And Governance 22:57 From Nice Search To Deep Analysis 29:59 Agents Writing The Integrations 31:45 The Cold Case Agent 35:02 The Anti-Network-Effect Proposition 38:40 Facial Recognition And Hard Decisions 40:48 Technology For The Underdogs 42:54 Trusting The Individual Contributor 48:50 Ten Thousand Cities
Sep 1
52 min
Parallel’s Parag Agrawal: Building a New Web for AI Agents
Parag Agrawal is making a bet that goes against two decades of web search: agents will query the web a thousand times more than humans ever have, and the infrastructure built around human clicks is wrong for them. The former Twitter CEO, now founder and CEO of Parallel Web Systems, explains why Parallel treats human click data as a bug and trains on agent feedback instead. He unpacks the counterintuitive choice to ship a search agent before a search engine, building an index incrementally, and how the new Turbo product cut agentic search to 200 milliseconds. But the problem Parag keeps returning to is economic: the ad-supported internet collapses when agents show up instead of people. His fix draws on Shapley values to pay content owners for the value their pages provide agents, with real dollars reaching publishers, he predicts, within 12 to 24 months. Hosted by Sonya Huang and Andrew Reed, Sequoia Capital 00:00 Introduction 03:25 What Is Web Search 05:17 Why Start a New Index 07:52 Search Agents First 10:17 Not a Neolab 13:14 Agents vs Google Search 19:38 Inside the Search Stack 28:59 Search Multipliers With Agents 30:21 Meeting Prep Agent Workflows 31:46 Quality Cost Latency And Turbo 32:42 Are Agents Overtaking Humans 34:28 Ads Model Meets Agent Web 37:20 New Incentives For Content 40:48 Shapley Values Attribution 47:46 Parallel Web And Future Vision
Aug 25
55 min
Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again
Rich Sutton, who helped pioneer reinforcement learning and wrote the seminal AI essay The Bitter Lesson, has now cofounded Oak Lab with his former student Khurram Javed. Their goal: to build agents that continuously learn from their own experience rather than from us. Rich doesn't think he holds a radical view: "I'm not weird. The field is weird." He says all learning is continual, and the field is the one that needed a new name for it. Rich and Khurram argue synthetic data is "a big mistake." Their "big world hypothesis" is that the world is massively more complex than any agent or simulator, so approximations have to be updated continuously rather than frozen at deployment. Rich calls LLMs an unanticipated scientific breakthrough, but says they represent roughly a quarter of intelligence. He says catastrophic forgetting is "totally curable" with the ideas behind their continual backprop algorithm. Khurram explains why the frontier labs can't follow: they sit in a local minimum where a new paradigm gets worse before it gets better. Their target, five to ten years out, is a trillion-parameter mind that keeps learning, stays coherent, and runs on 20 watts. Hosted by Sonya Huang and Alfred Lin, Sequoia Capital 00:00 Introduction 02:10 An AI winter, a cancer diagnosis, and the move to Alberta 07:07 Writing "The Bitter Lesson," and what people get wrong 09:53 Are LLMs a positive or a negative example of it? 11:03 Synthetic data is "just a big mistake," and the Big World Hypothesis 18:01 AlphaGo, human priors, and why prior knowledge and learning should be friends 22:37 "Their weights never change": do LLM assistants actually learn? 26:09 Babies, squirrels, and why no animal learns by supervised learning 32:02 Rockets, imagination, and where paradigm shifts come from 36:42 The Alberta Plan and its 12 steps 38:53 Catastrophic forgetting and the cure 43:43 Oak's biggest ambition: a self-maintaining mind 47:56 Why the big labs are stuck in a local minimum 49:13 If everything goes right: LLMs, many minds, and hiring
Aug 18
53 min
Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem
Most people treat biology as a bespoke, messy science. Josh Meier and Matt McPartlon, co-founders of Chai Discovery, treat it as an engineering problem. They make the case that drug design obeys the bitter lesson: scale data, models, and compute, and the model can learn what a hand-built pipeline simply couldn't capture. The results are concrete: Chai-2 pushed de novo antibody design from a sub 0.1% hit rate to 16%, turning a needle-in-a-haystack search into something more like designing a key to fit a lock. Josh argues, counterintuitively, that biology is more verifiable than code, and explains why the goal should be more lab experiments, not fewer. Their bet: a design suite that collapses drug discovery from nine months to nine days, and arms the pharma industry rather than competing with it. Hosted by Pat Grady and Sonali Singh, Sequoia Capital 00:00 Introduction 01:52 From Discovery to Design 03:25 Protein AI Breakthroughs Timeline 06:04 Why Start in 2024 10:13 Diffusion Models Intuition 11:41 Building the Avengers Team 15:22 Hit Rates and Scaling Laws 25:01 Molecular CAD Vision 25:24 Faster Design Loops 26:32 Future Drug Discovery 28:37 Platform Business Model 31:14 Partnering Reality Check 33:44 Data Flywheel Explained 37:16 Staying Ahead at Scale 39:44 Culture and What's Next
Aug 4
47 min
Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil
Jerry Tworek led reasoning at OpenAI, convinced that scaling reinforcement learning was the path to AGI. Rohan Anil co-led Gemini pre-training and built the Shampoo optimizer. Now they've teamed up at Core Automation on a contrarian premise: the transformer has carried us as far as it can, and the bottleneck to smarter systems is no longer scale — it's the architecture itself. The missing capability is continual learning, models that adapt at test time, which transformers can't do. In-context learning taps out fast (Codex needs compacting after ~20 minutes) and fine-tuning invites catastrophic forgetting. Rohan argues pre-training and RL should be optimized end-to-end, and that transformers spend computation inefficiently. They lay out why the largest labs won't chase alternatives while locked in the coding-agent race, and why building the world's most automated lab starts with automating kernel generation—the one place frontier models still lose to a high-taste human. Hosted by Sonya Huang and Pat Grady, Sequoia Capital
Jul 29
49 min
Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself
Factory started building fully autonomous coding agents in April 2023, two years before enterprises were ready. Matan Grinberg now says this is indistinguishable from being wrong. The Factory co-founder and CEO explains how the company survived its "journey in the desert," including the decision to hand nearly all of its revenue back to customers when the product wasn't making developers obsessed. Matan makes the contrarian technical case that a model-agnostic harness beats the model-and-harness co-design that labs like OpenAI and Anthropic favor, because exposing a harness to many models keeps it from overfitting to any single one. He argues open-weight models like GLM will capture the majority of tokens by staying one generation behind the frontier at a fraction of the cost, and that CIOs will soon justify every incremental token the way they justify headcount. Looking ahead, he predicts 90% of coding tokens will run asynchronously—the "dark factory" where software builds itself. Hosted by Sonya Huang and Pat Grady, Sequoia Capital
Jul 21
51 min
Anthropic's Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden
Katelyn Lesse and Angela Jiang lead the team building Anthropic's developer platform - the layer that both outside builders and Anthropic's own products run on top of. Angela frames the platform as a three-layer stack: knowledge, execution, and coordination. She argues the real leverage is what’s at the top: "strategies," or meta-harnesses that give each token a different job, from advising to executing to reflecting to memory. On the question of open ecosystem vs. walled garden, they say they aren't precious about owning the stack. Katelyn points to Anthropic's self-hosted sandboxes with partners like Modal, Vercel, and Cloudflare. Whether the work runs on Anthropic's infrastructure or someone else's, what really matters to them is that the architecture is sound. The deeper bet is standards: they hand skills and MCP to the whole industry, build connectors on the MCP spec, and help agents (Claude and non-Claude) work together. The one place they stay closed is model routing: they argue harnesses should be tuned to a model family, so they're designing for Claude rather than routing across models. Angela's frame for the ecosystem bet is electricity: transformative only because everyone could plug in, and no company wired it alone.Hosted by Sonya Huang and Lauren Reeder, Sequoia Capital 00:00 Introduction 01:49 Two North Stars 02:27 External Builders And Primitives 03:54 What To Externalize 06:00 From Messages To Agents 08:19 Managed Agents Adoption 09:07 Three Layer Cake 10:22 Execution Harnesses Explained 11:09 Coordination Strategies Roadmap 12:13 Ecosystem Standards And Safety 15:39 Open Ecosystem Not Walled 17:12 Vertical Products And Form Factors 22:26 Claude Tag Under The Hood 26:04 Harness Best Practices 38:13 Token Costs And Whats Next
Jul 14
48 min
Inside Zipline's Autonomous System: 140M Miles, Zero Incidents
The largest commercial autonomous system on earth isn't a robotaxi fleet — it's Zipline, which has flown 140 million autonomous miles with zero safety incidents. Co-founder Keller Rinaudo Cliffton and Eric Watson, who leads systems engineering and safety, explain why the drone itself is only 15% of the solution. The rest spans inventory management, air traffic integration, and engineering systems such as a dual flight computer failover protocol that recently saved a delivery mid-flight. They trace Zipline's path from launching blood delivery in Rwanda in 2016 (when drone delivery was illegal in the US) to a 51% reduction in maternal mortality in that country, a $550 million commercial diplomacy partnership with the State Department, and a cost curve that fell from $300 per delivery to $12. Zipline is now racing toward a million deliveries a day, and a quiet inflection point when autonomous delivery becomes cheaper than sending a car. Hosted by Alfred Lin and Pat Grady, Sequoia Capital
Jul 7
55 min
Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis
Dylan Patel, founder of SemiAnalysis, argues the biggest gains in AI don't come from faster chips, they come from software-hardware co-design. Optimizing the model, the kernels, and the silicon together turns a 2x here and a 2x there into 100x. He explains why DeepSeek's experts were shaped for Nvidia's Hopper (and why TPUs struggle to run it), why OpenAI's sparser models and Anthropic's denser ones pull them toward different hardware, and why the so-called CUDA moat was never really about CUDA. Dylan breaks down InferenceX, his living benchmark that runs the latest models on over $50M of donated hardware daily, tracking a roughly 60x annual drop in cost per unit of quality. He makes the case that inference will be a bigger market than oil, that the compute crunch persists because models expand the value of useful work faster than compute grows, and why Jensen Huang is bankrolling neoclouds to engineer a multipolar world. Hosted by Shaun Maguire and Sonya Huang, Sequoia Capital
Jun 30
1 hr 10 min
Memory and Continual Learning: Engram's Dan Biderman and Jessy Lin
Dan Biderman and Jessy Lin, co-founders of Engram, are building a neolab around memory and continual learning, which they call two sides of the same coin. Their contrarian premise: instead of stuffing ever-larger prompts into the context window or bolting on RAG, bake a team's knowledge directly into the model's weights, so it knows your company the way an employee of several years does.  The payoff: matching or beating frontier models while consuming up to 100x fewer tokens. Working with partners like Microsoft, Notion, and Harvey, the team draws on roots in computational neuroscience and state-space architectures to attack what they see as the real bottleneck in AI — not raw intelligence, but memory and continual learning. In contrast to the frontier labs' race toward one ever-bigger model and AGI, Dan and Jessy imagine a world where everyone has their own model — privately trained, always learning, and good at the things you actually care about. The real ChatGPT moment for memory, they argue, is the day your model feels like an intern that genuinely got smarter overnight. Hosted by Sonya Huang and Shaun Maguire, Sequoia Capital
Jun 24
44 min
Load more