
Shopify's distillation pipeline cuts production AI costs by up to 30x — and in some cases, the smaller model outperforms the frontier model on the narrow task. That's not a trade-off. That's a win on accuracy, latency, and cost simultaneously.
Farhan Thawar, VP & Head of Engineering at Shopify, runs AI across one of the largest commerce platforms on earth. In this episode, he breaks down the exact infrastructure decisions Shopify made to avoid being locked into any single model provider — and why 29% of enterprise AI projects die from token costs, not model failure.
Shopify built an internal LLM proxy that routes tokens across every major provider, enabling automatic failover when any one goes down. On top of that, their Universal Distillation Platform (UDP) lets any R&D team distill a frontier model (Opus 4, GPT-5+) down to a fine-tuned open source model (Qwen and others) for a specific subtask — in roughly a day, with evals baked in. Results range from 2x to 30x cheaper, faster, and more accurate than calling the frontier API for everything. Shopify currently runs roughly half a dozen of these distilled models in production, with more being added.
Farhan also details River, Shopify's internal agentic substrate — a public-only Slack agent that queries their data warehouse, reads their PM system, and improves its own answers when engineers jump in to correct it. Plus: how Shopify governs AI-generated code at scale, why they killed their token leaderboard, how UCP is positioning Shopify's catalog for agentic commerce, and what a two-to-three year horizon looks like when agents start holding spending budgets and buying autonomously.
🎙️ GUEST: Farhan Thawar | VP & Head of Engineering, Shopify
🎙️ HOST: Sam Witteveen | VentureBeat
__
If you enjoy these conversations, you need to be in Menlo Park this July.
VB Transform 2026 is VentureBeat's flagship enterprise AI event, built entirely around one question: How do you orchestrate AI autonomy at scale? July 14–15, Hotel Nia. Real projects, proprietary research, no fluff.
50% off for listeners with code BEYONDTHEPILOT: https://bit.ly/4fK4F6z
—
**CHAPTERS**
00:00 Intro — Infrastructure-first philosophy at Shopify
02:00 Episode overview: token economy, distillation, and the cost crisis killing AI projects
03:00 Toby's "AI reflexivity" mandate and what it actually means for engineers
05:20 Ecosystem strategy: when to build vs. leave room for third-party developers
06:20 Developer tooling stack: GitHub Copilot (2021), Claude Code, Cursor, Codex, and Shopify's own River
08:00 LLM proxy architecture: bulk token purchasing, multi-provider failover, and usage reporting
09:00 Model agnosticism: why Shopify lets engineers choose their own harness
10:00 AI adoption beyond R&D — finance, HR, sales, and the Qwik internal deployment platform
11:20 Code governance: who owns AI-generated code going to production
13:00 Token maxing, leaderboards, and the shift from AI reflexivity to AI leverage
14:20 Circuit breakers: how Shopify catches runaway token spend without hard limits
15:20 LLM proxy deep dive: uptime, insights, and cross-team learning from usage data
16:40 River: Shopify's agentic substrate, public-channel-only design, and emergent HITL behavior
18:40 Model distillation explained: teacher/student models, narrow tasks, and the trade-offs
21:00 Universal Distillation Platform (UDP): how any team submits a distillation job in ~one day
22:20 Tangle: open source pipeline visualization for distillation workflows
22:40 Who uses UDP today, and Farhan's vision for auto-selecting the distillation target model
24:00 Evals in practice: golden datasets, Toloka for data generation, threshold-based deployment
26:20 Sim Gym: simulating A/B tests for small merchants without enough traffic
27:40 Pulse: async AI insights on store performance and conversion
28:20 GPU infrastructure trade-offs: when running your own inference makes sense at scale
29:20 Frontier vs. distilled model split in production — and why dev tokens stay frontier
30:40 Should Shopify train its own coding model? Why Farhan says not yet
31:40 UCP protocol, agentic commerce, and how Shopify's catalog surfaces in every LLM
35:00 Early signals: agentic commerce growth rate and the shift away from SEO
36:00 What developers and entrepreneurs should build for now given multi-channel uncertainty
37:20 River as a hive-mind agent — and what "truly agentic" actually means
38:40 Two-to-three year forecast: agents with spending budgets, autonomous purchasing, proactive outreach
40:00 Project Glasswing, model access loss (Fable), and the case for multi-provider architecture
42:00 Why every company should have a backup plan — and how Shopify built theirs years ago
43:20 Wrap-up
---
Subscribe to VentureBeat: https://www.youtube.com/@VentureBeat
Apple Podcasts: https://podcasts.apple.com/us/podcast/venturebeat/id1839285239
Spotify: https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4
Website: https://venturebeat.com
LinkedIn: https://www.linkedin.com/company/venturebeat
Newsletter: https://venturebeat.com/newsletters
#EnterpriseAI #AIAgents #LLM #MLOps #AgenticAI
—
Learn more about your ad choices. Visit megaphone.fm/adchoices
Jun 24
44 min

If you enjoy these conversations, you need to be in Menlo Park this July.
VB Transform 2026 is VentureBeat's flagship enterprise AI event, built entirely around one question: How do you orchestrate AI autonomy at scale? July 14–15, Hotel Nia. Real projects, proprietary research, no fluff.
50% off for listeners with code BEYONDTHEPILOT: https://bit.ly/4fK4F6z
_
He negotiated seat-based (unlimited) licenses before token costs exploded — and projects only a 20–30% spend increase when that changes. MassMutual's CIO rebuilt a COBOL mainframe app into a working web prototype in 7 days — work that used to take a 15-person SI team 90 days. That's not a pilot. That's a new build-vs-buy equation.
Sears Merritt, CIO at MassMutual, runs AI inside one of America's most regulated legacy environments. In this conversation, he breaks down the architecture decisions, cost structures, and security posture behind real, production deployments — not roadmaps.
On infrastructure: MassMutual routes all agentic tool calls through centralized API gateways with identity and access controls, using Amazon Bedrock as a proxy layer. That multi-harness design preserves model optionality while enforcing FinOps discipline. On model selection: a trust score rubric drives every model decision, balancing cost against user experience. In their IT contact center, that rubric led them to choose the more expensive model after users said the quality gap was worth two extra seconds of inference time.
Productivity results are concrete: 30% boost in developer output across the SDLC, call resolution time dropping from 10 minutes to under 1 minute for specific call types, cost per interaction from dollars to cents. On the security side, MassMutual is embedding AI into its SDLC for vulnerability scanning and compressing cyber response cycles from days to hours — building agentic tier-one and tier-two capabilities to match the accelerated threat landscape that frontier models like Mythos have exposed.
🎙️ GUEST: Sears Merritt | Head of Enterprise Technology & Experience, MassMutual
🎙️ HOSTS: Matt Marshall | VentureBeat, Sam Witteveen | VentureBeat
—
If you enjoy these conversations, you need to be in Menlo Park this July.
VB Transform 2026 is VentureBeat's flagship enterprise AI event, built entirely around one question: How do you orchestrate AI autonomy at scale? July 14–15, Hotel Nia. Real projects, proprietary research, no fluff.
50% off for listeners with code BEYONDTHEPILOT: https://bit.ly/4fK4F6z
—
00:00 Intro & COBOL Modernization Preview
00:01:15 Guest Introduction: Sears Merritt, MassMutual CIO
00:02:00 Multi-Vendor Strategy & Avoiding Lock-In
00:02:30 How MassMutual Evaluates AI Tools (Cost vs. Experience Rubric)
00:03:15 12-Month Contracts and Switching Optionality
00:03:30 AI Standards Cycle: MCP, A2A, and the Early Internet Analogy
00:05:15 Measuring Developer Productivity: 30% SDLC Boost
00:06:15 Contact Center Results: 10 Minutes to 1 Minute, Dollars to Cents
00:07:00 Managing Token Cost Explosion
00:07:30 Seat-Based vs. Consumption Licensing Decision
00:08:45 Token Maxing While the All-You-Can-Eat Window Is Open
00:09:45 Building FinOps Infrastructure for Model Routing and Optimization
00:11:15 Outcome-First Model Selection: When to Pay for Opus vs. a Cheaper LLM
00:13:15 Trust Score Framework: How MassMutual Picks the Right Model
00:15:00 Sponsor: OutShift by Cisco
00:15:30 Claude, OpenAI Codex, and Multi-Harness Agentic Architecture
00:16:30 API Gateway Design: Identity, Access, and FinOps Controls
00:17:30 What the Usage Analytics Revealed (And What Merritt Was Afraid to Find)
00:18:15 Projected Token Cost Increase: 20–30% Off Unlimited Plan
00:19:45 Power Law Usage: Top 10% Consuming 80% of Tokens
00:20:45 COBOL Mainframe Modernization: The 7-Day Prototype Workflow
00:22:30 The Full AI-Assisted COBOL Migration Playbook
00:24:00 Implications for IBM and Mainframe-as-a-Service Providers
00:25:15 Open Source Models, DeepSeek, and the Cost Efficiency Question
00:27:45 Chinese Models in a Regulated Environment: Evaluation Criteria
00:29:30 Agentic Security: Identity Management and the Evolving Threat Landscape
00:30:00 How Frontier Models Changed the Threat Velocity (Not the Threat Types)
00:31:15 Fighting AI With AI: Agentic Tier-1 and Tier-2 Cyber Capabilities
00:31:30 Project Glasswing and the CISO Community Response
00:32:30 Embedding AI Into the SDLC for Security Scanning
00:34:00 Closing: When Will Agentic Standards Consolidate? Advice for Builders
---
Subscribe to VentureBeat: https://www.youtube.com/@VentureBeat
Apple Podcasts: https://podcasts.apple.com/us/podcast/venturebeat/id1839285239
Spotify: https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4
Website: https://venturebeat.com
LinkedIn: https://www.linkedin.com/company/venturebeat
Newsletter: https://venturebeat.com/newsletters
#EnterpriseAI #AIAgents #LLMInfrastructure #AIDeployment #MLOps
Learn more about your ad choices. Visit megaphone.fm/adchoices
Jun 10
35 min

Pinterest's open-source AI stack costs 90% less than frontier models — and their custom-trained recommender outperforms off-the-shelf alternatives by 30% in accuracy. Pinterest CTO Matt Madrigal breaks down exactly how they did it, and what enterprise AI teams can actually replicate.
Madrigal walks through the full architecture behind Navigator 1, Pinterest's conversational shopping assistant built on Qwen 3 VL — and the specific decision to rip out its native vision encoder and replace it with PinCLIP, Pinterest's proprietary multimodal embedding layer. That swap alone closes a 20x inference latency gap and makes the economics work at 620 million monthly active users. This is the clearest public explanation yet of how a scaled platform operationalizes the "core vs. context" principle for model selection: open-source and custom-built where it touches the user, frontier models where speed-to-prototype matters more than cost.
The conversation also covers the Taste Graph — Pinterest's knowledge graph across hundreds of billions of pins and 15 billion boards — and how post-training on that proprietary data lets a smaller, fit-for-purpose model beat a larger frontier model on production metrics. Madrigal details their eval framework: gold set benchmarks, product-level evals tied to engagement and merchant click outcomes, and a structured A/B test pipeline that runs from engineer PRs through to live user signal.
On the organizational side: how Pinterest manages a "default yes" multi-IDE policy (Cursor, Windsurf, Claude Code, Codex) without collapsing security posture, how they segment sandbox environments between ML engineers with Taste Graph access and general application developers, and why Madrigal measures AI coding ROI in token usage and experimentation velocity — not lines of code.
🎙️ GUEST: Matt Madrigal | CTO, Pinterest
🎙️ HOSTS: Matt Marshall | VentureBeat, Sam Witteveen | VentureBeat
00:00 Show Intro and Guest
01:17 Open Source Cost Breakdown
02:20 Pinterest Multimodal Roots
02:37 PinClip and Embeddings
05:46 Core vs Context Models
07:43 Navigator 1 Assistant Stack
11:52 Benchmarking and Evals
13:29 Accuracy from Proprietary Data
17:16 Taste Graph Explained
18:29 Taste Graph in Training
22:22 Fighting AI Slop
25:16 Developer Tools and Velocity
27:57 Tool Choice and Governance
28:56 Security Sandboxes and CICD
30:57 Wrap Up
If you enjoy these conversations, you need to be in Menlo Park this July.
VB Transform 2026 is VentureBeat's flagship enterprise AI event, built entirely around one question: How do you orchestrate AI autonomy at scale? July 14–15, Hotel Nia. Real projects, proprietary research, no fluff.
50% off for listeners with code BEYONDTHEPILOT: https://bit.ly/4fK4F6z
Subscribe to VentureBeat: https://www.youtube.com/@VentureBeat
Apple Podcasts: https://podcasts.apple.com/us/podcast/venturebeat/id1839285239
Spotify: https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4
Website: https://venturebeat.com
LinkedIn: https://www.linkedin.com/company/venturebeat
Newsletter: https://venturebeat.com/newsletters
#EnterpriseAI #OpenSourceAI #AIInfrastructure #LLM #MachineLearning
Learn more about your ad choices. Visit megaphone.fm/adchoices
May 27
33 min

Enterprise GPU hoarding is over. LinkedIn CTO Erran Berger and VentureBeat analyst Rob Strechay break down what comes next — and the infrastructure math most enterprises are only now being forced to confront.
VentureBeat's Q1 research shows GPU availability anxiety dropped from 20.8% to 15.4% among enterprise teams, while cost-per-inference and TCO concerns jumped from 34% to 41% — a number that's still climbing. The hoarding phase is giving way to an audit phase, and the companies that didn't build the instrumentation to understand their workloads are now paying for it.
Erran Berger explains how LinkedIn runs one of the few remaining at-scale applied ML shops outside the hyperscalers — owning the full stack from bare metal GPU clusters to member-facing products. That means LinkedIn engineers can optimize custom CUDA kernels, compress embeddings, prune models for throughput, and adapt networking and storage per workload — trade-offs that are simply unavailable on public cloud instance menus. The result: a rigorous ROI framework that evaluates not just current traffic costs, but the traffic shape agents will drive in 2–3 years.
On the market side, 72% of enterprises admit they lack sufficient control over their AI infrastructure. Open-source inference tools like vLLM and LLMD are seeing rapid adoption, while 17% of organizations have moved to full-stack ownership. Hyperscalers report 60–80% of workloads have already shifted from training to inference — and most enterprise teams are still figuring out how to staff and instrument for that reality.
🎙️ GUEST: Erran Berger | CTO, LinkedIn
🎙️ ANALYST: Rob Strechay | VentureBeat
🎙️ HOST: Matt Marshall | CEO, VentureBeat
---
00:00 Intro: The GPU Hoarding Hangover
00:10 Guest Introductions
02:00 VentureBeat Q1 Data: GPU Panic Fades, TCO Concerns Rise
03:00 LinkedIn's Early Shift to Inference ROI Discipline
04:00 Budget Moving Into Inference Optimization and Control
07:00 LinkedIn's Full-Stack Advantage: Kernels, Pruning, Embedding Compression
08:00 Private AI and Sovereign Stacks: What the Q1 Data Shows
09:00 Open Source Inference Tooling: vLLM, LLMD, RDMA
10:00 Data Sovereignty at LinkedIn Scale: Member Data and Board-Level ROI Framing
12:00 Why Instrumentation Beats GPU Hoarding
13:00 Planning for Ambient Agent Traffic — Not Just Today's Workloads
14:00 Closing Advice for the Enterprise CTO Staring at 5% GPU Utilization
---
Subscribe to VentureBeat: https://www.youtube.com/@VentureBeat
Apple Podcasts: https://podcasts.apple.com/us/podcast/venturebeat/id1839285239
Spotify: https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4
Website: https://venturebeat.com
LinkedIn: https://www.linkedin.com/company/venturebeat
Newsletter: https://venturebeat.com/newsletters
#EnterpriseAI #AIInfrastructure #MLOps #InferenceOptimization #GenerativeAI
---
Learn more about your ad choices. Visit megaphone.fm/adchoices
May 13
15 min

The CEO who built one of the most-starred RAG frameworks on GitHub (47,000 stars) just publicly declared that frameworks like his are becoming obsolete — and then pivoted his entire company around that conclusion.
Jerry Liu, CEO and co-founder of LlamaIndex, joins Matt Marshall and Sam Witteveen to explain exactly what broke in the AI stack, why 95% of his team's code is now AI-generated, and where the real defensibility in enterprise AI infrastructure actually lives in 2026.
The conversation covers the specific architectural shift that made RAG orchestration frameworks less central: agent reasoning has improved to the point where dumb tools plus smart agents outperform sophisticated retrieval pipelines, coding agents have collapsed the cost of custom integrations, and model providers like Anthropic are consolidating the harness layer around MCP, sandboxes, and session state. Jerry walks through Anthropic's managed agent diagram as a real architectural reference point and explains why engineering leaders should prioritize modular interfaces over implementation investment — because parts of your current stack will need to be thrown away in months, not years.
On SaaS survival, Jerry argues the companies that retain value are those becoming systems of record — and that the real opportunity is building AI agents that automate labor on top of their platforms, not defending UI/UX that agents are now bypassing. On LlamaIndex's own bet: document understanding — parsing PDFs, tables, charts, and forms at higher accuracy and lower cost than frontier models — is the context layer every agent stack needs regardless of which model wins the next benchmark cycle. LlamaParse and the newly released open-source ParseBench (April 13) are the commercial expression of that thesis.
If you're evaluating your AI stack architecture, deciding how much to build vs. buy, or trying to understand where horizontal tooling still has a moat, this episode is the conversation.
🎙️ GUEST: Jerry Liu | CEO & Co-Founder, LlamaIndex
🎙️ HOSTS: Matt Marshall | VentureBeat, Sam Witteveen | VentureBeat
---
**CHAPTERS**
00:00 Intro — LlamaIndex's origin and RAG framework origins
02:00 How LlamaIndex started: GPT-3, 4K context windows, and GPT Index
04:00 Why AI frameworks are becoming less useful in the agentic era
07:00 What changed in the stack: agent reasoning, coding agents, and RAG's evolution
09:00 How Anthropic's managed agent diagram reframes enterprise architecture
13:00 The lock-in question: managed agents, session state, and stack modularity
16:00 Should you build horizontal tooling? Why Jerry says probably not
18:00 Open vs. closed: the Apple/Android analogy applied to frontier labs
21:00 The abstraction level is rising — English is the new programming language
24:00 SaaS market cap destruction: who survives agents eating software
28:00 The "full stack builder" emergence and the future of SaaS seats
31:00 Buy vs. build for agents: the AI recruiter thought experiment
33:00 LlamaIndex's pivot: document understanding as defensible infrastructure
36:00 Why frontier models won't commoditize specialized document parsing
38:00 LlamaParse deep dive: zero-shot accuracy, tables, charts, handwriting
41:00 LightParse, ParseBench, and designing for agent consumers
44:00 Wrap-up and where to follow LlamaIndex
---
Subscribe to VentureBeat: https://www.youtube.com/@VentureBeat
Apple Podcasts: https://podcasts.apple.com/us/podcast/venturebeat/id1839285239
Spotify: https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4
Website: https://venturebeat.com
LinkedIn: https://www.linkedin.com/company/venturebeat
Newsletter: https://venturebeat.com/newsletters
#EnterpriseAI #AIAgents #LLMInfrastructure #RAG #AIArchitecture
---
Learn more about your ad choices. Visit megaphone.fm/adchoices
Apr 29
45 min

Cisco's OutShift deployed a multi-agent network configuration system that raised error detection from 10–15% to 100% and cut full change validation from 2–3 weeks to 6–7 minutes. The reason it worked — and why most enterprise multi-agent deployments still fail — comes down to a single gap nobody is talking about: agents can connect, but they cannot think together.
Vijoy Pandey, SVP and General Manager of OutShift by Cisco, joins Matt and Sam to explain why A2A, MCP, and existing agent protocols solve connectivity but leave out an entire layer: shared cognition. OutShift's research identifies this as a missing "Layer 9" — a semantic and cognitive communication stack above today's syntactic protocols — and they're already building it.
The conversation covers the four pillars of enterprise-grade multi-agent infrastructure (discovery, identity/access, communication, observability), why standard IAM models break when agents enter the picture, and how OutShift extended OpenTelemetry with Microsoft to cover multi-agent evaluation. Vijoy introduces three new cognition-state protocols — SSTP (Semantic State Transfer), LSTP (Latent Space Transfer), and CSTP (Compressed State Transfer) — and explains the staged rollout path for each, including a published MIT collaboration called the Ripple Effect Protocol.
The healthcare scheduling case study is particularly instructive: three independent third-party agents — insurance, diagnostics, scheduling — each with competing optimization functions and siloed context, and zero shared intent. That's the real multi-vendor, multi-org enterprise problem. Vijoy explains what an orchestrator can't fix, and what a cognitive fabric layer would.
🎙️ GUEST: Vijoy Pandey | SVP & General Manager, OutShift by Cisco
🎙️ HOSTS: Matt Marshall | VentureBeat, Sam Witteveen | VentureBeat
---
**CHAPTERS**
00:00 Intro & Cold Open: Agents Connect But Can't Think Together
00:03 Welcome & Guest Introduction: Vijoy Pandey, OutShift by Cisco
00:04 Do Agents Work Outside Coding & Customer Support? Challenging Amjad Masad's Diagnosis
00:05 What's Wrong With A2A and MCP? The Four Pillars of AGNTCY
00:08 Identity & Access Management for Agents: Why IAM Breaks and What TBAC Fixes
00:12 The Network Digital Twin: How OutShift Achieved 100% Error Detection in Production
00:13 From 2–3 Weeks to 6–7 Minutes: Real Results From Deployed Multi-Agent Networking
00:15 Agents Can Connect But Can't Think Together: The Core Thesis
00:20 The Cognitive Revolution Analogy: Shared Intent, Shared Context, Collective Innovation
00:25 The Healthcare Scheduling Case Study: Three Competing Agents, Zero Shared Intent
00:31 Why Orchestrators Fail in Multi-Vendor, Multi-Org Environments
00:36 Introducing Layer 9: SSTP, LSTP, and CSTP — The Cognition-State Protocol Stack
00:41 What OutShift Is Building Now: Protocols, Fabric, and Cognition Engines
00:44 MIT Collaboration: The Ripple Effect Protocol and Phase One Rollout
00:46 Cisco's 40-Year Networking Playbook Applied to the Internet of Cognition
00:49 Closing: Where to Find the Research, AGNTCY, and OpenClaw Integration
---
Subscribe to VentureBeat: https://www.youtube.com/@VentureBeat
Apple Podcasts: https://podcasts.apple.com/us/podcast/venturebeat/id1839285239
Spotify: https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4
Website: https://venturebeat.com
LinkedIn: https://www.linkedin.com/company/venturebeat
Newsletter: https://venturebeat.com/newsletters
#EnterpriseAI #AIAgents #MultiAgentSystems #AIInfrastructure #LLM
—
“Scaling Out Superintelligence” Vijoy Pandey, January 2026. The technical whitepaper detailing the Internet of Cognition architecture, three-layer stack, and cognition state protocols.
Internet of Cognition Interactive Demo Clickable walkthrough showing per-agent activity, intent, context, and collective reasoning across a multi-agent SRE system.
“A Layered Protocol Architecture for the Internet of Agents” Fleming, Muscariello, Pandey, Kompella. The OSI Layer 8/9 extension.
AGNTCY Open source multi-agent infrastructure under Linux Foundation governance. Covers discovery, identity, communication, observability.
Formative members: Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat.
Learn more about your ad choices. Visit megaphone.fm/adchoices
Apr 15
50 min

A QuickBooks customer discovered significant fraud by asking their AI assistant follow-up questions about transaction amounts that didn't add up. This isn't a demo — it's one of 3 million customers now using Intuit's AI agents in production, with 80.5% returning to use them again.
Marianna Tessel, EVP and GM of QuickBooks (formerly CTO of Intuit), walks through the architecture decisions behind one of the first enterprise AI deployments at true scale. Intuit's "done-for-you" agents now automate book closing, reconciliation, transaction categorization, and payroll — but the breakthrough came when they realized chatbots alone weren't enough. Businesses wanted human experts integrated directly into AI workflows, creating what Intuit calls the "AI + HI" model (artificial intelligence + human intelligence). The results: invoices paid 5 days faster, 90% more paid in full, 30% reduction in manual work, and 62% of users reporting bookkeeping is easier.
Tessel reveals the technical evolution: moving from monolithic agents to a dynamic orchestration layer that routes queries across multiple LLMs (including Intuit's proprietary FinLM built on open-source), 24,000 bank connections, and 600,000 customer attributes. The system now handles proactive anomaly detection, benchmarking against similar businesses, and even nascent vibe coding — all without requiring users to understand they're essentially programming workflows through natural language. She also addresses the "SaaS apocalypse" narrative head-on, explaining why QuickBooks saw 18% growth last quarter while competitors faced market pressure: durable data advantages and customer trust in financial accuracy matter more than ever when AI enters the mix.
For enterprise builders navigating agent architecture, data grounding, and human-in-the-loop design, this is a rare look inside a working system serving millions.
🎙️ GUEST: Marianna Tessel | EVP & GM, QuickBooks (Intuit)
🎙️ HOSTS: Matt Marshall | VentureBeat, Sam Witteveen | VentureBeat
00:00 Intro — Customer discovers fraud using QuickBooks AI
03:26 Intuit Intelligence: Agents, BI, and human expertise integration
05:20 First-time AI users and going beyond chatbots
08:02 How Intuit decides which workflows to automate
10:16 Sponsor: Outshift by Cisco
10:38 Human-in-the-loop: When to insert experts vs. full automation
13:00 The AI + HI model: Why customers want human verification
15:24 Human expertise as confidence layer, not just AI check
16:14 Proprietary data advantage: 24K bank connections, 600K attributes
18:39 Benchmarking: "Businesses like me" — using aggregate data for competitive insights
19:52 First-party vs. third-party data strategy
21:38 Addressing the "SaaS apocalypse" narrative — why Intuit grew 18% last quarter
24:39 Proactive AI: Anomaly detection for marketing expense spikes
25:20 Builder perspective: Leaning on LLM orchestration, not use-case-by-use-case builds
27:32 Architecture evolution: From monolithic agents to dynamic tools and skills
29:10 Composite UX: Chat side-by-side with traditional workflows
30:35 Multi-model strategy: Genos platform, FinLM, and model routing
31:16 Vibe coding and actions: Letting users automate without realizing they're coding
32:47 Personalization wave: Memory, persistence, and user-defined workflows
35:08 Docker background and primitives that survive disruption
36:00 Open Claw and agent automation: Real revolution or risky experimentation?
#EnterpriseAI #AIAgents #QuickBooks #Intuit #LLMOrchestration #AgenticAI
Presented by Outshift by Cisco Outshift is Cisco’s emerging tech incubation engine and driver of Agentic AI, quantum, and next-gen infrastructure. Learn more at outshift.cisco.com.
About VentureBeat: VentureBeat equips enterprise technology leaders with the clearest, expert guidance on AI – and on the data and security foundations that turn it into working reality.
🔗 CONNECT WITH US
Subscribe to our Newsletters for technical breakdowns: https://venturebeat.com/newsletters
Visit VentureBeat: Venturebeat.com
.
.
.
Subscribe to VentureBeat:
/ @VentureBeat
.
.
Subscribe to the full podcast here:
Apple: https://podcasts.apple.com/us/podcast/venturebeat/id1839285239
Spotify: https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4
YouTube: https://www.youtube.com/VentureBeat
Learn more about your ad choices. Visit megaphone.fm/adchoices
Apr 1
38 min

Major SaaS companies including Salesforce, Intuit, and ServiceNow saw stock drops of 45-50% as enterprises shift from bloated software suites to personalized AI agents that users can control directly. Microsoft just capitulated this week, opening Copilot to allow Claude Cowork-style functionality — a clear signal that the "build vs. buy" calculus for enterprise software has fundamentally changed.
Matt Marshall and Sam Witteveen break down why personalization is no longer optional for enterprise products. Companies like Zoom now offer personalized workflows that access your conversation history and profile context. Infrastructure decisions are moving fast: token budgets must account for per-user context, identity management has become the biggest technical challenge for agent deployments, and "skills" (not just MCP) are emerging as the key abstraction layer.
Zoom's Li Juan explains how their AI Companion moved beyond generic templates to user-controlled personalization: tracking opinion divergence in meetings, generating follow-up emails with specific context controls, and giving users explicit prompt examples instead of "good luck with your prompt." This is the new standard. If your product can't reason over which tools to use, which skills to apply, and which context to pull — all personalized to the individual user — you're competing with something that can be built in 10 days (Cowork's timeline).
The agents-are-taking-over reality is here: multi-user agent architectures require thinking about context contamination, security postures for computer-use capabilities, and whether you're building internal agents or buying SaaS that will adapt. Sam's take: "AGI is agentic, and we're well along that continuum now."
🎙️ HOSTS: Matt Marshall | CEO, VentureBeat & Sam Witteveen | VentureBeat
📺 CHAPTERS:
00:00 Intro — The SaaS Apocalypse
00:01:00 The Personalization Imperative
00:02:00 Microsoft Copilot Capitulates to Cowork
00:03:00 From Template Selection to Skill Generation
00:04:00 The Land Grab for User Context
00:05:00 Zoom's Li Juan on Personalized Meeting Intelligence
00:06:00 Why Context = Magic in Enterprise AI
00:07:00 Product-Market Fit in the Agent Era
00:08:00 Metrics That Matter: JP Morgan's 30,000 Agents
00:09:00 Build vs. Buy: The New Calculus
00:10:00 Why Slack Might Win on Agent Identity Management
00:11:00 Zoom's AI Companion: Control Over Randomness
00:13:00 Li Juan on Purposeful Prompts and Reference Control
00:15:00 Multi-Agent vs. Multi-User: The Critical Distinction
00:16:00 LinkedIn's GPU Optimization Strategy
00:17:00 AGI Is Agentic: Where
Presented by Outshift by Cisco Outshift is Cisco’s emerging tech incubation engine and driver of Agentic AI, quantum, and next-gen infrastructure. Learn more at outshift.cisco.com.
About VentureBeat: VentureBeat equips enterprise technology leaders with the clearest, expert guidance on AI – and on the data and security foundations that turn it into working reality.
🔗 CONNECT WITH US
Subscribe to our Newsletters for technical breakdowns: https://venturebeat.com/newsletters
Visit VentureBeat: Venturebeat.com
.
.
.
Subscribe to VentureBeat:
/ @VentureBeat
.
.
Subscribe to the full podcast here:
Apple: https://podcasts.apple.com/us/podcast/venturebeat/id1839285239
Spotify: https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4
YouTube: https://www.youtube.com/VentureBeat
Learn more about your ad choices. Visit megaphone.fm/adchoices
Mar 18
20 min

LangChain told employees they cannot install OpenClaw on company laptops due to "massive security risk" — yet this unhinged approach is exactly what makes it work. Harrison Chase unpacks why OpenClaw succeeds where AutoGPT failed, and why context engineering, not just smarter models, separates demo agents from production-ready systems.
The shift is architectural: Modern agent harnesses like Claude Code now dump 40,000-token API responses to file systems instead of cramming them into message history. LangChain's Deep Agents framework emerged from reverse-engineering Claude Code, Codex, and Deep Research — discovering they all use planning via to-do lists, subagents for focused work, file systems for context control, and 2000-line system prompts. Harrison explains why coding agents make surprisingly good general-purpose agents, how prompt caching creates accuracy trade-offs, and why "context engineering" — bringing the right information in the right format to the LLM at the right time — matters more than framework choice.
For enterprise teams: Harrison breaks down LangGraph (agent runtime with durable execution), LangChain (unopinionated agent framework), and Deep Agents (batteries-included harness). The conversation covers when to use graphs vs. loops, how skills differ from tools and subagents, and why nine months ago marked the inflection point where models could finally run reliably in autonomous loops.
🎙️ GUEST: Harrison Chase | Co-founder & CEO, LangChain
🎙️ HOSTS: Matt Marshall | CEO, VentureBeat | Sam Witteveen | VentureBeat
**CHAPTERS:**
00:00 Intro — OpenClaw security warning
01:00 LangChain's origin story: From open source library to company
03:00 Early LLM patterns: RAG and SQL agents before ChatGPT
05:00 Why OpenClaw works where AutoGPT failed
08:00 Step change in agent capability: The summer 2024 inflection
11:00 Deep Agents unpacked: Planning, subagents, file systems, prompting
14:00 Skills vs tools vs subagents
16:00 LangGraph, LangChain, and Deep Agents architecture
19:00 Context engineering: What the LLM sees vs what developers see
21:00 File systems for context management vs AutoGPT's approach
**LINKS:**
Subscribe to VentureBeat: https://www.youtube.com/@VentureBeat
Apple Podcasts: https://podcasts.apple.com/us/podcast/venturebeat/id1839285239
Spotify: https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4
Website: https://venturebeat.com
LinkedIn: https://www.linkedin.com/company/venturebeat
Newsletter: https://venturebeat.com/newsletters
#EnterpriseAI #AIAgents #LangChain #AgenticAI #LLMInfrastructure
Learn more about your ad choices. Visit megaphone.fm/adchoices
Mar 4
56 min

On February 2nd, a single plugin wiped nearly $800 billion off the enterprise software market. Wall Street is terrified that AI agents are about to eat the legal industry's lunch. But LexisNexis isn't scared—they're building the moat.
In this episode of Beyond the Pilot, Min Chen (Chief AI Officer, LexisNexis) reveals the sophisticated architecture they built to counter the "LLM wrapper" revolution. Moving beyond standard RAG, Min breaks down their move to "GraphRAG", their deployment of Agentic workflows (using Planner and Reflection agents), and why they created a proprietary "Usefulness Score" because standard accuracy metrics weren't good enough for lawyers.
AI Gets Real Here. No theory, just the execution roadmap for deploying AI in a zero-error environment.
In this episode, we cover:
The "Dangerous RAG" Problem: Why semantic search fails in professional domains (retrieving "relevant" but overruled cases) and how "Point of Law" knowledge graphs fix it.
The "Usefulness" Metric: The 8 sub-metrics LexisNexis uses (including Authority, Comprehensiveness, and Fluency) to grade AI quality.
Agentic ROI: How deploying a "Planner Agent" to break down complex questions increased answer usefulness by 20%.
The "Reflection Agent": Using a secondary agent to critique and refine drafts in real-time.
Hallucination Detection: Why you should never rely on an LLM to judge its own hallucinations (and the deterministic code they use instead).
⏱️ TIMESTAMPS
00:00 - Intro: The $800 Billion AI Threat to Legal Tech
02:18 - Min Chen’s Journey: From Feature Engineering to Chief AI Officer
05:55 - Why Standard RAG Fails in Law (and How GraphRAG Fixes It)
10:40 - "Accuracy" is a Vanity Metric: The 8-Point Usefulness Score
14:20 - The "Auto-Eval" Framework: Human-in-the-Loop at Scale
16:40 - The Secret Sauce: Don't Use LLMs to Detect Hallucinations
21:15 - Agentic AI: How "Planner Agents" Drove a 20% Gain
22:00 - The "Reflection Agent": Self-Critique Loops for Drafting
30:30 - Distillation: Balancing Cost, Speed, and Quality
32:45 - Min’s Advice: Don't Build the Product First (Build the Metrics)
Presented by Outshift by Cisco Outshift is Cisco’s emerging tech incubation engine and driver of Agentic AI, quantum, and next-gen infrastructure. Learn more at outshift.cisco.com.
About VentureBeat: VentureBeat equips enterprise technology leaders with the clearest, expert guidance on AI – and on the data and security foundations that turn it into working reality.
🔗 CONNECT WITH US
Subscribe to our Newsletters for technical breakdowns: https://venturebeat.com/newsletters
Visit VentureBeat: Venturebeat.com
.
.
.
Subscribe to VentureBeat:
/ @VentureBeat
.
.
Subscribe to the full podcast here:
Apple: https://podcasts.apple.com/us/podcast/venturebeat/id1839285239
Spotify: https://open.spotify.com/show/4Zti73yb4hmiTNa7pEYls4
YouTube: https://www.youtube.com/VentureBeat
Learn more about your ad choices. Visit megaphone.fm/adchoices
Feb 18
35 min
Load more
