vBrownBag
vBrownBag
vBrownBag
AI Observability: What Could Possibly Go Wrong?
1 hour 2 minutes Posted Jul 13, 2026 at 1:58 am.
Welcome & Introduction
Full Disclosure Observe, Snowflake, and How This Conversation Started
From Developer Concerns to Boss's Boss's Boss Spending Out of Control
What Actually Gets Measured Errors, Latency, Quality, and Cost
The Casino Chip Problem Confusing Token Pricing Models
Defining Quality When the Task Itself Is Nebulous
The New Role Engineers Who Just Build Testing Harnesses
Non-Determinism and Why Testing Agents Is Expensive
Trace Data, Tool Calls, and What Observability Tools Actually See
Prompt Injection, Zero-Width Characters, and Real World Failures
0:00
1:02:34
Download MP3
Show notes
Join us as John Mark Troyer and Rakesh Gupta break down what AI observability actually means once agents leave the demo and hit production - and why the old playbook for monitoring doesn't cut it anymore.
John Mark and Rakesh walk through why errors and latency are just the starting point for agents, how quality became a much harder thing to measure once bots went from answering questions to taking autonomous action, and why token-based costs are creating a confusing new economics problem for engineering teams. You'll learn the difference between online and offline evals, why a new engineering role has emerged just to build testing harnesses for agents, how trace data works differently when every prompt is its own trace, and what teams are doing to catch prompt injection and other AI-specific failure modes before they become expensive mistakes.
Timestamps
How to find John Mark:
https://www.linkedin.com/in/johnmarktroyer/
How to find Rakesh:
https://www.linkedin.com/in/rg0/
Links from the show: