Show notes
What happens after you’ve built your first MCP server and actually have to make it work in the real world?In this episode, James Ward dives into the more advanced side of MCP: observability, evals, tool design, code mode, authentication, and the challenges that appear once agents start using your server at scale.We also get into how AWS thinks about MCP across roughly 16,000 APIs, why inefficient tool design often gets blamed on MCP itself, and whether the future could involve more constrained, human-reviewable alternatives to full code mode.Along the way, James shares a great example of an AWS documentation change that accidentally triggered prompt injection warnings from agents, showing just how complicated testing across different models and harnesses is becoming.If you’re already building with MCP and want to understand what comes after the “hello world” stage, this one goes deep.Timestamps:[00:00] “This Code Is Gobbledygook”[00:44] What Happens After You Build an MCP Server?[02:11] The MCP Patterns You Actually Need in Production[03:48] Why Your Agent Is Making Too Many Tool Calls[05:29] How to Make MCP Use Fewer Tokens[08:44] Is MCP Actually Inefficient?[10:11] “We Built a Lot of Pretty Crappy MCP Servers”[11:25] The MCP 2.0 Migration Problem[14:02] Why AWS Is Rethinking Code Mode[15:35] The Problem With Letting Agents Write Python[19:31] How Do You Measure Agent Experience?[22:11] Finding Out Why an Agent Failed[23:42] Hundreds of Evals for Four Cents[24:36] The Evaluation Matrix Gets Massive[26:28] AWS Accidentally Triggered a Prompt Injection Warning[30:16] Should MCP Servers Expose Only Five Tools?[31:02] AWS Has 16,000 APIs. Now What?[34:09] One Super-Agent or Thousands of Specialized Agents?[36:55] The MCP Authentication Problem[40:15] What’s Coming at AgentCon

