Show notes
Goodfire co-founder and CTO Dan Balsam returns to discuss where interpretability research now stands and to introduce Silico, the $1,000-per-month research platform Goodfire built for itself. He and Nathan explore Predictive Data Debugging, including the idea that fine-tuning and RL often amplify behaviors already latent in pre-training, and that interpretability can identify the data and features driving unwanted updates. The conversation centers on concept manifolds: Dan argues that models do not store concepts as simple one-hot features, but as sparse mixtures of meaningful subspaces whose geometry determines what kinds of steering and control work. The stakes are practical as well as conceptual, from debugging training data and RL to understanding why steering can fail off-manifold and why modern interpretability may be moving beyond its reputation as a toy-model science.Silico: https://www.goodfire.com/silicoPredictive data debugging: https://www.goodfire.com/research/predictive-data-debugging#Neural Geometry: https://www.goodfire.com/research/the-world-inside-neural-networks#For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/thinking-in-silico-goodfire-cto-dan-balsam-on-concept-manifolds-a-1000-month-ml-research-agent/Sponsor:Claude: Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcrCHAPTERS:((((((((((((((((PRODUCED BY:https://aipodcast.ingSOCIAL LINKS:Website: https://www.cognitiverevolution.aiTwitter (Podcast): https://x.com/cogrev_podcastTwitter (Nathan): https://x.com/labenzLinkedIn: https://linkedin.com/in/nathanlabenz/Youtube: https://youtube.com/@CognitiveRevolutionPodcastApple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk



