Show notes
Join Prof. Subbarao Kambhampati and host Tim Scarfe for a deep dive into OpenAI's O1 model and the future of AI reasoning systems. * How O1 likely uses reinforcement learning similar to AlphaGo, with hidden reasoning tokens that users pay for but never see* The evolution from traditional Large Language Models to more sophisticated reasoning systems* The concept of "fractal intelligence" in AI - where models work brilliantly sometimes but fail unpredictably* Why O1's improved performance comes with substantial computational costs* The ongoing debate between single-model approaches (OpenAI) vs hybrid systems (Google)* The critical distinction between AI as an intelligence amplifier vs autonomous decision-makerSPONSOR MESSAGES:***CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. https://centml.ai/pricing/Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. Are you interested in working on reasoning, or getting involved in their events? Goto https://tufalabs.ai/***TOC:1. **O1 Architecture and Reasoning Foundations** [[[[2. **Monte Carlo Methods and Model Deep-Dive** [[[[[3. **System Design and Real-World Applications** [[[[[Transcript: https://www.dropbox.com/scl/fi/d0ef4ovnfxi0lknirkvft/Subbarao.pdf?rlkey=l3rp29gs4hkut7he8u04mm1df&dl=0REFS:[Witch trial scene: flawed logical reasoning.https://www.youtube.com/watch?v=zrzMhU_4m-g[Microsoft–OpenAI partnership evolution and control dynamics.https://www.nytimes.com/2024/10/17/technology/microsoft-openai-partnership-deal.html[Zero-shot chain-of-thought prompting ('Let's think step by step').https://arxiv.org/pdf/2205.11916[Multi-bot game solving with external and internal planning.https://deepmind.google/research/publications/139455/[AlphaGo's Monte Carlo Tree Search and Q-learning.https://www.nature.com/articles/nature16961[Evaluates O1's planning in "Strawberry Fields" benchmarks.https://arxiv.org/pdf/2410.02162[MARCO-O1: Chain-of-Thought + MCTS for improved reasoning.https://arxiv.org/html/2411.14405[Explores LLM "reasoning vs retrieval" debate.https://arxiv.org/html/2403.04121v2[Chain-of-thought prompting (introduces last-letter concatenation).https://arxiv.org/pdf/2201.11903[Transformer attention and "information over-squashing."https://arxiv.org/html/2406.04267v2[Influence functions to understand procedural knowledge in LLMs.https://arxiv.org/html/2411.12580v1(truncated - continued in shownotes/transcript doc)



