Programming Throwdown
Programming Throwdown
Patrick Wheeler and Jason Gauci
180: Reinforcement Learning
1 hour 52 minutes Posted Mar 17, 2025 at 3:00 pm.
PhD Throwdown
An AI Explains The Meaning of Hood
Grills vs Pizza Ovens
What is a Senior Engineer? (
The Art of Refactoring Code
Recraft Is the Most Powerful AI Image Platform I've Ever Used
Open Source vs. Closed Source
NASA's list of 10 rules for software development
AMD Radeon RX 9070 XT performance estimates leaked
What Kind of Jobs Are You Happy With?
Book of the Week
Basic Role Playing: The Universal Game Engine
Pandas vs. Pokemon
Foul AI: The Middleman Between Your AI Models and the
Reinforcement Learning
Reinforcement Learning vs Unsupervised Learning
Value Based Reinforcement Learning vs Policy Based Learning
Policy Algorithms for Complexity
Reward Shaping in Machine Learning
AlphaGo 2.8: Trust Region and Probability
AlphaGo: Inferring and Substitution
Model-based Reinforcement Learning
Post-supervised learning: What's the big deal?
Using Reinforcement Learning in the Store
ChatGPT: RLHF and Reinforcement Learning
The Deep Sequile Leap
DeepSeam Sniped OpenAI's AI
Facebook's Reinforcement Learning Explained
A Future of AI Is Not Scary
Reno on the Reno Hotel
0:00
1:52:22
Download MP3
Show notes
Patrick and Jason introduce reinforcement learning and place it alongside supervised and unsupervised learning. They cover Q-learning, SARSA, policy gradients, actor-critic methods, PPO, imitation learning, and why training and evaluating RL systems is so challenging.