Programming Throwdown
Programming Throwdown
Patrick Wheeler and Jason Gauci
172: Transformers and Large Language Models
1 hour 26 minutes Posted Mar 11, 2024 at 3:00 pm.
Programming Throwdown
Is Work From Home Really Work?
Working From Home
What Junior Developers Need to Know About Becoming Senior
Pure Pursuit: A Robotics Training Program
PID without a PhD
Google Launches Gemma on C#
Book of the Show
Book of the Day
Using a Stadia Controller as a Chromecast Controller
Fuse and sshfs
Neural Networks and Large Language Models
Attention layers in the Bayesian inference
Machine Learning: Direct Policy Optimization
The Large Language Model
Is Your Job Dead?
Programming Throwdown
0:00
1:26:08
Download MP3
Show notes

172: Transformers and Large Language Models


Intro topic: Is WFH actually WFC?

News/Links:


Book of the Show


Patreon Plug https://www.patreon.com/programmingthrowdown?ty=h


Tool of the Show


Topic: Transformers and Large Language Models

  • How neural networks store information
    • Latent variables
  • Transformers
    • Encoders & Decoders
  • Attention Layers
    • History
      • RNN
        • Vanishing Gradient Problem
      • LSTM
        • Short term (gradient explodes), Long term (gradient vanishes)
    • Differentiable algebra
    • Key-Query-Value
    • Self Attention
  • Self-Supervised Learning & Forward Models
  • Human Feedback
    • Reinforcement Learning from Human Feedback
    • Direct Policy Optimization (Pairwise Ranking)



★ Support this podcast on Patreon ★