EA Forum Podcast (Curated & popular)
EA Forum Podcast (Curated & popular)
EA Forum Team
“General capability - and capabilities generally - have no good y-axis” by Gregory Lewis🔸
1 hour 7 minutes Posted Aug 8, 2026 at 8:15 pm.
BLUF: To determine whether AI is ‘improving exponentially’, ‘hitting the wall’, or any other claim which involves a quantity or magnitude (e.g. ‘This model was a big leap/small increment’). We need a good y-axis: an interval scale of AI capability which means +1 unit always represents the same degree of ‘how much better’, in the same way +1 degree Celsius is always the same amount of ‘how much hotter’. Yet there is no good y-axis for AI capability. All our measures are of something related-to but clearly not identical-with it, thus ‘true’ AI capability can be a funhouse-mirror reflection of whatever was measured. Specifically: Benchmark score: One small step in benchmark score can be a giant leap in capability, or the opposite, or whatever else. (My 6/10 vs. your 4/10 ≠ I’m 50% better at maths than you). Elo et al: Can give a real y-axis in terms of winning chances, but doesn’t translate outside of beating others. (Going from 50% to 73% to 88% chance to get a higher score than you on a maths test ≠ gaining 0 → 1 → 2 units of maths ability over you) Epoch Capabilities Index: Analogous to IQ, so [...] ---Outline:
Introduction
Both a poor reflection and a dark glass
Human benchmarking also has a y-axis problem
A metrological elegy
The base case: benchmarks (cf. exams)
Elo et al.
(And maybe not quite 'game ability ≡ winning games', after all?)
ECI (cf. IQ)
Intervals Rarely True
Measure endogeneity
Forking IRT
(Dimensions of being, and beating, a bat)
Prediction (cf. chronometry)
Time horizons
Human "capability" is also exponential in time horizon
Intuitive/interpretative prelude
Time horizons and ECI share an axis kink
Perhaps money, as a measure, stinks the least
Finale: AI as normal epistemics
0:00
1:07:38
Download MP3
Show notes
BLUF: To determine whether AI is ‘improving exponentially’, ‘hitting the wall’, or any other claim which involves a quantity or magnitude (e.g. ‘This model was a big leap/small increment’). We need a good y-axis: an interval scale of AI capability which means +1 unit always represents the same degree of ‘how much better’, in the same way +1 degree Celsius is always the same amount of ‘how much hotter’. Yet there is no good y-axis for AI capability. All our measures are of something related-to but clearly not identical-with it, thus ‘true’ AI capability can be a funhouse-mirror reflection of whatever was measured. Specifically: Benchmark score: One small step in benchmark score can be a giant leap in capability, or the opposite, or whatever else. (My 6/10 vs. your 4/10 ≠ I’m 50% better at maths than you). Elo et al: Can give a real y-axis in terms of winning chances, but doesn’t translate outside of beating others. (Going from 50% to 73% to 88% chance to get a higher score than you on a maths test ≠ gaining 0 → 1 → 2 units of maths ability over you) Epoch Capabilities Index: Analogous to IQ, so [...] ---Outline:(03:05) Introduction(04:26) Both a poor reflection and a dark glass(07:51) Human benchmarking also has a y-axis problem(12:27) A metrological elegy(12:49) The base case: benchmarks (cf. exams)(13:30) Elo et al.(16:25) (And maybe not quite 'game ability ≡ winning games', after all?)(18:28) ECI (cf. IQ)(20:56) Intervals Rarely True(25:01) Measure endogeneity(31:12) Forking IRT(34:45) (Dimensions of being, and beating, a bat)(39:40) Prediction (cf. chronometry)(41:58) Time horizons(44:34) Human "capability" is also exponential in time horizon(47:38) Intuitive/interpretative prelude(52:56) Time horizons and ECI share an axis kink(57:13) Perhaps money, as a measure, stinks the least(01:01:07) Finale: AI as normal epistemics(01:06:48) Acknowledgements The original text contained 33 footnotes which were omitted from this narration. ---
First published:
August 4th, 2026
Source:
https://forum.effectivealtruism.org/posts/CQvdadxjCpd7i7kjA/general-capability-and-capabilities-generally-have-no-good-y
---
Narrated by TYPE III AUDIO.
---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.