80k After Hours
80k After Hours
The 80,000 Hours team
Highlights: #151 – Ajeya Cotra on accidentally teaching AI models to deceive us
25 minutes Posted Aug 2, 2023 at 7:07 pm.
Intro
How ML models might develop situational awareness
Why situational awareness makes safety tests less informative
What misalignment *doesn't* mean
Why it's critical to avoid training bigger systems
Why it's hard to negatively reinforce deception in ML systems
Can we require AI to explain its reasons for its actions
Ways AI is like and unlike the economy
0:00
25:25
Download MP3
Show notes

This is a selection of highlights from episode #151 of The 80,000 Hours Podcast.

These aren't necessarily the most important, or even most entertaining parts of the interview — and if you enjoy this, we strongly recommend checking out the full episode:

Ajeya Cotra on accidentally teaching AI models to deceive us

And if you're finding these highlights episodes valuable, please let us know by emailing [email protected].

Get this episode by subscribing to our podcast on the world’s most pressing problems and how to solve them: type ‘80,000 Hours’ into your podcasting app. Or read the transcript.

Highlights put together by Simon Monsour and Milo McGuire