AI Deep Dive is a show from The Information's TITV that gets to the bottom of the hardest technical problems in AI — the models, the research and the people building them. Hosted by our AI reporter, Rocket Drew.
Can an AI system do exactly what it was trained to do and still fail us? Rocket Drew joins UC Berkeley computer science professor Stuart Russell to explore where AI’s objectives come from, why human feedback can reward the wrong behavior, and the challenge of proving a powerful system is safe.
Related articles:
Exclusive: Anthropic Research Memo Shows Focus on Rogue Agents, Scheming Models: https://www.theinformation.com/articles/anthropic-research-memo-shows-focus-rogue-agents-scheming-models
Exclusive: OpenAI Technique in ‘Astra’ Model Sparks Security Concerns: https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns
AI Safety Push Sparks Demand for Watchdog Groups. Critics Doubt Their Independence: https://www.theinformation.com/articles/ai-safety-push-sparks-demand-watchdog-groups-critics-doubt-independence
Subscribe:
YouTube: https://www.youtube.com/@theinformation
The Information: https://www.theinformation.com/subscribe_h
Sign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agenda
Follow us:
X: https://x.com/theinformation
IG: https://www.instagram.com/theinformation/
TikTok: https://www.tiktok.com/@titv.theinformation
LinkedIn: https://www.linkedin.com/company/theinformation/
Chapters:
00:00 - Intro
00:43 - The alignment problem
07:18 - How AI misalignment shows up
16:07 - Training AI to imitate humans
26:05 - Can human feedback fix alignment?
37:47 - When humans become the obstacle
40:04 - Did AI take a wrong turn?
44:24 - AI & existential risk
51:09 - AI labs & safety evidence
58:06 - Assistance games & the off switch