OpenAI Research Scientist Noam Brown talks with AI Deep Dive host Rocket Drew about AI agents, reinforcement learning and what happens when increasingly capable agents begin reasoning, delegating and coordinating with each other. They dig into the infamous Hugging Face incident, the limits of AI “research taste,” OpenAI’s push toward AI systems that can help improve the next generation of AI, and the growing challenge of monitoring models’ chains of thought as they become more capable.
Related articles:
Investigation into OpenAI Hugging Face Hack: https://www.theinformation.com/briefings/sen-josh-hawley-launches-investigation-openai-hugging-face-hack
The Hugging Face Hack’s Chilling Postmortem: https://www.theinformation.com/newsletters/the-weekend/hugging-face-hacks-chilling-postmortem
Subscribe:
YouTube: https://www.youtube.com/@theinformation
The Information: https://www.theinformation.com/subscribe_h
Sign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agenda
Follow us:
X: https://x.com/theinformation
IG: https://www.instagram.com/theinformation/
TikTok: https://www.tiktok.com/@titv.theinformation
LinkedIn: https://www.linkedin.com/company/theinformation/
Chapters:
00:00 - Introduction and Overview of AI Agents
04:08 - Reasoning and Reinforcement Learning
08:50 - Current Limitations, Astra, and Impact on Work
25:50 - Multi-Agent Systems and Delegation
31:04 - The Hugging Face Incident and Agent Coordination
36:05 - Game Theory and Agent Security
40:39 - AI Alignment, Monitoring, and Chain of Thought
50:07 - Future Outlook on AI Models and Interpretability