Can We Stop AI Deception? Apollo Research Tests OpenAI's Deliberative Alignment, w/ Marius Hobbhahn

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis
18 September 2025 2h 8m
0:00 --:--
Episode Description
Today Marius Hobbhahn of Apollo Research joins The Cognitive Revolution to discuss their collaboration with OpenAI using "deliberative alignment" to reduce AI scheming behavior by 30x, exploring the safety challenges and concerning findings about models' growing situational awareness and increasingly cryptic reasoning patterns that emerge when frontier models like o3 and o4-mini operate with hidden chains of thought. Check out our sponsors: Fin, Linear, Oracle Cloud Infrastructure. Shownotes

Shared via Hopper