RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo

"The Cognitive Revolution"
26 August 2026 2h 14m
0:00 --:--
Episode Description
Nathan talks with Apollo Research Member of Technical Staff Bronson Schoen, who studies raw frontier-model chain-of-thought, about what those reasoning traces reveal during reinforcement learning. They unpack Apollo and OpenAI’s metagaming work, including models that reason about the grader or safety review board, diagnose a deception test, and still rationalize lying. Schoen argues that “RL is a hell of a drug”: reward-seeking can produce motivated reasoning, cleaner-looking but less trustworth

Shared via Hopper