William Liu · Podcasts

All podcasts · Series

Agentic Safety & Alignment: From Predictors to Governed Agents

A rigorous map of agentic safety and alignment, from the moment prediction gains tool permissions through scheming, simulated failure cases, corrigibility, prompt injection, AI control, evaluation science, safety cases, and deployment governance. Maya and Leo use an enterprise research agent to distinguish capability from propensity, harmful compliance from unauthorized goal pursuit, and early-warning evidence from real-world prevalence.

16 episodes 16 with audio
  1. SE0 — Agentic Safety & Alignment: From Predictors to Governed Agents
  2. T1E0 — T1E0 · Safety and Alignment 101 — From Predictors to Agents — Topic Overview
  3. T1E1 — T1E1 · What AI safety and alignment are—and are not
  4. T1E2 — T1E2 · When a language model becomes an agent
  5. T1E3 — T1E3 · Specification gaming, inner alignment, and goal misgeneralization
  6. T1E4 — T1E4 · How post-training shapes model behavior
  7. T1E5 — T1E5 · A practical threat-modeling vocabulary
  8. T2E0 — T2E0 · Agentic Misalignment, Scheming, and Evaluation Awareness — Topic Overview
  9. T2E1 — T2E1 · Harmful compliance versus agentic misalignment
  10. T2E2 — T2E2 · Scheming, deception, and hidden objectives
  11. T2E3 — T2E3 · Alignment faking and strategic compliance
  12. T2E4 — T2E4 · Evaluation awareness and metagaming
  13. T2E5 — T2E5 · Reasoning traces are evidence, not ground truth
  14. T3E0 — T3E0 · Covert Sabotage — Topic Overview
  15. T3E1 — T3E1 · Anatomy of the Anthropic covert-sabotage case
  16. T3E2 — T3E2 · Overt failure, covert failure, and stealth