
播客William Liu
播客
对谈式的深挖。每个系列围绕一个主题展开,每一期都配有音频、节目笔记与完整文字稿。



Agentic Safety & Alignment: From Predictors to Governed Agents
A rigorous map of agentic safety and alignment, from the moment prediction gains tool permissions through scheming, simulated failure cases, corrigibility, prompt injection, AI control, evaluation science, safety cases, and deployment governance. Maya and Leo use an enterprise research agent to distinguish capability from propensity, harmful compliance from unauthorized goal pursuit, and early-warning evidence from real-world prevalence.
