
Agentic Safety & Alignment: From Predictors to Governed Agents
A practical map from tool-using predictors to governed, evidence-bounded systems
节目笔记
A rigorous map of agentic safety and alignment, from the moment prediction gains tool permissions through scheming, simulated failure cases, corrigibility, prompt injection, AI control, evaluation science, safety cases, and deployment governance. Maya and Leo use an enterprise research agent to distinguish capability from propensity, harmful compliance from unauthorized goal pursuit, and early-warning evidence from real-world prevalence.
来源材料
- Agentic Misalignment in Summer 2026
- Artificial Intelligence Risk Management Framework (AI RMF 1.0)
- Model Evaluation for Extreme Risks
- AI Agents That Matter
- AI Control: Improving Safety Despite Intentional Subversion
















