
- SE0 — Agentic Safety & Alignment: From Predictors to Governed Agents
- T1E0 — T1E0 · Safety and Alignment 101 — From Predictors to Agents — Topic Overview
- T1E1 — T1E1 · What AI safety and alignment are—and are not
- T1E2 — T1E2 · When a language model becomes an agent
- T1E3 — T1E3 · Specification gaming, inner alignment, and goal misgeneralization
- T1E4 — T1E4 · How post-training shapes model behavior
- T1E5 — T1E5 · A practical threat-modeling vocabulary
- T2E0 — T2E0 · Agentic Misalignment, Scheming, and Evaluation Awareness — Topic Overview
- T2E1 — T2E1 · Harmful compliance versus agentic misalignment
- T2E2 — T2E2 · Scheming, deception, and hidden objectives
- T2E3 — T2E3 · Alignment faking and strategic compliance
- T2E4 — T2E4 · Evaluation awareness and metagaming
- T2E5 — T2E5 · Reasoning traces are evidence, not ground truth
- T3E0 — T3E0 · Covert Sabotage — Topic Overview
- T3E1 — T3E1 · Anatomy of the Anthropic covert-sabotage case
- T3E2 — T3E2 · Overt failure, covert failure, and stealth