AI Safety and Alignment

Failures that need no attacker. Misalignment, emergent behaviour, loss of control, interpretability, and the evaluation practice that tests for capability rather than for exploits.