AI Safety and Alignment
Failures that need no attacker. Misalignment, emergent behaviour, loss of control, interpretability, and the evaluation practice that tests for capability rather than for exploits.
-
Three Capability Scores for One Model, and METR Disowned All Three
METR produced three capability estimates for one model, ranging from 11 hours to over 270, and said none of them…
Read More » -
Marin’s Statement on AI Risk
The rapid development of AI brings both extraordinary potential and unprecedented risks. AI systems are increasingly demonstrating emergent behaviors, and…
Read More » -
The AI Alignment Problem
The AI alignment problem sits at the core of all future predictions of AI’s safety. It describes the complex challenge…
Read More » -
“Magical” Emergent Behaviours in AI: A Security Perspective
Emergent behaviours in AI have left both researchers and practitioners scratching their heads. These are the unexpected quirks and functionalities…
Read More » -
Explainable AI Frameworks in 2026: What Still Ships
This page recommended eight explainable AI frameworks in 2021. Only SHAP is still releasing. A census of what stopped, what…
Read More » -
Risks of AI – Meeting the Ghost in the Machine
Because it demands so much manpower, cybersecurity has already benefited from AI and automation to improve threat prevention, detection and…
Read More »