-
AI Security
Every Thumbs-Up Can Become a Training Label
Flipping 0.3% of preference labels reached a near-perfect attack success rate against safety alignment. The old defence advice assumes a…
Read More » -
AI Security
Testing Can Find Backdoors. It Cannot Prove They Are Not There.
A 2022 proof showed backdoors can be made computationally undetectable under stated assumptions. Scanners do find some of them. None…
Read More » -
AI Security
Your Guardrail Is a Text Classifier
The 2018 literature on fooling text classifiers reads like history until you notice that prompt injection detectors are text classifiers,…
Read More » -
AI Security
The Instruction Comes In Through the Camera
In 2022 a multimodal attack meant fooling a fusion model into misclassifying. In 2026 it usually means an instruction arriving…
Read More » -
AI Security
Query Attacks Buy Reconnaissance
NIST lists query access as an attacker capability, and three different attacks reach a model through it. Where the model…
Read More » -
AI Privacy
Differential Privacy Protects the Training Set, Not the System
What epsilon actually guarantees, why deployed values sit near 10, and what differential privacy leaves uncovered in a 2026 AI…
Read More » -
AI Safety and Alignment
Explainable AI Frameworks in 2026: What Still Ships
This page recommended eight explainable AI frameworks in 2021. Only SHAP is still releasing. A census of what stopped, what…
Read More » -
AI Security
Meta-Attacks: Using Machine Learning to Break Machine Learning
Four research groups have published different attacks under the name meta-attack, and they share no threat model. The machine-learning attack…
Read More » -
AI Security
Saliency Attacks Are Two Attacks, and the Second One Forges Evidence
Saliency attacks name two unrelated attacks. In one, the attacker uses saliency to find a cheap perturbation. In the other,…
Read More » -
AI Security
Attacks on Models That Learn While You Watch
Real-time inference and online learning are different architectures with different threat models. Where attacker-influenced data can reach a parameter update,…
Read More » -
AI Privacy
What Model Inversion Attacks Actually Recover
Model inversion recovers class representatives more often than individuals, its dominant evaluation framework counts adversarial examples as successful reconstructions, and…
Read More » -
AI Security
Data Spoofing Is Four Attacks, and Your Model Answers None of Them
Channel authenticity, measurement integrity and model correctness are three separate properties, and no one of them supplies the other two.…
Read More » -
AI Disinformation
Targeted Disinformation
Targeted disinformation poses a significant threat to societal trust, democratic processes, and individual well-being. The use of AI in these…
Read More » -
AI Security
Model Stealing Split Into Three Different Attacks
The attack the literature calls model stealing has split into three problems that share a name and little else. Query…
Read More » -
AI Disinformation
AI-Exacerbated Disinformation and Threats to Democracy
Recent events have confirmed that the cyber realm can be used to disrupt democracies as surely as it can destabilize…
Read More »