AI Security
Attacks against machine learning models and the pipelines that train them, and the defences that survive contact. Adversarial examples, poisoning, extraction, evasion and backdoors, assessed on what was demonstrated rather than what was announced.
-
When ML Bias Becomes a Security Failure
A demographic differential in false match rate is a demographic differential in exposure to impostors. NIST measured it in 2019…
Read More » -
Adversarial Attacks on AI: What Has Actually Been Demonstrated
Adversarial examples are a real and unsolved property of trained models. Almost every famous demonstration of one attacking a deployed…
Read More » -
Gradient-Based Attacks: Which Gradient, and Where It Stops Working
In the continuous case the gradient hands you the attack. Everywhere else it only ranks candidates, and something slower does…
Read More » -
The Poisoning Tool With Millions of Downloads
Nightshade corrupts a text-to-image concept with fewer than 100 samples, and it has been downloaded millions of times by the…
Read More » -
Every Thumbs-Up Can Become a Training Label
Flipping 0.3% of preference labels reached a near-perfect attack success rate against safety alignment. The old defence advice assumes a…
Read More » -
Testing Can Find Backdoors. It Cannot Prove They Are Not There.
A 2022 proof showed backdoors can be made computationally undetectable under stated assumptions. Scanners do find some of them. None…
Read More » -
Your Guardrail Is a Text Classifier
The 2018 literature on fooling text classifiers reads like history until you notice that prompt injection detectors are text classifiers,…
Read More » -
The Instruction Comes In Through the Camera
In 2022 a multimodal attack meant fooling a fusion model into misclassifying. In 2026 it usually means an instruction arriving…
Read More » -
Query Attacks Buy Reconnaissance
NIST lists query access as an attacker capability, and three different attacks reach a model through it. Where the model…
Read More » -
Meta-Attacks: Using Machine Learning to Break Machine Learning
Four research groups have published different attacks under the name meta-attack, and they share no threat model. The machine-learning attack…
Read More » -
Saliency Attacks Are Two Attacks, and the Second One Forges Evidence
Saliency attacks name two unrelated attacks. In one, the attacker uses saliency to find a cheap perturbation. In the other,…
Read More » -
Attacks on Models That Learn While You Watch
Real-time inference and online learning are different architectures with different threat models. Where attacker-influenced data can reach a parameter update,…
Read More » -
Data Spoofing Is Four Attacks, and Your Model Answers None of Them
Channel authenticity, measurement integrity and model correctness are three separate properties, and no one of them supplies the other two.…
Read More » -
Twitter API for Secure Data Collection in Machine Learning Workflows
While APIs serve as secure data conduits, they are not impervious to cyber threats. Vulnerabilities can range from unauthorized data…
Read More » -
Model Stealing Split Into Three Different Attacks
The attack the literature calls model stealing has split into three problems that share a name and little else. Query…
Read More »