AI Security
Attacks against machine learning models and the pipelines that train them, and the defences that survive contact. Adversarial examples, poisoning, extraction, evasion and backdoors, assessed on what was demonstrated rather than what was announced.
-
AI Model Security: How It Differs from Traditional Cybersecurity, and Where It Doesn’t
For most organisations, AI security is still conventional security applied to an unfamiliar stack. Where AI does add new failure…
Read More » -
Adversarial Training: What It Buys, What It Costs, and Where It Stops Working
Adversarial training is the one defence against adversarial examples that keeps surviving adaptive attack. It also costs an order of…
Read More » -
Robust Optimization Is the Only Defence That Held, and Only Inside Its Own Threat Model
Adversarial training as robust optimization is the only defence in fifteen years of adversarial machine learning that has survived attacks…
Read More » -
AI Deployment Security: Build the Boundary Outside the Model
A model in production will misread a document, accept a malicious premise or propose the wrong action. Securing the deployment…
Read More » -
White-Box vs. Black-Box Attacks: What Access Actually Buys
The attacker's access class used to predict how strong an attack would be. For LLM systems it no longer does…
Read More » -
Article 15 Requires Robustness Nobody Can Measure
Article 15(5) runs to three sentences and most summaries quote only the first. The other two set a proportionality standard…
Read More » -
Your Model Hub Scans for Malware, Not for Behaviour
European buyers are moving to open weights for reasons that make sense. The supply chain those weights arrive with has…
Read More » -
Feature Attribution Attacks: When Explanations Become the Vulnerability
An explanation is computed from the model it describes, which makes it both forgeable and informative. Seven years of published…
Read More » -
Your Model Explanations Are an Attack Surface
Explanation is deployed to build trust in a model and transfers more information to an attacker than to a defender.…
Read More » -
AI Security 101: Where Each Attack Belongs and What Contains It
The mechanisms behind AI attacks have not changed since the spam filters of 2008. What changed is the blast radius.…
Read More » -
Neural Trojans: Backdoors That Survive Safety Training
A backdoor is a specific trigger implanted in a model that stays dormant until it fires. The 2024 result that…
Read More » -
Model Fragmentation Is Model Sprawl, and Nobody Has the Inventory
A quantised derivative of a model you scanned is a different artefact that you did not scan. Every variant carries…
Read More » -
Model Evasion: Why Attackers Rarely Need Adversarial ML
Evasion research and evasion practice have been pointing in different directions for years. The attacks work in the lab, the…
Read More » -
What 250 Poisoned Documents Actually Proved
Two hundred and fifty documents planted the same backdoor in models from 600M to 13B parameters. What it planted was…
Read More » -
Semantic Adversarial Attacks: Leaving the Perturbation Budget Behind
A semantic attack changes the lighting, the hair colour or the phrasing. The result is enormous in pixel distance, obviously…
Read More »