AI Disinformation

The AI Act’s Transparency Code Relies on Watermarks That Break

Article 50(2) of the EU AI Act requires providers of AI systems generating synthetic audio, image, video or text to mark the outputs in a machine-readable format detectable as artificially generated or manipulated. The obligation applied on 2 August 2026 for systems placed on the market from that date. Systems already on the market received a transitional period ending 2 December 2026.

The AI Office published a Code of Practice on Transparency of AI-Generated Content on 10 June 2026, drawn up through a multi-stakeholder process. It endorses a layered approach: cryptographically signed and timestamped provenance metadata, plus imperceptible watermarking, with optional fingerprinting through a registry. Providers signing it must offer a watermark detection interoperability solution by 2 February 2027. The Commission reported that about 190 organisations across various sectors had signed by the end of July 2026, with named participants including Anthropic, Cohere, Google, Meta, Microsoft, Mistral, OpenAI and Synthesia. Meta signed this Code, having declined the separate General-Purpose AI Code of Practice, and the two are different instruments. The Code applies its watermarking measure to free-form text above a 200-token threshold, roughly 150 words, and exempts very short text below it. The Commission concluded on 8 July 2026 that the Code adequately covers the Article 50 obligations it addresses, and the AI Board adopted its own adequacy assessment the following day.

Article 50 itself prescribes no particular technique. It requires marking that is effective, interoperable, robust and reliable as far as this is technically feasible, which is a technology-neutral standard. The Code is what selects the techniques, and published attacks defeat both of them. Ordinary re-encoding removes provenance metadata. Diffusion-based regeneration removes imperceptible image watermarks. Separate work has demonstrated forgery against specific schemes.

Here is the sentence for your CISO. The legal duty is technology-neutral, the Commission-endorsed route to satisfying it depends on watermarking, and published research both removes and forges those marks, so an absent mark proves nothing and a present one can be planted. Do not build a detection or authentication control on Article 50 compliance.

What Article 50 requires of whom

The provision contains four separate duties addressed to different parties.

Article 50(1) requires providers of systems intended to interact directly with people to design them so users are informed they are interacting with an AI system, unless that is obvious.

Article 50(2) requires providers of systems generating synthetic content to mark outputs in a machine-readable format. This is a provider obligation and it covers the technical marking.

Article 50(3) requires deployers of emotion recognition or biometric categorisation systems to inform the people exposed to them.

Article 50(4) requires deployers who generate or manipulate image, audio or video content constituting a deepfake to disclose that the content is artificially generated or manipulated. For text published to inform the public, deployers must disclose artificial generation where the subject is a matter of public interest, subject to exceptions including editorial review by a natural person holding editorial responsibility.

The split has operational consequences that most summaries get wrong. Article 50(1) and 50(2) are provider duties. Article 50(3) and 50(4) are deployer duties. An organisation that deploys a third-party chatbot does not thereby acquire the 50(1) design duty, though it can become a provider under Article 25 by putting its name on the system or changing its intended purpose. For most deployers the exposure that applies today is 50(4) disclosure of deepfakes and public-interest text, and 50(3) notification for emotion recognition.

What happens to metadata in normal handling

The Code’s first layer is provenance metadata, a signed and timestamped record attached to the file and detectable as altered. C2PA Content Credentials is the mature implementation family and the obvious practical choice. The Code stays technology-neutral and confers no formal safe harbour on any specification.

Signed metadata answers one question well. Intact credentials establish who signed a provenance assertion and whether the signed manifest has been tampered with. They do not establish that every asserted fact is true, and a provenance chain can carry later edits as further signed assertions rather than breaking.

Ordinary handling removes it. Many transcoding, re-encoding and screenshot workflows strip embedded provenance unless preservation is deliberately supported, and platform handling is inconsistent. Absent credentials therefore tell a verifier very little in open distribution, because absence is the expected state for a large share of content that has passed through any platform.

Provenance metadata is therefore close to a positive-only signal in open distribution. Present and valid establishes what was signed. Absent establishes very little. A control treating absence as evidence of either authenticity or synthesis will be wrong at scale. Inside a controlled pipeline that preserves metadata end to end, absence does carry information, which is why this technique works internally and fails on the open internet.

Why watermarks break

The Code’s second layer is imperceptible watermarking, applied at generation and detectable afterwards, and required for free-form text above a length threshold as well as for images, audio and video.

Two attack classes matter and they are not equally understood.

Removal

Diffusion-based regeneration removes invisible image watermarks. The attacker adds noise to a watermarked image and then denoises it with a diffusion model, producing an output perceptually close to the original with the watermark gone, and needs no knowledge of the scheme or its key. Zhao and colleagues established this at NeurIPS 2024 against four pixel-level invisible watermarking schemes, removing 98 per cent of the marks from the most resilient of the four, RivaGAN, while holding image quality close to the unwatermarked original. The same paper proposes semantic-preserving watermarks as an alternative, so the result does not generalise to every scheme.

Statistical text watermarking fails more easily. Paraphrasing with a second model removes the signal and preserves meaning. Shorter passages retain less signal to begin with, which is why the Code applies its text measure above a length threshold.

Forgery

Forgery is the more serious failure and receives far less attention. The Warfare work demonstrates removal and forging in a single framework, using a pre-trained diffusion model to process the content and a generative adversarial network to manipulate the watermark, and reports high success rates on both across datasets and embedding setups while preserving content quality. Forging is the half that matters here: it applies a scheme’s mark to content the provider never generated, which turns a provenance signal into a way of attributing material to someone who did not produce it.

Forgery defeats the regime in a way removal does not. Removal allows synthetic content to pass as authentic, which is the harm Article 50 addresses. Forgery allows authentic content to be marked as synthetic, which gives anyone a way to deny a genuine recording by claiming it contains a generator’s watermark. Forgery also allows an attacker to attribute content to a provider that did not produce it. The provider can rebut through issuance logs or signed provenance. The watermark alone can no longer establish attribution, which is the property the whole scheme was supposed to supply.

securing.ai/ has covered the surrounding problem in AI disinformation and democratic process and targeted disinformation, with the background in the introduction to AI disinformation.

The claim-evidence gap in the Code

The Code of Practice was produced through a process including the organisations that build the generation systems and the provenance standards. It endorses the layered approach on the reasonable basis that two imperfect layers are better than one, and combining a signature with a watermark does raise the cost of an attack.

It does not state the residual risk. A reader of the Code would not learn that regeneration attacks against image watermarks are published and reproducible, that researchers have demonstrated forgery without key access, or that platforms remove metadata as a matter of routine transcoding. The 2 February 2027 detection interoperability requirement commits providers to make detection available and says nothing about detection being reliable.

That gap has a practical consequence. Organisations reading the Code as a technical specification will build detection pipelines and trust their outputs. Those pipelines will produce false negatives on any content that has been through a platform. They will produce false positives wherever forgery is attempted.

What to do with this

Treat Article 50 as a labelling obligation and not as a detection capability. Compliance means marking your own outputs and disclosing your own use. It does not give you a way to determine whether content you received is synthetic.

Confirm your deployer obligations before your provider duties. Article 50(3) and 50(4) apply now to most organisations. An undisclosed deepfake in a marketing video exposes the deployer to up to 15 million euro or 3 per cent of worldwide annual turnover, whichever is higher, under Article 99.

Check the 2 December 2026 date against your product inventory. Any system you placed on the market before 2 August 2026 that generates synthetic content needs machine-readable marking by then.

Design high-consequence authentication so it does not depend on watermarks. For executive payment authorisation, evidentiary video and journalistic verification, use an independently authenticated channel, a call-back on a known number, transaction confirmation or separation of duties, and do not treat watermark detection as proof of origin. Provenance metadata works as a positive signal inside a controlled pipeline that preserves it, which describes an internal workflow and excludes the open internet.

Keep the obligation

Marking synthetic content at generation increases the cost of undetected large-scale synthesis, and no obligation at all would be worse. Provenance metadata inside controlled distribution chains gives verifiers something they would otherwise lack.

The error is reading a legal requirement as a technical guarantee. Article 50 requires providers to apply a mark. It does not require that the mark defeats an adversary, it could not require that, and the Code of Practice claims no such thing. Anyone assuming otherwise has misread both documents.

The other four provisions of the AI Act that touch a security function are covered in the EU AI Act for security teams.

222fb9d292e3d0111656a33900e24a27cfb6a36eb7b202a94a66bb84766154b4?s=120&d=mp&r=g
[email protected] | About me |  Other articles

In the early 2000s, running emerging-technology risk labs at CyberAgency, a defence client asked my team to break the AI systems they planned to put into weapons. We did. That is where my work on AI security started, two decades before the current wave of attention. I kept at it through risk labs at IBM, Accenture, PwC and KPMG. In 2016 I co-wrote a book on AI and leadership. My commercial work today is quantum, at Applied Quantum, which is why this site sells nothing.