What the law is actually asking for
Article 4(3) of the CDSM Directive requires that a rightsholder wishing to reserve their work from text and data mining does so in an appropriate manner, such as machine-readable means in the case of content made publicly available online. Recital 18 elaborates only slightly, pointing at metadata and terms and conditions of a website or a service.
That is the entire statutory specification. There is no named protocol, no schema, no conformance test. Two consequences follow and both matter commercially.
- Almost anything machine-parseable is arguable, and nothing is certain. Until the Commission's list lands or a court draws a line, every implementation is a bet on interpretation.
-
Being human-readable is not enough. The Amsterdam
District Court held on 30 October 2024
(
ECLI:NL:RBAMS:2024:6563) that a reservation which was not properly machine-readable did not bind. The decision is under appeal, but it sets the direction of travel: terms buried in prose in a website's T&Cs are a weak reservation.
The mechanisms, compared
Every option below is a way of saying no, or of saying “here are my terms.” They differ in what a crawler has to do to see them, in how widely they are honoured, and in how much they can express.
| Mechanism | How it works | What it can and cannot express |
|---|---|---|
robots.txt user-agent rules |
Per-crawler Disallow directives at the site root, parsed per RFC 9309. The GPAI Code of Practice commits signatories to reading and complying with robots.txt. |
Blanket allow or disallow per named crawler and path. Cannot express purpose (train vs. search vs. inference), scope, duration, or price. Requires you to know every crawler's user-agent string in advance. |
| Cloudflare Content Signals | A structured policy served through robots.txt and HTTP headers, separating search, ai-input and ai-train as distinct signals. |
Purpose-separated reservation — the important advance over plain robots.txt, because “index me but don't train on me” is finally expressible. Still a standing declaration, not a per-counterparty grant. |
| TDM Reservation Protocol (TDMRep) | An EDRLab-developed convention expressing reservations via HTTP header, a well-known file, or in-page metadata, designed explicitly against Article 4(3). | Purpose-built for the legal test and granular to the asset level. Adoption outside publishing is thin, which matters because the Commission's criteria include wide adoption. |
X-Robots-Tag / meta robots |
Response header or in-page tag carrying directives such as noai. |
Cheap and per-URL. No standardised AI vocabulary, so honouring is inconsistent between crawlers. |
| RSL — Really Simple Licensing 1.0 | A License: directive in robots.txt pointing to a machine-readable license.xml declaring terms across paths. Backed by Cloudflare, Akamai, Fastly and Creative Commons. |
The most expressive of the standing-declaration options: terms, not just refusal, and a server= endpoint where terms resolve. Still a standing offer published before any specific request — not a record that a named party was granted anything. |
ai.txt |
A convention placing AI-usage terms at a well-known path. | Simple, but not a formal standard and unevenly honoured. Treat as supplementary, not load-bearing. |
| C2PA Content Credentials | A tamper-evident provenance manifest embedded in the media file at creation, in production across Adobe and Microsoft tooling and in-camera on flagship Sony, Nikon and Leica bodies. | Answers what is this and where did it come from, and can carry a training and data-mining assertion. It is not a license: one asset has one provenance but many licenses, and you cannot embed a set of counterparty-specific revocable grants into a single manifest. |
| Paywalls, authentication, IP blocking | Access control rather than signalling. | Effective against compliant and non-compliant crawlers alike, and relevant to whether access was “lawful” at all. Expresses nothing machine-readable about terms, and blocks the search traffic you may want. |
One thing this table deliberately omits.
llms.txt is sometimes listed alongside these as a
permissions mechanism. It is not one. It is a readability and
discovery convention that helps a model find and parse your content.
Citing it as a rights reservation is a category error and will not
help you in a dispute.
What every mechanism on that list has in common
Read down the right-hand column and the pattern is hard to miss. Every one of these is reservation-side. They are increasingly sophisticated ways of expressing a default — no, or not for training, or here are my standing terms. Not one of them produces a record that a specific counterparty was granted a specific set of rights on a specific date, revocable, and checkable by a third party afterwards.
That asymmetry is structural, not accidental. Reservation is a broadcast: one publisher, one signal, everybody reads it. A grant is a transaction: two named parties, a scope, a term, and a state that can change. Broadcast formats are the wrong shape for transactions, which is why every attempt to inline license terms into a reservation file runs into the same four walls — no counterparty binding, no revocation, no signature that survives rehosting, and a frozen vocabulary. We wrote that argument up separately in Pointer, not payload.
The practical upshot for a publisher: the reservation stack tells the world what you refuse. It cannot tell an AI lab's counsel what you permitted, and it cannot tell a regulator, four years later, what you permitted then.
A defensible configuration today
- Set a purpose-separated default. Content Signals in robots.txt, distinguishing search from AI input from AI training, so you are not choosing between discoverability and reservation.
- Reinforce it in a second channel. Response headers or in-page metadata, so a crawler that skips robots.txt still meets the signal. Redundancy is cheap and the legal test is about whether the reservation was reasonably discoverable.
-
Declare terms, not just refusal. An RSL
license.xmlturns a dead end into an invitation to license. A block that offers no path to a deal converts nothing. -
Keep the grant layer separate and signed. When you do
license, issue the grant as a signed, revocable credential referenced
by a stable URI — not as terms inlined into a file every crawler reads.
RSL 1.0 anticipates exactly this: §3.13 defines
<legal type="proof">and names verifiable credentials as its cryptographic-evidence layer. - Date and version your configuration. If a dispute turns on what your site declared in March 2027, you need to be able to show what it declared in March 2027.
Check where you actually stand — free, no login
Point the rights checker at any URL and it reads what that content declares to AI right now: reserved, granted, verified — or silent. ASSESS scores a rights declaration against the public compliance framework and returns the specific gaps.
Check what your content allows Run a free ASSESS scoreChangelog
- 2026-07-31 — Hub pillar expanded: added the Digital Omnibus (Regulation (EU) 2026/1744) compliance timeline, the Article 88 exclusive-competence distinction, the three-tier penalty detail, a 'what counts as evidence' section, and a 'where to start' checklist. Dates re-confirmed against primary sources.
- 2026-07-30 — Page published. Enforcement table, Article 53 / Article 4(3) plain-language summary, and Measure 1.3 consultation status current as of this date.
LicenseFoundry is not a law firm and this page is not legal advice. It is a plain-language technical reference maintained against primary sources — the AI Act text, the CDSM Directive, and published European Commission process documents. Where a question is genuinely unsettled, this page says so rather than resolving it. Verify against the primary sources before relying on any of it in a filing.