Two of those obligations are about content rights: publish a summary of what you trained on, and operate a policy that identifies and complies with rightsholders' reservations. Both of them eventually resolve to the same question — what can you actually show?
This page is a plain-English guide, not legal advice. LicenseFoundry is not a law firm. Consult qualified counsel on how these obligations apply to your organisation.
What the AI Act says about training data
The EU AI Act (Regulation (EU) 2024/1689) is structured in tiers. Most of the attention has gone to the high-risk regime — the one that governs AI in hiring, credit, medical devices and law enforcement. That regime was significantly rescheduled by the Digital Omnibus on AI, now Regulation (EU) 2026/1744, in force since 27 July 2026. Stand-alone high-risk obligations now land in December 2027.
The obligations on general-purpose AI models were not rescheduled. They have applied since 2 August 2025, and from 2 August 2026 the Commission can enforce them.
For anyone whose business touches licensed content, two provisions matter, and they sit next to each other in Article 53:
- Article 53(1)(c) — a GPAI provider must “put in place a policy to comply with Union law on copyright and related rights, and in particular to identify and comply with, including through state-of-the-art technologies, a reservation of rights expressed pursuant to Article 4(3) of Directive (EU) 2019/790.”
- Article 53(1)(d) — a GPAI provider must publish a sufficiently detailed summary of the content used to train the model, using the template published by the AI Office.
Read together, these say something specific: it is no longer enough to have licensed content. You have to be able to demonstrate which content you were permitted to use, and show your working on how you honoured the reservations of everyone who said no. That is a documentation problem before it is a legal one.
Article 53: the training-data obligations
The transparency summary — Article 53(1)(d)
On 24 July 2025 the Commission's AI Office published the mandatory template for the public summary of training content. It applies to every provider of a GPAI model placed on the EU market — including providers of free and open-source models.
The template is deliberately not a work-by-work manifest. It asks for aggregated, narrative disclosure: the general characteristics of the training data, the large publicly available datasets used, and the licensed and third-party sources relied on. The design intent was to give third parties a meaningful view without forcing providers to expose trade secrets.
The practical consequence is the part most teams underestimate. To describe your licensed sources accurately and publicly, you need an internal record of what you licensed that is complete, current, and defensible. A public summary is a claim. If your internal records can't substantiate it, you have published an inaccurate regulatory disclosure — and Article 101 lets the Commission fine providers for supplying incorrect, incomplete or misleading information, separately from the underlying breach.
The copyright policy — Article 53(1)(c)
This is the provision that quietly reshapes the content-licensing market. A GPAI provider must operate a policy to identify and comply with rights reservations expressed under Article 4(3) CDSM. The phrase doing the work is “including through state-of-the-art technologies.” The obligation is not “check once at acquisition.” It is to maintain an ongoing, technically current capability to determine what a rightsholder has reserved — and to honour it.
Reservations change. Licenses are granted, scoped, and withdrawn. A policy that reads a robots.txt once and caches the answer indefinitely is not complying with a reservation that was made afterwards.
Article 4 CDSM: machine-readable rights reservations
Article 4 of the CDSM Directive created a broad text-and-data-mining exception — and then attached a condition. The exception applies only where use has not been expressly reserved by the rightsholder in an appropriate manner, which for content made publicly available online means machine-readable means.
This is the legal foundation under every
Content-Signal: ai-train=no, every RSL License:
directive, and every TDMRep record. A properly expressed reservation
removes the exception. Training on that content without a license is then
simply infringement.
Two things follow, and they cut in opposite directions depending on which side of the market you sit on.
If you hold rights: a reservation is only worth what a
machine can read. A Dutch court has already held that a reservation which
is not properly machine-readable does not bind (Rechtbank Amsterdam,
30 October 2024, ECLI:NL:RBAMS:2024:6563 — a TDM and
media-monitoring case, under appeal). Reserving rights in your terms of
service, in prose, is not the same as reserving them in a form a crawler
can parse.
If you train models: the reservation is the default, and a license is the override. Which means that when you do have permission, you need to be able to prove it — because the baseline legal position on reserved content is that you had no exception to rely on.
We wrote at length about how a reservation should be structured — pointing at a signed record rather than inlining terms — in Pointer, not payload. On the specific question of expressing an opt-out so it binds, see TDM opt-out under Article 4(3).
Who must comply, and by when
Dates reflect the AI Act as amended by the Digital Omnibus on AI — Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026 and in force since 27 July 2026.
| Date | What applies | Who it hits |
|---|---|---|
| 1 Aug 2024 | AI Act enters into force | — |
| 2 Feb 2025 | Prohibited practices; AI literacy obligations (Article 4 subsequently amended — see note) | All operators |
| 24 Jul 2025 | AI Office publishes the training-data summary template | GPAI providers |
| 2 Aug 2025 | GPAI obligations apply (Chapter V), including Article 53. National penalty regimes under Article 99 also become applicable | New GPAI models |
| 2 Feb 2026 | Commission guidelines on Article 6 classification | High-risk providers |
| 27 Jul 2026 | Digital Omnibus on AI — Regulation (EU) 2026/1744 — enters into force | — |
| 2 Aug 2026 | The Commission gains enforcement and fining powers over GPAI providers under Article 101. Article 50 transparency obligations begin to apply. | GPAI providers; most AI systems |
| 2 Dec 2026 | Article 50(2) machine-readable marking duty for generative AI systems already on the EU market before 2 Aug 2026 (3-month transition added by the Omnibus) | Existing gen-AI systems |
| 2 Aug 2027 | GPAI models placed on the market before 2 Aug 2025 must be compliant | Legacy GPAI models |
| 2 Dec 2027 | Stand-alone Annex III high-risk obligations (deferred from 2 Aug 2026) | High-risk providers |
| 2 Aug 2028 | Annex I high-risk — AI embedded in regulated products (deferred from 2 Aug 2027) | Product manufacturers |
Penalties. For providers of GPAI models, Article 101 permits fines of up to 3% of total worldwide annual turnover or €15 million, whichever is higher, imposed by the Commission directly — including for supplying incorrect, incomplete or misleading information. Elsewhere in the Act, Article 99 sets three tiers: 7% / €35M for prohibited practices, 3% / €15M for most other breaches, and 1% / €7.5M for supplying incorrect information to authorities. For SMEs and start-ups the formula inverts to whichever is lower, and Regulation (EU) 2026/1744 added further reduced caps for SMEs and small mid-caps.
What the Omnibus did and didn't move. It deferred the high-risk regime — and it did not defer the GPAI obligations or the 2 August 2026 enforcement date. Article 50 is the nuance: it applies from 2 August 2026 as planned, but generative AI systems already on the market before that date get until 2 December 2026 for the machine-readable marking duty in Article 50(2). If your planning assumed a blanket delay, that assumption is wrong.
Watch for stale secondary sources on two points. Some still cite “2 February 2027” for the Article 50(2) transition — that was the Commission's November 2025 proposal, not the adopted text. And the Omnibus also rewrote Article 4 (AI literacy) from a duty to ensure literacy into a duty to take measures to support it, applicable since 27 July 2026 — anyone quoting the 2024 wording is quoting superseded text.
The Commission's protocol list (Measure 1.3)
The most consequential open process for anyone building in this space is Measure 1.3 of the GPAI Code of Practice — “Identify and comply with rights reservations when crawling the World Wide Web.” Under it, the Commission is assembling a published list of generally agreed machine-readable reservation-of-rights solutions: the set of protocols a GPAI provider can be expected to read and honour.
| Element | Status |
|---|---|
| Stakeholder consultation and call for expression of interest | Opened 1 December 2025; information session 9 December 2025; closed 23 January 2026 (extended from 9 January). |
| Selection criteria for the list | Solutions must be state-of-the-art, technically implementable, and widely adopted. |
| Review cadence | The list is to be reviewed at least every two years, with Commission-facilitated stakeholder workshops. |
| Workshop participation | Signatories to the GPAI Code of Practice are invited by default; other stakeholders express interest through the consultation form. Entry on the EU Transparency Register is not a prerequisite. |
| Publication of the list | Targeted for late 2026. Not yet published as of 31 July 2026. |
One distinction inside Measure 1.3 is worth holding onto, because it is
regularly flattened: instructions expressed through the Robot Exclusion
Protocol (robots.txt) must be respected in any case, while
the duty to honour other appropriate machine-readable protocols is
framed as a best-efforts obligation, limited to those adopted by
international or European standardisation bodies or otherwise
state-of-the-art. The list is what will convert “best
efforts” into something concrete. Deeper on one such protocol:
does robots.txt
satisfy the AI Act?
The scope nuance is the whole story for licensing. Measure 1.3 standardises the reservation — the machine-readable “no.” It does not standardise the grant — the machine-verifiable “yes.” A provider who honours every listed reservation protocol perfectly still has no standard way to evidence the content it did license.
How this will actually be enforced
The AI Office has stated it will not perform content-level audits of training corpora. Its posture is reactive: it responds to complaints and to alerts from the scientific panel. That has a practical consequence most compliance planning misses.
Who enforces what. This distinction is regularly got wrong. Under Article 88, the Commission has exclusive competence to supervise and enforce the GPAI obligations in Chapter V — national market surveillance authorities do not have power to sanction GPAI model providers. National authorities enforce the rest of the Act against other operators, and have been able to do so since August 2025.
If enforcement is complaint-driven, the operative risk is not a scheduled inspection you can prepare for. It is an unscheduled demand, triggered by a third party, asking a specific question about a specific corpus — probably years after ingestion, quite possibly after the staff who ran that pipeline have left. What survives that is documentation that reconstructs itself: per-source records that an outside party can verify without trusting your systems or your account of them.
In the Netherlands, the supervisory picture is further along than most member states: a single multi-sectoral regulatory sandbox coordinated by the RDI and the Autoriteit Persoonsgegevens, spanning eight market-surveillance authorities, becomes operational in August 2026.
What counts as evidence
Here is the gap that sits underneath both obligations.
Almost every organisation licensing content today can produce an artifact. A signed PDF. A purchase order. A row in a rights-management system. An email from a rightsholder saying yes. None of those answer the question a regulator, an auditor, or an opposing counsel actually asks:
Show me that this specific content was licensed to this specific party, for this specific use, on this date — and that the permission was still valid at the moment the content was used.
An invoice proves payment, not permission. A contract proves what two parties agreed, but not that it covered the asset in front of you, and not whether it was later withdrawn. A database row proves what your own system asserts about itself — which is exactly the evidentiary weight a hostile reader assigns to it: your word.
Under adversarial reading, self-produced records are the weakest category of evidence available. That is a general principle, not an AI-specific one — it's why financial statements are audited by someone other than the company, and why a certificate authority is not the website it vouches for.
What holds up better is a record with four properties:
- Signed by an identifiable party — so its origin can be established without trusting the holder.
- Bound to a counterparty and a date — so it evidences a specific grant, not a standing offer.
- Independently verifiable — by a regulator or a court, without depending on any vendor still being in business.
- Revocation-tracked — so “was this still valid at the time of use?” has an answer.
That is what a verifiable credential is. LicenseFoundry issues them as W3C
Verifiable Credentials 2.0, signed with Ed25519, resolvable through
did:web, revocation-tracked via W3C Bitstring Status List
v1.0, and verifiable offline in milliseconds by anyone holding the bytes.
For AI labs, that turns the Article 53(1)(d) summary from an assertion into a position you can substantiate on demand, and gives the Article 53(1)(c) policy a technical mechanism — the “state-of-the-art technologies” the provision asks for.
For publishers and rightsholders, it turns a reservation into something with a companion: the ability to grant specific, scoped, revocable permission to specific labs, and prove later exactly what you granted.
See how credentials are issued and verified · Read the compliance framework
Where to start
- Segment the corpus by rights basis. Which portions rest on a license, which on public-domain status, which on a TDM exception you would have to argue, and which on nothing you have written down. Most organisations cannot produce this split, and it is the first thing any substantiation request forces.
- For the licensed portion, test whether you could produce per-asset evidence. Not the master agreement — evidence, per source, of what was granted, by whom, for which uses, at what scope, and that the grant was valid at the moment of use.
- Check what your own reservations actually say. If you are also a rightsholder, a reservation that is not machine-readable in a form a crawler recognises may not bind at all.
- Change what new deals require. An audit trail that needs to hold up in 2030 has to start accumulating now. Retrofitting evidence onto a 2026 handshake in 2029 is not a project; it is a negotiation with a counterparty who no longer needs you.
Check what your content declares to AI — free →
Do you actually need a credential layer?
Often, no — and we would rather say so here than waste your time.
If your position is “no AI training on my content, ever, no exceptions”: you don't need licensing infrastructure. Express the reservation properly in machine-readable form and enforce it at your CDN. Cloudflare's AI Crawl Control, or equivalent content-signals enforcement, is sufficient. We are not the right answer for that.
If you are an AI lab training only on data you own outright or that is public domain: the Article 53(1)(d) summary still applies to you, but the rights-provenance problem largely doesn't. Your work is documentation, not attestation.
Where a credential layer earns its place is the middle case — and it is a large one. You grant some rights to some counterparties under specific terms, and refuse everything else. That is what a licensing market is. It's what the Leiden Declaration's call for “licensing agreements that prevent use of published work as training data without consent” describes. Blanket blocking cannot express it, and a PDF cannot prove it.
If that's your situation, the question isn't whether you need a record. It's whether the record you have will survive someone reading it adversarially in 2031.
Common questions
Does the EU AI Act apply to me if I'm not in the EU?
The Act applies to providers placing GPAI models on the EU market regardless of where they are established. Territorial scope is one of the more contested areas of the Act and is worth specific legal advice for your situation.
Was the AI Act delayed?
Partly. The Digital Omnibus on AI — Regulation (EU) 2026/1744, in force since 27 July 2026 — deferred the high-risk regime: stand-alone Annex III systems moved to 2 December 2027 and embedded Annex I systems to 2 August 2028. The GPAI obligations and the 2 August 2026 enforcement date were not deferred. Article 50 also applies from 2 August 2026, with one transitional exception: generative AI systems already on the EU market before that date have until 2 December 2026 to meet the machine-readable marking duty in Article 50(2).
What happens on 2 August 2026 under the EU AI Act?
The Commission gains its enforcement and penalty powers over GPAI providers under Chapter V, and Article 50 transparency obligations begin to apply. The underlying GPAI obligations have applied since 2 August 2025 — what arrives is the ability to enforce them. Note that enforcement against GPAI model providers sits exclusively with the Commission under Article 88; national market surveillance authorities do not have competence over Chapter V.
What fines apply to GPAI providers under the EU AI Act?
Under Article 101, the Commission may fine providers of general-purpose AI models up to 3% of total worldwide annual turnover or €15 million, whichever is higher — including for supplying incorrect, incomplete or misleading information.
Has the Commission published its list of approved reservation protocols?
Not yet. The stakeholder consultation under Measure 1.3 of the GPAI Code of Practice ran from 1 December 2025 to 23 January 2026, and publication of the list is targeted for late 2026. Until then, robots.txt is the one protocol that must be respected in any case, and the duty regarding other machine-readable protocols is a best-efforts obligation.
Do open-source model providers have to publish a training-data summary?
Yes. The AI Office's template applies to all providers of GPAI models, including those released under free and open-source licenses. (Some other AI Act obligations do have open-source carve-outs — this one does not.)
Is a C2PA manifest enough to show I had rights?
No. C2PA establishes provenance — what an asset is and where it came from. It does not express who is permitted to use it, for what, or whether that permission still stands. The two compose; neither substitutes for the other. We explain the structural reason in the FAQ.
Does a credential prove I paid for the license?
No, and it doesn't claim to. A credential proves the permission — which rights, on what content, still valid. Payment is a separate fact with separate proof. A credential can reference a settlement ID, but it never vouches for the money.
What about my content — can I check what it currently declares?
Yes, free and without an account. The Rights Checker reads what any URL or file declares to AI systems: robots.txt, Content-Signals, license files and verifiable credentials. Most content turns out to declare nothing at all.
Reserved is the default. Prove your exceptions.
The Act's architecture assumes rightsholders can say no in a machine-readable way, and that anyone who trained anyway can show why they were allowed to. Enforcement of that assumption starts on 2 August 2026.
Check where you actually stand — free, no login
Point the rights checker at any URL and it reads what that content declares to AI right now: reserved, granted, verified — or silent. ASSESS scores a rights declaration against the public compliance framework and returns the specific gaps.
Check what your content allows Run a free ASSESS scoreChangelog
- 2026-07-31 — Hub pillar expanded: added the Digital Omnibus (Regulation (EU) 2026/1744) compliance timeline, the Article 88 exclusive-competence distinction, the three-tier penalty detail, a 'what counts as evidence' section, and a 'where to start' checklist. Dates re-confirmed against primary sources.
- 2026-07-30 — Page published. Enforcement table, Article 53 / Article 4(3) plain-language summary, and Measure 1.3 consultation status current as of this date.
LicenseFoundry is not a law firm and this page is not legal advice. It is a plain-language technical reference maintained against primary sources — the AI Act text, the CDSM Directive, and published European Commission process documents. Where a question is genuinely unsettled, this page says so rather than resolving it. Verify against the primary sources before relying on any of it in a filing.