The obligation, and the dates attached to it
Article 53(1)(d) of the AI Act obliges providers of general-purpose AI models to draw up and make publicly available a sufficiently detailed summary of the content used for training, according to a template provided by the AI Office. It sits alongside Article 53(1)(c), the copyright-compliance policy — the two are designed to work together, and in practice they fail together.
| Date | What happens |
|---|---|
| 2 August 2025 | Article 53 GPAI obligations apply Providers of general-purpose AI models placed on the EU market must have a policy to comply with Union copyright law — including honouring Article 4(3) CDSM rights reservations — and must publish a sufficiently detailed summary of training content using the Commission's template. The obligations are legally in force from this date. |
| 2 August 2026 | AI Office enforcement powers and fines begin The Commission gains the power to request documentation, conduct evaluations, compel corrective measures and impose fines on GPAI providers of up to 3% of global annual turnover or €15 million, whichever is higher. The obligation did not change on this date; the ability to enforce it did. |
| 2 August 2027 | Legacy models must comply Models placed on the EU market before 2 August 2025 have until this date to bring themselves into compliance with the Article 53 obligations. |
What “sufficiently detailed” is doing
The recitals frame the summary as serving parties with legitimate interests — including rightsholders — in exercising their rights. That framing is the interpretive key. The summary is not primarily a research transparency instrument; it is an enablement instrument. It exists so that a rightsholder can work out whether their material is plausibly in a model and decide what to do about it.
Read that way, the standard for “sufficiently detailed” becomes less mysterious. A summary is sufficient if it lets an interested rightsholder form a view. A list of three broad categories does not; an itemised manifest of every URL is not required either. The template is the Commission's attempt to fix the resolution somewhere in between — naming major datasets and data sources, describing scraped web data and the crawlers used, identifying licensed sources and categories of other data.
Check the current template before you build against it. This page describes the shape and purpose of the obligation, not a field-by-field reproduction of the template, which the Commission maintains and revises. Work from the AI Office's published version. Note also that a coalition of rightsholder organisations — including the European Publishers Council — has formally rejected the GPAI Code of Practice and the training-data template as inadequate, so pressure for revision is live.
Publishing a summary is the easy half
A summary is a document. Most providers can produce one. What the obligation quietly creates is something harder: a public, dated, attributable statement about your corpus, made under penalty, that anyone can read and that any rightsholder can test against their own knowledge of where their content sits.
The AI Office has said it will not conduct content-level audits of training corpora. Its enforcement posture is reactive — responding to complaints and to alerts from the scientific panel. So the realistic sequence is not an inspection. It is:
- You publish a summary naming your data sources and categories.
- A rightsholder reads it, recognises their material in a described category, and believes they never licensed it.
- They complain — to the AI Office, to a national authority, or in court.
- Someone asks you, for that source, what rights position you relied on and what evidence you hold.
Step four is where the summary stops helping. A summary describes; it does not substantiate. And a crawl log — the artefact most pipelines actually retain — records what you fetched, not what you were permitted to fetch. Those are different facts, and only one of them is a defence.
What substantiation would look like
For each source in the licensed portion of a corpus, the questions an outside party will ask are narrow and consistent:
| Question | What answers it |
|---|---|
| What was granted? | An enumeration of permitted uses — training, retrieval, embedding, display, derivatives — not a single undifferentiated “licensed” flag. |
| At what scope? | Territory, term, model family, and any conditions attached to each granted use. |
| By whom? | An identified issuer whose authority to grant can be checked independently of your say-so. |
| When? | A time anchor that fixes the grant to a date, so validity at the moment of ingestion is a checkable fact rather than a recollection. |
| Was it still valid? | A revocation path the verifier can query — because a license, unlike a paid invoice, can be withdrawn. |
| Can a third party check all of the above without trusting you? | A cryptographic signature verifiable against a published key, offline, by the regulator's own technical staff. |
Nothing in Article 53 requires that these questions be answered with a verifiable credential. It is worth being exact about that, because the opposite claim is made often and it is wrong: the AI Act does not mandate any evidence format, and it does not require providers to obtain verifiable licenses from suppliers. What it creates is a foreseeable demand for substantiation, arriving through a complaint-driven channel, probably years after the fact.
The question for a provider is therefore commercial rather than legal: what does it cost to be able to answer, and when does that cost have to be paid? Evidence gathered at licensing time is nearly free. Evidence reconstructed in 2030, from a counterparty who no longer needs your business, is not.
Three things worth doing now
- Segment the corpus by rights basis — licensed, public domain, exception-dependent, and undocumented. The fourth category is the one that determines your exposure, and most organisations have never sized it.
- Test one licensed source end to end. Pick a single supplier and try to produce per-asset evidence of the grant. The exercise usually reveals within an hour whether the answer is “we have this” or “we have a PDF and a hope.”
- Change the evidence clause in new deals. Requiring a machine-verifiable grant record costs a supplier almost nothing at signature and is close to impossible to obtain retroactively.
Check where you actually stand — free, no login
Point the rights checker at any URL and it reads what that content declares to AI right now: reserved, granted, verified — or silent. ASSESS scores a rights declaration against the public compliance framework and returns the specific gaps.
Check what your content allows Run a free ASSESS scoreChangelog
- 2026-07-31 — Hub pillar expanded: added the Digital Omnibus (Regulation (EU) 2026/1744) compliance timeline, the Article 88 exclusive-competence distinction, the three-tier penalty detail, a 'what counts as evidence' section, and a 'where to start' checklist. Dates re-confirmed against primary sources.
- 2026-07-30 — Page published. Enforcement table, Article 53 / Article 4(3) plain-language summary, and Measure 1.3 consultation status current as of this date.
LicenseFoundry is not a law firm and this page is not legal advice. It is a plain-language technical reference maintained against primary sources — the AI Act text, the CDSM Directive, and published European Commission process documents. Where a question is genuinely unsettled, this page says so rather than resolving it. Verify against the primary sources before relying on any of it in a filing.