The case for yes
- It is explicitly in scope. The GPAI Code of Practice commits signatories to deploying crawlers that read and follow instructions expressed in robots.txt in accordance with RFC 9309, and to identifying and complying with other appropriate machine-readable protocols on top of that.
- It is unambiguously machine-readable. Article 4(3) asks for an appropriate manner, such as machine-readable means. robots.txt is a standardised, parseable file at a well-known location. Whatever else is arguable, its readability is not.
- It is universally adopted. The Commission's Measure 1.3 criteria include wide adoption. Nothing else in this space comes close.
The case for not-so-fast
| Limitation | Why it matters |
|---|---|
| It is not a safe harbour | No provision states that a robots.txt reservation is sufficient. It is one recognised mechanism among several, and no European court has ruled on its adequacy in a specific dispute. |
| It is a blocklist, so it degrades silently | Per-crawler Disallow rules require you to know the user-agent string in advance. New AI crawlers appear constantly. A reservation expressed as a list of names quietly stops covering the traffic that matters. |
| It cannot separate purposes | Plain robots.txt has one verb. “Index me for search but do not train on me” is not expressible, so you choose between discoverability and reservation. Content Signals exists precisely to fix this, splitting search, ai-input and ai-train. |
| Compliance is voluntary | It guides compliant crawlers. It stops nobody else. It is a legal signal, not an access control — for that you need authentication, paywalls or rate limiting, which also bear on whether access was lawful under Article 4(1) at all. |
| It offers no path to a deal | A Disallow is a dead end. As IAB Tech Lab's own bot guidance has been quoted: a block should be an invitation to license, not a dead end. Refusal alone converts nothing. |
| It has no memory | robots.txt tells the world what your site says today. If a dispute turns on what it said at ingestion two years ago, the live file is not evidence. |
The question underneath the question
Most people asking whether robots.txt satisfies the AI Act are really asking one of two different things, and the answers diverge sharply.
“Have I done enough to reserve my rights?”
Probably a reasonable start, and materially better with three additions: a purpose-separated policy such as Content Signals rather than a blanket refusal; a second channel — an HTTP response header or in-page metadata — so a crawler that skips the file still meets the signal; and a dated, versioned record of the configuration so you can later prove what was in force when.
“Have I done enough to be compliant as an AI provider?”
No, and not close. Reading robots.txt discharges part of one limb of Article 53 — the copyright-compliance policy. It says nothing about the training-content summary under Article 53(1)(d), and nothing at all about the content you affirmatively licensed. Honouring every reservation you encounter establishes that you did not take what was withheld. It establishes nothing about whether you were entitled to what you took.
The asymmetry worth naming
robots.txt, Content Signals, TDMRep, meta directives, RSL, paywalls — the entire permission stack that has grown up around AI crawling — is reservation-side. Every layer is a more precise way of saying no. Not one of them can prove a yes.
That is not a criticism of the stack; it is a description of what broadcast formats can do. A reservation is one publisher signalling to everybody. A grant is two named parties, a scope, a term, and a state that changes. You cannot encode the second in the first, which is why attempts to inline license terms into reservation files keep hitting the same wall — no counterparty binding, no revocation, no signature that survives being copied elsewhere.
The consequence is that both sides of every AI content deal are currently well-equipped to document a refusal and badly equipped to document a permission. Which is the wrong way round, given that the permission is the part somebody paid for.
Check where you actually stand — free, no login
Point the rights checker at any URL and it reads what that content declares to AI right now: reserved, granted, verified — or silent. ASSESS scores a rights declaration against the public compliance framework and returns the specific gaps.
Check what your content allows Run a free ASSESS scoreChangelog
- 2026-07-31 — Hub pillar expanded: added the Digital Omnibus (Regulation (EU) 2026/1744) compliance timeline, the Article 88 exclusive-competence distinction, the three-tier penalty detail, a 'what counts as evidence' section, and a 'where to start' checklist. Dates re-confirmed against primary sources.
- 2026-07-30 — Page published. Enforcement table, Article 53 / Article 4(3) plain-language summary, and Measure 1.3 consultation status current as of this date.
LicenseFoundry is not a law firm and this page is not legal advice. It is a plain-language technical reference maintained against primary sources — the AI Act text, the CDSM Directive, and published European Commission process documents. Where a question is genuinely unsettled, this page says so rather than resolving it. Verify against the primary sources before relying on any of it in a filing.