licensefoundry / study
The Rights Check Study

Every piece of content answers to AI now.
Most of it says nothing.

A quarterly, reproducible index of the machine-readable rights signals AI systems actually see — robots.txt AI directives, Content Signals, RSL declarations, and verifiable credentials — across the domains of the world's leading publishers.

Edition: · Scanned: · Cohort: domains · Method open & re-runnable · Methodology

—% of scanned publisher domains publish no machine-readable AI rights signal at all — no reservation, no grant, no proof. Silence isn't consent. But it isn't protection, or proof, either.

The four verdicts

⚪ SILENT

No AI-relevant signal found. Every AI system decides for itself what the silence meant.

🔴 RESERVED

An opt-out is on the record — it blocks compliant use, but proves nothing about what the owner would allow.

🟡 GRANTED

Machine-readable terms are declared — on the owner's say-so alone. Self-attested, not independently verifiable.

🟢 VERIFIED

A signed, revocable credential proves exactly what's allowed — checkable years later by anyone, with no account and no permission.

Key findings

Look up a domain

Every domain in the cohort, with the signals we saw. Not in the cohort? Check any site live instead.

Methodology — open, respectful, reproducible

What we read: only public, machine-readable signals — /robots.txt (AI user-agents such as GPTBot, ClaudeBot, Google-Extended, PerplexityBot; Content-Signals; RSL License: discovery and the referenced license.xml), response headers (X-Robots-Tag: noai), and homepage <head> license/credential links. Where a signal points to a verifiable credential, we resolve and verify it — signature via did:web, revocation via status list — using the open-source content-license-verify reference verifier. The scanner itself is open source and dependency-free: a present-but-failing credential is recorded, never promoted, and a domain we couldn't check is recorded as uncertain, never as silent.

What we don't do: nothing behind logins or paywalls, no article content (metadata only), no natural-language ToS interpretation, no ranking of named "offenders." Aggregates are what we publish; individual results go to the domain owner.

Respect: identified user-agent, metadata-only fetches, rate-limited. Every cohort domain is notified with its result and evidence before publication and can fix and re-scan — fixes before publication count as fixed.

Reproduce it: the scanner, tier rules, crawler list, cohort, and scan date are published. Run the same check on your own domain at licensefoundry.com/#rights-tool.

What does your content tell AI?

Ten seconds, no login, nothing uploaded.

Check your site →   Press & data inquiries