CrawlLedger
Append-only observatory · live pilot

A public change log for the machine-readable web.

CrawlLedger preserves what domains publish in robots.txt, llms.txt, and RSL licensing files—plus when it was observed, how it parsed, and how each record links to the one before it.

20watched domains
392observations
57unique artifacts
2parsed changes

Evidence that stays inspectable

A normal checker tells you what a file says now. CrawlLedger keeps the observation metadata needed to distinguish a current result, an unchanged response, a parse failure, and a fetch skipped by robots policy.

01 / Observe narrowly

One queued domain at a time, limited to public policy files in a reviewed 20-domain pilot.

02 / Preserve provenance

Content hashes, bounded response metadata, parser versions, and per-domain signal chains remain attached to every event.

03 / Report carefully

Pages use factual states such as observed, not observed, unverified, parse failed, or skipped—not legal conclusions.

Recent parsed artifact changes

Only post-baseline byte changes that produced a parsed result are shown here. Changing error pages and routine 304 responses are excluded.

usatoday.com

robots.txt · Observed and parsed

9bf8ba4edf17…artifact hash
usatoday.com

robots.txt · Observed and parsed

b0d291be5914…artifact hash

Recently observed domains

Each domain page is a canonical, crawlable landing page for the public metadata already available from the JSON API.

oreilly.com

2 signal types · last observed Aug 13, 2026

20records
quora.com

1 signal types · last observed Aug 13, 2026

10records
medium.com

2 signal types · last observed Aug 13, 2026

20records
yahoo.com

2 signal types · last observed Aug 13, 2026

20records
reddit.com

1 signal types · last observed Aug 13, 2026

10records

Built for publishers, researchers, and agents

The archive is useful whether you need to audit a crawler rule, cite a historical observation, monitor emerging licensing signals, or consume structured metadata from an API.

Free policy checker

Paste a robots.txt file and evaluate the tracked AI crawler identifiers without triggering a network request or storing the input.

Open the checker →

Source-backed guide

Compare robots.txt, llms.txt, and RSL without treating guidance, crawl rules, and licensing terms as interchangeable.

Read the guide →

CrawlLedger 0.2.0 · first observed Aug 4, 2026 · latest observation