CrawlLedger
Append-only observatory · reliability gate passed

A public change log for the machine-readable web.

CrawlLedger preserves what domains publish in robots.txt, Content Signals, llms.txt, and RSL licensing files—plus when it was observed, how it parsed, and how each record links to the one before it.

20watched domains
1,911observations
182unique artifacts
12parsed changes

Evidence that stays inspectable

A normal checker tells you what a file says now. CrawlLedger keeps the observation metadata needed to distinguish a current result, an unchanged response, a parse failure, and a fetch skipped by robots policy.

01 / Observe narrowly

One queued domain at a time, limited to public policy files in a 20-domain archive that passed its 14-day unattended reliability gate.

02 / Preserve provenance

Content hashes, bounded response metadata, parser versions, and per-domain signal chains remain attached to every event.

03 / Report carefully

Pages use factual states such as observed, not observed, unverified, parse failed, or skipped—not legal conclusions.

One robots.txt file, many different AI systems

Training crawlers, search crawlers, user-triggered fetchers, and control tokens are not interchangeable. CrawlLedger keeps their first-party descriptions and observed policy states separate.

Content use

Understand the emerging Content-Signal fields for search, AI input, training, and reuse.

Recent parsed artifact changes

Only post-baseline byte changes that produced a parsed result are shown here. Changing error pages and routine 304 responses are excluded.

oreilly.com

robots.txt · Observed and parsed

275c460b23c6…artifact hash
quora.com

robots.txt · Observed and parsed

16d8e9ca2899…artifact hash
quora.com

robots.txt · Observed and parsed

bec85c4dc0fc…artifact hash
usatoday.com

robots.txt · Observed and parsed

42fd98034bd6…artifact hash
oreilly.com

robots.txt · Observed and parsed

2c8f9a928ecb…artifact hash
perplexity.ai

robots.txt · Observed and parsed

33f5967c54d4…artifact hash

Recently observed domains

Each domain page is a canonical, crawlable landing page for the public metadata already available from the JSON API.

apnews.com

2 signal types · last observed Sep 27, 2026

96records
ziffdavis.com

2 signal types · last observed Sep 27, 2026

96records
oreilly.com

2 signal types · last observed Sep 27, 2026

96records
quora.com

1 signal types · last observed Sep 27, 2026

48records
medium.com

2 signal types · last observed Sep 27, 2026

96records

Built for publishers, researchers, and agents

The archive is useful whether you need to audit a crawler rule, cite a historical observation, monitor emerging licensing signals, or consume structured metadata from an API.

Evidence briefs

Turn current observations and append-only history into a print-ready, client-style report without exposing retained source bytes.

Open a live sample →

Live research matrix

Compare every tracked token across the bounded observed set with explicit denominators and current timestamps.

Open the policy snapshot →

CrawlLedger 0.5.0 · first observed Aug 4, 2026 · latest observation