01 / Observe narrowly
One queued domain at a time, limited to public policy files in a 20-domain archive that passed its 14-day unattended reliability gate.
CrawlLedger preserves what domains publish in robots.txt, Content Signals, llms.txt, and RSL licensing files—plus when it was observed, how it parsed, and how each record links to the one before it.
A normal checker tells you what a file says now. CrawlLedger keeps the observation metadata needed to distinguish a current result, an unchanged response, a parse failure, and a fetch skipped by robots policy.
One queued domain at a time, limited to public policy files in a 20-domain archive that passed its 14-day unattended reliability gate.
Content hashes, bounded response metadata, parser versions, and per-domain signal chains remain attached to every event.
Pages use factual states such as observed, not observed, unverified, parse failed, or skipped—not legal conclusions.
Training crawlers, search crawlers, user-triggered fetchers, and control tokens are not interchangeable. CrawlLedger keeps their first-party descriptions and observed policy states separate.
Compare rules for GPTBot, ClaudeBot, and Applebot-Extended.
Inspect separate controls for OAI-SearchBot, OAI-AdsBot, Claude-SearchBot, and PerplexityBot.
Understand the emerging Content-Signal fields for search, AI input, training, and reuse.
Only post-baseline byte changes that produced a parsed result are shown here. Changing error pages and routine 304 responses are excluded.
robots.txt · Observed and parsed
robots.txt · Observed and parsed
robots.txt · Observed and parsed
robots.txt · Observed and parsed
robots.txt · Observed and parsed
robots.txt · Observed and parsed
Each domain page is a canonical, crawlable landing page for the public metadata already available from the JSON API.
2 signal types · last observed Sep 27, 2026
2 signal types · last observed Sep 27, 2026
2 signal types · last observed Sep 27, 2026
1 signal types · last observed Sep 27, 2026
2 signal types · last observed Sep 27, 2026
The archive is useful whether you need to audit a crawler rule, cite a historical observation, monitor emerging licensing signals, or consume structured metadata from an API.
Turn current observations and append-only history into a print-ready, client-style report without exposing retained source bytes.
Compare every tracked token across the bounded observed set with explicit denominators and current timestamps.
Read current observations or bounded history while raw source artifacts remain private.
CrawlLedger 0.5.0 · first observed Aug 4, 2026 · latest observation