CrawlLedger
Crawler identity · RFC 9309-conscious

A small crawler with a long memory.

CrawlLedger fetches only public policy and agent-guidance files. It does not crawl site content, bypass access controls, solve challenges, or pay for content.

Identifiable on every request

CrawlLedgerBot/0.1.0 (+https://crawlledger.net/crawler)

Operating limits

The pilot is intentionally slower and smaller than the platform can support.

Default files
/robots.txt and, only when permitted, /llms.txt plus explicitly linked RSL documents
Rate
One queued domain job at a time; at least two seconds between same-host policy requests; published crawl-delay honored within the operating window
Caching
ETag and If-Modified-Since conditional requests
Response cap
512 KiB with bounded streaming and one-hop-at-a-time redirect validation
Raw bytes
Retained privately and content-addressed; public pages expose hashes and metadata
Pay Per Crawl
Unverified until authenticated Cloudflare Discovery API access is active

Fail-closed behavior

CrawlLedger stops when its user agent is disallowed. Authentication failures, rate limits, server failures, unclear cross-host permissions, and overlong crawl delays do not become silent permission.

Missing robots.txt

An explicit 404 or 410 permits only the narrow policy-file pilot under the crawler's conservative operational rules.

Cross-host RSL

The linked host's robots.txt is checked independently before an RSL document is requested.

Corrections

New observations append to history. Existing evidence rows and artifact keys cannot be overwritten or deleted by application code.