CrawlLedger
Common Crawl · Open web archive

AI crawler directory / CCBot

CCBot robots.txt reference.

Common Crawl's automatic crawler for building its openly accessible web crawl repository.

20observed domains
6disallowed
9partial policies
4allowed by matched rules

What CCBot is documented to do

This summary is tied to a first-party operator page reviewed on Aug 13, 2026. It should be rechecked when the registry version changes.

Operator
Common Crawl
Documented purpose
Open web archive
Separate HTTP fetcher
Documented as a fetcher or crawler
Official documentation
https://commoncrawl.org/ccbot

Common full-site robots.txt block pattern

This example states a crawl preference for the named token. It does not authenticate requests, enforce blocking, or automatically control another token from the same operator.

User-agent: CCBot
Disallow: /

Check the official operator source before publishing a policy. User-requested fetchers and non-fetcher control tokens can have different behavior from automatic crawlers.

Latest observed pilot policies

These states are calculated from each domain's latest available robots.txt record using the versioned CrawlLedger parser. They are observations of published text, not proof of crawler behavior.

DomainParsed stateMatched groupLast observed
anthropic.comallowed*
apnews.comdisallowedCCBot
apple.compartial*
cloudflare.comallowedCCBot
commoncrawl.orgpartial*
google.compartial*, Yandex
medium.compartial*
openai.compartial*
oreilly.comdisallowedGPTBot, anthropic-ai, Google-Extended, cohere-ai, CCBot, AI2Bot, Amazonbot, Amazonbot-Video, Bytespider, meta-externalagent, Diffbot, omgili, TimpiBot, SeznamBot, Exabot, YandexBot, Sogou, 360Spider, YisouSpider, Baiduspider-Render/2.0, MJ12bot, DataForSeoBot, AliyunSecBot, ArchiveTeam, ArchiveTeam ArchiveBot, AwarioBot, ZoominfoBot, ali-implementer, Blueno, BIGO-baiguoyuan, WorksOgCrawler, OnPageBot, DoCoMo, SAMSUNG-SGH-E250, TA-Googlebot
perplexity.aipartial*
quora.compartial*
reddit.comdisallowed*
rslcollective.orgallowed*
rslstandard.orgallowed*
stackoverflow.comunverifiedNo matching group observed
theguardian.comdisallowedNewsNow, CCBot, TurnitinBot, PetalBot, MoodleBot, FacebookBot, Bytespider, Mojeek, JenkersBot, Seekr, YouBot, Arquivo-web-crawler, coccocbot-web, SeznamBot, PerplexityBot, yacy, anthropic-ai, ClaudeBot, Claude-SearchBot, Claude-User, AwarioRssBot, AwarioSmartBot, SentiOne, ImageSift, Applebot-Extended, YandexAdditional, YandexAdditionalBot, scalepostAI, Buck, meta-externalagent, Amazonbot, amazon-QBusiness, DuckAssistBot, Google-CloudVertexBot, Amzn-SearchBot, AhrefsBot, AhrefsSiteAudit
usatoday.comdisallowedCCBot
voxmedia.compartial*
yahoo.comdisallowedADmantX, AlphaBot, anthropic-ai, AwarioRssBot, AwarioSmartBot, BLEXBot, Buzzbot, Bytespider, CCBot, ChatGPT-User, claritybot, Claude-Web, ClaudeBot, cohere-ai, Diffbot, FacebookBot, FriendlyCrawler, Google-Extended, GPTBot, huggingface, ImagesiftBot, img2dataset, magpie-crawler, Meltwater, Neevabot, news-please, NewsNow, Nutch, omgili, omgilibot, panscient.com, Perplexity-ai, PerplexityBot, PetalBot, PiplBot, scoop.it, Scrapy, Seekr, SentiBot, SeznamBot, TurnitinBot, YouBot, ZumBot
ziffdavis.compartial*

Interpret the state carefully

“Allowed” means a selected group was observed without a blocking rule for the tested policy shape. “Partial” means at least one non-empty disallow path was observed. “No matching group” is silence, not an affirmative grant. None of these states authenticate the requester or establish legal permission.