ClaudeBot
Model training. Anthropic's automatic crawler for web content that could contribute to model training.
AI crawler directory / ClaudeBot
Anthropic's automatic crawler for web content that could contribute to model training.
This summary is tied to a first-party operator page reviewed on Sep 20, 2026. It should be rechecked when the registry version changes.
Anthropic documents these tokens for different purposes. A rule for one does not automatically control the others.
Model training. Anthropic's automatic crawler for web content that could contribute to model training.
AI search. Anthropic's automatic crawler for improving the relevance and accuracy of Claude search results.
User-requested fetch. Anthropic's user-directed fetcher for retrieving web content in response to a Claude user's request.
This example states a crawl preference for the named token. It does not authenticate requests, enforce blocking, or automatically control another token from the same operator.
User-agent: ClaudeBot
Disallow: /Check the official operator source before publishing a policy. User-requested fetchers and non-fetcher control tokens can have different behavior from automatic crawlers.
These states are calculated from each domain's latest available robots.txt record using the versioned CrawlLedger parser. They are observations of published text, not proof of crawler behavior.
| Domain | Parsed state | Matched group | Last observed |
|---|---|---|---|
| anthropic.com | allowed | * | |
| apnews.com | disallowed | ClaudeBot | |
| apple.com | partial | * | |
| cloudflare.com | allowed | * | |
| commoncrawl.org | partial | * | |
| google.com | partial | *, Yandex | |
| medium.com | disallowed | Amazonbot, Applebot-Extended, Bytespider, ClaudeBot, FacebookBot, GoogleOther, GPTBot, meta-externalagent | |
| openai.com | partial | * | |
| oreilly.com | partial | ClaudeBot, claude-web, Claude-SearchBot, YouBot, Applebot-Extended, MistralAI-User, meta-webindexer, GoogleAgent-Mariner, Google-Extended, DotBot, archive.org_bot, Pinterestbot, TelegramBot, kakaotalk-scrap, getstream.io/opengraph-bot, RamblerMail, Brave, OAI-SearchBot | |
| perplexity.ai | partial | * | |
| quora.com | disallowed | ClaudeBot | |
| reddit.com | disallowed | * | |
| rslcollective.org | allowed | * | |
| rslstandard.org | allowed | * | |
| stackoverflow.com | unverified | No matching group observed | |
| theguardian.com | disallowed | NewsNow, CCBot, TurnitinBot, PetalBot, MoodleBot, FacebookBot, Bytespider, Mojeek, JenkersBot, Seekr, YouBot, Arquivo-web-crawler, coccocbot-web, SeznamBot, PerplexityBot, yacy, anthropic-ai, ClaudeBot, Claude-SearchBot, Claude-User, AwarioRssBot, AwarioSmartBot, SentiOne, ImageSift, Applebot-Extended, YandexAdditional, YandexAdditionalBot, scalepostAI, Buck, meta-externalagent, Amazonbot, amazon-QBusiness, DuckAssistBot, Google-CloudVertexBot, Amzn-SearchBot, AhrefsBot, AhrefsSiteAudit | |
| usatoday.com | disallowed | ClaudeBot | |
| voxmedia.com | partial | * | |
| yahoo.com | disallowed | ADmantX, AlphaBot, anthropic-ai, AwarioRssBot, AwarioSmartBot, BLEXBot, Buzzbot, Bytespider, CCBot, ChatGPT-User, claritybot, Claude-Web, ClaudeBot, cohere-ai, Diffbot, FacebookBot, FriendlyCrawler, Google-Extended, GPTBot, huggingface, ImagesiftBot, img2dataset, magpie-crawler, Meltwater, Neevabot, news-please, NewsNow, Nutch, omgili, omgilibot, panscient.com, Perplexity-ai, PerplexityBot, PetalBot, PiplBot, scoop.it, Scrapy, Seekr, SentiBot, SeznamBot, TurnitinBot, YouBot, ZumBot | |
| ziffdavis.com | partial | * |
“Allowed” means a selected group was observed without a blocking rule for the tested policy shape. “Partial” means at least one non-empty disallow path was observed. “No matching group” is silence, not an affirmative grant. None of these states authenticate the requester or establish legal permission.