CCBot: what it is and how to allow or block it
The crawler behind Common Crawl, a free public web archive that many AI models have been trained on.
Last checked 2026-10-08
- Company
- Common Crawl
- Used for
- The Common Crawl open web archive
- Job
- Collects pages to train AI models
- Follows robots.txt
- Yes
- Name in robots.txt
- CCBot
What CCBot does
CCBot belongs to Common Crawl, a non-profit that builds “an open repository of web crawl data” anyone can download. Many AI developers use that archive as training data.
Blocking CCBot is therefore an indirect training opt-out: it covers any AI company that relies on Common Crawl rather than crawling for itself.
Should you allow it?
A training choice. If you are opting out of training generally, include CCBot, since it reaches more AI developers than any single company’s crawler.
What blocking it changes: Your pages stop being added to future Common Crawl archives. Copies already in past archives stay there.
Allow or block CCBot in robots.txt
Common Crawl documents a robots.txt group for CCBot as the way to block it.
To allow it
User-agent: CCBot Allow: /
To block it
User-agent: CCBot Disallow: /
robots.txt lets every bot in by default, so you only need the allow lines if another rule in your file would otherwise block it. Our robots.txt guide explains how rules for different bots combine.
How it appears in your server logs
Common Crawl publishes this as the bot’s user-agent string, the label it sends with every visit:
CCBot/2.0 (https://commoncrawl.org/faq/)
How to check a visit is really CCBot
CCBot runs on dedicated IP ranges with reverse DNS (hostnames end in crawl.commoncrawl.org). Common Crawl warns that other crawlers falsely use its name.
Our guide to server log checks shows where to find these visits on common hosting setups.
Source: Common Crawl’s CCBot page, checked 2026-10-08. Bots and their rules change often, so check the company’s own page before making a decision that matters.
See which AI bots your site lets in
The free AI readiness check reads your robots.txt and reports which of the main AI crawlers are allowed or blocked, along with everything else that decides whether AI can understand your site.
Check your site free