← Fix-it library

Unblock CCBot and get into AI training data

Published· by

Your robots.txt blocks CCBot, the crawler behind Common Crawl - the open web dataset that many AI labs use as training data. Sites absent from Common Crawl are less likely to be known by the long tail of AI models, including new ones that haven’t crawled the web themselves.

Why does Common Crawl matter for AI visibility?

Because it’s upstream of everything. When a lab trains or fine-tunes a model, Common Crawl is frequently in the mix. Being present means models across the ecosystem - not just the big four - have some baseline knowledge of your business. Being absent means starting from zero with every new AI product.

Is there a downside to allowing it?

For content publishers whose text is the product, keeping it out of training sets is a legitimate stance. For businesses that want to be found and recommended, blocking CCBot mostly means being less known. Decide deliberately - this check exists because most blocks aren’t deliberate.

How do you unblock CCBot?

Find and remove the disallow, or add:

User-agent: CCBot
Allow: /

Deploy the updated robots.txt to your web root, then rescan your site to confirm the check passes. And since you’re auditing bot rules anyway, make sure GPTBot made it through the same cleanup.

See where your site stands. The free scan takes about fifteen seconds and shows every fix.

Run a free AI visibility scan

Written by

Abdul Jaafar is the founder of AIOScan and runs Mason, a marketing agency focused on search and AI visibility for local businesses. He built AIOScan because most AI visibility scores are made up, and he wanted one that isn't. More on the about page.