Should You Block AI Crawlers?

Should You Block AI Crawlers?

  • September 4, 2026
  • SEO
No Comments

Every AI assistant that answers a buying question is reading somebody's website to do it...and yours is on the menu whether you ever decided anything or not. The decision exists, it is just being made by default.

Robots.txt now carries a second job. For decades it told search engines where they could go, and the same plain text file now decides whether AI crawlers can read, store and repeat what you publish.

Blocking them is not hygiene and allowing them is not surrender...each choice has a price, and the right default depends on which business you run.

Two Kinds Of AI Crawler

The word crawler is hiding two different visitors, and the whole decision turns on telling them apart.

  • Training: Bots collecting pages to build the model itself, months and/or years before anyone asks a question.
  • Retrieval: Fetches made while a person waits, so the assistant can read a current page and cite it in the answer.

Blocking the first shapes what future models know about you. Blocking the second removes you from answers being written right now...along with the visit that a citation can send.

The Names In Your Log File

These are the agents worth recognizing when they show up in a server log.

  • GPTBot: OpenAI's training crawler, separate from the fetches ChatGPT makes while it browses for a user.
  • ClaudeBot: Anthropic's crawler, with user-triggered fetches arriving under related agent names.
  • PerplexityBot: The index behind an answer engine built almost entirely on citing sources.
  • CCBot: Common Crawl, the open web archive many early training sets were built from.
  • Google-Extended: Not a crawler at all...a robots token that tells Google not to use your pages for AI training, while ordinary Googlebot keeps crawling as before.

The vendors document these names publicly, and the list keeps growing. A quarterly skim of your own logs beats any published roster.

A Request, Not A Lock

Robots.txt is voluntary. The large companies publish their agent names and honor the file...smaller and hungrier operations have been widely reported to ignore it.

So the file is a policy statement, not a fence. If a block genuinely has to hold, it gets enforced at the server or the CDN (content delivery network) by matching the agent and refusing the request. That is more work. Most sites never need it.

The Server Bill Nobody Mentions

Crawling is not free for the crawled. AI bots have been widely reported hitting some sites hard enough to strain hosting, and a block written purely for load reasons is legitimate...that one is capacity management rather than strategy.

There is a middle path worth knowing about. Rate limiting slows an aggressive bot without disappearing from it, and most CDNs can hold a crawler to a crawl...the content stays readable while the bill stays flat.

What Blocking Actually Buys

For a publisher whose words are the product, blocking training crawlers keeps new work out of the next round of models from compliant companies. That is a real position...licensing deals between AI companies and publishers exist because the material has value, and withholding it is how a seller negotiates.

What blocking does not do is claw anything back. Pages crawled in years past are already inside models shipping today, and no robots rule reaches into a finished model...the file only governs the next crawl.

The Licensing Backdrop

The block-versus-allow debate is happening inside a market that is still forming. Some large publishers have signed licensing deals with AI companies, others have filed lawsuits, and both groups are betting on what access will be worth once the dust settles.

A small site is not going to negotiate its own deal...that is not the reason to watch this. The reason is that the biggest holders of content treat access as something with a price, which is the strongest available evidence that an open robots.txt gives away something of value. Give it away on purpose, for a return you can name...that is a fine trade.

What Blocking Costs

Assistants that cite sources send visitors, and those visitors arrive with the question already half answered. The referrals show up in analytics under sources like chatgpt.com and perplexity.ai...small counts for most local sites today, moving in one direction.

A site that blocks retrieval fetches cannot be read when the assistant goes looking, which means someone else's page gets cited in the answer. But the deeper cost is quieter than a lost click...the model recommends what it can verify, and a business it cannot read is easy to leave out entirely.

What A Citation Is Worth

A visitor from an AI answer arrives late in the decision...the comparing already happened inside the conversation, and the click sits closer to the end of it. Fewer visitors doing more is the shape of this channel.

Citations also compound quietly. An answer that names you gets read aloud, quoted and pasted into a group chat, and none of that registers as a session...word of mouth ran on the same invisible accounting long before models did.

Questions Are Changing Shape

Some share of the questions that used to be typed into a search box are now asked in a chat window. Nobody outside the AI companies knows the split, and any number quoted this quarter is stale by the next...the direction, though, is not in dispute.

That moves the robots file from a technical footnote to part of being findable. A decade ago the question was whether your site could be crawled at all. The question underneath this article is the same one wearing new agent names.

The Google Catch

Google-Extended looks like the switch that removes you from Google's AI answers. It is not...as things stand it governs training for Google's models, while AI Overviews are assembled from the ordinary search index that Googlebot already built.

The levers that do touch Overviews are the snippet controls, and they price the choice steeply, because they restrict your regular search listings at the same time. You cannot leave the AI answer and keep the search result...for a business that lives on Google traffic, that settles it.

A Default By Business Type

No single rule fits every site, and the split falls along what the words on the page are for.

  • Publishers: The words are the inventory, so there is a case for blocking training bots while the licensing market settles.
  • Stores: Product pages an assistant can read can end in a purchase, so a block mostly hides the shelf.
  • Service Businesses: Being named and cited when a neighbor asks for a recommendation is marketing...staying open is the default that pays.

I think the referral column settles it for most local operators. A plumber's how-to page was never the product...it exists to make the phone ring, and an assistant repeating it with the business name attached is doing the page's job for it.

How To Write The Block

The mechanics are one stanza per agent in robots.txt, and precision matters more than reach.

  • Agents: One user-agent line per bot, spelled exactly as the vendor documents it.
  • Scope: Disallow everything or a single directory...keeping paid content closed while the marketing pages stay open is a legitimate shape.
  • Others: A stanza for one bot changes nothing for the rest, so Googlebot reads on unaffected.

Test it before trusting it. Each vendor documents how its agent reads the file, and a typo in an agent name produces a rule that blocks nobody. That part is dull. Check it anyway.

Check Before You Choose

The policy should follow the numbers, and the numbers take under an hour to pull.

  1. Search a month of server logs for the agent names above and count the visits from each.
  2. Filter analytics referrers for chatgpt.com, perplexity.ai and copilot.microsoft.com, and note the trend rather than the total.
  3. Ask the assistants your customers actually use who they would recommend for what you sell, and record whether you appear.

Run the checks before touching the file and the decision mostly makes itself. A site nobody's assistant reads has little to protect, and a site already earning citations has something to lose.

What To Do This Week

If you run a local service business, leave the crawlers open and spend the effort making pages worth citing instead...the readiness checklist covers what a model can actually repeat. An open robots.txt in front of an unreadable site earns nothing.

If you publish for a living, block the training agents by name, keep the retrieval fetches you want, and enforce the block at the CDN if it matters commercially. Half measures here signal a policy without holding one.

Either way, write the decision down with a date on it. The crawler list will change, the referral numbers will change...the site that rereads its own policy once a year is the one that never gets surprised by it.

No Comments

Categories

A Little About Me

I am a business consultant with a ton of digital experience. I help companies achieve success with a focus on technology and the Web.

Request a callback

I'd be more than happy to discuss your project to see how I can help!

More from my blog

See all posts