Skip to content
Popproxx brand logo in stylized cursive font.
A tall wooden gate in an old stone archway, with a small door in it standing open onto a sunlit garden path

Should You Let AI Crawlers Read Your Site?

Jamin Giersbach

In my post on AI visitors, I sorted the bots AI companies send to websites into three jobs: collecting pages to train models, indexing pages for AI search, and fetching a page because a person asked. This is the practical follow-up: deciding which of them get in.

"Should AI read my site?" turns out to be three questions, and you can answer each one differently.

What robots.txt can and can't promise

robots.txt is the plain text file at yoursite.com/robots.txt. It's made of groups: a User-agent line naming a bot, then Allow and Disallow lines for the addresses that bot may or may not crawl. A group headed User-agent: * is for every bot you haven't named.

Google says the file is "used mainly to avoid overloading your site with requests," and that its instructions "cannot enforce crawler behavior" (Google Search Central). The standard behind it, RFC 9309, says it "is not a substitute for valid content security measures." So it's a request, not a lock. Three more details matter here:

  • Names can be faked. Google warns that a user agent string "can be spoofed" (Google), and Common Crawl says it knows of crawlers "falsely identifying themselves as CCBot" (Common Crawl).
  • A bot you name reads only its own group. Under RFC 9309, a crawler obeys the group that matches its name and falls back to the * group only "if no matching group exists."
  • Each address needs its own file. Rules in example.com/robots.txt don't cover subdomains such as m.example.com (Google Search Central). Anthropic asks you to add its opt-out "for every subdomain that you wish to opt out from."

One more from the standard: listing paths in robots.txt "exposes them publicly." Don't use it to hide a private page.

Three questions, not one

Training: may my pages help build AI models?

Training is how an AI model is built, from enormous amounts of text. If your answer is no, these are the names to block, each from the company's own page:

  • GPTBot (OpenAI). Disallowing it "indicates a site's content should not be used in training" OpenAI's foundation models (OpenAI).
  • ClaudeBot (Anthropic). Restricting it "signals that the site's future materials should be excluded" from training (Anthropic).
  • Google-Extended (Google). It isn't a crawler, just a name Google reads in robots.txt. It covers training future Gemini models, and grounding (feeding your pages to the model while it answers) in Gemini Apps and Google's Vertex AI (Google).
  • Applebot-Extended (Apple). It also "does not crawl webpages." Disallowing it opts you out of training the models behind Apple Intelligence and other Apple features (Apple).
  • CCBot (Common Crawl). Common Crawl isn't an AI company. It's a nonprofit that keeps "a free, open repository of web crawl data that can be used by anyone" (Common Crawl), and the research its homepage features includes language models. Blocking CCBot stops it crawling your site.

Perplexity says its crawler "is not used to crawl content for AI foundation models" (Perplexity), so it has no line here.

Saying no to training doesn't take you out of these companies' search, by their own account. OpenAI says a site can allow its search bot "while disallowing GPTBot." Google says Google-Extended "does not impact a site's inclusion in Google Search." Apple says pages that disallow Applebot-Extended "can still be included in search results."

AI search: may assistants find and link to my pages?

These crawlers build the indexes AI assistants search, and link to, when they answer:

  • OAI-SearchBot (OpenAI). Sites that block it "will not be shown in ChatGPT search answers," though they can still appear as plain navigational links (OpenAI).
  • Claude-SearchBot (Anthropic). Blocking it "may reduce your site's visibility and accuracy in user search results" (Anthropic).
  • PerplexityBot (Perplexity), which is "designed to surface and link websites in search results on Perplexity" (Perplexity).
  • Applebot (Apple), which powers search in Spotlight, Siri and Safari. Apple says its data may also feed AI answers to broad questions in Siri and Search, and that a nosnippet tag keeps specific content out of them (Apple).

Google has no separate AI search bot: its AI Overviews and AI Mode use what Googlebot crawls for ordinary Search. Its switch for them is in Search Console, which the earlier post covers. Two details to add. That switch "doesn't affect AI training," which is Google-Extended's job (Search Console Help). And Search Console's Generative AI performance report shows how often links to your site appear in those features, so you can see what you'd give up before switching them off.

Fetches a person asked for

ChatGPT-User, Claude-User, Perplexity-User and Google-Agent visit when someone asks an assistant to open a page. OpenAI, Perplexity and Google say robots.txt may not apply to these, since a person made the request. Anthropic says blocking Claude-User does stop those fetches. The earlier post has the details; the point here is that each of these visits starts with someone's question.

The trade-off for a local business

For most local businesses, I think the search question matters more than the training one. Your site's job is to put the facts people ask about in front of them accurately: what you do, where, when you're open, what it costs. Search crawlers are how an AI answer gets those facts from you and links back. Anthropic's own warning is that blocking its search crawler may cost you "visibility and accuracy."

Training is a different deal. It sends you no visitor and no link; the question is whether you're comfortable with your words helping build a model. For a plumber whose pages list services and hours, that may not weigh much. For a photographer, writer or consultant whose pages are the product, it's a bigger question.

Either way, the companies' docs above say you don't have to trade one for the other.

A robots.txt for yes to search, no to training

If that's where you land, this file lets search crawlers and user fetches in and asks the training bots to stay out. Each name is spelled exactly as its company documents it:

User-agent: *
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: CCBot
Disallow: /

It extends the shorter example in my earlier post with Apple and Common Crawl. A few notes:

  • Search and user bots aren't named, so they follow the * group. Name one only if you want to treat it differently, and then copy in any rules it should still obey.
  • Keep your existing rules. If your file already blocks an admin area under User-agent: *, leave those lines there.
  • Some site builders won't let you edit the file. Google says platforms such as Wix or Blogger may offer a search settings page instead (Google Search Central).
  • Changes take a day or so. OpenAI says its search can take about 24 hours to reflect a robots.txt update, and Perplexity says up to 24 hours.

Then open yoursite.com/robots.txt in a browser and read what's actually there.

If your site sits behind Cloudflare

Cloudflare is a network service that can sit between a website and its visitors. It has its own AI settings, and these block bots outright instead of asking. Under Security Settings, "Configure AI bot policies" sets Search, Agent and Training bots separately: block everywhere, block only on pages with ads, or allow. Cloudflare's docs say that from September 15, 2026, new domains start with Training and Agent bots blocked on pages that display ads, and Search allowed (Cloudflare). It can also write matching robots.txt lines for you, while noting that "robots.txt compliance is voluntary" (Cloudflare). Its AI Crawl Control dashboard, on all plans, shows which AI services visit and whether they follow your robots.txt (Cloudflare).

Make the two agree. Google's list for its AI features includes "Ensuring that crawling is allowed in robots.txt, and by any CDN or hosting infrastructure" (Google Search Central). And Anthropic says blocking its IP addresses instead of using robots.txt "may not work correctly or persistently guarantee an opt-out," because it keeps Anthropic from reading your robots.txt.

What I do on this site

The robots.txt for popproxx.com lets every crawler in, AI bots included. It names none of them. The only addresses it asks bots to skip are /api/, the site's behind-the-scenes endpoints, and /dashboard, where signed-in clients follow their project.

What this means for you

Open your robots.txt this week. If it names no AI bots, every one that honors the file is welcome. That may be exactly what you want, but it should be a choice, not an accident.

Answer the three questions separately. By the companies' own account, you can ask to keep your pages out of training and still be found and linked in their search. And keep the limits in mind: robots.txt is a sign on the gate, not the gate.

The names also change. Google's list of user-triggered fetchers notes that Gemini Notebook's old name, Google-NotebookLM, was supported only until August 2026 (Google). It's worth rereading the vendors' pages now and then rather than trusting a list copied once, this one included.

Jamin Giersbach, who runs Popproxx
Written by

Jamin Giersbach

Jamin Giersbach is one of the three designers who started Popproxx in New York City in 2000. Early clients included WebMD, MSD Capital and Bookmans. In 2007 he took Popproxx to Oregon, and he has run it on his own ever since.

Read the full story →