Skip to content
Popproxx brand logo in stylized cursive font.
A person in a dark coat and knit cap stands at night in front of a lit bakery window full of pastries

Your Next Website Visitor Might Be an AI Agent

Jamin Giersbach

I listened to Eric Bigelow of Goodfire talk with Nathan Labenz on The Cognitive Revolution about how AI agents make decisions. Giving tips for working with agents, Bigelow said that "if you're just an agent waking up in this fresh context where you don't know how the world works, like, you do need context." He was talking about agents doing research, not websites.

But it's a fair picture of an AI tool landing on a small business's site. It arrives knowing nothing about you, and your pages are the context it gets.

Google says people are "increasingly gravitating to generative AI experiences to help them find information" (Google Search Central). Some of that software reads your pages to answer a question. Some opens them and clicks around for someone. Here's who they are, what you can control, and a checklist built from the companies' own documentation.

Three kinds of AI visitors

Each bot identifies itself by a name, its user agent. The AI companies publish theirs, sorted into three jobs:

  • Training crawlers collect pages that may be used to train AI models. (A crawler is a program that fetches pages automatically.)
  • Search crawlers index pages so an AI assistant can find them and link to them in its answers.
  • User-initiated fetches happen when an assistant opens your page right then, for the person asking it a question.
CompanyTrainingSearch and answersFetches a user asked for
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
PerplexityNone listedPerplexityBotPerplexity-User
GoogleGoogle-ExtendedGooglebotGoogle-Agent and others

Perplexity says PerplexityBot is not used to crawl content for AI foundation models. Google-Extended isn't a separate crawler; it's a name you use in robots.txt to tell Google how it may use what its existing crawlers collect (Google). Google says Google-Agent is used by agents "to navigate the web and perform actions upon user request" (Google).

The three jobs are separate choices. OpenAI says each of its settings "is independent of the others": you can allow OAI-SearchBot so you appear in ChatGPT's search results while blocking GPTBot to signal that your pages shouldn't be used for training (OpenAI). Anthropic says blocking ClaudeBot signals that your future pages should be left out of training, while blocking Claude-SearchBot "may reduce your site's visibility and accuracy in user search results" (Anthropic).

What robots.txt can and can't do

robots.txt is a plain text file at the root of your site (yoursite.com/robots.txt) that tells bots which addresses they may crawl. The rules date to 1994; the Internet Engineering Task Force wrote them up as RFC 9309 in September 2022. The standard calls them rules "that crawlers are requested to honor," and says plainly: "These rules are not a form of access authorization."

It's a request. Google says "it's up to the crawler to obey them," and some crawlers might not (Google Search Central). Two more limits:

  • It doesn't keep a page out of Google. Google says robots.txt is "not a mechanism for keeping a web page out of Google"; use a noindex tag or a password for that (Google Search Central).
  • User-initiated fetches play by different rules. For ChatGPT-User, OpenAI says "robots.txt rules may not apply." Perplexity says Perplexity-User "generally ignores robots.txt rules," and Google says the same of its user-triggered fetchers. Anthropic is the exception: it says blocking Claude-User stops Claude from retrieving your content in response to a user's question.

If you'd rather keep your pages out of AI training but still be found in AI search, a robots.txt naming only the training bots looks like this:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

The search crawlers aren't named, so this leaves them alone. And robots.txt isn't the only gatekeeper: Google's AI features page lists "Ensuring that crawling is allowed in robots.txt, and by any CDN or hosting infrastructure" (Google Search Central).

Google's AI answers have their own switch

Google works differently. It says AI is built into Search, so robots.txt rules for ordinary Googlebot are what control crawling for its AI Overviews and AI Mode (Google Search Central). Blocking Google-Extended doesn't take you out of them: Google says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal" (Google). It governs whether Google may use your pages to train future Gemini models, and to ground answers in Google's Gemini apps (feed them in as source material while they answer).

The switch for AI Overviews and AI Mode is in Search Console, Google's free tool for site owners, under Settings, then Search generative AI. Google had rolled it out to all websites by August 31, 2026, and "include" is the default. If you exclude your site, its links and content stay out of AI Overviews, AI Mode and generative features in Discover, and you get no traffic from them. Google says the setting isn't used as a ranking or inclusion signal elsewhere in Search (Search Console Help).

As for getting in, Google's answer is ordinary SEO. "There are no additional technical requirements" beyond being indexed and eligible for a snippet, and "all existing SEO fundamentals continue to be worthwhile" (Google Search Central).

How an agent reads a page

Agents go further than reading. Google describes them as systems that "can perform tasks on behalf of people, such as booking a reservation or comparing product specifications." To do that, Google says, browser agents may analyze screenshots, inspect the DOM structure (the page as the browser has built it from your HTML) and interpret the accessibility tree (Google Search Central).

The accessibility tree is the browser's summary of a page for screen readers and other assistive technology: each button, link and form field, its name, and its state, such as checked or expanded. Google's developer guide on web.dev calls it "a high-fidelity map" for an agent. It also shows what breaks that map: a <div> styled to look like a button. On screen it's a button; in the code, an agent "might see a <div> without knowing" it is one.

I like that the advice for agents turns out to be the advice for screen readers. The same web.dev guide says: "Everything we suggest to make a site 'agent-ready' also makes sites better for humans."

What about llms.txt?

llms.txt is a proposal, not a standard. Jeremy Howard proposed it in September 2024 (llmstxt.org): a Markdown file at /llms.txt that gives AI tools a short summary of a site and links to its key pages. (LLM means large language model, the kind of AI behind ChatGPT.) The proposal says it's used most heavily for software documentation.

Does search use it? Google says no. You don't need AI text files or Markdown versions of your pages to appear in Google Search or its AI features, "as Google Search itself doesn't use them," and making one neither helps nor hurts your visibility there (Google Search Central). Lighthouse, Google's page-testing tool, has experimental agent checks that look for one, but they mark a missing file "Not Applicable" because "providing the file is optional at the moment" (Chrome for Developers). The crawler documentation from OpenAI, Anthropic and Perplexity says nothing about their bots reading one, so I found no confirmation that their search tools use it.

For a small business site, I'd do everything below before thinking about one.

A checklist for your site

1. Put the important words in text

Google's list for its AI features includes "Making sure that important content is available in textual form" (Google Search Central). Hours on a photo of the door sign, or prices in a scanned menu, fail that test. Some agents can read a screenshot, but web.dev says that "can be slow and expensive" and is better as a backup (web.dev).

2. Use clear headings and sections

Google says people appreciate pages "organized by paragraphs and sections, along with headings that provide a clear structure to navigate content." It also says you don't need to chop pages into tiny pieces for AI (Google Search Central).

3. State your hours, address, phone and prices plainly

Google says its AI answers can include information about local businesses, and points to Google Business Profile (Google Search Central). Its AI features page lists keeping that profile up to date (Google Search Central). Make your site say the same thing; my local SEO checklist covers the profile.

4. Add structured data, without expecting magic

Structured data is code, using the shared schema.org vocabulary, that labels a page's facts for machines. With Google's LocalBusiness markup you can tell Google about business hours, departments and more. But Google says it "isn't required for generative AI search, and there's no special schema.org markup you need to add" (Google Search Central). It's still worth having for rich results, and Google says it should match the visible text on the page.

5. Keep pages fast and steady

Google lists a good page experience among the basics for its AI features (Google Search Central). Chrome's Lighthouse docs add a reason agents care: layout shifts "can move elements between the time an agent identifies them and the time it attempts an interaction" (Chrome for Developers). My Core Web Vitals guide covers how to measure speed and layout shift.

6. Make buttons, links and forms real HTML

Google's web.dev guide says to use real <button> and <a> elements instead of styled boxes, and to connect each form label to its field. Google's own crawler can generally only follow a link that's an <a> element with an href address (Google Search Central). Ask whoever built your site whether your menu, booking button and contact form are built this way.

7. Decide which bots to allow, on purpose

Open yoursite.com/robots.txt and read it. Decide separately about training and search crawlers, make sure your host or firewall agrees, and check the Search generative AI setting in Search Console.

What this means for you

Google says its AI features in Search work by retrieving pages from its index, then reviewing what those pages say to write the answer, with links (Google Search Central). A page that states your facts plainly gives that process something accurate to work with.

None of this guarantees you'll appear in an AI answer. Google says indexing and serving aren't guaranteed even when a page meets every requirement, and that many suggested AI search "hacks" "aren't effective or supported by how Google Search actually works" (Google Search Central).

The encouraging part is that the list is short and familiar: real text, clear structure, accurate details, fast pages and honest HTML. A site built that way works for the customer at the counter, the one using a screen reader, and the software sent on someone's behalf.

Jamin Giersbach, who runs Popproxx
Written by

Jamin Giersbach

Jamin Giersbach is one of the three designers who started Popproxx in New York City in 2000. Early clients included WebMD, MSD Capital and Bookmans. In 2007 he took Popproxx to Oregon, and he has run it on his own ever since.

Read the full story →