Which AI Crawlers Can Read Your Website, and Which Matter
Published
Whether an assistant can read your website starts with whether your robots.txt lets its crawler in, and the crawler that answers live questions is not the one most blocklists name. Each AI company runs several, for different jobs. Here is the documented list, vendor by vendor, and what blocking each one costs you.
This is a five-minute check on a file you already own, and on most small-business sites it comes back clean. Every quote below is from the company that runs the crawler, because these four have published the answer themselves and almost nobody in this subject cites them.
Three jobs, one file
The crawlers split by purpose, and that split is the whole thing to understand.
Training. Collects pages that may be used to build a future model. Blocking it is a choice plenty of owners make on principle, and it does nothing to today's answers either way.
Search indexing. Builds the index the assistant searches when a person asks it something now. Blocking this one is how a business disappears from live answers by accident.
User-initiated fetching. Goes and gets a specific page because a person just asked a question that needs it. Two vendors say outright that robots.txt may not govern this case at all, which matters later.
One robots.txt file addresses all of them, by name, in separate groups. So a file can welcome the training crawler and turn away the search crawler, which is exactly backwards from what most owners intend.
The documented list, vendor by vendor
OpenAI (ChatGPT)
OpenAI's crawler documentation names four agents:
- OAI-SearchBot is the one to care about: "OAI-SearchBot is used to surface websites in search results in ChatGPT's search features." OpenAI's own instruction is plain: "To help ensure your site appears in search results, we recommend allowing OAI-SearchBot in your site's robots.txt file."
- GPTBot is the training crawler. "It is used to crawl content that may be used in training our generative AI foundation models," and "Disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models."
- ChatGPT-User handles "certain user actions in ChatGPT and Custom GPTs." OpenAI adds a caveat worth reading twice: "Because these actions are initiated by a user, robots.txt rules may not apply."
- OAI-AdsBot "is used to validate the safety of web pages submitted as ads on ChatGPT," so it only concerns you if you advertise there.
Note the address. That documentation moved: the old platform.openai.com/docs/bots now redirects, and the current page is developers.openai.com/api/docs/bots. If a checklist you were handed points at the old URL, it is old enough to be worth re-reading.
Anthropic (Claude)
Anthropic documents three crawlers, and it is the only vendor of the four that states the cost of blocking in plain sentences:
- Claude-SearchBot "navigates the web to improve search result quality for users." Block it and, in Anthropic's words, that "prevents our system from indexing your content for search optimization, which may reduce your site's visibility and accuracy in user search results."
- Claude-User fetches a page when someone asks: "When individuals ask questions to Claude, it may access websites using a Claude-User agent." Blocking it "prevents our system from retrieving your content in response to a user query, which may reduce your site's visibility for user-directed web search."
- ClaudeBot is the training crawler, "collecting web content that could potentially contribute to their training." Restricting it "signals that the site's future materials should be excluded from our AI model training datasets."
Anthropic also states its position on the file itself: "Anthropic's Bots respect 'do not crawl' signals by honoring industry standard directives in robots.txt." That is Anthropic describing Anthropic's crawlers, and it is not a claim any vendor makes on behalf of the others.
Perplexity
Perplexity documents two, and draws the training line harder than anyone:
- PerplexityBot "is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models." Perplexity also says: "we recommend allowing PerplexityBot in your site's robots.txt file."
- Perplexity-User "supports user actions within Perplexity. When users ask Perplexity a question, it might visit a web page to help provide an accurate answer and include a link to the page in its response." And then the caveat: "this fetcher generally ignores robots.txt rules."
Google (AI Overviews, AI Mode, Gemini)
Google splits the control differently, and the difference trips up a lot of advice.
- Googlebot is the control for everything in Search, the AI features included: "robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search" (Google Search Central). The same page gives the gate for being linked in an AI answer: "To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements."
- Google-Extended is a separate token for training and grounding, something "web publishers can use to manage whether content Google crawls from their sites may be used for training future generations of Gemini models." Google states the consequence directly: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search" (Google).
So blocking Google-Extended does not remove you from AI Overviews, and allowing it does not put you there. For Google, being crawlable and indexable by Googlebot is the whole lever.
What the popular recipe gets wrong
"Block the training bots, allow the search bots" is the advice going around, and it is closer to right than the blanket block it replaced. It is still not the clean switch it is sold as, for two documented reasons.
The first is the user-initiated agents. OpenAI says robots.txt rules "may not apply" to ChatGPT-User, and Perplexity says Perplexity-User "generally ignores robots.txt rules." Both of those fetch pages to answer live questions. Your file is worth getting right. It is not a set of permissions you fully control.
The second is that the categories do not line up across vendors. Perplexity runs one search crawler and says it never feeds a foundation model. Google runs one crawler for Search and a separate token for Gemini training. OpenAI and Anthropic each run three. A recipe written for one vendor's tokens does not translate, which is why the list above is by vendor rather than by rule.
Google's version: indexable, not specially marked up
The upsell in this subject is a file: add an AI text file, add a new kind of markup, get read by the assistants. Google refuses that in its own documentation, which is a better source than we could be: "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add" (Google Search Central).
What Google does document are controls that limit what it shows, and they are worth knowing because they can be on by accident: nosnippet, data-nosnippet, max-snippet and noindex, all used "to limit the information shown from your pages in Search." A page carrying noindex is not eligible for that supporting-link slot at all, whatever your robots.txt says. If a developer added those during a staging build and nobody removed them, the file check below will not find it, so check your page source too.
Our page on why ChatGPT may not know your business covers the llms.txt question with the research behind it, so this page leaves it there.
Your host may have made this decision for you
The other place this gets decided is a dashboard, not a file. Network-level bot controls sit in front of your site and can turn a crawler away before it ever reads your robots.txt, and their categories are not the vendor tokens above.
Cloudflare, which sits in front of a large share of small-business sites, published new defaults for domains onboarding after September 15, 2026: "Bots classified as Training or as Agent are blocked on pages that display ads, while Search remains allowed" (Cloudflare, posted July 1, 2026). Cloudflare also describes the tradeoff its older blanket setting carried, that it "did not apply to mixed-use crawlers because blocking them could also affect search discoverability," and its newer option "lets you easily stay indexed for search while refusing to let that same crawler train on your content" (Cloudflare, September 15, 2026).
Read past the marketing and the practical point is small and useful: if someone clicked a one-button AI block for you, that choice lives in a control panel, in categories, and it is a separate thing to check from the file on your server.
Reachable is not the same as readable
A crawler that gets in still has to find your facts in the HTML it receives. Text that only appears after a browser runs your JavaScript is not there for a reader that does not run it.
There is now one measured number on this for local businesses, and it is somebody else's. TiltStack fetched 62 Atlanta local business sites, 58 of them analyzable, and published the results on July 29, 2026. Its method, in its own words: "We fetched the raw HTML of each site (what a non-JavaScript agent or crawler sees) and checked five machine-readable signals: business structured data, hours, phone, address, and whether content was server-rendered." It found 10% were JavaScript-only shells and 33% published machine-readable hours. That is Atlanta, not Sacramento, and it is TiltStack's measurement rather than ours.
Whatever the local figure is, the two halves of the question are separate and both have to pass. Crawler allowed, facts in the HTML. A site can do the first and fail the second, which is the case that looks fine in a browser and comes back nearly blank to the thing answering your customer.
How to read your own robots.txt in two minutes
Open yourdomain.com/robots.txt in a browser. The file lives at the root of the site and it is public, which means you can read yours, and so can your competitor. Google's documentation is specific about the location: "You must place the robots.txt file in the top-level directory of a site, on a supported protocol."
Four things to look for:
- A Disallow: / under User-agent: *. That is the whole site closed to every crawler that honors the file. It is rare on a live site and usually left over from a staging build.
- Any of the names above in a User-agent: line. OAI-SearchBot, Claude-SearchBot, Claude-User, PerplexityBot, Googlebot. Read what follows it. A Disallow: / under one of those is a specific door closed on a specific assistant.
- Which group actually applies to a crawler. Rules are grouped by user-agent, and per Google, "Only one group is valid for a particular crawler," the one matching most specifically, while "the order of the groups within the robots.txt file is irrelevant" (Google Search Central). So a permissive * group does not rescue a crawler that has a restrictive group of its own, and a line further down the file is not automatically the winner.
- A block you did not write. A security plugin, a hosting default, or a well-meaning developer following a 2023 blog post will happily add a list of AI user-agents. If you find one, the question is which of the named crawlers is which job, and the list above answers it.
If you find nothing, that is the normal result, and it is worth knowing rather than assuming. One caution on what the file does: blocking crawling is not the same as removing a page from results. Google says it "can't index the content of pages which are disallowed for crawling, but it may still index the URL and show it in search results without a snippet."
What a clean file does not buy you
Allowing a crawler removes a blocker. That is all it does, and anyone telling you otherwise is guessing. The vendors do not promise the next step either: Google's own documentation says "Just because a page meets all requirements, best practices, and complies with the policies, doesn't mean that Google will crawl, index, or serve its content. Indexing and serving isn't guaranteed."
Being read is the floor. Being named in an answer also depends on your facts agreeing with each other across the web and on other sources corroborating them, which is the ground our page on the five levers you control covers. And if what you want to know is what the assistants currently say about you, that is the other side of this question: how to check what ChatGPT says about your business is the free method for the output side, where this page is about the input side.
How RevivedLocal fits
We run both of these checks on every business we measure, which is why this page is written from practice rather than from reading. The AI Visibility Monitor scores your site's readability 0 to 100 check by check, and two of those checks are the two halves above: whether your robots.txt leaves the crawlers that answer live questions free to read you, and whether your text loads without JavaScript, the way a crawler reads it. It also captures what the assistants say about your business word for word, each quote tagged with the assistant and the date, including the runs where you are not named. One honest note on scope: we query four assistants: Claude, ChatGPT, Gemini and Perplexity.
Five dollars runs it once and includes a second run 30 days later, so you have two dated observations instead of one. It is $99 a month per location to keep it running. If you want the fixing done as well as the measuring, the SEO Crew is $399 a month with the Monitor included, working on your existing site with a person approving every change before it goes live. When the site itself is the blocker, a rebuild from us ships with a correct robots.txt, a real sitemap, and your business published as structured data from day one.
We are in Orangevale and work with businesses across the Sacramento region in person.
Common questions
How do I know if my website is blocking AI crawlers?
Open yourdomain.com/robots.txt in a browser and read it. Look for Disallow: / under User-agent: *, and for any crawler named by name: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot. Check your CDN or security plugin separately, because a network-level block never appears in that file.
Does blocking GPTBot stop ChatGPT from mentioning my business?
No. GPTBot is OpenAI's training crawler, and OpenAI says disallowing it signals that your content should not be used in training. The crawler behind ChatGPT's search answers is OAI-SearchBot, and OpenAI recommends allowing that one in your robots.txt.
Do AI crawlers obey robots.txt?
The search and training crawlers are documented as honoring it. The user-initiated fetchers are the exception, in the vendors' own words: OpenAI says robots.txt rules "may not apply" to ChatGPT-User, and Perplexity says Perplexity-User "generally ignores robots.txt rules."
My robots.txt is clean. Is my website readable by ChatGPT now?
Not necessarily. A clean file means nothing is turning the crawler away. Your facts still have to be in the HTML that comes back rather than drawn in later by a script, and hours, address, phone and services still have to be stated in plain text.
Do I need an llms.txt file or special AI markup for crawlers to read me?
Google says no: you don't need to create new machine readable files, AI text files, or markup to appear in these features, and there's no special schema.org structured data that you need to add. Ordinary structured data can help Google understand a page, but no AI company says its assistant reads it, and we don't score it.
Can AI assistants read a website built in JavaScript?
Only what arrives in the HTML. TiltStack's July 2026 study of 62 Atlanta local business sites found 10% were JavaScript-only shells, which come back nearly blank to a reader that does not run scripts. Server-render your key facts, or publish them as structured data as well.
Find out what they say about you
Five dollars runs the Monitor on your business and quotes the answers word for word, with a second run 30 days later.
