How to Block AI Crawlers in robots.txt Without Hurting Search

In this article
- Which robots.txt user-agent tokens control AI training and which control search?
- How to block AI crawlers in robots.txt, step by step
- What a safe Google-Extended and GPTBot robots.txt looks like
- Why robots.txt is a request, not a lock
- How to check which AI bots still visit your site
- Mistakes that cost search visibility when you opt out of AI training
- FAQ
Want to block AI crawlers without hurting search? Give each AI training token, like GPTBot and Google-Extended, its own robots.txt group with Disallow: /, and leave search crawlers such as Googlebot and OAI-SearchBot allowed. That’s the short version. One caveat before we start: robots.txt is a request that compliant bots honour. It doesn’t enforce anything. Below I go through the tokens that matter, the setup steps, a sample file, where the method falls short, how to verify the result and the mistakes that cost rankings.
Which robots.txt user-agent tokens control AI training and which control search?
AI training and search crawling run on separate user-agent tokens. So you can allow or block each one on its own. A robots.txt user-agent token is the name you put in the User-agent: line, and it’s shorter than the full user agent string a bot sends with its HTTP requests (people mix these two up all the time). At Google, Googlebot handles Search. Google-Extended is the AI training opt-out. Google’s list of common crawlers says Google-Extended doesn’t affect inclusion in Google Search and isn’t used as a ranking signal. And Google-CloudVertexBot? It only covers crawls that site owners request for building Vertex AI Agents. No effect on Search.
OpenAI splits its bots the same way. Disallow the GPTBot user agent and you’re telling OpenAI that crawled content shouldn’t be used to train its foundation models. OAI-SearchBot is a different matter: it decides whether you show up in search results. The OpenAI crawler documentation says each setting is independent of the others. My advice: copy tokens exactly as the vendors list them, and recheck those pages before every edit. The lists change.
How to block AI crawlers in robots.txt, step by step
One group per AI training token with Disallow: /. Your existing search groups stay as they are. That’s really all there is to it.
- Open the current robots.txt at the root of your site.
- Keep the existing Googlebot and
User-agent: *rules unchanged. - Add a group for GPTBot with
Disallow: /. - Add a group for Google-Extended with
Disallow: /. - Save and deploy the file, then fetch the live version in a browser to confirm the server actually returns it.
Don’t put Disallow: / under User-agent: * to stop AI bots. Ever. That rule shuts out search crawlers too, and then you’ve got a much bigger problem than AI training. If you only want a partial AI training opt out, add Allow and Disallow paths inside the AI group instead. Google’s own example groups do exactly this: they allow one archive page and disallow the rest of the archive.
What a safe Google-Extended and GPTBot robots.txt looks like
A safe file keeps search crawlers allowed and gives each AI training bot its own group:
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xmlA crawler follows the group that names its own token and ignores the general one. Simple enough. Some Google crawlers carry more than one token, though, and Google says a rule applies when any one of them matches. The OAI-SearchBot group is optional. Why keep it? It spells out your intent if you want to appear in ChatGPT search results while opting out of training. And your Google-Extended robots.txt entry stays separate from Googlebot, so indexing carries on as before.
Why robots.txt is a request, not a lock
robots.txt tells well-behaved crawlers what you’d prefer. It can’t stop a bot that decides to ignore it. Google states that its common crawlers always obey robots.txt when crawling automatically. Other operators write their own policies, and some skip the file entirely. Tempted to work around that by serving bots different content from what visitors see? Don’t. That’s cloaking, and our guide to cloaking rules for bots explains why search engines penalise it. I’m sticking to robots.txt here, since access controls are a whole separate topic.
How to check which AI bots still visit your site
Server logs tell you which crawlers actually request pages after the change. Our walkthrough on identifying bots in logs covers the filtering. But anyone can fake a user agent string, so confirm what a bot claims to be before you trust it. Reverse DNS lookups handle that, and the guide to PTR checks for bot IPs shows how. In the logs, look for:
- page requests from blocked tokens after the update went live,
- whether those bots fetched robots.txt itself,
- Googlebot crawling at its usual pace.
Be patient with search visibility, it can lag. According to OpenAI’s notes on its bots, its systems can take ~24 hours to adjust search results after a robots.txt update.
Mistakes that cost search visibility when you opt out of AI training
Most lost search traffic comes down to two things: blocking the wrong token, or writing a rule that’s too broad. Block Googlebot-Image or Googlebot-Video by accident and your images or videos drop out of Google Images, Discover and other Search features that show media. Block OAI-SearchBot when you only meant GPTBot, and you’re gone from ChatGPT search results. The reverse happens too. Some owners assume Google-Extended affects rankings and skip the opt-out just to be safe, even though Google says it doesn’t.
So yes, you can block AI crawlers with a few precise groups and leave search alone, as long as you recheck the vendors’ token lists now and then. The rules don’t change if you run your origins behind an NGINX proxy for your sites either, because the file is served from your root domain either way. More crawling guides live in the Jalvo blog archive.
FAQ
Does blocking Google-Extended remove my site from Google Search?
No. Google states that Google-Extended doesn’t affect whether your site is included in Google Search and isn’t used as a ranking signal. Googlebot keeps crawling and indexing your pages as before. The token only controls whether your content is used for Google’s AI training.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot is about training: disallowing it tells OpenAI not to use your content for its foundation models. OAI-SearchBot controls whether your pages show up in ChatGPT search results. You set each one separately, so blocking the first and allowing the second is perfectly fine.
Will robots.txt stop every AI crawler?
No. robots.txt is a voluntary standard, and only compliant bots follow it. Check your server logs to see which crawlers still request pages, and confirm their identity with reverse DNS before you treat them as genuine.


