September 30, 2026

/ AEO/Head

10 min read

What should robots.txt block in 2026

One bad Disallow line can hide your whole site. Here is the 2026 list of what robots.txt should block, what it must never block, and the crawler table.

What should robots.txt block in 2026

In 2026, your robots.txt should block, in this order: (1) private and low value URL paths such as /admin/, /cart/, /checkout/ and internal search results, (2) staging and duplicate parameter URLs, and (3) optionally the training-only bots GPTBot, ClaudeBot, Google-Extended and Applebot-Extended; it should never block Googlebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot or the CSS and JavaScript files your pages need to render. Google’s documentation says Googlebot reads only the first 500 KiB of a robots.txt file and ignores the rest, so the file has to stay short, and Google states plainly that robots.txt “is not a mechanism for keeping a web page out of Google.” That second fact is the one most site owners get wrong.

Most robots.txt files fail in one of two ways. They block too much, and a search or answer engine can no longer fetch the pages that earn traffic. Or they block the wrong thing for the wrong reason, trusting a Disallow line to hide a page that Google will still index by URL. This post gives you the ordered list, the rules behind it, a crawler table, and a copy and paste file.

What is robots.txt actually for?

Robots.txt manages crawl traffic. It is a plain text file at the root of your host that tells crawlers which URLs they may request. Google’s Search Central documentation says its main use is to avoid overloading your site with requests. The format is standardized as RFC 9309, the Robots Exclusion Protocol, and Google’s crawlers follow it.

Four facts from Google’s own specification shape every decision below.

First, the file is per host, protocol and port. A robots.txt at www.example.com does not cover blog.example.com. Second, Google caches the file for up to 24 hours, so a change is not instant. Third, if the file returns a 4xx error (except 429), Google treats it as if no file exists and crawls everything. If it returns a 5xx error, Google pauses crawling, retries for 12 hours, then falls back to the last cached copy for up to 30 days. Fourth, Google supports two wildcards: * matches any run of characters and $ marks the end of a URL.

The important limit: robots.txt controls crawling, not indexing. If another site links to a blocked URL, Google can still index that URL and show it without a description. To keep a page out of Google, use a noindex meta tag or header, or put the page behind a login. And a noindex tag only works if the page is not blocked, because a crawler that cannot fetch the page never sees the tag.

Want to know which of your pages search engines and AI assistants can actually reach today? Run the free AI visibility audit and see where you stand.

What are the five things to block and not block?

Block private paths and duplicate URL spaces first, decide on training bots second, and leave every search and answer crawler alone. Here are the five buckets in order of how much they matter, using the crawlers from Google, OpenAI, Anthropic, Perplexity and Apple as reference points.

1. Block private and transactional paths

Carts, checkout, account pages, admin panels and thank you pages have no search value and burn crawl requests. Disallow them. Remember this is about crawl efficiency, not secrecy: a Disallow line publishes the path to anyone who reads your robots.txt. Anything truly private needs authentication.

2. Block infinite URL spaces

Internal search results, faceted filters, session IDs and calendar archives can generate thousands of near identical URLs. Wildcards handle these cleanly, for example Disallow: /*?sort= or Disallow: /search. Googlebot, Bingbot and others respect these rules.

3. Decide on training only bots

GPTBot, ClaudeBot, Google-Extended and Applebot-Extended exist to control training data. OpenAI says disallowing GPTBot signals your content should not be used to train its generative AI foundation models. Anthropic says ClaudeBot collects web content that could contribute to training. Blocking these is a legitimate policy choice about your content. Whether it changes how often AI assistants cite you is not established by any primary source I could find, so treat it as a licensing decision, not a visibility tactic.

4. Never block search and answer crawlers

OAI-SearchBot, Claude-SearchBot and PerplexityBot are the crawlers that let those products surface your pages. OpenAI states that sites that disallow OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. Anthropic says blocking Claude-SearchBot reduces indexing and visibility in its search responses. Perplexity says PerplexityBot surfaces and links websites in its results and is not used to crawl content for foundation models. If you want to be found, allow all three.

5. Never block rendering resources

Googlebot needs your CSS, JavaScript and images to render a page the way a visitor sees it. A blanket Disallow: /assets/ or Disallow: /wp-includes/ can leave crawlers looking at a broken page. If AI crawlers matter to you, see our post on whether AI crawlers can read JavaScript before you block any script directory.

Which crawlers belong in a robots.txt table?

Group crawlers by what they feed, then apply the recommendation to the group. All names below come from each vendor’s own documentation.

User agentOperatorWhat it feedsRecommendation
GooglebotGoogleSearch, Images, Video, News, DiscoverAllow
BingbotMicrosoftBing SearchAllow
OAI-SearchBotOpenAIChatGPT search resultsAllow
ChatGPT-UserOpenAILive fetches when a user asksAllow (robots.txt may not apply)
GPTBotOpenAIFoundation model trainingYour call, block if you want out of training
Claude-SearchBotAnthropicClaude search result qualityAllow
Claude-UserAnthropicLive fetches when a Claude user asksAllow
ClaudeBotAnthropicModel trainingYour call
PerplexityBotPerplexityPerplexity search resultsAllow
Perplexity-UserPerplexityLive fetches when a user asksAllow (generally ignores robots.txt)
Google-ExtendedGoogleGemini training useYour call, no Search effect
Applebot-ExtendedAppleApple generative model trainingYour call, no Search effect
GoogleOtherGoogleInternal research, one off crawlsAllow unless load is an issue

For a longer engine by engine breakdown, our complete list of AI crawlers covers verification and IP ranges.

Does blocking Google-Extended hurt your Google rankings?

No. Google describes Google-Extended as a standalone product token that publishers use to manage whether content Google crawls may be used to train future generations of Gemini models. Google states it does not impact inclusion in Search or ranking. It also has no separate HTTP user agent string: Google crawls with its existing user agents and reads the token only as a permission signal.

Apple works the same way. Apple’s documentation says Applebot-Extended does not crawl webpages, that pages disallowing it can still appear in search results, and that its rules are not considered in ranking for Search. Disallowing it only opts your content out of training Apple’s general purpose foundation models.

The practical rule: a -Extended token is a training permission switch, not a visibility switch. Blocking Googlebot, by contrast, removes you from Google. Confusing the two is the most expensive robots.txt mistake we see. If you block by category through a CDN or firewall setting, confirm it is not catching Googlebot or the search crawlers listed above.

Which AI fetchers ignore robots.txt?

The user triggered fetchers can. OpenAI’s documentation says that because ChatGPT-User actions are initiated by a user, robots.txt rules may not apply. Perplexity says Perplexity-User generally ignores robots.txt because a user requested the fetch. Anthropic lists Claude-User as the fetcher for user initiated access.

This matters for two reasons. If your goal is to keep content away from AI products entirely, robots.txt will not do it. You would need authentication or a paywall. And if your goal is visibility, there is nothing to fix here: these fetchers visit when a person asks about your page.

Also note that IP blocking backfires. Anthropic warns that blocking by IP address may not reliably prevent crawling and can stop a bot from reading your robots.txt at all, which defeats the purpose. Anthropic publishes its crawler IP ranges so you can verify traffic instead. OpenAI notes that robots.txt changes take about 24 hours to propagate to its systems, so test and wait before you judge the result. Our guide on how to check if ChatGPT can see your website walks through the test.

What should a 2026 robots.txt file look like?

Start from allow everything, then add narrow blocks. This sample blocks transactional paths and URL traps for all crawlers, opts out of training, and leaves search and answer crawlers untouched.

# Default: everyone may crawl, except private and low value paths
User-agent: *
Disallow: /admin/
Disallow: /cart/
Disallow: /checkout/
Disallow: /account/
Disallow: /search
Disallow: /*?sessionid=
Disallow: /*?sort=

# Training only opt outs (optional, a content policy choice)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

Sitemap: https://www.example.com/sitemap.xml

Three details make this file safe. Crawlers follow the most specific matching group, so GPTBot obeys its own block and ignores the wildcard group. Search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot fall under the * group, which blocks only private paths. And the Sitemap line points every crawler to your canonical URLs.

If you would rather stay out of the training decision, delete the four training blocks. The first group alone is a solid default for most sites.

How do you fix the mistakes that hide pages?

Most damage comes from a handful of patterns. Check for them in this order.

A leftover staging rule. A Disallow: / under User-agent: * blocks the entire site. It often ships from a staging environment. Search your live file for it today.

Blocking a page you want deindexed. If a page is already indexed and you add a Disallow, Google can no longer see a noindex tag on it. Remove the Disallow, add noindex, wait for recrawl, then decide whether to block.

Trusting Disallow for privacy. Blocked URLs can still be indexed if linked elsewhere, and the file itself lists your sensitive paths. Use a login.

An oversized file. Anything past 500 KiB is ignored by Google. Consolidate rules with wildcards or restructure directories.

A file that errors. A 5xx response makes Google pause crawling. Monitor the URL like any critical page.

Wrong host. Each subdomain and protocol needs its own file.

To confirm the live state, fetch yoursite.com/robots.txt in a browser, then use the robots.txt report in Google Search Console. If you also care how AI systems read your site, our AI search ranking factors post covers what else gets you cited. Our services page shows how Subscribe PR handles the technical layer for clients who prefer not to do it themselves.

Frequently asked questions

Should I block GPTBot in robots.txt?

Block GPTBot only if you do not want OpenAI to use your content for training its foundation models. OpenAI documents GPTBot as the training crawler and separate from OAI-SearchBot, which handles ChatGPT search. Blocking GPTBot does not remove you from ChatGPT search results. Whether it changes citation rates is not established by any primary source, so decide on your content policy.

Does robots.txt remove a page from Google?

No. Google says robots.txt is not a mechanism for keeping a page out of Google, and a blocked URL can still be indexed if other sites link to it, usually without a description. Use a noindex meta tag or header on a crawlable page, require a login, or remove the page. Then confirm in Search Console.

How big can a robots.txt file be?

Google enforces a limit of 500 KiB. Content beyond that limit is ignored, so rules placed late in an oversized file may never apply. Keep the file short by using wildcards, consolidating repeated paths, and moving excluded content into shared directories.

Will blocking Google-Extended hurt my SEO?

No. Google states Google-Extended is a standalone token controlling whether crawled content may be used to train future Gemini models, and that it does not affect inclusion in Search or ranking. It uses no separate user agent string. Blocking Googlebot, however, removes your pages from Google Search, so keep those two separate.

How long does a robots.txt change take to work?

Expect up to 24 hours. Google generally caches robots.txt for up to 24 hours and may hold it longer if refreshing fails, and OpenAI says changes take about 24 hours to propagate to OAI-SearchBot. Existing indexed pages do not vanish immediately either. Verify the live file, wait a day, then review crawl stats.

Do all AI bots obey robots.txt?

No. OpenAI says robots.txt rules may not apply to ChatGPT-User because a person initiates the fetch, and Perplexity says Perplexity-User generally ignores robots.txt for the same reason. Training and search crawlers such as GPTBot, ClaudeBot and PerplexityBot are documented to respect it. For real access control, use authentication.

What is the short answer to what robots.txt should block?

Block the paths that waste crawl requests: admin, cart, checkout, account, internal search and parameter traps. Decide on training bots as a content policy, knowing Google-Extended and Applebot-Extended never touch your rankings. Never block Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, PerplexityBot or your rendering files. And when the goal is hiding a page, skip robots.txt: use noindex or a login.

Keep the file under 500 KiB, one host per file, and test it after every change. If you want a second set of eyes on what search engines and AI assistants can reach on your site, request your AI visibility audit and we will show you exactly what is blocked.

Tagged

robots.txt ai crawlers technical seo googlebot gptbot