HomeToolsSEOrobots.txt generator
robots.txt generator and tester with AI crawlers
Build a robots.txt from a template, block private sections, decide which AI crawlers may access your site, and test any URL: is it allowed for Googlebot, Bingbot or GPTBot, and which group and rule decided it.
—
No Crawl-delay or Host. Google does not support Crawl-delay, and Host was a Yandex-only directive retired in 2018. The generator does not add them.
—
Rules of the applied group
| Line | Rule | Length | Match |
|---|
This URL for different crawlers
| Crawler | Group | Access |
|---|
File check
Everything is calculated in your browser — nothing you enter is sent anywhere.
How to use it
- Pick a template“SaaS” blocks the app area and API; “E-commerce” blocks the cart, filters and sorting. Add your own paths and separate groups for Googlebot or Bingbot.
- Decide on AI crawlersBots are grouped by purpose: model training, search with citations, user requests. You can block training and keep search.
- Test URLs“Test this file” moves the result to the “Test URL” tab, which shows the group and rule that decide a URL’s fate for each crawler.
- Publish the fileCopy the text and put it at the domain root: https://example.com/robots.txt. Every subdomain and protocol has its own file.
robots.txt syntax: groups, rules and precedence
robots.txt is a text file at the site root that tells search crawlers which URLs they may crawl. Since 2022 the parsing rules are standardized in RFC 9309; Google and Bing document the same core principles plus a few extensions of their own.
The file is made of groups. A group starts with one or more User-agent lines followed by Disallow and Allow rules. Sitemap lines do not belong to groups and can appear anywhere.
How a crawler picks a group
A crawler looks for the group with its token — case-insensitively — and follows only that group. If there is none, it uses the * group; if that is missing too, everything is allowed. Groups with the same token are merged. Hence a common mistake: creating a separate User-agent: Googlebot group and forgetting to repeat the * rules in it.
Search engines document fallback groups: Googlebot-Image and Googlebot-News use the Googlebot rules when they have none of their own.
How a rule is chosen
A rule value is compared with the start of the path including the query string. * matches any sequence of characters, $ at the end marks the end of the URL. If several rules match, the longest wins; on a tie, Allow wins. Line order does not matter.
This is the precedence table from Google’s documentation; the tool checks it in automated tests. An empty Disallow: blocks nothing. Paths are compared in normalized form: non-ASCII characters are UTF-8 percent-encoded, so /café/ and /caf%C3%A9/ are the same. Paths are case-sensitive: /Admin/ and /admin/ are different rules.
robots.txt for Google and Bing: what differs
The core is shared: groups, Allow and Disallow, the * and $ wildcards, longest-match precedence. The differences are in extensions.
- Crawl-delay. Google ignores it. Bing honors it, but Bing Webmaster Tools’ crawl control is the better place to manage crawl rate.
- Fallback groups. Googlebot-Image, -Video and -News fall back to the Googlebot group when they have no group of their own.
- Size. Google processes the first 500 KiB of the file and ignores the rest.
- Host and Clean-param. These are Yandex directives. Google and Bing ignore them; use 301 redirects for the main host and
rel="canonical"for URL parameters. - Noindex. Google stopped honoring
Noindex:in robots.txt in 2019. Use the robots meta tag or theX-Robots-Tagheader.
To see how a specific search engine reads your live file, use the robots.txt report in Google Search Console or the robots.txt tester in Bing Webmaster Tools. This tool is handy before publishing: it works with any text, not only the file on your site.
How to block AI crawlers, and whether you should
AI companies crawl the web for different purposes, and the big players use a separate token for each. Decide by purpose rather than “block them all”.
| Token | Owner | Purpose |
|---|---|---|
| GPTBot | OpenAI | model training |
| OAI-SearchBot | OpenAI | ChatGPT search |
| ChatGPT-User | OpenAI | user request |
| ClaudeBot | Anthropic | model training |
| Claude-SearchBot | Anthropic | search answers |
| Claude-User | Anthropic | user request |
| PerplexityBot | Perplexity | search and citations |
| Perplexity-User | Perplexity | user request |
| Google-Extended | token: Gemini training and grounding | |
| Applebot-Extended | Apple | token: Apple model training |
| CCBot | Common Crawl | open web archive |
| meta-externalagent | Meta | training and products |
| Amazonbot | Amazon | Amazon services and AI |
| Bytespider | ByteDance | data collection |
Three groups, three decisions
- Model training (GPTBot, ClaudeBot, CCBot and others). Blocking keeps new pages out of future training sets. These bots bring no traffic, so they are the most commonly blocked — especially for paid and original content.
- Search and linked answers (OAI-SearchBot, Claude-SearchBot, PerplexityBot). These bots send visitors: your site is cited with a link. Blocking them means leaving a new channel. SaaS and media sites usually keep them open.
- User requests (ChatGPT-User, Claude-User, Perplexity-User). The page is opened because a person explicitly asked. This is closer to a browser than a crawler; some companies state such requests may not follow robots.txt.
Google-Extended and Applebot-Extended are not crawlers. Pages are still crawled by the regular Googlebot or Applebot; the token only forbids using them to train models. Blocking Google-Extended does not affect Google Search rankings. Blocking Googlebot to stay out of Google’s AI answers would take you out of search too.
Above all, robots.txt is an agreement, not protection. Reputable companies honor it, but you cannot verify that from outside, and some bots ignore the rules. If content must not be served, you need authentication or server or CDN restrictions.
Common robots.txt mistakes
- Blocking CSS and JavaScript. Rules like
Disallow: /*.js$orDisallow: /assets/stop search engines from rendering the page, which may then look empty or not mobile-friendly. - Hiding pages with robots.txt instead of noindex. Blocking crawling does not remove a URL from the index: Google may show it without a description if other pages link to it. To remove a page from search, allow crawling and add
<meta name="robots" content="noindex">or anX-Robots-Tagheader. - Treating robots.txt as security. The file is public. A line like
Disallow: /secret-admin-panel/advertises the path to anyone who reads it. Protect private areas with a login; a generic/admin/in robots.txt is enough — or nothing. - A dedicated group without the shared rules. You add
User-agent: Googlebotfor one exception, and Googlebot stops seeing the*rules. - Rules before User-agent.
Disallowlines at the top of the file, before the first group, are ignored. - A relative Sitemap. Use a full URL:
Sitemap: https://example.com/sitemap.xml. - A leftover
Disallow: /from staging. The staging site was blocked entirely, the file shipped to production, and the site drops out of search. Check robots.txt after every release. - One file for all subdomains. robots.txt applies only to its own host and protocol:
app.example.comandexample.comhave separate files.
A typical SaaS setup: the marketing site is open, the /app/ area is blocked from crawling and protected by login, sign-up and log-in pages are open. If the product lives on a separate subdomain, it has its own robots.txt.
Sources
- RFC 9309. Robots Exclusion Protocol. IETF, 2022 — rfc-editor.org/rfc/rfc9309.
- Google Search Central. How Google interprets the robots.txt specification — developers.google.com/search/docs/crawling-indexing/robots/robots_txt.
- Google Search Central. Overview of Google crawlers and fetchers (user agents) — developers.google.com/search/docs/crawling-indexing/overview-google-crawlers.
- Google Search Central Blog. A note on unsupported rules in robots.txt, 2019.
- Bing Webmaster Guidelines and Bing Webmaster Tools Help: robots.txt and crawl control — bing.com/webmasters.
- OpenAI. Overview of OpenAI Crawlers — platform.openai.com/docs/bots.
- Anthropic. Does Anthropic crawl data from the web, and how can site owners block the crawler? — support.claude.com.
- Perplexity. Perplexity Crawlers — docs.perplexity.ai.
- Apple. About Applebot — support.apple.com/119829.
- Common Crawl. CCBot — commoncrawl.org/ccbot.
FAQ
How do I create a robots.txt file?
Pick a template in the generator (open site, private sections, SaaS or e-commerce), adjust the paths, add your sitemap and optionally block AI crawlers. Copy the result and save it as robots.txt at your site root so it opens at https://your-domain/robots.txt.
How do I check whether a page is blocked by robots.txt?
Open the “Test URL” tab, paste the file contents and the page URL, and choose a crawler. The tool shows whether the URL is allowed, which group applied and which rule decided — plus the result for Googlebot, Bingbot and AI crawlers.
How do I block AI crawlers?
Add groups with their tokens and Disallow: /, for example User-agent: GPTBot and User-agent: ClaudeBot. Separate bots by purpose: training bots bring no traffic, while search bots like OAI-SearchBot and PerplexityBot send visitors through links in answers. Remember that robots.txt is followed voluntarily.
Does blocking Google-Extended remove my site from Google?
No. Google-Extended is a control token, not a crawler: it forbids using pages for training and grounding Gemini models, but does not affect Googlebot crawling or Google Search.
Do I need Crawl-delay or Host?
Usually not. Google ignores Crawl-delay, and Host is an obsolete Yandex directive. Bing honors Crawl-delay, but its Webmaster Tools crawl control is the better option.
Will Disallow remove a page from search results?
Not reliably. Disallow blocks crawling, but Google can index the URL without its content if other pages link to it. To remove a page from results, keep it crawlable and add a robots meta tag with noindex.
Which wins: Allow or Disallow?
The rule with the longer path pattern, regardless of line order. If the lengths are equal, Allow wins. That is how Google and RFC 9309 work.
Where should robots.txt be placed?
At the root of each host: https://example.com/robots.txt. A file in a subfolder is ignored, and each subdomain and protocol needs its own file.
More in SEO
- SERP snippet preview
How will the title and description look in Google, and will they be cut off?
- Open Graph generator
Which Open Graph tags does the page need, and how will the link preview look?
- Readability checker
How hard is this text to read, and which sentences are too long?
- hreflang generator
Which hreflang tags connect the language versions of a page?
- UTM builder
How do I build consistent UTM links for every campaign and channel?
- UTM QR code generator
Which booth, flyer or poster brought visitors — and how do I tag a QR code for each one?