Clew

HomeToolsSEOrobots.txt generator

robots.txt generator and tester with AI crawlers

Build a robots.txt from a template, block private sections, decide which AI crawlers may access your site, and test any URL: is it allowed for Googlebot, Bingbot or GPTBot, and which group and rule decided it.

● Free, no sign-upUpdated:

Template
A template replaces the groups; edit by hand afterwards.

A crawler follows one group — its own or, if there is none, *. So a Googlebot or Bingbot group must repeat the shared rules.

AI crawlers: block access

Model training

Search and linked answers

User-initiated requests

robots.txt is a request, not a lock. Reputable companies follow it, but some bots ignore it. Only server-side controls — authentication or blocking by IP and user agent — reliably restrict access.

Full sitemap URLs, one per line.

—

 

    No Crawl-delay or Host. Google does not support Crawl-delay, and Host was a Yandex-only directive retired in 2018. The generator does not add them.

    Everything is calculated in your browser — nothing you enter is sent anywhere.

    How to use it

    1. Pick a template“SaaS” blocks the app area and API; “E-commerce” blocks the cart, filters and sorting. Add your own paths and separate groups for Googlebot or Bingbot.
    2. Decide on AI crawlersBots are grouped by purpose: model training, search with citations, user requests. You can block training and keep search.
    3. Test URLs“Test this file” moves the result to the “Test URL” tab, which shows the group and rule that decide a URL’s fate for each crawler.
    4. Publish the fileCopy the text and put it at the domain root: https://example.com/robots.txt. Every subdomain and protocol has its own file.

    robots.txt syntax: groups, rules and precedence

    robots.txt is a text file at the site root that tells search crawlers which URLs they may crawl. Since 2022 the parsing rules are standardized in RFC 9309; Google and Bing document the same core principles plus a few extensions of their own.

    The file is made of groups. A group starts with one or more User-agent lines followed by Disallow and Allow rules. Sitemap lines do not belong to groups and can appear anywhere.

    User-agent: * # group for all crawlers Disallow: /app/ # block everything starting with /app/ Allow: /app/signup # except the sign-up page User-agent: GPTBot # two User-agent lines — User-agent: CCBot # one shared group Disallow: / Sitemap: https://example.com/sitemap.xml

    How a crawler picks a group

    A crawler looks for the group with its token — case-insensitively — and follows only that group. If there is none, it uses the * group; if that is missing too, everything is allowed. Groups with the same token are merged. Hence a common mistake: creating a separate User-agent: Googlebot group and forgetting to repeat the * rules in it.

    Search engines document fallback groups: Googlebot-Image and Googlebot-News use the Googlebot rules when they have none of their own.

    How a rule is chosen

    A rule value is compared with the start of the path including the query string. * matches any sequence of characters, $ at the end marks the end of the URL. If several rules match, the longest wins; on a tie, Allow wins. Line order does not matter.

    Allow: /p + Disallow: / → /page allowed (2 chars beat 1) Allow: /folder + Disallow: /folder → /folder/page allowed (equal length, Allow wins) Allow: /page + Disallow: /*.htm → /page.htm blocked (6 chars beat 5) Allow: /$ + Disallow: / → / allowed; /page.htm blocked

    This is the precedence table from Google’s documentation; the tool checks it in automated tests. An empty Disallow: blocks nothing. Paths are compared in normalized form: non-ASCII characters are UTF-8 percent-encoded, so /café/ and /caf%C3%A9/ are the same. Paths are case-sensitive: /Admin/ and /admin/ are different rules.

    robots.txt for Google and Bing: what differs

    The core is shared: groups, Allow and Disallow, the * and $ wildcards, longest-match precedence. The differences are in extensions.

    • Crawl-delay. Google ignores it. Bing honors it, but Bing Webmaster Tools’ crawl control is the better place to manage crawl rate.
    • Fallback groups. Googlebot-Image, -Video and -News fall back to the Googlebot group when they have no group of their own.
    • Size. Google processes the first 500 KiB of the file and ignores the rest.
    • Host and Clean-param. These are Yandex directives. Google and Bing ignore them; use 301 redirects for the main host and rel="canonical" for URL parameters.
    • Noindex. Google stopped honoring Noindex: in robots.txt in 2019. Use the robots meta tag or the X-Robots-Tag header.

    To see how a specific search engine reads your live file, use the robots.txt report in Google Search Console or the robots.txt tester in Bing Webmaster Tools. This tool is handy before publishing: it works with any text, not only the file on your site.

    How to block AI crawlers, and whether you should

    AI companies crawl the web for different purposes, and the big players use a separate token for each. Decide by purpose rather than “block them all”.

    AI crawlers by purpose
    TokenOwnerPurpose
    GPTBotOpenAImodel training
    OAI-SearchBotOpenAIChatGPT search
    ChatGPT-UserOpenAIuser request
    ClaudeBotAnthropicmodel training
    Claude-SearchBotAnthropicsearch answers
    Claude-UserAnthropicuser request
    PerplexityBotPerplexitysearch and citations
    Perplexity-UserPerplexityuser request
    Google-ExtendedGoogletoken: Gemini training and grounding
    Applebot-ExtendedAppletoken: Apple model training
    CCBotCommon Crawlopen web archive
    meta-externalagentMetatraining and products
    AmazonbotAmazonAmazon services and AI
    BytespiderByteDancedata collection

    Three groups, three decisions

    • Model training (GPTBot, ClaudeBot, CCBot and others). Blocking keeps new pages out of future training sets. These bots bring no traffic, so they are the most commonly blocked — especially for paid and original content.
    • Search and linked answers (OAI-SearchBot, Claude-SearchBot, PerplexityBot). These bots send visitors: your site is cited with a link. Blocking them means leaving a new channel. SaaS and media sites usually keep them open.
    • User requests (ChatGPT-User, Claude-User, Perplexity-User). The page is opened because a person explicitly asked. This is closer to a browser than a crawler; some companies state such requests may not follow robots.txt.

    Google-Extended and Applebot-Extended are not crawlers. Pages are still crawled by the regular Googlebot or Applebot; the token only forbids using them to train models. Blocking Google-Extended does not affect Google Search rankings. Blocking Googlebot to stay out of Google’s AI answers would take you out of search too.

    Above all, robots.txt is an agreement, not protection. Reputable companies honor it, but you cannot verify that from outside, and some bots ignore the rules. If content must not be served, you need authentication or server or CDN restrictions.

    Common robots.txt mistakes

    • Blocking CSS and JavaScript. Rules like Disallow: /*.js$ or Disallow: /assets/ stop search engines from rendering the page, which may then look empty or not mobile-friendly.
    • Hiding pages with robots.txt instead of noindex. Blocking crawling does not remove a URL from the index: Google may show it without a description if other pages link to it. To remove a page from search, allow crawling and add <meta name="robots" content="noindex"> or an X-Robots-Tag header.
    • Treating robots.txt as security. The file is public. A line like Disallow: /secret-admin-panel/ advertises the path to anyone who reads it. Protect private areas with a login; a generic /admin/ in robots.txt is enough — or nothing.
    • A dedicated group without the shared rules. You add User-agent: Googlebot for one exception, and Googlebot stops seeing the * rules.
    • Rules before User-agent. Disallow lines at the top of the file, before the first group, are ignored.
    • A relative Sitemap. Use a full URL: Sitemap: https://example.com/sitemap.xml.
    • A leftover Disallow: / from staging. The staging site was blocked entirely, the file shipped to production, and the site drops out of search. Check robots.txt after every release.
    • One file for all subdomains. robots.txt applies only to its own host and protocol: app.example.com and example.com have separate files.

    A typical SaaS setup: the marketing site is open, the /app/ area is blocked from crawling and protected by login, sign-up and log-in pages are open. If the product lives on a separate subdomain, it has its own robots.txt.

    Sources

    1. RFC 9309. Robots Exclusion Protocol. IETF, 2022 — rfc-editor.org/rfc/rfc9309.
    2. Google Search Central. How Google interprets the robots.txt specification — developers.google.com/search/docs/crawling-indexing/robots/robots_txt.
    3. Google Search Central. Overview of Google crawlers and fetchers (user agents) — developers.google.com/search/docs/crawling-indexing/overview-google-crawlers.
    4. Google Search Central Blog. A note on unsupported rules in robots.txt, 2019.
    5. Bing Webmaster Guidelines and Bing Webmaster Tools Help: robots.txt and crawl control — bing.com/webmasters.
    6. OpenAI. Overview of OpenAI Crawlers — platform.openai.com/docs/bots.
    7. Anthropic. Does Anthropic crawl data from the web, and how can site owners block the crawler? — support.claude.com.
    8. Perplexity. Perplexity Crawlers — docs.perplexity.ai.
    9. Apple. About Applebot — support.apple.com/119829.
    10. Common Crawl. CCBot — commoncrawl.org/ccbot.

    FAQ

    How do I create a robots.txt file?

    Pick a template in the generator (open site, private sections, SaaS or e-commerce), adjust the paths, add your sitemap and optionally block AI crawlers. Copy the result and save it as robots.txt at your site root so it opens at https://your-domain/robots.txt.

    How do I check whether a page is blocked by robots.txt?

    Open the “Test URL” tab, paste the file contents and the page URL, and choose a crawler. The tool shows whether the URL is allowed, which group applied and which rule decided — plus the result for Googlebot, Bingbot and AI crawlers.

    How do I block AI crawlers?

    Add groups with their tokens and Disallow: /, for example User-agent: GPTBot and User-agent: ClaudeBot. Separate bots by purpose: training bots bring no traffic, while search bots like OAI-SearchBot and PerplexityBot send visitors through links in answers. Remember that robots.txt is followed voluntarily.

    Does blocking Google-Extended remove my site from Google?

    No. Google-Extended is a control token, not a crawler: it forbids using pages for training and grounding Gemini models, but does not affect Googlebot crawling or Google Search.

    Do I need Crawl-delay or Host?

    Usually not. Google ignores Crawl-delay, and Host is an obsolete Yandex directive. Bing honors Crawl-delay, but its Webmaster Tools crawl control is the better option.

    Will Disallow remove a page from search results?

    Not reliably. Disallow blocks crawling, but Google can index the URL without its content if other pages link to it. To remove a page from results, keep it crawlable and add a robots meta tag with noindex.

    Which wins: Allow or Disallow?

    The rule with the longer path pattern, regardless of line order. If the lengths are equal, Allow wins. That is how Google and RFC 9309 work.

    Where should robots.txt be placed?

    At the root of each host: https://example.com/robots.txt. A file in a subfolder is ignored, and each subdomain and protocol needs its own file.

    More in SEO

    Open the collection →

    Tours, tooltips and checklists whose impact shows up in the numbers.

    Try Clew for freeHow it worksFree plan forever · no credit card

    Product tours, popups and onboarding in the age of AI

    7 patterns with step counts, copy rules and what to measure, what AI changes, and a checklist before you publish. PDF, 2 pages. What is inside →