SEO & web Site files

robots.txt generator

Built in your browser · nothing is uploaded

Local · a public file, never a security measure

Generates one robots.txt group: a user agent, its disallow and allow rules, an optional crawl delay and any sitemap lines. The distinction the file exists on is that it controls crawling and not indexing, and Google says so in the first paragraph of its own documentation.

How to use the robots.txt generator

1 Set the user agent. Leave it at * for every crawler, or name one to write rules that apply to it alone.
2 List the paths to disallow, one per line, each starting with a slash. Allow lines carve exceptions out of them.
3 Add your sitemap URLs. Those apply to the whole file rather than to the group above them.
4 Copy the result to a file named robots.txt at the root of the site, so it answers at https://example.com/robots.txt.
5 Check it in Search Console’s robots.txt report, and confirm the pages you meant to keep are still fetchable.

Crawling and indexing are two different permissions, and confusing them causes more damage than anything else in technical SEO. Google’s introduction to the file puts it plainly: robots.txt "is not a mechanism for keeping a web page out of Google". Disallowing a URL stops the fetch. It does not stop the URL appearing in results, because a link from anywhere else on the web is enough to put it there, listed without a description because Google was never allowed to read one.

The failure that follows from this is specific and common. Google’s own guidance states that "for the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file", and that if it is blocked "the crawler will never see the noindex rule, and the page can still appear in search results". So blocking a page you want removed prevents its removal, and the fix is the opposite of the instinct: allow crawling, serve noindex, and wait for a recrawl, which Google warns can take months on a page it visits rarely.

One group per crawler, and only one applies. Google’s guidance is that a crawler obeys "the group with the most specific user agent that matches" it and that "the order of the groups within the robots.txt file is irrelevant", with multiple groups for the same agent merged before processing. Position decides nothing; specificity decides everything. This tool writes a single group, so the moment you paste a User-agent: Googlebot block beside an existing User-agent: * block, Googlebot obeys the named one wherever it sits in the file and ignores every rule in the wildcard group. Any disallow you still want to apply to Googlebot has to be repeated inside its own group.

Within a group, order is cosmetic. This generator prints all the disallows and then all the allows, and it makes no difference: crawlers "use the most specific rule based on the length of the rule path", and where two rules of equal length conflict, Google takes "the least restrictive rule". RFC 9309, which standardised the protocol in September 2022, says the same in stricter language: "the most specific match is the match that has the most octets". A four-character Allow: /a/b therefore loses to a longer disallow, whatever line each sits on.

Four fields are supported and one is not. Google names user-agent, allow, disallow and sitemap, adding that "other fields such as crawl-delay aren’t supported". The crawl delay box is here because Bing and several other crawlers do read it; put a value in and Googlebot will pass over it without complaint. Google’s crawl rate is managed in Search Console, or by serving a 503 when you genuinely need it to back off.

The file is scoped more narrowly than most people assume: the rules "apply only to the host, protocol, and port number where the robots.txt file is hosted". A file at https://example.com/robots.txt says nothing about http://example.com, nothing about shop.example.com, and nothing about a service on another port. Every host you serve needs its own copy, and a subdomain with no robots.txt is fully crawlable no matter what the parent domain says.

Two smaller mechanics from the same spec. Google parses the first 500 kibibytes and stops, so a generated file with tens of thousands of disallow lines can be silently truncated mid-rule. And an empty Disallow: with no path is the explicit "nothing is blocked": this tool emits exactly that when you leave both boxes empty, which is a valid and deliberate allow-everything file rather than an incomplete one.

Finally, treat the file as published copy. It is world-readable by design, RFC 9309 states outright that these rules "are not a form of access authorization", and a tidy list of the directories you would rather nobody visited is a map for anyone curious. Anything genuinely private needs authentication. And avoid disallowing the CSS and JavaScript a page needs to render, because a crawler that cannot fetch them judges the page on a broken version of itself.

What people use it for

  • Keeping a staging path out of crawl
  • Pointing crawlers at a sitemap from one file
  • Writing rules for one crawler without affecting the rest
  • Producing a deliberate allow-everything file for a new site
  • Adding a crawl delay for the crawlers that honour one
  • Checking whether a path you blocked is the reason a page will not drop out of the index

Questions

No. It stops crawling. Google states the file "is not a mechanism for keeping a web page out of Google", and a blocked URL linked from elsewhere can still appear, listed without a description.

Google Search Central, robots.txt introductionGoogle Search Central — robots.txt specificationGoogle Search Central — block indexing with noindexRFC 9309 — Robots Exclusion Protocol
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding