robots.txt generator
Built in your browser · nothing is uploaded
Generates one robots.txt group: a user agent, its disallow and allow rules, an optional crawl delay and any sitemap lines. The distinction the file exists on is that it controls crawling and not indexing, and Google says so in the first paragraph of its own documentation.
How to use the robots.txt generator
Crawling and indexing are two different permissions, and confusing them causes more damage than anything else in technical SEO. Google’s introduction to the file puts it plainly: robots.txt "is not a mechanism for keeping a web page out of Google". Disallowing a URL stops the fetch. It does not stop the URL appearing in results, because a link from anywhere else on the web is enough to put it there, listed without a description because Google was never allowed to read one.
The failure that follows from this is specific and common. Google’s own guidance states that "for the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file", and that if it is blocked "the crawler will never see the noindex rule, and the page can still appear in search results". So blocking a page you want removed prevents its removal, and the fix is the opposite of the instinct: allow crawling, serve noindex, and wait for a recrawl, which Google warns can take months on a page it visits rarely.
One group per crawler, and only one applies. Google’s guidance is that a crawler obeys "the group with the most specific user agent that matches" it and that "the order of the groups within the robots.txt file is irrelevant", with multiple groups for the same agent merged before processing. Position decides nothing; specificity decides everything. This tool writes a single group, so the moment you paste a User-agent: Googlebot block beside an existing User-agent: * block, Googlebot obeys the named one wherever it sits in the file and ignores every rule in the wildcard group. Any disallow you still want to apply to Googlebot has to be repeated inside its own group.
Within a group, order is cosmetic. This generator prints all the disallows and then all the allows, and it makes no difference: crawlers "use the most specific rule based on the length of the rule path", and where two rules of equal length conflict, Google takes "the least restrictive rule". RFC 9309, which standardised the protocol in September 2022, says the same in stricter language: "the most specific match is the match that has the most octets". A four-character Allow: /a/b therefore loses to a longer disallow, whatever line each sits on.
Four fields are supported and one is not. Google names user-agent, allow, disallow and sitemap, adding that "other fields such as crawl-delay aren’t supported". The crawl delay box is here because Bing and several other crawlers do read it; put a value in and Googlebot will pass over it without complaint. Google’s crawl rate is managed in Search Console, or by serving a 503 when you genuinely need it to back off.
The file is scoped more narrowly than most people assume: the rules "apply only to the host, protocol, and port number where the robots.txt file is hosted". A file at https://example.com/robots.txt says nothing about http://example.com, nothing about shop.example.com, and nothing about a service on another port. Every host you serve needs its own copy, and a subdomain with no robots.txt is fully crawlable no matter what the parent domain says.
Two smaller mechanics from the same spec. Google parses the first 500 kibibytes and stops, so a generated file with tens of thousands of disallow lines can be silently truncated mid-rule. And an empty Disallow: with no path is the explicit "nothing is blocked": this tool emits exactly that when you leave both boxes empty, which is a valid and deliberate allow-everything file rather than an incomplete one.
Finally, treat the file as published copy. It is world-readable by design, RFC 9309 states outright that these rules "are not a form of access authorization", and a tidy list of the directories you would rather nobody visited is a map for anyone curious. Anything genuinely private needs authentication. And avoid disallowing the CSS and JavaScript a page needs to render, because a crawler that cannot fetch them judges the page on a broken version of itself.
What people use it for
- Keeping a staging path out of crawl
- Pointing crawlers at a sitemap from one file
- Writing rules for one crawler without affecting the rest
- Producing a deliberate allow-everything file for a new site
- Adding a crawl delay for the crawlers that honour one
- Checking whether a path you blocked is the reason a page will not drop out of the index
Questions
No. It stops crawling. Google states the file "is not a mechanism for keeping a web page out of Google", and a blocked URL linked from elsewhere can still appear, listed without a description.
Allow crawling and serve a noindex rule. Google is explicit that a page blocked in robots.txt never has its noindex seen, so blocking it prevents the removal you wanted.
As long as a recrawl takes, which Google warns "may take months" for a page it visits rarely. The URL Inspection tool can request a faster one, and the Removals tool hides a URL temporarily while you wait.
Not usefully in that order. Allow the crawl first so the noindex is read, and only add a disallow later, once the page has actually dropped out.
At the root of each host, so it answers at /robots.txt. Google says the rules "apply only to the host, protocol, and port number where the robots.txt file is hosted".
No. A shop or blog subdomain needs its own file. Without one it is fully crawlable regardless of what the main domain disallows.
Technically yes, because the protocol is part of the scope. In practice most sites redirect http to https, which resolves it.
The longer path. Google uses "the most specific rule based on the length of the rule path", and RFC 9309 phrases the same rule as the match with the most octets. Equal-length conflicts go to the least restrictive rule.
Within a group, no. Specificity decides, not position, so this tool grouping the disallows before the allows changes nothing.
Each crawler obeys the most specific group that names it and ignores the rest. Google states that the order of the groups in the file is irrelevant, so moving a block up does nothing: a Googlebot group means Googlebot stops reading the wildcard group entirely, and anything you still want enforced has to be repeated inside it.
Google combines multiple groups for the same user agent into one before processing, so duplicates are tolerated. Writing them once is still easier to audit.
No. Google lists user-agent, allow, disallow and sitemap as supported and says other fields such as crawl-delay are not. Bing and some other crawlers do read it, and the field is here for them.
Through the crawl rate settings in Search Console, or by returning 503 or 429 for a short period. Google treats a sustained run of those as a signal to reduce its rate.
Nothing is blocked. It is the explicit allow-everything form, and this generator emits it when you leave both path boxes empty.
Google supports * for any run of characters and $ to anchor the end of a path, so Disallow: /*.pdf$ blocks PDFs. Not every crawler implements them, so keep critical rules literal.
Yes. /Admin/ and /admin/ are different rules, and matching is by prefix, so /admin also covers /administrator unless you anchor it.
Google parses the first 500 kibibytes and ignores the rest. A generated file of tens of thousands of rules can be cut off partway through, so prefer a few wildcard rules over an enumerated list.
It does not belong to a group at all. Sitemap directives apply to the whole file regardless of which user-agent block they sit near, and this tool prints them after a blank line for readability.
The opposite. RFC 9309 says outright that these rules "are not a form of access authorization", the file is public, and listing your private paths tells a curious visitor exactly where to look. Use authentication.
No. A crawler that cannot fetch the assets a page needs renders a broken version of it and evaluates that. Blocking /assets/ is a frequent and expensive mistake.
Everything is treated as crawlable, which is usually what you want for a small site. A 500 is different: a persistent server error can make Google stop crawling the site as a precaution.
Only the ones that choose to read it. It is a convention rather than an enforcement mechanism, and a crawler that ignores it faces no obstacle. Rate limiting and authentication are the enforcement layer.