Robots.txt: what it actually blocks
A robots.txt file tells search engine crawlers which URLs on your site they may crawl. It does not keep a page out of Google: a disallowed URL can still show up in results, without a description. And it binds no one: compliant crawlers follow it, others ignore it.
The file is short and rarely reread, so a badly written line can sit there for years with no visible effect. It is one of the things our SEO audit checks on every site.
What the file can and cannot do
Since 2022 the protocol has been an IETF standard, RFC 9309. The file sits at the root of the site and groups rules by crawler: each group names one or more crawlers, then the paths they may or may not crawl.
Google sums it up in its introduction to robots.txt: the file is mainly there to keep a site from being overloaded with requests, and it is not a mechanism for keeping a page out of Google. Reputable crawlers follow its rules, but each one may interpret them differently.
The standard adds that the protocol is no substitute for security measures. The file is public: anyone can read it, including the paths it disallows.
Blocking crawling is not the same as deindexing
A disallowed URL is no longer read, but it can still be known. If other pages link to it, Google may list it in results without a description, since it never read the page.
To keep a page out of the index, Google asks for a noindex rule or a password. And its noindex documentation states that the rule only works if the page is not blocked by robots.txt. A crawler that is not allowed to read the page never sees the request to drop it.
Combining the two is the mistake that shows the least: the page is disallowed, it carries the right rule, and its URL stays in the results.
The lines a compliant crawler ignores
A crawler that cannot parse a line does not report it. It skips it. A badly written disallow rule protects nothing, while the file looks as if it does.
- A line with no colon, usually a typo. No compliant crawler reads it.
- A rule placed before the first line that names a crawler. It belongs to no group, so it applies to no one.
- Fields Google does not support. According to its reading of the standard, Google reads only four fields: User-agent, Allow, Disallow and Sitemap. A Crawl-delay or Noindex line does nothing for Google, even if some other crawlers honor one of them.
- Letter case in paths. The standard matches paths character by character: /Admin/ and /admin/ are two different paths.
- Anything past 500 KiB. Google ignores everything after that size.
None of these mistakes breaks anything on screen. They surface when a crawler visits what you thought was closed, or when a page you wanted indexed stays invisible.
A named crawler ignores the rules written for everyone
When a crawler finds a group with its own name, it follows that group and nothing else. Google puts it this way: only one group is valid for a given crawler, the one whose name matches it most specifically. The rules in the "*" group, the one that applies to all crawlers, are not added on top. They are replaced.
Our own file is a case in point. It allows the whole site, and it names nine crawlers from AI assistants and AI search engines, including GPTBot, ClaudeBot and PerplexityBot, to allow them explicitly. Our audit measures how readable sites are to these assistants, so blocking them on our own site would make no sense.
The file makes one exception: the folder of sample reports published as PDFs. These reports are meant to be read from the page that introduces them, not served on their own as search results. Written once in the "*" group, the rule would have stopped only anonymous crawlers and let the nine named ones through. So it is repeated in each of the ten groups.
And it has the limit described above: if another site links to one of these PDFs, a search engine can list the URL without having read it.
Allowing these crawlers in robots.txt and offering them a summary of the site are two separate things. Our guide to the llms.txt file explains what the second one does.
What we find on the sites we audit
15%have no robots.txt file, or no sitemap that actually lists pages
0%have a robots.txt with a line that compliant crawlers ignore
These percentages cover about 30 websites we audited between August 14, 2026 and September 24, 2026. Each one counts only the sites where the point could be checked. Many were audited because a defect showed up quickly, so these figures describe our audits, not websites in general.
The second figure describes each file as it stood on the day of the audit. Files are edited by hand, and the same check can give a different result after the next edit.
What the audit checks, and why
The audit reads the file at the start of its crawl of the site, in a browser, the same way it reads the pages.
- The server response. A missing file (404 error) is valid: the standard reads it as everything allowed. An access denial or an anti-bot challenge page is not a missing file, and the audit does not count it as one. A site that returns its home page instead of the file has no robots.txt, even if the response says success.
- The syntax, line by line, against RFC 9309. Ignored lines and rules outside any group count as a defect. Fields outside the standard are flagged but not counted, since some crawlers read them.
- Broad blocks: the whole site closed to all crawlers, to Googlebot, or to Bingbot, the crawler behind Bing.
- The choice made for AI assistant crawlers, recorded as it is. Blocking them and allowing them are both legitimate choices: the report only flags a block that contradicts a goal the site states.
- Whether there is a sitemap that actually lists pages. A missing robots.txt or sitemap counts as a defect in this pillar.
Staging subdomains get their own check: a copy of the site left public, with no disallow rule and no noindex, is reported.
What is quick to fix, and what takes a project
A malformed line, a rule outside any group, or a disallow rule missing from a named group are fixed in the file itself, without touching the pages.
It becomes a project when robots.txt is being used to hide what should be protected or removed: a customer area, a staging site, duplicate pages. Each set of pages then needs its own decision between noindex, a password and a redirect. A redesign is when these decisions get lost: our guide to SEO site migrations covers what the audit checks after a launch.
What this check doesn't tell you
We read the file as your server serves it on the day of the audit, and we check its syntax against the standard. We do not know how each crawler interprets it: each one decides for itself whether to follow it, and how to read it.
We cannot see which URLs Google has crawled or keeps on record. That information is in the site's Search Console.
Read next
Redirect chains: why one hop is enough
A 301 redirect does its job in one hop. Why redirect chains cost you, where they come from, and what the audit checks on your site.
Read the guide →SEO migration: before and after launch
A redesign can change every URL on your site. What to settle before launch, and what the audit remeasures once the new site is live.
Read the guide →Structured data: which types still work
Organization, WebSite, BreadcrumbList, Article, Product: the structured data that still does something, and why it has to say what the page says.
Read the guide →