Noindex vs Disallow: Keeping Pages Out of Google Correctly

In this article
- Noindex vs Disallow: What Each One Actually Controls
- Why Can a Page Blocked by robots.txt Still Show Up in Google?
- The Trap: Noindex Behind a Disallow Rule Is Never Seen
- How to Add a Robots Meta Tag Noindex or an X-Robots-Tag Header
- Which Should You Use for Thin, Test and Duplicate Pages?
- How to Check That Google Has Seen Your Noindex
- Where robots.txt Disallow Is Still the Right Tool
- FAQ
Noindex vs disallow? If you want a page out of Google, use noindex. A robots.txt disallow rule only stops Googlebot from crawling a URL, and the page can still get indexed. Say you run a handful of small sites and want the thin, test or duplicate pages gone from results. A noindex tag on a crawlable page does that. A robots.txt rule doesn’t.
Noindex vs Disallow: What Each One Actually Controls
Disallow controls crawling. Noindex controls indexing. People mix these up all the time. According to Google’s introduction to robots.txt, the file tells crawlers which URLs they’re allowed to access. It exists mostly to keep bots from hammering your server, and Google says outright that it’s not a way to keep a page out of its results. Noindex works differently. Once Googlebot crawls the page and reads the rule, Google drops the page from Search, even if other sites link to it.
- Disallow stops fetching, manages crawl traffic and keeps image, video and audio files out of results.
- Noindex removes the page from results, but only works if the page can be crawled.
Why Can a Page Blocked by robots.txt Still Show Up in Google?
Because Google knows the URL exists from links pointing at it, even though it can’t fetch the content. So the page shows up with no description. That’s what “indexed though blocked by robots.txt” in Search Console means: Google knows the address but has never read the page. Google’s advice splits into two cases. If you want to fix a result like that, remove the robots.txt entry that blocks the page. If you want the page hidden from Search completely, use something else, like noindex.
The Trap: Noindex Behind a Disallow Rule Is Never Seen
This one catches a lot of people. If robots.txt blocks a URL, Googlebot never fetches it, so it never reads the noindex meta tag or header on that page. Google has to crawl a page to see its meta tags and HTTP headers. Makes sense, right? So when a disallow rule and a noindex sit on the same URL, the disallow wins and the noindex does nothing at all. And putting noindex into robots.txt won’t save you either, because Google doesn’t support that rule there. The order that works is simple. Allow crawling first. Then add noindex.
How to Add a Robots Meta Tag Noindex or an X-Robots-Tag Header
You’ve got two options: a <meta name="robots" content="noindex"> tag in the HTML head, or an X-Robots-Tag HTTP response header. They do the same thing. The meta tag is fine for ordinary HTML pages. The header is what you want for PDFs and other non-HTML files, and it can cover a whole path in one go (handy when a test directory has dozens of URLs). On your own NGINX origin, a location block for a test directory looks like this:
location /test/ {
add_header X-Robots-Tag "noindex" always;
}
location ~* \.pdf$ {
add_header X-Robots-Tag "noindex" always;
}Reload NGINX after you edit the config, then check that the header actually comes back in the server response. Don’t just assume it does.
Which Should You Use for Thin, Test and Duplicate Pages?
Pages that have to disappear from results get noindex. Robots.txt disallow is only for URLs you just don’t want crawled. Here’s how I’d handle the usual mix on a small site:
- Thin pages: noindex. Google drops them after the next crawl.
- Test pages: noindex, or password protection, which Google itself lists as the other option.
- Duplicate or similar low-value URLs that were never indexed: disallow can save crawling, which matters if you care about crawl limits on small sites.
Staging and origin hosts are a different problem, and they have their own guide on origin hosts left in search. One more thing: don’t block the CSS, JavaScript or images a page needs to render. Without them Google can’t analyse the page properly.
How to Check That Google Has Seen Your Noindex
Open the URL Inspection tool, look at the HTML Googlebot received, and check the noindex is actually in it. Per Google’s guide to blocking indexing, if the page still shows up, the usual reason is that Google hasn’t recrawled it since you added the rule. Nothing’s broken, it just needs a recrawl. You can request one from the same tool. For a site-wide view, the Page Indexing report in Search Console lists every page where Googlebot found a noindex rule.
Where robots.txt Disallow Is Still the Right Tool
Disallow still has a place. Use it to manage crawl traffic, keep image, video and audio files out of Google, and skip unimportant or near-identical URLs. It won’t stop other pages or users from linking to blocked media files, though. Crawler opt-outs are another use, but opting out of AI training is a separate topic with its own tokens. So, noindex vs disallow in one line: to remove a page, allow crawling and add noindex. To save crawling on URLs that don’t matter, disallow them.
FAQ
Can I put noindex in robots.txt?
No. Google doesn’t support a noindex rule in robots.txt. Put it in a robots meta tag in the page head or send it as an X-Robots-Tag header.
Should I use both disallow and noindex on the same page?
No. The disallow rule stops Googlebot from fetching the page, so it never sees the noindex. Leave the page crawlable until it has dropped out of results.
How do I remove a page from Google that is blocked by robots.txt?
Remove the robots.txt entry that blocks the page, then add a noindex meta tag or header. After that, request a recrawl in the URL Inspection tool so Google can read the new rule.


