Skip to content
SEO

Keeping Your Origin Server Out of Google’s Index

Keeping Your Origin Server Out of Google's Index
In this article
  1. Why Does Google Index Your Origin IP or Staging Host?
  2. How to Check Whether Your Origin Server Is Indexed
  3. Block Direct IP Access and Restrict the Origin to Your Proxy
  4. Setting the Canonical Host Behind a Proxy
  5. Removing Origin URLs Google Has Already Indexed
  6. Summary: The Fix in Order
  7. FAQ

Found your origin server indexed in Google? The fix comes in two parts. Lock the origin down so the only way in is through your proxy, and make every canonical signal point at the public domain, so Google folds everything into that one host. Here’s how it usually goes. The site sits behind a reverse proxy, and yet the raw origin IP (or some staging hostname everyone forgot about) turns up in results as a full copy of the real site. Below: why it happens, how to confirm it, how to block direct access, which canonical signals actually count, and how to clean up URLs that are already in the index.

Why Does Google Index Your Origin IP or Staging Host?

Google will index any address that serves your pages as a normal, successful response. So if the origin answers requests for its bare IP or an old hostname, that’s a second copy of your site. Simple as that. And the causes are almost always boring:

  • a web server default vhost that serves the main site no matter what Host header comes in,
  • staging subdomains left wide open, with no password and no allowlist,
  • links, sitemaps or feeds that leaked the raw IP,
  • old DNS records still pointing straight at the origin.

What you end up with is origin IP duplicate content. It competes with your canonical domain and eats crawl time. Google’s guidance on consolidating duplicate URLs says it’s better for Googlebot to spend its time on new or updated pages than on duplicate versions. Fair enough. But there’s a quieter cost too, and I’d argue it’s the worse one: an indexed origin gives away the exact address the proxy was supposed to hide. It also turns into one of the signals that tie sites together.

How to Check Whether Your Origin Server Is Indexed

Run site: searches in Google for the origin IP and your staging hostnames, then look at what the origin sends back when someone hits it directly. Go in this order:

  1. Do a site: search for the bare IP and for every staging or old hostname you can think of.
  2. Curl the origin IP twice: once with no Host header, once with the production hostname.
  3. Open Search Console and look for hosts you don’t recognise, plus duplicate pages where Google picked a canonical other than yours.
  4. Go through sitemaps, feeds and internal links for raw IPs or staging addresses.

The responses are what matter. Say a direct request to the IP comes back as a successful page with the full content. Then the origin is serving your site to the public, and Google can crawl it just as easily as you did. No magic involved.

Block Direct IP Access and Restrict the Origin to Your Proxy

The most reliable fix is an origin that refuses every request not coming through the proxy. Then there’s nothing public left for Google to crawl. On a server you run yourself, that usually means:

  • Firewall rules that accept HTTP and HTTPS traffic only from your proxy’s IP addresses.
  • An NGINX default_server block that closes or rejects requests for unknown hosts and bare IPs.
  • Separate server_name blocks for your real domains only, so no hostname falls through to the main site.
  • A password or IP allowlist on every staging environment. Every one. The one you forget will be the one that gets indexed.

Refusing direct connections to the origin is plain access control. Serving crawlers something different from what people see on the public domain is cloaking, and that’s never an acceptable fix (no matter how tempting it looks at 2 a.m.). Putting a proxy layer before hosting, with EU or US IP addresses, keeps the public address separate from the origin, and the origin’s own configuration stays in your hands. A proxy won’t repair a leaky setup on its own, though. It helps to know when a proxy hides problems instead of solving them.

Setting the Canonical Host Behind a Proxy

Behind a proxy, every signal the origin sends out has to name the public domain as the canonical host. Why? Because the origin often has no clue which hostname the visitor actually typed. Put rel=canonical in the HTML source, pointing at the public HTTPS URL, and make sure no JavaScript rewrites it later. Google’s canonicalization documentation says a Link HTTP header with rel=canonical works too, and that includes non-HTML files like PDFs.

Set your application to build absolute URLs from the forwarded host and protocol headers your proxy sends. That way canonicals, sitemaps and internal links never contain the IP. Internal links point only at canonical URLs. Permanent redirects on a duplicate host? Only when you’re retiring that host. And stay away from bad TLS certificates and HTTPS-to-HTTP redirects, because both push Google hard toward HTTP versions. Certificates can give away your setup as well, so go through the footprints SSL and CDN add while you’re at it.

Removing Origin URLs Google Has Already Indexed

For Google to drop indexed origin URLs, it has to be able to crawl them and see either a noindex signal or a redirect to the canonical domain. Robots.txt is the wrong tool here. A lot of people reach for it first, but Google’s documentation on blocking indexing explains that disallowed URLs can still get indexed without their content. You’ve got two options that actually work:

  • Option A: temporarily send an X-Robots-Tag noindex header on every response from the origin or staging host, then shut off access once the URLs have dropped out.
  • Option B: if the origin hostname has to stay reachable, permanently redirect it to the public domain.

If it’s urgent, Google’s removals tool hides URLs fast. But only for a while. It’s a band-aid, not a replacement for fixing the server config.

Summary: The Fix in Order

Confirm the leak. Restrict the origin to the proxy. Align your canonical signals. Then clean up whatever’s already indexed. Bots and people get the same content on the public domain; the only thing you refuse is direct access that skips the proxy. Getting an origin server indexed out of search results is a server configuration job, not a robots.txt one.

FAQ

Can I just disallow the origin in robots.txt?

No. URLs blocked in robots.txt can still be indexed without their content. And Google can’t see a noindex on pages it isn’t allowed to crawl, so a disallow can actually keep the duplicate in the index longer.

Is blocking direct IP access considered cloaking?

No. Refusing connections that bypass your proxy is access control, and it applies to every visitor the same way. Cloaking means showing crawlers different content from what people see on the same URL.

Should staging servers use noindex or authentication?

Authentication or an IP allowlist is the stronger choice, since it stops crawling completely. While the existing staging URLs are still being removed, add noindex on top.

ShareLinkedInX
Web Systems team

The team that builds and runs Jalvo.

Give each site its own IP.

Add your first domain in a few minutes.

Start now

Choose which cookies we may use. Necessary cookies keep the site and your login working and cannot be switched off.