Google Crawler IP Ranges Moved: Updating Your NGINX Allowlists

In this article
- What changed with the Google crawler IP ranges?
- Which JSON files cover which crawlers?
- How do you update scripts that fetch Googlebot IP ranges JSON?
- Building the NGINX Googlebot allowlist from the files
- Refresh on a schedule and test with a dry run
- How to verify Googlebot IP when a request looks suspicious
- FAQ
The JSON files listing Google crawler IP ranges are moving. Old home: /search/apis/ipranges/. New home: developers.google.com/crawling/ipranges/. So every script that feeds an NGINX or firewall allowlist needs the new path. Google published the announcement about the new location on March 31, 2026, and says the old address keeps working for now and will be redirected within six months. Do you allowlist Googlebot by IP and pull those files automatically? Then this one is yours to fix.
What changed with the Google crawler IP ranges?
The URL. That’s it. The files leave the /search/apis/ipranges/ directory on developers.google.com and now live under developers.google.com/crawling/ipranges/. Google’s reasoning: the ranges apply to more than Search crawlers, so they deserve a more general home. Fair enough. Requests to the previous path still succeed today, and later they will be answered with a redirect. Nothing in the announcement touches the file contents or the CIDR notation, so your parser can stay exactly as it is.
Which JSON files cover which crawlers?
Google splits its crawlers into three groups, and each group gets its own file. How they treat robots.txt is spelled out on the page about verifying requests from Google:
- Common crawlers such as Googlebot, listed in
common-crawlers.json, always respect robots.txt rules for automatic crawls. - Special-case crawlers such as AdsBot do specific jobs for Google products and may or may not respect robots.txt.
- User-triggered fetchers such as Google Site Verifier ignore robots.txt, because a person asked for the fetch.
That last group is the messy one. It spans user-triggered-fetchers.json, user-triggered-fetchers-google.json and user-triggered-agents.json. Addresses in the -google file resolve to a google.com hostname, while those in user-triggered-fetchers.json resolve to gae.googleusercontent.com. My advice: pick the groups your site really needs and skip the rest, instead of importing everything just in case. And knowing where Googlebot crawls from matters as much as knowing its addresses.
How do you update scripts that fetch Googlebot IP ranges JSON?
Replace the old path wherever it shows up, and make the download step picky about what it accepts. I’d go in this order:
- Grep cron jobs, configuration management, CI pipelines and firewall tooling for
/search/apis/ipranges/. - Swap in the new
/crawling/ipranges/URL for each file you use. - Tell the HTTP client to follow redirects, so a reference you forgot about survives the switch.
- Parse the response as JSON before anything else reads it.
- Exit with an error when the file is empty, malformed or contains no prefixes.
Keep the last good copy next to the script. A failed download must never wipe the allowlist. Never. The addresses come in CIDR format, and your code has to handle both IPv4 and IPv6 prefixes (the second one is easy to forget).
Building the NGINX Googlebot allowlist from the files
Don’t type ranges by hand. Generate an include file from the JSON. The main configuration should hold no prefixes at all - it only points to a separate file that the script overwrites on each run. A geo block fits nicely here:
geo $google_crawler {
default 0;
include /etc/nginx/generated/google-crawlers.conf;
}
location /restricted/ {
if ($google_crawler = 0) {
return 403;
}
}Each line of the generated file is one prefix followed by 1;. Simple. One catch, though: if an NGINX reverse proxy with EU or USA addresses, such as Jalvo, sits in front of your hosting, apply the rule on the origin, where the real client IP is visible. And is this cloaking? No. The variable decides access only, and every visitor who gets through receives identical content. The same generated map comes in handy when you set up hotlink protection for Google Images and want genuine crawlers exempted.
Refresh on a schedule and test with a dry run
Put the fetch-and-rebuild job on a schedule, so the allowlist follows whatever Google publishes. But every run should start as a dry run:
- write the new include to a temporary file,
- diff it against the live one,
- run
nginx -t, - reload only when the test passes.
Log every added or removed prefix. Send an alert when the download fails or the diff is unexpectedly large. And once you have switched to the new path, scan the access logs for denied requests carrying a Googlebot user agent. Those give away a missing group or a stale file.
How to verify Googlebot IP when a request looks suspicious
Google documents two methods: a manual DNS check for single cases, and automatic matching against the published lists at scale. For a one-off, run a reverse DNS lookup on the address and confirm the hostname ends in googlebot.com, google.com or googleusercontent.com. Then do a forward lookup and compare the result with the original IP. What about the user agent string? It proves nothing, because anyone can send it. The address is what counts.
So the whole job with Google crawler IP ranges boils down to four things: point your scripts at the new path, let them follow redirects, rebuild the allowlist on a schedule and keep prefixes out of hand-written configuration.
FAQ
Will my allowlist break if I keep the old URL?
Not right away. The previous path still responds for now and will be redirected to the new one, so clients that follow redirects keep getting the files. Update the reference anyway. A tool that ignores redirects would end up with an empty or invalid response.
Do I need all the JSON files or only common-crawlers.json?
Depends on which Google products reach your site. common-crawlers.json covers Googlebot and the other common crawlers, special-case crawlers such as AdsBot have their own list, and user-triggered fetchers are split across separate files. Import only the groups whose traffic you expect and want to let in.
Do user-triggered fetchers obey robots.txt?
No. Google’s verification page states that these fetchers ignore robots.txt rules, because the fetch was requested by a user and not started by an automatic crawl. So access control for them has to happen at the IP level.


