Googlebot’s 2MB Limit: What Large HTML Pages Lose in Crawling

In this article
- What does the Googlebot 2MB limit actually say?
- 2MB vs the 15MB crawl limit: which figure applies to your pages?
- What do large HTML pages lose past the cutoff?
- How to measure the uncompressed HTML size with curl
- Where the bloat comes from: inline CSS, JavaScript and base64 images
- Put important content and links early in the HTML
- FAQ
Googlebot fetches only the first 2MB of a supported file type and the first 64MB of a PDF. Whatever sits past that cutoff is never sent for indexing consideration, as stated in Google’s Googlebot documentation. Who should care? Mostly WordPress and page-builder sites whose templates spit out very heavy HTML. That’s where the Googlebot 2MB limit bites. Most pages are far smaller, so don’t panic: measure your own documents first and fix only the ones that come close.
What does the Googlebot 2MB limit actually say?
Googlebot stops downloading a file once it hits the cutoff and passes on only the part it already has. On February 3, 2026 Google moved the default file size limits of its crawlers into the crawler documentation and clarified Googlebot’s own values. The threshold is applied to uncompressed data, so the transferred size you see with gzip or Brotli is not the figure that counts. And no, nothing in the documentation describes the limit as a ranking signal. It only defines how much of a file reaches indexing.
2MB vs the 15MB crawl limit: which figure applies to your pages?
For Google Search it’s the smaller Googlebot file size limit, not the general default. The 15MB crawl limit is the baseline for Google’s crawlers and fetchers as a whole, and content beyond it is ignored, according to the overview of Google’s crawlers. That same page explains that individual projects may set different thresholds per crawler and per file type, with Googlebot as the example.
What do large HTML pages lose past the cutoff?
Everything after the cutoff. On a long document that can mean late body text, footer navigation, internal links and structured markup placed at the end of the source. And since Googlebot finds new URLs mainly through links on crawled pages, a cut-off footer can leave you finding orphan pages that looked well connected in the browser.
CSS and JavaScript referenced in the HTML are fetched separately, and each of those fetches is bound by the same cap. An oversized bundle can get truncated too, which affects how the page renders for Google. In my view, large HTML pages SEO boils down to completeness: is the whole document, and every file it depends on, actually reaching the indexing pipeline?
How to measure the uncompressed HTML size with curl
Download the raw document without compression and read its byte count. Here is how to check HTML page size in a way that matches what the limit counts:
- Request the page with curl and don’t ask for compression (leave out the -compressed flag).
- Save the output and print the size:
curl -s -o page.html -w '%{size_download}' https://example.com/page/ - Repeat the check for your largest CSS and JavaScript files, because each one has its own allowance.
Look at the HTML document itself, not the total page weight reported by browser tools. Start with the heaviest templates: long landing pages assembled in a page builder, category archives, pages with embedded data.
Where the bloat comes from: inline CSS, JavaScript and base64 images
Rarely from text. Most oversized documents are inflated by code and data embedded directly in the markup. The usual suspects behind inline CSS and JavaScript bloat in WordPress and page builders: per-widget style blocks, inlined scripts, serialized JSON state, inline SVG and base64-encoded images. The fixes happen in your CMS, theme or origin server:
- Move inline CSS and JS into external files.
- Replace base64 images with regular image files.
- Remove unused builder modules and duplicated sections.
- Split very long pages into several shorter ones.
Put important content and links early in the HTML
Source order beats visual order. The main content and the internal links you care about should appear before heavy scripts and decorative blocks. Keep the head lean so the body starts early, and push large inline scripts and data blobs to the end of the document or into files. If a page gets truncated anyway, the part that survives is then the part that matters.
Check as well that crawlers can fetch the CSS, JavaScript and images the page relies on, for example by writing hotlink rules that spare crawlers on your origin. An NGINX reverse proxy such as Jalvo, with EU and USA IPs in front of existing hosting, does not change this: HTML size is still decided by the origin and the CMS. For heavy pages that have no reason to appear in search at all, first get clear on when noindex beats disallow.
Measure the uncompressed HTML of your heaviest templates, move inline code and base64 images into files, then place the main content and links early in the source. Handled that way, the Googlebot 2MB limit becomes a routine template check. Not a hidden reason why part of a page never reaches the index.
FAQ
Does the Googlebot file size limit apply to compressed or uncompressed size?
Uncompressed, as Google’s documentation states. The transfer size shown for a gzip or Brotli response is smaller than the number that counts.
Do CSS and JavaScript files count toward the HTML page’s limit?
No. Each resource referenced in the HTML is fetched separately and is bound by the same limit on its own. Moving inline code into external files lightens the document, but a very large bundle can still be cut by itself.
Is the PDF crawl limit the same as for HTML?
No, the PDF crawl limit is higher than the one for other supported file types when Googlebot crawls for Search. Resources referenced from HTML follow the regular threshold, not the PDF one.


