Orphan Pages: How to Find Them and Link Them Back In

In this article
Orphan pages are published URLs on your site that no other page links to. Simple as that. To find them, you compare the full list of URLs you publish with the URLs a crawler reaches when it follows links from the home page. And if you publish often, you’ve almost certainly got some. Older posts slowly slide off category pages, archives and related-post blocks until nothing points to them anymore. The method has five steps: build a URL list from your sitemap or CMS, run a crawl that follows links only, compare the two lists, check the gaps against your server logs, and relink what’s worth keeping.
What are orphan pages and why do they matter?
An orphan page is a published URL with no internal link pointing to it. Readers can’t reach it by clicking around. Neither can crawlers. Search engines discover new URLs by following links on pages they’ve already crawled, and the Google guide to sitemaps describes exactly this process. It also admits that on large sites it gets harder to make sure every page has a link from at least one other page.
Content sites don’t create orphans through some dramatic mistake. It’s the boring stuff. Paginated archives push old posts further and further back. Categories get deleted or renamed, someone removes a widget, the menu changes. Edit a slug and you might break the only link a post ever had. Now, orphans aren’t the same thing as deep pages. Deep pages are linked, they just sit many clicks from the home page. Both can end up as deep pages not indexed, but the fix differs: deep pages need shorter paths, orphans need a first link, full stop. (And no, fixing either one doesn’t guarantee better indexing or rankings.)
What counts as a crawlable link?
Only one kind of link takes a page out of orphan status: a standard <a> element with an href attribute that Google can follow. Google’s page on making links crawlable for search says anchor text belongs between <a> elements Google can crawl, and that every page you care about should get a link from at least one other page on your site.
So what about links that live only in a script or a click handler, with no proper <a href>? They may not count. My advice: open the rendered HTML of the linking page and confirm the element is actually there before you tick an orphan off as fixed. Takes a minute. Saves you a false “done”.
How to find orphan pages: sitemap vs crawl comparison
Export every URL you consider published. Crawl the site from the home page, following links only. Anything in the first list but missing from the second is an orphan candidate. That sitemap vs crawl comparison is the heart of any internal linking audit:
- Export URLs from your XML sitemap or straight from the CMS, covering published posts and pages.
- Run a generic site crawler from the home page with sitemap discovery turned off.
- Normalise both lists: same protocol, consistent trailing slash, no tracking parameters.
- Diff the lists to get the URLs the crawler never reached.
- Remove intentional exclusions, like thank-you pages or noindexed URLs.
One gotcha people trip over all the time. The crawler has to ignore the sitemap. Let it read the sitemap and it’ll happily find your orphans there, and the problem quietly vanishes from your results. For the diff itself, a spreadsheet lookup or a short script does the job. Nothing fancy.
Check your server logs before you fix anything
Why bother with logs? Because they show which orphan candidates Googlebot still requests and which ones it never asks for, and that tells you what to relink first. A post that still gets crawler visits is in a very different spot from one that’s dropped out of view entirely. Our guide to finding URLs Googlebot never requests walks through how to filter logs for this.
While you’re in there, look for URLs that show up in neither list: old slugs, parameter variants, pages that never made it into the sitemap in the first place. Oh, and verify the traffic really is Googlebot. Don’t trust the user agent string alone - anyone can fake it.
How to link orphan pages back into your site
Every page worth keeping should get at least one contextual link from a related page that’s already well linked. The rest? You decide case by case. In practice, each candidate lands in one of these buckets:
- Relink: the page is still useful and just needs a way in.
- Merge and redirect: it overlaps a stronger post, so combine the content and add a redirect on your own origin or NGINX.
- Update and relink: it’s outdated but still valuable.
- Exclude or remove: it has no value for readers.
Where do the new links go? Related existing posts, topic hubs or category pages, series navigation, and older posts you’re already updating anyway (that last one is the cheapest win). Google’s guidance asks for anchor text that’s descriptive, concise and relevant, and it points out that the words around a link matter too. So don’t cram several links right next to each other. Our notes on internal link anchors for navigation put these rules into practice. Just keep expectations honest: a new link makes a page reachable, but it doesn’t guarantee indexing or rankings.
Keep orphan pages from coming back
Here’s the thing. Orphans come back whenever you publish faster than you link. So the link check has to be part of every new post and every structural change, not a once-a-year cleanup. A short routine covers most of it:
- Link each new post from at least one existing post on the day it goes live.
- Recheck links after changing menus, categories or slugs.
- Rerun the sitemap vs crawl comparison on a regular schedule.
A sitemap helps discovery. It doesn’t replace internal links. Google’s sitemap overview describes a comprehensively linked site as one where Googlebot can find all the important pages by following links from the home page. Clear paths also help with focusing crawling on key pages rather than stray URL variants.
So it really is one loop you keep repeating: list what you publish, crawl what’s linked, compare the two, check the gaps in your logs, and relink the pages that still serve readers. Boring? A bit. Works, though.
FAQ
Is a page in the sitemap still an orphan page?
Yes, if no internal link points to it. The sitemap helps search engines discover the URL, but readers and crawlers that follow links still can’t get there. It stays an orphan until another page on your site links to it.
Should I delete orphan pages or relink them?
Relink the ones that are still useful. Merge or update content that overlaps with other posts, and redirect the merged URLs. Only remove pages that have no value for readers.
Will linking an orphan page get it indexed?
Linking makes the page reachable by following links, which is what Google asks for. Indexing still isn’t guaranteed, though, and there’s no set timeframe for when it might happen.


