To find all the URLs of a website for 301 redirects, do not rely on a single source, because none of them is complete on its own. Combine three: a crawl that follows every internal link, the XML sitemap, and your analytics plus Search Console exports. The crawl finds linked pages, the sitemap adds pages nothing links to, and analytics and Search Console catch orphaned URLs that still get traffic or sit in the index. The pages you miss are the ones that 404 after launch, so completeness is the whole job. WPBuildAI combines all three into one URL inventory, so the redirect map covers every page.

Extract from the variants too, not just the canonical host, because Google’s June 2026 guidance now expects the move to be declared for all subdomains and both the www and non-www forms of the old domain, and a crawl seeded from one hostname will happily miss URLs that only ever existed on another.

Why no single source is enough

Each source has a blind spot, which is why one alone always leaves gaps. A link crawl only sees what something links to, so it misses orphans, pages no longer linked from anywhere on the current site. The XML sitemap lists declared pages but is often stale or incomplete, missing recently changed URLs or never including certain types. Analytics shows what gets traffic but only pages that had visitors in the window you export. Search Console shows what Google has indexed and what returns errors, including old URLs that vanished from the site but stay in the index. The 2024 Web Almanac documents how many URLs a real site accumulates across archives, tags, and generated pages, far more than anyone lists by hand. Only the union of the three sources approaches complete, because each catches what the others miss.

The four sources and what each catches

It helps to know precisely what each source contributes, so you know why all are needed. The link crawl, starting from your homepage and following internal links, finds everything reachable through your navigation and content, the bulk of the site. The XML sitemap adds pages the crawl could not reach because nothing links to them, plus it is the list you think you have. Analytics (export your top landing pages over a long window, a year or more) reveals URLs that still receive visits, including ones the crawl and sitemap missed. Search Console (the Pages/coverage report and the Performance report) reveals what Google has indexed and what it is requesting, including old URLs returning 404s and orphaned pages still getting impressions. A backlink export adds externally-linked URLs. Merge and de-duplicate these and you have a list far more complete than any single tool produces.

The orphans are the risk

The dangerous URLs are the ones nobody links to anymore but that still hold value, because they are invisible to the obvious source, the crawl, yet costly to lose. Ahrefs found in its search traffic study that a small share of pages carries most organic traffic, and some of those are old, orphaned URLs still ranking or holding backlinks from years ago, an old guide that earned links, a discontinued product page people still search for. A link crawl cannot see them because the current site does not link to them. Without the sitemap and the Search Console and analytics exports, they simply vanish in a migration, 404ing silently and taking their rankings and backlinks with them. Pulling those exports is the specific step that catches the orphans, and the orphans are precisely where the unmanaged risk concentrates, which is why the multi-source approach is not optional thoroughness but the core of the method.

A worked example: the site that “had 200 pages”

Picture a site the owner believes has about 200 pages. The link crawl confirms roughly 200 reachable pages. But the work is not done. The XML sitemap lists 240, adding 40 the crawl could not reach. The Search Console coverage report shows 90 more URLs Google has indexed, old blog posts, paginated archives, and tag pages no longer linked. The analytics export, over two years, surfaces another 30 URLs still receiving the occasional visit, including a couple of old landing pages with backlinks. De-duplicated, the real inventory is over 300 URLs, not 200. Had the owner built the redirect map from the crawl alone, more than a third of the site, including some backlinked, still-ranking pages, would have 404’d after the migration. The multi-source discovery turned a confident-but-wrong 200 into a complete 300+, which is the difference between a clean migration and a quiet traffic loss.

From list to redirects

Once the inventory is complete, the rest follows mechanically. Add a new-URL column, map each old URL to its closest relevant new page, and that becomes your 301 rules, per Google’s site move guidance and redirects guidance. Mark genuinely retired pages, ones with no equivalent that you meant to remove, for a clean 404 or 410 rather than a forced redirect. Prioritise by the traffic and links each URL had, so the high-value pages get precise destinations. The complete list is the hard, error-prone part; the mapping is comparatively straightforward once you have it. This is the discovery step behind mapping a site’s structure, the URL mapping template, and bulk redirects for a store, and it feeds straight into setting up the redirects.

How far back to look

A practical question that affects completeness: how long a window to pull from analytics and how much of the index to include. The answer is to err generous. Export analytics over at least a year, ideally more, because a page that got a handful of visits last quarter may have ranked steadily for years and still hold value, and a short window misses seasonal or evergreen URLs. Include everything Search Console has indexed, even URLs you do not recognise, because Google requesting a URL means it considered it real. The cost of including a URL that turns out not to matter is small (you map it or 410 it), while the cost of excluding one that did matter is a 404 on a ranking page. So when in doubt, include it: a slightly over-broad inventory is far safer than a tidy but incomplete one, since the whole point of the exercise is to miss nothing.

When the old site is already gone

Sometimes you need the URL list but the old site is down, replaced, or you have lost access, which removes the crawl as a source. The multi-source approach still works, leaning on the sources that persist independently of the live site. Search Console retains the URLs it indexed, so its reports are the backbone of a reconstruction. Analytics keeps its history of which URLs received traffic. The Wayback Machine holds snapshots of the old site you can mine for URLs. A backlink export shows externally-linked pages. And an old sitemap, if archived, lists declared pages. Together these can rebuild a substantial URL inventory even with no live site to crawl. So the method degrades gracefully: the live crawl is the easiest source when available, but the others are independent enough to reconstruct the list when it is not, which matters precisely when a migration went wrong and you are recovering after the fact.

Common mistakes finding URLs

The recurring errors all stem from trusting one source. Building the list from a crawl alone misses every orphaned URL, the ones most likely to hold stranded value. Trusting the sitemap as complete overlooks that it is often stale. Exporting analytics over a short window drops URLs that ranked for years but were quiet recently. Ignoring Search Console’s index leaves out URLs Google still considers real. And assuming you “know” your page count, rather than discovering it, guarantees surprises, since real sites almost always have more URLs than their owners think. Each is avoided by the same discipline: combine the crawl, the sitemap, a long analytics window, the full Search Console index, and a backlink export, then de-duplicate, which is what produces an inventory complete enough that the redirect map misses nothing.

Key points to remember

Find all of a website’s URLs for 301 redirects by combining four sources, a link crawl, the XML sitemap, a long analytics export, and the Search Console index (plus a backlink export), because each has a blind spot and only their union is complete. The orphans, old URLs nothing links to but that still rank or hold backlinks, are invisible to a crawl and are exactly the pages that 404 after a migration, so the sitemap and the analytics and Search Console exports that surface them are the most important sources, not optional extras. Err on the side of a long window and including everything indexed, and reconstruct from Search Console, analytics, and the Wayback Machine when the old site is gone. The complete list is the hard part; the redirect map follows from it. WPBuildAI merges all sources into one inventory and builds the redirect map, so no page is missed; send your site URL for a fixed quote.

Not affiliated with Google, WordPress, or Lovable.