Skip to content
Rankora
Free account · 3 live checks/day

Sitemap vs Crawl

Your sitemap says one thing, your site shows another. Compare the two to find the gaps.

Privacy and cost

This tool runs a live check on our servers. Free accounts get 3 live checks per day in total across all such tools. We keep your report in your account so you can delete it any time.

How it works

  1. Enter your site address; the sitemap is found via robots.txt or /sitemap.xml.
  2. The crawler explores the site by following links and reads the sitemap URLs.
  3. Three lists come out: sitemap URLs not reached, crawled pages outside the sitemap and sitemap URLs that do not return 200.

Example

The sitemap lists 120 URLs and the crawl reaches 85: 35 sitemap URLs are linked from no page and 4 crawled pages are missing from the sitemap.

What the result means

A sitemap URL the crawl did not reach is often an orphan page. A crawled page missing from the sitemap may be forgotten or deliberately excluded. A sitemap URL that does not return 200 should be removed from it. The comparison depends on the crawled page cap.

Common problems this tool helps with

  • A sitemap listing deleted or redirected pages.
  • Important pages missing from the sitemap.
  • Pages that exist only in the sitemap and are never linked.
  • An outdated sitemap that is no longer regenerated.

Frequently asked questions

Why are some sitemap URLs not reached?

Either no page links to them (orphans) or they fall beyond the crawled page cap of your plan.

Must the sitemap contain every page?

It should list the important indexable pages, without pages in error, redirected or excluded from the index.

Where is the sitemap looked up?

In the Sitemap lines of robots.txt, then at /sitemap.xml, following one level of sitemap index.