XML sitemap generator: Which of your pages belong in your XML sitemap?
This XML sitemap generator crawls up to 100 pages of your site like a search engine, keeps only the live, indexable, canonical ones and hands you a valid sitemap.xml to upload.
Switches on with donations
The crawler is built and tested. A crawl costs more server time than a single check, so it switches on for everyone, free, as soon as donations cover it. Until then, the example below is a real crawl of our sample shop.
- Free, no account
- Up to 100 pages in a few minutes
- Any public http or https site
example-shop.hr
23 URLs crawled · 18 in the sitemap · 5 left out
sitemap.xml: generated code
Upload it to https://example-shop.hr/sitemap.xml and add "Sitemap: https://example-shop.hr/sitemap.xml" to robots.txt.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example-shop.hr/</loc>
<lastmod>2026-09-20</lastmod>
</url>
<url>
<loc>https://example-shop.hr/about-us/</loc>
<lastmod>2026-05-01</lastmod>
</url>
<url>
<loc>https://example-shop.hr/blog/</loc>
<lastmod>2026-09-15</lastmod>
</url>
<url>
<loc>https://example-shop.hr/care/</loc>
<lastmod>2026-07-01</lastmod>
</url>
<url>
<loc>https://example-shop.hr/contact/</loc>
<lastmod>2026-05-01</lastmod>
</url>
<url>
<loc>https://example-shop.hr/privacy-policy/</loc>
<lastmod>2026-01-15</lastmod>
</url>
<url>
<loc>https://example-shop.hr/returns/</loc>
<lastmod>2026-08-10</lastmod>
</url>
<url>
<loc>https://example-shop.hr/shipping/</loc>
<lastmod>2026-08-10</lastmod>
</url>
<url>
<loc>https://example-shop.hr/shop/</loc>
<lastmod>2026-09-20</lastmod>
</url>
<url>
<loc>https://example-shop.hr/size-guide/</loc>
<lastmod>2026-08-10</lastmod>
</url>
<url>
<loc>https://example-shop.hr/terms/</loc>
<lastmod>2026-01-15</lastmod>
</url>
<url>
<loc>https://example-shop.hr/blog/break-in-new-sneakers/</loc>
<lastmod>2026-08-22</lastmod>
</url>
<url>
<loc>https://example-shop.hr/blog/how-we-make-a-pair/</loc>
<lastmod>2026-09-15</lastmod>
</url>
<url>
<loc>https://example-shop.hr/blog/resoling-service/</loc>
<lastmod>2026-07-30</lastmod>
</url>
<url>
<loc>https://example-shop.hr/shop/gift-card/</loc>
<lastmod>2026-06-01</lastmod>
</url>
<url>
<loc>https://example-shop.hr/shop/kupa-runner/</loc>
<lastmod>2026-09-02</lastmod>
</url>
<url>
<loc>https://example-shop.hr/shop/sava-high/</loc>
<lastmod>2026-09-12</lastmod>
</url>
<url>
<loc>https://example-shop.hr/shop/sava-low/</loc>
<lastmod>2026-09-18</lastmod>
</url>
</urlset>
Your robots.txt already lists a sitemap. Validate what is there now before you replace it.
Left out, and why
| URL | Detail |
|---|---|
| https://example-shop.hr/cart/ |
| URL | Detail |
|---|---|
| https://example-shop.hr/shop/?sort=price | → https://example-shop.hr/shop/ |
| URL | Detail |
|---|---|
| https://example-shop.hr/shop/drava-classic/ | HTTP 404 |
| URL | Detail |
|---|---|
| https://example-shop.hr/tag/sale/ |
| URL | Detail |
|---|---|
| https://example-shop.hr/old-journal/ | → https://example-shop.hr/blog/ |
What this tool checks
Every page we find is followed like a search engine would, and only the ones Google should index go into the file.
Found by crawling
We start at the page you enter and follow every same-site link, the way a search engine discovers pages, up to 100 pages and 10 clicks deep.
Only pages worth indexing
Pages that redirect, answer 4xx or 5xx, say noindex or name another URL as canonical are left out, each with the reason.
Respects robots.txt
Paths your robots.txt disallows are not fetched and not listed, and its Crawl-delay is honoured up to 5 seconds.
lastmod from your server
When your server sends a Last-Modified header, the date goes into the file. No made-up priority or changefreq values; Google ignores both.
Copy, download or CSV
Copy the XML, download sitemap.xml, or download a CSV of every URL we found with its status and verdict.
Gentle on servers
Two requests at a time, 500 ms apart, through the same SSRF guard as our reports. A 429 or 503 answer stops the crawl at once.
Show 2 more checks
Copy, download or CSV
Copy the XML, download sitemap.xml, or download a CSV of every URL we found with its status and verdict.
Gentle on servers
Two requests at a time, 500 ms apart, through the same SSRF guard as our reports. A 429 or 503 answer stops the crawl at once.
How it works
Enter your home page
Enter the address of your home page. We read robots.txt first, then follow links from that page.
We crawl your site
Up to 100 pages are requested with GET, two at a time, 500 ms apart. Most sites finish in one to three minutes.
Upload the file
Upload sitemap.xml to your site root, add a Sitemap line to robots.txt and submit it in Google Search Console.
One crawl covers up to 100 pages, and pages only reachable through JavaScript, forms or search are not found, because we read the HTML your server sends.
Questions
What is a sitemap generator?
A tool that finds the pages of your site and writes them into a sitemap.xml file for search engines. This one crawls your site the way a search engine does, starting at your home page and following links, then keeps only the pages Google should index: live, not noindex, and canonical to themselves. Everything it leaves out is listed with the reason.
Is it free, and why is there sometimes no Generate button?
It is always free, with no account. Crawling costs more server time than a single check, so the crawler is a funded unlock: it runs for everyone while donations cover that time and pauses for everyone when they do not. While it is paused you see a real example crawl of our sample shop instead, and the funding page shows what switches on next.
How do I create an XML sitemap?
Check first whether your platform already makes one: WordPress publishes /wp-sitemap.xml, SEO plugins such as Yoast and Rank Math replace it with their own, and Shopify, Wix and Squarespace build /sitemap.xml automatically. If yours does not, generate the file here, upload it to your site root, add a Sitemap line to robots.txt and submit it in Google Search Console.
Do I need a new sitemap when I add pages?
Yes, because a sitemap only lists the pages that existed when it was made. A CMS such as WordPress, Shopify or Wix keeps its own sitemap up to date automatically, so use that one if you have it. For a static site, run the generator again after adding pages and upload the new file; Google reads it on its next visit.
Which pages are left out of the sitemap, and why?
Pages that redirect (the target is listed instead), pages that answer 404, 410 or 5xx, pages with a noindex robots tag or header, pages whose canonical names another URL, paths robots.txt disallows, and files that are not web pages, such as PDFs and images. Each is listed with its reason in the results and the CSV, so you can fix what should be indexed.
Why only 100 pages?
100 pages keeps each crawl to a few minutes and polite to your server: two requests at a time, 500 ms apart. Crawls of up to 500 pages are planned for owners who verify their site, so nobody can make us crawl a big site they do not run. For larger sites, your CMS or SEO plugin is the better source of the sitemap.
Do I still need a sitemap if Google crawls my site anyway?
It helps most for new sites, large sites and pages with few internal links. A sitemap tells Google which URLs you want indexed and when they changed, so new and deep pages are found sooner. Small, well-linked sites get crawled fine without one. A clean file with only live, canonical pages is worth more than a long one full of redirects.
Where do I put the sitemap file?
Upload it to the root of your site, so it answers at https://yoursite.com/sitemap.xml. Then add the line "Sitemap: https://yoursite.com/sitemap.xml" to robots.txt and submit the address in Google Search Console and Bing Webmaster Tools. Check it any time with our XML sitemap validator, which also samples the listed URLs to confirm they answer 200.
Related free tools
All 43 tools →Free, funded by the people who use it
Donations keep getReport free and switch on site crawl up to 500 pages + weekly re-check next, for everyone.