AI Search Strategy

HTML Sitemap vs XML Sitemap: What Google Just Said

Google says an HTML sitemap won't replace your XML sitemap, but it still helps people and crawlers. Here's how to build one that also works for AI search.

Core takeawayKeep your XML sitemap as the complete, accurate list for search engines, and add a small HTML sitemap that helps people reach your main categories, built with plain links that every crawler, including AI crawlers, can read.

Overview

An HTML sitemap is a page of links that helps visitors find the main sections of a website. Google's John Mueller says it doesn't replace an XML sitemap, which is a structured file made for search engines. But crawlers can still follow its links. Use both: XML for machines, HTML for people.

Mueller made the point on Google's Search Off the Record podcast with Martin Splitt, as Search Engine Journal reported. It's a small comment. But it settles a question many site owners still ask, and it matters more now that AI crawlers from OpenAI, Anthropic, Perplexity and others are reading the web alongside Google and Bing.

**A note on our position:** Vanaxity sells AI search services, so we have a stake in how sites get crawled. We stick to what Google and Bing have said, and we label our own advice as ours.

Key Takeaways

  • An HTML sitemap is a page of links for people. An XML sitemap is a file for search engines.
  • Google can't process an HTML sitemap the way it processes an XML sitemap, Mueller said, because it lacks a strict structure.
  • Crawlers can still follow the links on an HTML sitemap, so it helps discovery a little.
  • Mueller suggested listing categories, not every product, so people can find the right section fast.
  • Most AI crawlers don't run JavaScript, so plain HTML links help them reach your pages.
Book a free fit check

Map your SEO, GEO and AEO workflow before you build.

Van avatar
Chat with Van

What Did Google Say About HTML Sitemaps?

Splitt asked whether site owners are stuck with XML sitemaps or can use HTML ones instead. Mueller's answer had three parts.

  • **It's for people.** An HTML sitemap is "basically almost like a map of your website for users," he said.
  • **It's not a replacement.** "It's not something that replaces an XML sitemap file." You can't submit it as a sitemap, because "it doesn't have this strict structure to it."
  • **It still helps crawling.** If a crawler finds an HTML sitemap, it can follow those links like any other links on the site.

Then he added the practical part. On an online store, "you wouldn't list all of your products in an HTML sitemap file." You'd list the categories, so people can go to the right place and pick a product from there.

HTML Sitemap vs XML Sitemap: What's the Difference?

They share a name, but they do different jobs. Here's how they compare:

HTML sitemapXML sitemap
Made forPeople browsing your siteSearch engines
FormatA normal web page of linksA structured XML file
Can you submit it to Google?NoYes, in Search Console or robots.txt
What it listsMain sections and categoriesEvery page you want indexed
Extra dataNoneOptional last-modified dates
Size limitNone set, but keep it useful50,000 URLs or 50MB per file

The limits and formats come from Google's sitemap documentation, which lists XML, RSS or Atom, and plain text files as supported formats. HTML pages aren't on that list.

Why Did HTML Sitemaps Fall Out of Fashion?

Because something built for machines came along.

Years ago, site owners linked new pages from older ones and kept a big HTML sitemap in the footer, so search engines could find everything. Many SEOs packed it with every page they wanted to rank.

Then Google, Yahoo and Microsoft agreed on a shared XML sitemap standard. It gave search engines a clean list of URLs, with optional dates for when each page last changed. Once that existed, giant HTML sitemaps lost their main job, and many sites dropped them or let them rot.

Mueller's comment suggests the old idea isn't dead. It just has a different job now: helping people, with crawling as a side benefit.

Where Does It Fit in Internal Linking?

Think of it as one piece of your internal linking, not the whole plan.

Your menu, breadcrumbs, related-article links and links inside your content do most of the work of guiding people and crawlers through a site, and they also show which pages matter most. A sitemap page adds one more route. It's a safety net.

So if a key page is only reachable from the sitemap page, that's a warning sign. Fix it.

When Does an HTML Sitemap Still Help?

Not every site needs one, and a badly kept one can do more harm than good, but it clearly earns its place in a few common cases:

  • **Large sites.** Stores, publishers and directories with deep sections where visitors get lost.
  • **Menus built with JavaScript.** If your main menu only appears after scripts run, an HTML sitemap gives every crawler a plain path in.
  • **Deep or orphaned pages.** Important pages that are many clicks from the home page, or linked from nowhere at all.
  • **Accessibility.** A simple list of links is easy to use with a screen reader or keyboard.
  • **Content hubs.** Guides and resources that need one clear place to start.

A small site with five pages and a clear menu doesn't need one. The menu already does the job, and a second list of the same five links would only add clutter for visitors who can already see everything from the top of the page.

How Do You Build a Useful HTML Sitemap?

Follow Mueller's lead and build it for people first. Six steps:

  • **List sections, not everything.** Link your main categories, key guides and important pages, and let people drill down from there.
  • **Use real links.** Each item should be a normal link with an href, which is the kind Google says it can crawl reliably.
  • **Render it on the server.** Make sure the links are in the page's HTML, not added later by scripts.
  • **Group by topic.** Use clear headings, such as Products, Guides and Company, so people can scan the page in a few seconds and jump straight to the part of the site they came for.
  • **Write clear link text.** "Running shoes" beats "Click here."
  • **Link to it from the footer.** That way people and crawlers can find it from any page.

Keep it short enough to scan. As a rule of thumb, if the page needs a search box, it's probably too long. That's our guideline, not Google's.

What Should Your XML Sitemap Get Right?

The HTML sitemap is the extra. The XML sitemap is still the core. Google's documentation sets out the basics:

  • **Use full URLs,** such as https://example.com/page, not short paths.
  • **Save it as UTF-8.**
  • **Stay under the limits:** 50,000 URLs or 50MB per file. Split bigger sites into several files with an index.
  • **Keep lastmod honest.** Google uses it only if it's consistently accurate, and a changed copyright date doesn't count.
  • **Skip priority and changefreq.** Google ignores both.
  • **Tell search engines where it is,** in Search Console or with a Sitemap line in robots.txt.

Google also notes that a sitemap is only a hint. It doesn't guarantee that every listed page will be crawled or indexed.

Is llms.txt the New HTML Sitemap?

In a way, yes. llms.txt is a short, hand-picked list of a site's key pages, written for AI tools. That's close to what Mueller describes for HTML sitemaps.

But Google doesn't use it for search. Google's guide to AI features lists llms.txt among things you don't need, and Mueller has compared it to the old keywords meta tag, Search Engine Journal reported. Meanwhile, Google's Lighthouse tool added an experimental llms.txt check for AI agents, so the signals are mixed.

**Our view:** a good HTML sitemap does most of what llms.txt promises, and it helps human visitors too. If you add llms.txt anyway, treat it as a small extra for AI agents and coding tools, not as a ranking tactic that will change how often Google or ChatGPT cite you.

What HTML Sitemap Mistakes Should You Avoid?

Most problems come from treating the page as a dumping ground rather than a guide. The common ones:

  • **Dumping every URL.** Thousands of links help nobody, and Mueller suggested listing categories instead.
  • **Letting it go stale.** Broken links and removed sections make the page useless.
  • **Hiding it.** If nothing links to it, people and crawlers won't find it.
  • **Using it instead of an XML sitemap.** Google can't process it that way.
  • **Building it with scripts.** Links that only appear after JavaScript runs may be invisible to AI crawlers.

Fix these, and the page becomes a quiet, useful part of your site rather than a forgotten list.

One more tip: review it every quarter. It takes ten minutes.

What Does a Good HTML Sitemap Look Like?

Here's a simple layout for a mid-sized B2B or ecommerce site:

SectionWhat to link
Products or servicesEach main category, not each product
Guides and resourcesKey guides, the blog home and topic hubs
CompanyAbout, team, careers and contact
HelpSupport home, FAQ, shipping and returns
LegalPrivacy policy and terms

That's often fewer than 50 links. Small and clear works better than big and complete, because people can actually use it.

What Should You Do This Week?

  • **Check your XML sitemap** in Search Console for errors and stale lastmod dates.
  • **Map your internal linking.**
  • **View your site with JavaScript off,** and see whether your main menu still shows real links.
  • **Decide if you need one.**
  • **If you do, build a small one** around your main sections and link it from the footer.
  • **Check AI answers** for your key topics to see which of your pages get cited.

None of this takes long, and it makes your site easier to explore for people, search engines and AI crawlers alike.

How Vanaxity Helps Sites Get Found in AI Search

We help brands make sure search engines and AI assistants can reach, read and cite their best pages. That means checking crawl paths, fixing sitemaps, testing how pages look to crawlers that don't run JavaScript, and tracking where AI answers cite you.

A good HTML sitemap is a small fix with a clear job. To see how your site looks to AI crawlers, see our services, read our guide to technical SEO debt or browse more insights.

Frequently asked questions

What is the difference between an HTML sitemap and an XML sitemap?

An HTML sitemap is a normal web page of links that helps people find the main sections of a site. An XML sitemap is a structured file that lists the pages you want search engines to index, with optional last-modified dates. Google says an HTML sitemap doesn't replace an XML sitemap.

Does Google still use HTML sitemaps?

Google can follow the links on an HTML sitemap like any other links, which helps crawling. But it can't submit or process one as a sitemap file, because it lacks the strict structure of an XML sitemap, Google's John Mueller said on the Search Off the Record podcast.

Should an HTML sitemap list every page?

No. Mueller suggested listing categories rather than every product, so people can go to the right section and choose from there. Your XML sitemap is the place for the complete list of pages.

Do HTML sitemaps help AI search?

Indirectly. Vercel found that major AI crawlers, apart from Google's, don't run JavaScript, so plain HTML links help them reach your pages. There's no evidence an HTML sitemap boosts AI citations by itself, but it makes sure crawlers can find your key pages.

Is llms.txt a replacement for a sitemap?

No. llms.txt is a short list of key pages for AI tools, but Google says it doesn't need it for search or AI features. Keep an accurate XML sitemap, and consider a small HTML sitemap, which also helps people.

Tran Tien VanFounder, Van Data Team - builds Vanaxity, the AI content agent for SEO, GEO and AEO, and leads data engineering delivery for B2B teams.Connect on LinkedIn