Skip to content
Article

How to Add llms.txt and Markdown Pages Without Maintaining a Second Website

The llms.txt convention asks for a curated Markdown index plus clean .md versions of your pages. Here's how to generate that layer from the pages you already publish, so it never drifts, never leaks private routes, and never competes with your HTML in search.
TLDR
  • Treat llms.txt Markdown pages as a rendering of the page you already publish, not a second copy of the site. Generate each .md version on request from the same server render.
  • Keep one disallow list that matches robots.txt, so private paths like cart, checkout, and account return 404 instead of leaking through the Markdown layer.
  • Send noindex and rel=canonical on Markdown responses, and rel=alternate and rel=describedby on HTML, so search engines keep ranking the HTML and AI systems know what to cite.
  • Keep the index short and built for change: stable collection URLs over churning SKUs, policies under Optional, and format validation in code.
  • llms.txt is still a proposed standard, so keep the build small and check your server logs for AI crawler traffic before investing more.

AI assistants, answer engines, and shopping agents are starting to read websites differently than people do. They don't want your mega-menu, cookie banner, or carousel markup. They want the facts on the page, in a format that's cheap to parse. The proposed llms.txt convention offers one answer: a small, curated Markdown index at /llms.txt, plus clean .md versions of important pages.

The usual objection is a fair one: "We are not going to maintain a second copy of our website." You shouldn't. What we've learned, first on our own site and then on two sister storefronts for a multi-location retail client, is that the Markdown layer works only when it is a rendering of the page you already publish, not a separate set of content. If you get that one decision right, most of the risks (stale copies, duplicate URLs, policy drift, maintenance cost) go away.

Know what llms.txt is (and what it isn't)

Three files tend to get lumped together, but they do different jobs:

  • robots.txt grants or denies crawler access. It's a policy file.

  • sitemap.xml is an exhaustive list of every URL you want indexed. It's an inventory.

  • llms.txt is neither. It's a short, hand-picked reading list for a reader that will only look at a handful of pages: what the organization is, where the primary pages are, which listings are stable, and where the policies live.

Keep in mind that llms.txt is still a proposal, not a widely adopted standard. When the client's merchandising team asked whether it would pay off, the honest answer was "maybe." That's why the implementation had to be cheap to build and close to free to maintain. If the work only pays off once the content team is updating a parallel site, it isn't worth doing yet.

Make every Markdown page a view of an existing URL

The core pattern is simple: every public page gets a Markdown twin at the same path with .md appended (/ maps to /index.md). No one writes those files. When a .md path is requested, the server:

  1. Maps the Markdown path back to its HTML path.

  2. Renders the normal HTML page internally, through the same request handler real visitors use.

  3. Pulls out the main content region, strips non-content blocks (scripts, styles, iframes, SVG, templates, comments), and converts what's left to Markdown with a library like Turndown, with a simple regex fallback if conversion throws.

  4. Returns the result as text/markdown.

Because the Markdown is generated from the live render on each request (and cached like the page itself), it can't drift from the HTML. Change a product description, publish a blog post, or fix a typo, and the .md version updates with it.

// Simplified: route .md requests through the normal page renderer
const htmlPath = htmlPathFromMarkdownPath(url.pathname); // "/about.md" -> "/about"
if (!htmlPath) return notFound();

const html = await renderPage(new Request(origin + htmlPath, {
  headers: {...request.headers, 'x-markdown-companion': '1'},
  redirect: 'manual',
}));

const markdown = htmlToMarkdown(mainContent(html));
return new Response(markdown, {headers: markdownHeaders(htmlPath)});

Handle the rendering edge cases on purpose

On a modern streaming or server-rendered frontend, "render the page" hides a few gotchas we had to deal with directly:

  • Streaming responses. Many React frameworks stream HTML and finish rendering after the first bytes go out. The server already waited for the full render when it detected a known bot. We extended that rule to the internal companion request (flagged with a request header) so the Markdown never comes from half-rendered HTML.

  • Redirects. Old URLs that redirect in HTML should also resolve as Markdown. The handler follows a few redirect hops manually (including platform-managed URL redirects) and gives up after a fixed limit instead of looping.

  • Titles. The <title> tag is written for search results and often ends with a brand suffix. The page's visible main heading usually describes the content better, so the converter uses it as the document title and removes it from the body so it doesn't appear twice. It falls back to <title> only when there's no heading.

  • Collapsed and client-side content. Accordions that are closed by default often don't render their answers into the HTML at all, and widgets that fetch data in the browser (reviews, for example) aren't in the server render. You can force accordions your team controls to render open for companion requests. For anything else, make a deliberate call. Ours was simple: if the content isn't available server-side, the Markdown version doesn't chase it. If some piece of that content really matters to AI readers, the fix is to move it into the server render, which helps HTML visitors and search engines too.

Keep one access policy, not two

A Markdown layer can quietly become a back door. If /checkout.md or /account/orders.md renders anything useful, you've exposed pages you never meant to expose. The fix is one shared list of disallowed path prefixes (cart, checkout, account, orders, API routes, CMS preview routes, internal tools) that does three things:

  • Any .md request under a disallowed prefix returns 404 without rendering the page.

  • Pages under those prefixes don't advertise a Markdown alternate.

  • The list stays aligned with robots.txt.

That last point is the one that slips. During review, the team noticed that one storefront's disallow list still blocked a path that its robots.txt now allowed, so the fix was to make the two match. The general rule is that whenever crawler policy changes, the Markdown policy changes in the same commit.

Tell crawlers which version is canonical

Serving the same content at two URLs is a classic duplicate-content risk, so the relationship has to be stated explicitly in both directions:

  • Markdown responses carry X-Robots-Tag: noindex and a Link header with rel="canonical" pointing to the HTML URL. Search engines keep ranking the HTML page, and AI systems know which URL to cite.

  • HTML responses advertise the Markdown alternate with <link rel="alternate" type="text/markdown"> in the head (and the same thing as a Link header), plus rel="describedby" pointing to the most relevant llms.txt index.

Link: <https://example.com/about>; rel="canonical",
      <https://example.com/llms.txt>; rel="describedby"
X-Robots-Tag: noindex
Content-Type: text/markdown; charset=utf-8

The root llms.txt also says it in plain language: AI systems may discover, summarize, and cite the listed public pages, should prefer canonical HTML URLs when citing, and should not use account, cart, or checkout URLs.

Curate the index for change

A good llms.txt is short. On a catalog site, the hard part is deciding what belongs in it when the catalog changes every week. A few rules held up well:

  • Split by section. The root index lists primary pages (home, about, contact, store locator, FAQs), stable collections, and links to sub-indexes. Products and blog content each get their own index (for example /products/llms.txt and /blogs/llms.txt), generated from the same queries the storefront uses.

  • Prefer stable URLs over churning ones. A "Sale" or "Clearance" listing changes daily, but its URL doesn't, so it can go in the root index with a note that the assortment rotates. Individual SKUs get added, discontinued, and re-priced, so the product index is a capped, best-selling slice, and it tells readers to start with collections.

  • Put policies in an Optional section. Shipping, returns, warranty, financing, and privacy pages are high-value context for an AI answering a shopper's question. Listing them in the spec's "Optional" section signals they're secondary to the primary reading list.

Validate the index in code

Because the indexes are generated, a CMS edit or a bad query can quietly break the format. We added a small assertion that runs every time an index is rendered. It fails loudly if the output stops matching the spec's shape:

  • Exactly one top-level title, followed immediately by a blockquote summary.

  • Only second-level section headings, with no deeper nesting.

  • Every list item is a Markdown link, and every link is an absolute http(s) URL.

  • The optional section is named exactly Optional.

It's a tiny guard, but it turns "AI readers have been getting a malformed index for a month" into an error you see the same day.

Measure before you invest more

We treat llms.txt as a way to prepare for where things are heading, not as a ranking tactic. The build was deliberately small: about an hour with AI assistance on our own simpler site, and roughly a day or two of developer time per storefront on the client side, mostly spent on edge cases. The next step is evidence: watch server logs for AI crawlers and agent user agents fetching /llms.txt and .md paths, and only then decide whether richer Markdown (server-rendered reviews, expanded FAQs) is worth the extra work.

A checklist you can reuse

  • Generate Markdown from the live render. Never hand-maintain a copy.

  • Use the same path plus .md, and run it through the same request handler as the HTML page.

  • Wait for the full render, follow redirects with a limit, and use the visible main heading as the title.

  • Decide explicitly what to do about collapsed and client-only content.

  • Keep one disallow list that matches robots.txt, and return 404 for private paths.

  • Send noindex and rel="canonical" on Markdown, and rel="alternate" and rel="describedby" on HTML.

  • Keep the index short, split by section, and focused on stable URLs, with policies under Optional.

  • Validate the index format in code.

  • Check your logs before you invest further.

If you'd like help making your storefront or marketing site readable for AI systems without adding a second content workflow, talk to our team.

Drag to pan. Use +/− or Ctrl/Cmd + scroll to zoom. Pinch to zoom on touch devices.

llms.txt and Markdown Pages Without a Second Website