- On a headless storefront, redirects, canonicals, sitemaps, robots.txt, and noindex are code you own, not platform defaults.
- Framework redirect helpers often default to 302; choose 301 vs 302 in one shared helper that knows the business intent.
- Collapse duplicate URLs (like trailing slashes) with a 301 that preserves query strings, and list only canonical 200 URLs in sitemaps.
- In robots.txt a named user-agent group replaces the wildcard group, so repeat private-path disallows in every named group.
- noindex only works if crawlers can fetch the page: ship noindex before a robots.txt block for anything already indexed.
When a Shopify theme renders your store, the platform quietly answers a lot of questions for search engines: which old URLs moved for good, which version of a page is the real one, what belongs in the sitemap, and what should never be crawled. Move to a headless storefront and those answers don't disappear. They move into your route loaders, your redirect helpers, and a handful of text files you now write by hand.
That shift is easy to miss because nothing breaks. Pages render, customers check out, and the site looks finished. The problems show up months later as an SEO audit item that "has been sticking around forever." The lesson is simple to state and harder to practice: on a headless site, every HTTP response is a statement to crawlers, so each one should be a deliberate decision with an owner, not a framework default.
Below are the six crawl signals we tightened during a recent SEO cleanup across two sister Hydrogen storefronts and an internal sales app for a multi-brand retail client. The details are Shopify Hydrogen and Remix, but the checklist applies to any headless or framework-rendered site.
1. Redirects: Match the Status Code to the Intent
Shopify lets merchants manage URL redirects in the admin, and a headless storefront can read them through the Storefront API's urlRedirects query. The usual pattern is a shared "not found" helper: if a route can't find its product or page, check for a configured redirect, otherwise fall back somewhere sensible.
The trap is that the helper returned a URL string, and each route wrapped it in redirect(url). In Remix, as in many frameworks, that helper defaults to a 302 (temporary). So every redirect a merchandiser created to say "this product URL has permanently moved" went out as "this is temporary, keep the old URL indexed." Ranking signals never consolidated onto the new URLs, and the issue survived several audits because it was spread across dozens of route files.
The fix was to make the helper return the full response, with the status chosen where the intent is known:
// Admin-configured redirects are permanent decisions
if (redirectTarget) {
return redirect(redirectTarget, 301);
}
// Soft fallbacks are not a statement that the page moved
if (isStoreLocator) {
return redirect(`${origin}/store-locator`, 302);
}
return pathname !== '/' ? redirect(`${origin}/`, 302) : null;Callers now just return whatever the helper gives them, so no route can quietly downgrade a permanent redirect again. Two design points carry over to any stack:
Decide the status code in one place, next to the business rule. If call sites build the response, the status will drift.
Be honest about fallbacks. A temporary redirect to the homepage or a locator page is kinder to shoppers than a dead end, but search engines often treat "missing page sends you to the homepage" as a soft 404. When a URL truly has no close successor, a real 404 or 410 is the clearer signal. Use fallbacks deliberately, not as a catch-all.
2. Canonical URLs: Collapse Duplicates at the Edge
One high-traffic interactive page was indexed twice: once at its normal path and once with a trailing slash. Both rendered the same content, so search engines had to guess which one to rank. The loader now normalizes the URL before doing any work:
const canonicalPath = '/pages/product-quiz';
if (url.pathname.endsWith('/') &&
url.pathname.replace(/\/+$/, '') === canonicalPath) {
return redirect(`${canonicalPath}${url.search}`, 301);
}Note the url.search: a canonical redirect that drops query parameters breaks campaign tracking and "retake" style flows. A canonical link tag is still worth having, but a 301 settles the question for crawlers and users alike. If duplicates show up on more than one route, move the normalization into a shared request handler instead of patching pages one by one.
3. Sitemaps: List Only URLs You Want Ranked
A sitemap is a list of promises: each entry says "this URL is canonical, returns 200, and is worth indexing." Headless sitemaps are usually hand-assembled from several sources, so they collect leftovers. The audit found two:
A retired content section that still had routes and a sitemap entry, even though the brand had stopped publishing there.
Two different URLs for the same store locator, both listed.
We removed both entries and deleted the unused routes entirely, rather than just hiding them from the sitemap. Dead routes are a liability: they keep rendering thin pages that can be crawled through old links. A quick test to add to your release checklist is to fetch every URL in the sitemap and confirm each returns 200, is self-canonical, and isn't marked noindex. Anything that fails is either a sitemap bug or a page bug.
4. robots.txt: Treat It as Access Policy, and Mind the Group Rules
The storefront's robots.txt already blocked cart, checkout, account, and order paths for all crawlers, with a separate group for Google's ad crawler (which ignores the wildcard group unless named). The new requirement was to welcome search engines and AI assistants explicitly, so a decision that used to be implicit became visible in the file.
That's a reasonable goal, but there is a subtlety that catches many teams. Under the robots exclusion standard, a crawler follows only the most specific group that matches its user agent and ignores the rest. If you add:
User-agent: *
Disallow: /cart
Disallow: /account
User-agent: Googlebot
Allow: /then Googlebot no longer sees the cart and account disallows at all. The "explicit allow" has silently reopened your private paths to the crawler you care about most. If you name bots, repeat the private-path rules inside each named group, or generate the file from one list of disallowed prefixes so every group stays in sync. The same shared list can drive any other access rules your app enforces, such as the paths excluded from AI-readable Markdown, so robots.txt and runtime behavior never disagree.
5. Hiding a Whole App: noindex Needs Crawl Access
The client also runs a tablet sales app for in-store associates on its own subdomain. It was built on the same storefront starter, so it inherited a full sitemap and an indexable SEO configuration. The cleanup made it unmistakably private to crawlers:
noindex, nofollowin the SEO defaults for every page type (home, product, collection, page, policy, blog).An
X-Robots-Tag: noindex, nofollowheader on every HTML response, so routes that don't use the SEO helper are covered too.An empty sitemap instead of a generated product list.
Disallow: /for all crawlers in robots.txt.
Layering signals is sensible, but order matters. A noindex directive only works if the crawler can fetch the page and see it. If robots.txt blocks crawling, URLs that are already indexed can linger in results as "indexed, though blocked by robots.txt," because the crawler never gets to read the instruction to drop them. The safer sequence for an app that has already leaked into search is:
Ship
noindex(meta tag and header) while still allowing crawling.Watch Search Console until the URLs drop out, and use the removals tool if you need them gone faster.
Then add the robots.txt block, or better, put the app behind authentication so there is nothing public to index.
If the app was never indexed, blocking crawling from day one is fine. The point is to know which situation you're in before choosing the order.
6. Structured Data: Generate It From Content You Already Maintain
The last item was adding FAQ structured data to product listing pages. The easy path would have been another CMS field for merchandisers to fill in. Instead, the team generated the schema from the accordion blocks editors already use for FAQs on those pages: walk the page's content blocks, extract each question and answer, strip the HTML, and emit FAQPage JSON-LD only when real items exist.
This keeps schema and visible content identical by construction, which is what search engines expect, and it costs editors nothing. It's the same principle as the rest of this list: derive crawl signals from a single source of truth instead of maintaining a parallel copy that can drift.
A Short Crawl-Signal Checklist for Headless Sites
Because these issues hide behind pages that look fine, they're best caught with a small, repeatable check after each release:
Redirects: request a few known moved URLs with
curl -Iand confirm 301s for permanent moves and 302s only where the move is truly temporary.Canonicals: try trailing-slash, uppercase, and parameter variants of key pages and confirm they resolve to one URL.
Sitemap: every listed URL returns 200, is self-canonical, and is indexable.
robots.txt: every named user-agent group still contains your private-path disallows.
Private apps and staging: HTML responses carry
X-Robots-Tag: noindex, and you know whether crawling should still be allowed while they drop out.Structured data: schema is generated from on-page content and validates in Google's Rich Results Test.
When two storefronts share an architecture, ship each fix to both at once. In this case the same redirect helper change went to both sister sites in the same cleanup, so neither drifted back to the old behavior.
Key Takeaways
Going headless moves SEO decisions out of the platform and into your code. Treat redirects, canonicals, sitemaps, robots rules, and indexing headers as features with owners.
Framework redirect helpers usually default to temporary. Choose the status code in one shared place, next to the business rule that knows the intent.
Sitemaps should contain only canonical, indexable, 200-status URLs. Delete dead routes instead of just hiding them.
In robots.txt, a named user-agent group replaces the wildcard group for that crawler. Repeat your disallows or generate the file from one list.
noindex only works when crawlers can read it. Sequence noindex before robots blocks for anything already indexed.
If you're planning a headless build or auditing one that has been live for a while, our overview of Hydrogen's built-in SEO features covers the defaults you start with, and our guide to launching a new website without losing search visibility covers the migration side.
