Why traffic drops after a headless migration, and how to stop it
Client-side rendering, dropped canonicals and a slow origin. The three ways a headless migration quietly costs organic traffic — and how to close each one.
A headless CMS has no opinion about your HTML. It returns content over an API;
everything a search engine sees is produced by code you wrote.
That is why "is headless bad for SEO?" is the wrong question. The right one is:
what did the rebuild change about the HTML? In practice the answer is one of
three things, and all three are fixable.
#Leak 1 — Content that only exists after JavaScript runs
The most common and the most expensive.
You moved to a framework with server rendering and then fetched the article body
in a client effect, because that is what the tutorial did. Now:
Googlebot gets an empty shell on first pass, and rendering is queued for
later — sometimes much later.
Other crawlers (Bing, AI crawlers, social preview bots) mostly do not render
at all. Your links unfurl blank.
Largest Contentful Paint measures the content, and your content arrives one
round trip after the shell.
0 means your content is not in the HTML. Everything else in this article is
secondary until that returns 1.
Fix — fetch on the server. In the Next.js App Router that is the default;
the mistake is usually a 'use client' boundary placed too high in the tree.
Move it down to the interactive component instead of wrapping the page.
#Leak 2 — Canonical, meta and structured data lost in the rebuild
The old CMS had an SEO plugin that produced canonical tags, meta descriptions,
Open Graph tags and Article schema for every page. Nobody explicitly decided to
remove that. It just was not on the rebuild checklist.
Symptoms: duplicate-content clustering (parameter URLs, trailing-slash variants,
paginated archives all indexing separately), missing rich results, and social
shares that unfurl with the wrong title.
Fix — make these fields part of the content model, not the theme. In
corpusctl the SEO module stores per-entry title, description, canonical and
social image alongside the content, so they travel with the entry rather than
living in whatever renders it.
Then assert them in CI. A test that fetches ten representative URLs and checks
that each has exactly one self-referencing canonical, a non-empty description
under 155 characters, and valid JSON-LD costs an hour and catches this class of
regression permanently.
#Leak 3 — A slow origin behind a cache you have not thought about
Headless usually makes performance better. It makes it worse when the front end
fetches from the CMS on every request and the CMS is a database query away.
Your TTFB is now the CMS's response time plus your render time, on every cache
miss, and every cache miss is a real user waiting.
Fix — decide where content lives at request time. Three options, in
increasing order of how well they hold up:
Fetch per request. The CMS is on the critical path. Fine for a
dashboard, wrong for content you want ranked.
ISR with a time window. Content is at most N seconds stale; every window
expiry is a cache miss someone pays for.
Pre-build and serve from cache. Rendered at publish time, served as a
static artefact. Fastest and cheapest — provided invalidation is handled.
The hard part of option three is invalidation, because the page is a copy of a
computation that could change at any moment.
corpusctl removes the problem rather than managing it. Publishing writes content
to an immutable, hash-addressed file in object storage and refreshes a
manifest; the client resolves the slug to a manifest shard locally, then fetches
that address from the nginx edge cache. The address is a content hash, so the
same address can never return different content — it is cacheable for thirty
days, and publishing writes a new address. There is nothing to purge.
For Core Web Vitals this matters directly: the document arrives from a cache
rather than from a computation, so TTFB stops varying with database load and LCP
stops depending on whether you got a cache hit.
If you run several languages — and a headless CMS is often chosen precisely
because you do — this is the fourth leak.
Rules that are easy to state and easy to get wrong:
Every language version links to every other version, including itself.
There is an x-default pointing at the version for unmatched visitors.
The links are reciprocal. A one-way hreflang is ignored entirely.
Language codes are valid BCP 47. Simplified Chinese is zh-Hans, not zh-CN,
when you mean the script rather than the region.
The URL in hreflang matches the canonical on the page it points to.
Do not verify this by hand. Fetch each URL and assert the full set:
Related: the language-negotiating root URL should return 302, not 301. The
response varies by Accept-Language, so it is not permanent; a 301 gets cached
in the browser and the next person on that machine cannot change their language.
Send Vary: Accept-Language, Cookie with it, or an intermediate cache will
serve the first visitor's language to everybody.
Two weeks minimum after cutover. Rankings move on their own schedule and a
three-day panic tells you nothing.
Search Console → Coverage. A rise in 404s means a redirect you missed. A
rise in "Crawled – currently not indexed" often means thin or duplicate output
from a template.
Search Console → Performance, segmented by page type. If one template lost
more than the others, the bug is in that template, not in the migration.
Server logs, not just Search Console. The crawl log tells you what Googlebot
actually requested and what it received. A 404 there is free to fix today and
expensive to fix in a month.
Core Web Vitals field data, not lab scores. Lab scores measure your laptop.
Draft mode, ISR, the App Router and the render layer. What actually matters when you wire a headless CMS to Next.js — and the three ways it goes wrong.