first โ it argues,
honestly, that most WordPress sites should stay put.
This article assumes the decision is made. It is about doing it without waking
up to a traffic graph shaped like a cliff.
#Step 1 โ Inventory what the plugins actually do
Do this before anything else, because it is where migrations get their real
budget.
Go through the plugin list and sort each into one of three buckets:
Bucket
Example
Migrates?
Stores content
Advanced Custom Fields, custom post types
Yes โ it is in the export
Transforms content
Shortcodes, page builders, related-posts blocks
Partly โ output is in the body, logic is not
Provides behaviour
Forms, bookings, memberships, e-commerce, caching
No โ rebuild or replace
The third bucket is the project. A form plugin becomes a form service. A
membership plugin becomes auth plus gating in your front end. Nobody exports
this, and it is the reason migrations run long.
The second bucket has a specific trap: page builders. If your posts are
built with Elementor, WPBakery or similar, the exported body is a soup of
shortcodes and wrapper markup, not clean HTML. Plan to clean it, or to rewrite
the affected pages by hand. Count them now.
Tools โ Export โ All content produces a WXR file: XML containing posts, pages,
custom post types, taxonomies, comments and references to media.
References, not files. The images are still sitting in wp-content/uploads and
have to move separately.
Two things worth checking before you trust the file:
Size. Very large sites hit PHP limits mid-export and produce a truncated
file that still looks valid. Compare the post count in the XML against
wp_posts.
Encoding. Check a post with non-ASCII characters end to end. Encoding bugs
discovered after import are miserable to unpick.
The instinct is to import and then tidy. Do it the other way round.
A WordPress post with eight custom fields is usually two or three proper types
in a structured CMS. Decide that shape before the import, or you will carry
WordPress's model into a system you chose specifically to escape it.
In corpusctl there are no built-in content types; you define them from the panel
or the API, and fields carry roles โ title, slug, body. The role is
how the system knows which field is the heading regardless of what you named it,
which is what makes computed fields like reading time and table of contents work
without configuration.
A typical mapping:
WordPress
Becomes
Post title
field with role title
Post slug
field with role slug
Post content
field with role body, blocks
Featured image
media reference
Categories, tags
taxonomy relation, or a tags field
Author
separate author type with a relation
ACF field group
its own fields, or a nested object
Localisation is a field property, not a duplicate site. If you were running
WPML or Polylang, this is the step where that structure gets simpler.
corpusctl's importer module reads WXR and has a dry-run mode: it produces a
plan without writing anything.
The response contains the entries it would create, plus skipped with a
reason for each and warnings. Read both. Silent skipping is how you
discover in month two that 400 posts never arrived.
Import limits exist for a reason โ the parser rejects XML containing
DOCTYPE/ENTITY declarations outright, because that is the XML-bomb vector, and
files above the size ceiling are rejected rather than streamed into memory. If
your export trips either, split it.
Then run it for real and reconcile counts: posts in, posts out, per type.
This is the largest task and it is a normal web project, so the only migration-
specific advice is about URLs.
Preserve the URL structure if you possibly can. Every URL you keep is a
redirect you do not have to write and a ranking signal you do not have to
transfer. If the current structure is /2019/03/some-post/ and you hate it,
changing it is a separate project โ do not bundle a URL restructure into a CMS
migration. If both go wrong you will not know which one caused it.
On the read side, corpusctl's client does not call an API: published content is
written to an immutable hash-addressed file at publish time, and the client
resolves the slug to a manifest shard locally, then fetches that address from
the nginx edge cache. Two requests, both cacheable, neither touching the
database.
Block renderers ship with zero default styling, so your design work is
your design work โ the CMS contributes no CSS to argue with.
#Step 6 โ Redirects, including the ones you forget
Every old URL needs a 301 to its new home. The list people write:
Posts and pages
The list they forget:
Attachment pages (/some-post/image-name/) โ WordPress generates one per
media item and Google has indexed them
Uppercase and trailing-slash variants, if the old server accepted them
Export the URL list from Search Console (Pages report) rather than from the
sitemap. The sitemap is what you meant to publish; Search Console is what Google
actually found.
Go live at a low-traffic hour, then watch for two weeks:
Search Console โ Coverage for a spike in 404s or "excluded" pages;
Performance segmented by page type to see if one template lost more than the
others.
Server logs โ the crawl log tells you what Googlebot is hitting and what it
gets. A 404 in the log is a redirect you missed, and fixing it the same day
costs nothing while fixing it in a month costs rankings.
Core Web Vitals โ field data, not lab scores. This is usually where a
headless migration shows a genuine improvement, because pre-built content
served from cache does not have a slow origin to hide.
Do not delete the WordPress install for at least a month. Keep it running,
unlinked and noindex, so you can check what a page used to look like.
Editors will complain about preview. WordPress previews inside the CMS.
Headless preview goes through your front end via draft mode, and if you have not
wired it up before cutover, the editorial team will notice on day one. Wire it
first. (corpusctl's draft mode helper validates the secret with a constant-time
comparison and refuses any slug that is not a confirmed relative path, which
closes the open-redirect hole most hand-rolled preview endpoints have.)
Shortcodes leave residue. Anything a plugin was rendering appears as raw
text in the imported body. Grep the corpus for [ after import.
Comments. WXR exports them; most headless CMSs have nowhere to put them. If
comments matter, plan a third-party service before cutover, not after.
The first month is slower. People know where everything is in WordPress
after ten years. Budget for that, and do not schedule a launch in the same week.