Testing a Sitemap Generated from a Headless CMS
A headless CMS can contain drafts, scheduled articles, archived records and published snapshots.
TL;DR
- Decision: Acceptance tests should ensure the sitemap reflects only the intended public content from the CMS, producing a truthful inventory of public URLs rather than every backend record.
- Useful method: Define the eligible article set in ordinary language before coding, and use a consistent canonical route policy so sitemap entries and pages resolve to the same public URLs.
- Limit / success check: Verify completeness and correctness by exercising pagination with a dataset larger than one response page and by validating listed URLs, not just XML syntax or database counts.
The sitemap should reflect public content
A headless CMS can contain drafts, scheduled articles, archived records and published snapshots. The website’s sitemap should be generated from the intended public set, not from every record returned by a convenient database query.
This guide focuses on the acceptance tests for a CMS-backed sitemap. It is not a replacement for a general sitemap reference. RankWin publishes it as a content-workflow provider, and the exact implementation depends on the site’s delivery contract and publication model.
Define the eligible article set
Write the rule in ordinary language before implementing it. For example, include articles that have a public published version and a valid canonical route; exclude private drafts and records that have been deliberately withdrawn. The rule should match what the website actually serves.
For a fictional SaaS blog, an article can have a published version while an editor works on a new draft. That article may remain eligible for the sitemap because the current public snapshot still exists. A new scheduled draft with no public version is a different case.
| CMS state | Sitemap question | Test evidence |
|---|---|---|
| Draft only | Is anything public yet? | Excluded if no public page exists |
| Published with newer draft | Which version remains public? | Existing public URL retained appropriately |
| Scheduled future article | Has publication actually occurred? | No premature exposure |
| Archived or removed | What public response is intended? | Listing matches the decision |
| Changed slug | Which URL is canonical now? | New route and redirect agree |
Related reading: A Website Migration SEO Checklist With a URL-Level Rollback Plan.
Use the canonical route builder
The sitemap and article pages should use a consistent route policy. If one assumes /blog and another serves /blogs, the sitemap can list invalid or redirected URLs. Centralize or carefully coordinate the host, path and slug rules.
Check trailing slashes, encoding and unusual slugs. A title containing punctuation should not produce a malformed URL merely because it was assembled differently in the sitemap code.
Google’s sitemap documentation provides the official format and submission guidance. Use it to validate the output, while separately verifying that each listed URL represents the intended public page.
Test pagination and completeness
A content API may return only the first page of records. A sitemap generator that ignores pagination can silently omit older articles while appearing valid. Create or use a test set larger than one response page and verify the full eligible count.
Do not equate database count with sitemap count without considering publication eligibility. The comparison should use the same rule. Keep a small list of known included and excluded articles as a regression fixture.
For a growing blog, also verify the strategy for splitting sitemap files when needed. The implementation should follow current protocol limits rather than relying on an unbounded response forever.
Related reading: XML Sitemap Best Practices for a Growing Blog.
Treat modification dates as facts
A sitemap modification date should correspond to a meaningful page change under the site’s chosen policy, not simply the time the sitemap was requested. Updating every date on every fetch can make the signal less informative.
Decide whether the date comes from the public snapshot, publication event or another verified field. A draft edit that has not changed the public page should not automatically imply that the public content was modified.
Document the choice so future developers do not replace it with a convenient current timestamp without understanding the consequence.
Verify the response and the listed pages
Check that the sitemap endpoint returns the intended XML and content type, not an HTML error page with a successful status. Parse it with an appropriate validator and fetch representative listed URLs.
Inspect the page canonical, public content and relevant robots signals. A well-formed sitemap can still list a page that returns an error or points its canonical elsewhere. Those are separate checks.
RankWin’s CMS delivery may provide the content source, but the destination website owns the final sitemap response. Test that boundary through the actual frontend rather than assuming the CMS record proves it.
Handle upstream failures deliberately
Decide what happens if the CMS is temporarily unavailable. A site may use a valid cached sitemap, return an explicit failure or follow another documented strategy. It should not silently produce an empty successful sitemap that looks like the entire blog disappeared.
Test the chosen behavior in a controlled environment. Make sure operational monitoring can distinguish a valid empty site from a failed content fetch.
Keep credentials server-side where required by the delivery integration. A public sitemap does not justify exposing a private CMS key in browser code or generated output.
Review after publication changes
Test the sitemap after a new article goes live, an article is withdrawn and a slug changes. Verify that the expected update occurs and that caching does not preserve a stale set indefinitely.
Record the public endpoint and the relevant verification results in the release checklist. Submission to a search engine is a separate action from proving the sitemap is correct, and neither guarantees indexing.
A reliable CMS-backed sitemap is a truthful inventory of intended public URLs. Its value comes from matching publication state, routing and modification evidence—not merely from producing XML that passes a syntax check.
