RankWin

Reconciling a Partial Article Import Before Retrying the Batch

Reconciling a Partial Article Import Before Retrying the Batch

An import can fail after creating some articles but before returning a complete response.

RankWin Team

TL;DR

  • Treat an interrupted import as an unknown result: preserve the source batch and inspect destination records before retrying to avoid duplicates, overwrites or misattached assets.
  • Use stable identities and durable receipts to reconcile items: include project context, fingerprints and a manifest so each source maps unambiguously to a destination during recovery.
  • Declare success only after explicit reconciliation and checks: verify every source has an explained destination outcome, test retries with small simulated interruptions first.

Treat an interrupted import as an unknown result

An import can fail after creating some articles but before returning a complete response. A timeout therefore does not prove that nothing was written. Retrying the entire folder without checking the destination can create duplicate drafts, overwrite later edits or attach assets to the wrong record.

The first recovery action is to preserve the intended batch and inspect what actually arrived. Keep the source files unchanged while reconciling them with destination records. The goal is to determine the result for each article, not merely make the import command finish with a successful exit code.

Give every source article a stable identity

Use an identity that does not depend only on the title or slug. Both can change during editorial work. A source identifier such as a project-specific article ID should remain attached to the same content item throughout import, revision and publication.

Include the destination project in the identity boundary. Two projects may legitimately use the same slug, while two separate articles in one project may accidentally request it. The importer needs enough context to distinguish those cases before writing.

For a hypothetical batch of fifty articles, preserve a manifest containing each source ID, intended project, content fingerprint and requested slug. This is a proposed operational contract, not a claim that every CMS automatically supplies these fields.

Keep a receipt for each attempted item

A useful receipt records the source identity, destination record ID, attempted operation and outcome. It should also identify the imported version or fingerprint so an operator can tell whether the destination contains the intended text.

Receipt stateRecovery question
Confirmed createdDoes the destination match the source version?
Confirmed updatedWhich prior version was replaced?
RejectedWhat specific validation failed?
UncertainDid the write complete before the response was lost?

Do not convert uncertain into failed merely to simplify a progress bar. That distinction determines whether a retry is safe. Store receipts somewhere durable enough to survive the process interruption that made them necessary.

Reconcile by identity and content evidence

Suppose the importer reports thirty successful records and then disconnects while processing article thirty-one. Inspect the destination for that source identity before issuing another create operation. If the record exists with the expected fingerprint, record the recovered success. If it exists with different content, investigate the version history.

A matching title alone is weak evidence. Another editor may have created a different article with that title, or a prior import may have used an earlier draft. Compare the stable identity, project and relevant content version together.

Keep editorial changes made after the import attempt. A retry should not silently overwrite a newer destination revision simply because the local file is older and easier to access.

Resolve slug collisions explicitly

A slug collision requires a decision, not an automatic numeric suffix in every case. The existing record might be the same source article from a previous attempt, an unrelated article or a deliberately reserved destination. Each situation implies a different action.

For the same identity, reconcile or update through the supported version path. For a genuinely different article, ask the editorial owner to choose a distinct public destination or reconsider the content plan. Preserve the final mapping so links inside other imported articles can be resolved correctly.

Google’s canonicalization guidance is relevant to the eventual public URL policy, but a canonical tag does not repair an importer that confused two records. Resolve identity before publication signals are generated.

Related reading: Canonical Tags vs Redirects: Use the Signal That Matches the Move.

Retry only the operations that remain necessary

If the destination supports idempotency keys or equivalent stable-write semantics, use the documented mechanism consistently. Do not assume that repeating the same request is harmless unless that behavior has been established. A locally generated key offers no protection if the server ignores it.

Construct the retry set from reconciliation results. Confirmed successes should be skipped unless a deliberate revision is required. Rejected records need corrected inputs. Uncertain records need inspection before mutation. Record each new attempt against the same source identity so the history remains understandable.

Test this behavior with a small controlled batch before relying on it for a large release. Simulate an interruption and verify that recovery neither duplicates records nor erases a newer edit.

Related reading: SEO Automation: A Workflow With Clear Review Points.

Check references and assets before declaring completion

A partially imported batch can leave links pointing to articles that do not yet exist. After reconciliation, resolve source references through the final destination mapping. Check that images belong to the intended article and that metadata did not retain a temporary or rejected slug.

The import is complete when every intended source item has an explained destination outcome, not when the counts happen to match. Fifty destination records could still contain a duplicate and a missing article. Keep unresolved exceptions separate from the successful set and leave affected articles out of publication scheduling.

Finish with a reconciliation report containing created, updated, skipped, rejected and unresolved items, plus the mapping needed for subsequent review. This turns recovery into an accountable process and protects editorial work from the most dangerous assumption in an interrupted import: that the last response accurately describes every write that already happened.