Published November 12, 2025
Recognize ambiguous completion
A timeout does not prove that the receiving system rejected the request. Processing may have completed before the response was lost. Conversely, an acceptance response may only indicate that work entered a queue. Understand the interface’s documented semantics and the evidence available for checking final state. This distinction is essential when the event creates receipts, adjustments, shipments, or other records that must not be duplicated.
Preserve the identity of the operation
Use a stable reference that distinguishes the original event from a genuinely new transaction. Where supported, duplicate protection should recognize repeated attempts and return a consistent result. Verify how long that protection applies and what happens after partial processing. Replacing the identifier on every retry can defeat the mechanism intended to protect the operation. The recovery design must be agreed across both sides of the exchange.
Establish state before intervention
Check the relevant business record and its linked downstream events. Determine whether nothing happened, everything happened, or only part of the process completed. Avoid relying exclusively on a technical log line when the business state can be inspected through a supported method. Preserve evidence and use authorized access. If the outcome remains unclear, escalate through the agreed recovery process rather than guessing and hoping a second attempt is harmless.
Control backlog recovery
A large queue can create pressure to replay everything quickly. Correct the underlying cause first, then test a bounded representative event and verify its result. Consider ordering, dependencies, duplicate protection, and downstream capacity before widening the recovery. Record the scope and responsible operator. A replay that overwhelms another system or reintroduces rejected data can extend the incident even if the original connection problem has been fixed.
Make recovery an ordinary design requirement
Include uncertain responses, repeated messages, and partial completion in integration testing. Provide support instructions that explain what to check and which actions are permitted. Review incidents for missing evidence or unsafe manual steps. Dependable integration is not defined by how rarely it fails alone; it also depends on whether the team can recover without creating a second, harder-to-explain problem for the warehouse.
Use with your team
Working checklist
Checks are temporary and are not saved or submitted. Use Print / save PDF for your working copy.
Educational guidance. Apply it to your product, version, operating conditions, and agreed controls; it is not a project-specific solution design.
Explore development & integration