Transient dependency
network timeout · temporary 5xx · service unavailable
Retry with bounded attempts and backoff when the operation is safe to repeat.
D2 Automation Knowledge · Failure recovery
Reliable recovery starts by asking what kind of failure happened and whether the operation is safe to repeat. Blind retries amplify incidents; no retries turn temporary dependency problems into lost work. Production workflows need bounded recovery states between those extremes.
Direct answer
Retry only when the failure has a realistic chance of succeeding later and the operation is safe to repeat. Bound the attempts, increase delay when the dependency is under pressure, preserve the original event and error context, and move exhausted or terminal work into an inspectable dead-letter state with a controlled recovery path.
Failure classification
The same retry policy should not be applied to a temporary network timeout, an expired credential and an invalid payload. Error semantics determine whether waiting can actually change the outcome.
network timeout · temporary 5xx · service unavailable
Retry with bounded attempts and backoff when the operation is safe to repeat.
429 · quota pressure · downstream saturation
Respect provider timing, reduce pressure and retry later rather than immediately amplifying load.
expired token · invalid key · missing scope · forbidden
Stop blind retries, surface ownership and repair credentials or permissions first.
missing field · invalid type · unsupported state
Quarantine or route for correction because the same payload is unlikely to succeed unchanged.
invalid transition · duplicate-sensitive action · policy rejection
Resolve the business condition explicitly instead of treating it like transport instability.
timeout after write · connection lost after remote acceptance
Reconcile remote state by business or idempotency key before deciding whether another write is safe.
Recovery flow
Capture the failure with correlation ID, business object, dependency and exact error context.
Decide whether the error is transient, rate-limited, terminal, data-related, business-related or ambiguous.
Determine whether the operation is safe to repeat and whether idempotency or reconciliation is required first.
Run only the bounded retries justified by the failure class, with appropriate delay or backoff.
Persist terminal or exhausted work as an inspectable state with the original event and attempt history.
Correct credentials, data, dependency state or business conditions before another attempt where needed.
Re-enter processing through the same validation and duplicate-protection controls as live events.
Confirm the intended business outcome instead of stopping at a technically green replay.
Backoff model
Every automatic retry policy needs a terminal condition. Infinite retry hides failure state and can consume capacity indefinitely.
Staged or exponential backoff gives a recovering dependency time to become healthy instead of receiving another immediate request.
Small timing variation reduces synchronized retry waves when many executions fail against the same dependency at once.
Dead-letter state
Once automatic recovery stops, the system should preserve enough context to explain what failed, what may already have happened and how an operator can safely continue.
Keep the source payload or a durable reference to the event that failed.
Connect the failure to the exact order, lead, ticket, document or other affected object.
Know which implementation processed the event when the failure happened.
Record whether the cause was transient, terminal, validation, auth, rate-limit, business rule or ambiguous state.
Persist attempt count, timestamps and the last error instead of losing context after a final retry.
Record whether downstream effects are not attempted, confirmed, failed or uncertain before replay.
Identify who or what is expected to correct, retry, replay, reconcile or escalate the item.
Anti-patterns
Permanent credential, schema and business-rule failures become noise and can amplify incidents.
A dependency under pressure receives more traffic precisely when it is least able to handle it.
Work never reaches an explicit terminal state, hiding operational debt and consuming capacity indefinitely.
Operators know a workflow failed but cannot reconstruct the event or determine a safe recovery action.
Manual recovery becomes a second path capable of creating duplicate records, messages or writes.
The execution succeeded, but the actual business object may still be missing, duplicated or in the wrong state.
Production checklist
Define failure classes before configuring retry behavior.
Separate transient failures from authentication, validation and business-rule failures.
Set a maximum attempt count for every automatic retry path.
Use staged or exponential backoff where repeated requests could increase dependency pressure.
Add jitter when many executions may fail and retry together.
Protect duplicate-sensitive side effects with stable event or business identity.
Reconcile ambiguous write outcomes before repeating irreversible actions.
Persist terminal failure context outside ephemeral execution memory.
Make dead-letter state visible to an owner with a defined recovery action.
Route replay through the same validation and idempotency controls as live processing.
Verify the final business outcome after recovery.
Related systems
Protect duplicate-sensitive side effects so retries and replay remain safe under repeated delivery.
ExploreMake retry state, dead-letter backlog and affected business outcomes visible to operators.
ExploreSee durable event state, recovery paths and operating controls applied to an event-driven customer operations system.
ExploreFAQ
Retry failures that have a realistic chance of succeeding later and are safe to repeat. Temporary network errors, rate limits and some server-side failures may qualify. Invalid credentials, unsupported schemas, forbidden access and deterministic business-rule failures usually require intervention rather than repeated attempts.
There is no universal retry count. Choose a bounded number of attempts and delay from the downstream recovery pattern, request cost, latency expectation, rate limits and duplicate-side-effect risk. The policy should end in a visible terminal state rather than retry forever.
Backoff spaces repeated requests so a temporary dependency problem is not amplified by immediate retries. Jitter adds variation to retry timing so many failed executions are less likely to retry at exactly the same moment and create another traffic spike.
Preserve enough durable context to diagnose and replay safely: the original event or business reference, workflow version, correlation ID, failure category, attempt count, timestamps, last error, known side-effect state and the recovery action or owner.
Not for important workflows. A notification should point to durable failure context and a defined recovery path. Without preserved event state, operators may know something failed but still be unable to determine what happened or replay safely.
Replay should use the same validation, idempotency and side-effect protections as live processing. Operators should know whether the previous attempt performed no side effect, definitely completed it or ended in an uncertain state before another write is allowed.
Recovery architecture
Authorship & accountability
D2 AI & Automation TeamProduction automation, APIs, data pipelines and AI-assisted systems
D2 keeps claims, assumptions and evidence separate. Citations are attached only when a relevant source or evidence asset is available; unresolved material is not automatically presented as a verified fact.
Review D2's evidence methodology →