Transient dependency
network timeout · temporary 5xx · service unavailable
Retry bằng bounded attempts + backoff khi operation thực sự safe-to-repeat và dependency có khả năng recover.
D2 Automation Knowledge · Failure recovery
Reliable recovery bắt đầu bằng việc phân loại failure và xác định operation có safe-to-repeat hay không. Blind retry có thể khuếch đại incident; không retry lại biến temporary dependency problem thành lost work.
Direct answer
Retry khi failure có khả năng thành công sau đó và operation safe-to-repeat. Bound attempts, tăng delay khi dependency chịu pressure, preserve original event/error context và đưa exhausted hoặc terminal work vào một dead-letter state có thể inspect và recover.
Failure classification
Temporary timeout, expired credential và invalid payload không thể dùng cùng một retry policy. Failure semantics quyết định waiting có thể thay đổi outcome hay không.
network timeout · temporary 5xx · service unavailable
Retry bằng bounded attempts + backoff khi operation thực sự safe-to-repeat và dependency có khả năng recover.
429 · quota pressure · downstream saturation
Tôn trọng provider timing, giảm pressure/concurrency và defer thay vì retry tức thì làm dependency quá tải hơn.
expired token · invalid key · missing scope · forbidden
Dừng blind retry, surface ownership và sửa credentials/permissions trước khi thử lại.
missing field · invalid type · unsupported state
Quarantine hoặc route để correction vì cùng payload sẽ không tự nhiên thành hợp lệ chỉ nhờ retry.
invalid transition · policy rejection · duplicate-sensitive action
Giải quyết business condition explicit thay vì coi đây là transport instability.
timeout after write · connection lost after remote acceptance
Reconcile remote state bằng business/idempotency key trước khi quyết định repeat write có an toàn hay không.
Recovery flow
Capture correlation ID, business object, dependency và exact error context.
Phân transient, rate-limited, auth, validation, business-rule hoặc ambiguous failure.
Xác định operation có safe-to-repeat không và có cần idempotency/reconciliation trước hay không.
Chỉ chạy bounded retries phù hợp failure class, với delay/backoff thích hợp.
Persist terminal/exhausted work thành inspectable state có original event và attempt history.
Sửa credentials, data, dependency state hoặc business condition trước attempt tiếp theo nếu cần.
Re-enter processing qua cùng validation và duplicate-protection controls như live path.
Xác minh intended business outcome thay vì dừng ở một replay execution xanh.
Backoff model
Mọi automatic retry policy cần terminal condition. Infinite retry che operational debt và tiêu tốn capacity vô hạn.
Staged hoặc exponential backoff cho dependency thời gian recover thay vì bị đánh thêm request ngay lập tức.
Timing variation giảm synchronized retry waves khi nhiều executions cùng fail trên một dependency.
Dead-letter state
Giữ source payload hoặc durable reference đến event đã fail.
Nối failure với exact order, lead, ticket, document hoặc affected object.
Biết implementation version nào xử lý event khi failure xảy ra.
Record cause là transient, rate-limit, auth, validation, business-rule hay ambiguous.
Persist attempt count, timestamps và last error sau final retry.
Giữ not-attempted, confirmed, failed hoặc uncertain state trước replay.
Xác định ai hoặc system nào có trách nhiệm correct, retry, replay, reconcile hoặc escalate.
Anti-patterns
Credential, schema và business-rule failures bị biến thành noise hoặc incident amplification.
Dependency đang chịu pressure lại nhận thêm traffic đúng lúc yếu nhất.
Work không bao giờ đi vào explicit terminal state và operational debt bị ẩn.
Operator biết workflow fail nhưng không reconstruct được event hay recovery action an toàn.
Manual recovery trở thành path thứ hai có thể tạo duplicate records/messages/writes.
Execution thành công nhưng business object vẫn có thể missing, duplicated hoặc ở sai state.
Recovery checklist
Khai báo failure taxonomy trước khi cấu hình retry behavior.
Auth, validation và business-rule failures không dùng chung retry policy với temporary dependency errors.
Mỗi automatic retry path phải có bounded attempt count.
Áp staged/exponential backoff khi repeated requests có thể tăng downstream pressure.
Giảm synchronized retry storms khi nhiều executions fail cùng dependency.
Dùng stable event/business identity và idempotency/reconciliation controls.
Không repeat irreversible action khi chưa biết previous write đã xảy ra hay chưa.
Dead-letter context không được chỉ nằm trong ephemeral execution memory.
Mỗi terminal item cần owner và explicit next action.
Replay phải đi qua cùng validation/idempotency controls như live processing.
Recovery chỉ hoàn tất khi authoritative business state đạt intended result.
Claim boundaries
Retry chỉ là một recovery action cho failure có khả năng transient; terminal/data/business failures cần path khác.
Backoff giảm pressure và tăng chance dependency recover nhưng không chứng minh request cuối sẽ thành công.
Jitter giảm synchronized retries nhưng không thay rate limits, queue/backpressure hay downstream capacity design.
Dead-letter là durable terminal/review state để preserve context và controlled replay, không phải nơi silently drop work.
Notification không đủ nếu operator thiếu original event, failure class, attempt history, side-effect state và recovery action.
Hết automatic attempts chỉ nói policy đã dừng; business object vẫn cần classify, resolve hoặc reconcile.
Replay phải respect validation, state checks và idempotency; manual execution không được bypass controls.
Workflow success không chứng minh downstream authoritative state đúng nếu chưa verify business outcome.
Có retry/backoff/dead-letter không chứng minh recovery time, success rate, uptime hay incident reduction nếu chưa có production evidence.
Related reading
Bảo vệ duplicate-sensitive side effects để retries/replay an toàn dưới repeated delivery.
Đọc tiếpLàm retry state, terminal failures, dependencies và business outcomes visible cho operators.
Đọc tiếpĐặt retry, timeout ambiguity và integration recovery trong production API contract.
Đọc tiếpQuay lại knowledge hub về API, webhook, queue mode, production readiness và RAG reliability.
Đọc tiếpFAQ
Retry những failure có realistic chance thành công sau đó và operation safe-to-repeat, ví dụ một số temporary network error, rate limit hoặc server-side failure. Invalid credentials, unsupported schema, forbidden access và deterministic business-rule failure thường cần intervention thay vì repeated attempts.
Không có universal retry count. Chọn bounded attempts và delay theo downstream recovery pattern, request cost, latency expectation, rate limits và duplicate-side-effect risk. Policy phải kết thúc ở visible terminal state thay vì retry mãi.
Backoff tăng khoảng cách giữa các attempts để temporary dependency problem không bị khuếch đại bởi immediate retries. Nó giảm pressure chứ không cam kết lần sau sẽ thành công.
Jitter thêm variation vào retry timing để nhiều executions không cùng retry đúng một thời điểm. Nó hữu ích khi nhiều jobs có thể fail đồng thời trên cùng dependency.
Là durable terminal/review state cho work đã exhausted retry hoặc cần human/system resolution. Nó giữ original event, business ID, workflow version, failure class, attempts, side-effect state và recovery ownership để replay có kiểm soát.
Không. Nếu chỉ drop work thì đó là data loss. Dead-letter đúng nghĩa phải preserve enough evidence và có explicit resolve/replay/escalation path.
Không mặc định. Timeout chỉ chứng minh caller không nhận usable response. Với duplicate-sensitive write, cần query/reconcile destination bằng business hoặc idempotency key trước khi repeat.
Không blind-retry. Nếu token hết hạn hoặc thiếu scope, việc lặp cùng request không sửa được root cause. Restore credentials/permissions trước rồi mới reprocess.
Có thể quarantine hoặc dead-letter nếu cần correction/review. Quan trọng là không repeatedly retry unchanged invalid payload và phải preserve rejected context để sửa mapping/data.
Không. Recovery phải đi qua cùng validation, state checks và duplicate protection như live path. Nếu không, operator replay có thể tạo duplicate side effects.
Chưa chắc. Cần verify intended business outcome ở authoritative system; green execution chỉ là technical signal.
Không. Đây là recovery-control architecture. Uptime, recovery time, success rate hoặc incident reduction chỉ nên claim khi có measured production evidence.
Need a recovery architecture?
Tác giả & trách nhiệm
Đội ngũ D2 AI & AutomationAutomation production, API, data pipeline và hệ thống có AI hỗ trợ
D2 tách claim, giả định và evidence. Citation chỉ được gắn khi có nguồn hoặc evidence asset phù hợp; nội dung chưa kiểm chứng không được tự động trình bày như fact đã xác nhận.
Xem phương pháp evidence của D2 →