Event intake
Know whether expected triggers are arriving at all. A workflow with zero failures can still be broken if upstream events silently stop.
volume · event ID · timestamp
D2 Automation Knowledge · Observability
A green workflow run does not prove the business outcome happened. Production observability has to connect event intake, execution telemetry, dependencies, retries, infrastructure health and the final downstream state operators actually care about.
Direct answer
Monitor both technical execution and business delivery. Execution success rate alone is insufficient: track trigger intake, execution outcome and latency, dependency errors, retries and dead-letter state, queue and worker health where applicable, and the final acknowledgement that proves the expected business outcome actually happened.
Observability model
These layers answer different questions. Monitoring becomes useful when operators can trace one business event across all of them instead of opening isolated workflow executions and guessing what happened.
Know whether expected triggers are arriving at all. A workflow with zero failures can still be broken if upstream events silently stop.
volume · event ID · timestamp
Track execution outcome, duration and error class so operators can distinguish healthy processing from repeated or slow failure patterns.
success/failure · duration · execution ID
External APIs, databases and services fail independently. Preserve status, timeout, rate-limit and authentication context instead of collapsing every dependency problem into one generic workflow error.
2xx/4xx/5xx · timeout · rate limit
For queue-mode deployments, observe whether work is waiting, workers are available and execution capacity is keeping up with intake.
queue depth · worker health · backlog
Retries, terminal failures and replayable dead-letter records need to be visible as operational states rather than hidden inside individual executions.
attempts · retry state · dead-letter
Verify the outcome the automation exists to create: the lead stored, ticket acknowledged, message delivered, document indexed or business state updated.
business ID · expected state · acknowledgement
Correlation first
Connect the trigger, n8n execution, downstream requests, retries and final business state with the same traceable identifier wherever the systems allow it.
Choose where timing starts and ends, then monitor the distribution. Average latency can hide slow-tail events, so percentiles such as P95 are useful when the process has time expectations.
Core signals
How many events entered the system, and did expected event streams unexpectedly stop?
Detect missing work before execution metrics can even see it.
Success, failure, cancelled or other relevant execution state by workflow and failure class.
Separate isolated failures from systemic patterns.
Time from defined business intake boundary to the final usable outcome; use distributions where useful.
Expose long-tail delay that averages can hide.
Timeouts, authentication failures, rate limits, server errors and schema failures by dependency.
Show whether the workflow or an external system is the actual bottleneck.
Attempt count, retry age, terminal failures and replay status.
Make recovery backlog visible and bounded.
Queue depth, worker availability and execution pressure where queue mode applies.
Identify capacity pressure without confusing it with downstream slowness.
Evidence that the destination reached the expected business state.
Close the gap between a green run and a correct outcome.
Actionable alerts
Alerting is not a copy of the execution log. It is an operational handoff with enough evidence to classify impact and start recovery safely.
Which operating surface is affected?
Which exact business event should the operator inspect?
Transient dependency, authentication, validation, business rule, timeout or other known class.
Is this the first failure, an active retry or a terminal condition?
What customer, order, ticket, lead, document or downstream state may be incomplete?
Retry, replay, reconcile, correct data, rotate credentials, escalate or stop processing.
Incident loop
Alert on an actionable technical or business condition, not every noisy execution detail.
Use the event or business key to connect intake, workflow execution, dependency calls and downstream state.
Identify whether the problem is transient, terminal, data-related, dependency-related or a business-rule exception.
Bound retries, pause harmful side effects or isolate the failing dependency when the incident can amplify.
Replay or retry from durable context through the same idempotency and validation controls.
Confirm the intended downstream business outcome instead of stopping at a green rerun.
Update the alert, runbook, threshold or workflow control so the same failure is easier to diagnose next time.
Common anti-patterns
You miss silent trigger loss, partial success and business outcomes that never materialize.
Long-tail delays disappear inside the average even when users or downstream processes experience them.
Operators know something failed but cannot reliably find the affected record or customer outcome.
Authentication, schema and business-rule failures become noise or incident amplification instead of actionable exceptions.
An alert becomes a notification feed rather than an operational control.
A workflow may finish while a downstream write, delivery or state transition is missing or incorrect.
Production checklist
A stable event or correlation ID connects intake to outcome.
Execution outcome and duration are visible without opening every run manually.
Dependency failures are classified rather than flattened into generic errors.
Retries and terminal failures have durable, inspectable state.
Queue depth and worker health are monitored where queue mode applies.
The expected downstream business acknowledgement is observable where possible.
Each actionable alert has an owner and a runbook or recovery instruction.
Replay and retry paths preserve idempotency and validation controls.
Related evidence & guidance
See how webhook ingress, Redis queues, workers and PostgreSQL create separate infrastructure failure domains.
ExploreSee event persistence, durable conversation state, SLA logic and recovery paths applied to customer operations.
ExploreDesign bounded retry and replay paths so monitoring connects directly to recoverable operational states.
ExploreFAQ
No. Execution success is one technical signal. Production observability also needs intake volume, end-to-end latency, downstream dependency evidence, retries or dead-letter state, and verification that the expected business outcome actually happened.
At minimum, identify the workflow or system, correlation or business event ID, failure category, attempt count, timestamp, affected business outcome, current recovery state, and the operator action or runbook needed to diagnose or replay safely.
Log enough context to trace one business event from intake to final outcome: stable event or correlation ID, execution ID, timestamps, status, dependency calls and error classes, retry state, important state transitions, and the final downstream acknowledgement where available.
Define the end-to-end boundary first, then monitor the latency distribution rather than only an average. Percentiles such as P95 can expose long-tail delays that average duration hides, especially for workflows with response-time expectations.
Where queue mode is used, monitor queue depth, backlog age where available, worker availability or health, execution duration, failure rate and dependency limits together. More workers do not fix slow APIs, blocked databases or serial workflow bottlenecks.
Need production visibility?
Authorship & accountability
D2 AI & Automation TeamProduction automation, APIs, data pipelines and AI-assisted systems
D2 keeps claims, assumptions and evidence separate. Citations are attached only when a relevant source or evidence asset is available; unresolved material is not automatically presented as a verified fact.
Review D2's evidence methodology →