Most automations feel easy at the start: call an API, transform some data, write a record, notify someone. The trouble begins when the automation spans multiple systems and multiple steps, and a failure can happen between any two of them.
When that happens, teams often discover an uncomfortable gap between “each step works” and “the whole workflow is reliable.” You end up with orphaned records, duplicate actions, confused customers, and a backlog of manual cleanup.
This post explains a practical, small-team approach based on the “saga” idea: treat your automation as a series of steps with checkpoints, and define what to do when a step fails after earlier steps already succeeded. You do not need a distributed transaction manager or a complex new platform. You need clarity, a little structure, and discipline.
Why multi-step automations fail (even when each API is fine)
In a multi-step automation, the main risk is not that an API call returns an error. It is that the automation stops in the middle, leaving the world in a half-updated state.
Common causes include:
- Time gaps: Step 1 succeeds, then step 2 fails later because of a transient outage or rate limit.
- Ambiguous outcomes: The request times out. Did the provider create the resource or not?
- Non-atomic side effects: Sending an email, charging a card, or creating a ticket cannot be “rolled back” by default.
- Retries without context: A simple “retry the whole job” may repeat already-successful steps and create duplicates.
- Human interventions: Someone edits a record mid-run, changing assumptions the automation was built on.
A reliable design accepts that partial failure is normal and plans for it. The goal is not zero failures. The goal is controlled recovery and fast, predictable cleanup.
The saga idea in plain language
A saga-style workflow is a sequence of “forward” steps. Each step has a recorded outcome and, optionally, a “compensating” step that can undo or neutralize its effects if later steps fail.
Instead of trying to make the whole workflow atomic, you make it traceable and repairable. That is often the best fit for integrations, where you cannot lock multiple external systems into one transaction.
Forward steps and compensating steps
Forward steps do the normal work: create a shipment, update a CRM, write an invoice. Compensating steps reverse or counteract earlier steps when needed: cancel a shipment, void an invoice, restore a previous status.
Not every action has a perfect undo. In those cases, “compensation” can mean creating a follow-up task, issuing a correcting record, or marking a state so people do not act on it.
What makes this work is a durable workflow record that answers two questions at any time:
- Where did we get to? Which steps completed, with what IDs and timestamps?
- What should happen next? Continue forward, retry a step, run compensation, or escalate to a human.
{
"workflowId": "order_10492",
"state": "CREATED_SHIPMENT",
"steps": [
{"name":"reserve_inventory","status":"ok","ref":"inv_res_7781"},
{"name":"create_shipment","status":"ok","ref":"ship_55219"},
{"name":"capture_payment","status":"failed","error":"timeout"}
],
"nextAction": "retry:capture_payment"
}
A concrete example: order to shipment workflow
Imagine a small ecommerce operation with a custom storefront, a shipping provider, and an accounting system. The automation is supposed to:
- Reserve inventory in the warehouse system.
- Create a shipment label with the shipping provider.
- Capture payment.
- Record the invoice in accounting.
- Email the customer tracking information.
A partial failure scenario: inventory is reserved and a shipping label is created, but payment capture times out. If you blindly retry the full workflow, you might create a second label. If you do nothing, you have a paid-later order with a label that might ship anyway.
A saga-style approach makes the recovery explicit:
- If payment capture fails, retry only that step with a bounded policy (for example, a few attempts over a short period).
- If it still fails, run compensation: cancel the shipping label (if possible) and release inventory, or mark the order as “hold” with a task for staff.
- If the label cannot be canceled, compensate by preventing pickup (status update) and notifying staff. The key is to stop the workflow from quietly “completing” in an inconsistent state.
This is also where you decide what “done” means. For example, you might treat the email as best-effort and allow the workflow to complete even if email fails, as long as tracking is stored in your system for later resend.
Design checklist: make each step recoverable
You can implement saga-style reliability in many ways, but the design questions are consistent. Use this checklist when designing or reviewing an automation.
- Write down the steps in order, including the side effects (records created, messages sent, statuses changed).
- Define step boundaries so each step is small enough to retry safely and easy to observe.
- Persist a workflow record with a stable workflow ID (often derived from the business object, like an order ID) and store external references returned by APIs.
- Make “already done” detectable by storing step results. On retry, the system should skip completed steps or confirm their result rather than repeat them.
- Classify errors per step: transient (retry), permanent (stop), ambiguous (verify before retry).
- Choose a retry policy per step: max attempts, spacing (backoff), and a “give up” state.
- Define compensation for each step that creates an external side effect. If there is no true undo, define a neutralizing action or a human task.
- Decide which steps are optional and can fail without blocking completion (for example, non-critical notifications).
- Log with context (workflow ID, step name, external IDs) so you can reconstruct what happened quickly.
- Create an operator view even if it is simple: a list of failed workflows with next action and error message.
Key Takeaways
- Reliability comes from tracking where you are in the workflow, not from hoping retries fix everything.
- Store external IDs from successful steps so you can continue safely without duplicating work.
- Define compensating actions for side effects, even if the “undo” is a human task and a clear status.
- Different steps need different retry rules. Treat timeouts and rate limits differently from validation errors.
Common mistakes to watch for
Most painful incidents in automations come from a few predictable design gaps. Here are the ones that show up repeatedly in small-team systems.
- Retrying the whole workflow by default. This is how duplicates are born. Retry at the step level using recorded progress.
- No durable workflow ID. If you cannot correlate logs and external records to one workflow instance, recovery becomes guesswork.
- Assuming timeouts mean failure. A timeout often means “unknown.” Build a verification step: look up by idempotency key, search by reference, or query status before retrying creation.
- Not capturing external references. If the shipping provider returns
ship_55219and you do not store it, you cannot cancel, query, or resume safely. - Compensation is an afterthought. Even a simple “release inventory” step can save hours of manual cleanup.
- Blending business logic with plumbing. Keep your step definitions and state transitions readable. The more tangled it is, the more likely operators will bypass it.
A good test is to ask: “If the server stops right after this line, what is the state of each external system, and how do we recover?” If you cannot answer quickly, the workflow is not yet operationally safe.
When not to use a saga-style workflow
Saga-style workflows are a great fit for cross-system automations, but they are not universal. Consider other approaches when:
- You need strict atomicity. If a business rule demands “all or nothing” with no intermediate visibility, you may need a single system of record, a tighter architecture, or a different product decision.
- There is no acceptable compensation. If you cannot undo or neutralize a side effect and it is high impact, reduce the number of steps before that side effect or require manual approval.
- The workflow is tiny and local. If everything happens inside one database transaction, keep it simple and transactional.
- Human judgment dominates. If a person must review most cases anyway, the better investment may be a queue with clear statuses and templates rather than complex automation.
Sometimes the best “reliability improvement” is changing the process: move the irreversible action later, add an approval gate, or redesign to eliminate a step.
Conclusion
Multi-step automations fail in the seams between systems. A saga-style workflow makes those seams visible and manageable by recording progress, retrying intentionally, and compensating when later steps cannot complete.
If you want a simple place to start, pick one existing automation from your archive of scripts and jobs, write down its steps, add a durable workflow record with step status, and define compensation for the first irreversible action. You will feel the operational burden drop quickly.
FAQ
Do I need a workflow engine to do this?
No. Many small teams implement the core idea with a single “workflow table” (or document) and a worker that advances one step at a time. The engine can come later if complexity grows.
What if an external API does not support cancellations or rollbacks?
Compensation can be a neutralizing action rather than a true undo: mark the record as voided, prevent downstream steps, create a human task, or create a correcting transaction. The key is to make the system’s next action explicit.
How do I handle “unknown” results after a timeout?
Treat timeouts as ambiguous, not failed. Add a verification step before repeating the action: check for an existing resource using the workflow ID, a stored request key, or provider search fields. Only create again if you are confident it did not happen.
How granular should steps be?
Small enough that a step is easy to retry and observe, but large enough to avoid excessive state overhead. A good rule is “one external side effect per step” or “one clear business milestone per step.”
What is the minimum set of data I should store?
At minimum: workflow ID, current state, step statuses, timestamps, external IDs returned by APIs, and the last error for any failed step. This is enough to resume safely and to support manual intervention.