Most automations fail in boring ways: a request times out, a worker restarts, a webhook gets delivered twice, or someone clicks “run” again because they are not sure it worked. The failure is annoying, but the real damage is often the retry. Without safeguards, the same event can create two invoices, two shipments, or two tickets.
Idempotency is a reliability concept that turns retries from a risk into a feature. Instead of asking “Did this run already?”, you structure the automation so running it twice produces the same external result as running it once.
This post explains an evergreen, system-agnostic pattern you can apply in scripts, serverless functions, integration platforms, or full backend services. You will leave with a simple design that prevents duplicates and makes incidents easier to unwind.
Why duplicates happen
Duplicate side effects appear when your automation crosses a boundary where you cannot rely on exactly-once delivery. In practice, most distributed systems are “at least once” at some layer, meaning repeats are normal.
- Webhook redelivery: Many providers retry webhooks when they do not receive a timely 2xx response.
- Client retries: Your job runner or integration tool retries on 429s, 5xx responses, and timeouts.
- Ambiguous outcomes: You send a “create invoice” call, the network drops, and you cannot tell if it succeeded.
- Concurrency: Two workers pick up the same queued message, or the same schedule overlaps.
- Human replays: Someone reruns yesterday’s workflow to “be safe.”
If the side effect is benign, duplicates are annoying. If the side effect is expensive or user-facing, duplicates become a trust problem: customers get double charged, agents see multiple tickets, inventory counts drift, and auditors start asking questions.
What idempotency means in practice
An operation is idempotent if repeating it does not change the final state beyond the first application. “Set status to Paid” can be idempotent. “Add $10 to balance” usually is not.
For API automations, the goal is narrower and more practical: given the same business intent, ensure the external side effect happens at most once, and on retries you return or reuse the original result.
Key Takeaways
- Assume duplicates will happen; design for safe retries instead of hoping they do not.
- Use a stable idempotency key that represents the business intent, not the attempt.
- Store outcomes so retries can return the same result without repeating the side effect.
- Idempotency is not just a header; it is a data model decision.
The core pattern: idempotency key plus outcome store
The most reliable approach has two pieces:
- Idempotency key: A unique identifier for the business action you are trying to perform.
- Outcome store: A place to record what you did the first time, so you can reuse it on retries.
This can live in a database table, a durable key-value store, or even a spreadsheet in low-volume cases. The essential requirement is that you can do an atomic “check then write” operation, or a single “insert if absent” that only one worker wins.
Choosing a good idempotency key
A good key is stable across retries and specific to the intent. Common sources include:
- A webhook event ID provided by the sender.
- Your own internal job ID, if the job is generated once per business action.
- A composite like
customerId + invoiceNumberwhen those values are truly unique.
Avoid keys based on timestamps, random UUIDs generated per attempt, or “current run ID.” Those identify attempts, not intent, and will not prevent duplicates.
What to store and why
Your outcome store should capture enough to make retries boring:
- Status: in-progress, succeeded, failed (and optionally expired).
- External references: the created invoice ID, ticket ID, or payment ID.
- Normalized request summary: what you intended to do (not necessarily the whole payload).
- Timestamps: createdAt, updatedAt, and optionally a TTL/expiry.
- Error info: a short code/message for failed attempts so you can decide whether to retry.
Conceptually, a record can be as simple as:
{
"idempotencyKey": "evt_12345",
"status": "succeeded",
"externalId": "inv_987",
"requestFingerprint": "hash(...)",
"createdAt": "…",
"updatedAt": "…"
}
When a duplicate arrives, your automation checks the store by key:
- If it already succeeded, return the stored external ID and do nothing else.
- If it is in-progress, either wait, back off, or return “try later” depending on your architecture.
- If it failed, decide whether to retry automatically or require manual review.
- If it does not exist, create an in-progress record, then perform the side effect, then mark succeeded.
Real-world example: webhook to invoicing
Imagine a small SaaS business that receives a “subscription renewed” webhook from a billing platform. The automation creates an invoice in an accounting system and then opens a “receipt sent” ticket in a helpdesk tool for auditing.
One afternoon, the webhook endpoint becomes slow. The billing platform retries the same webhook event three times. Without idempotency, you get three invoices and three tickets for a single renewal.
With the pattern above:
- The webhook includes an event ID like
evt_123. Use that as the idempotency key. - The first time you see
evt_123, you create an in-progress record, then create the invoice and storeinv_555. - You create the helpdesk ticket and store
tkt_901, then mark the record succeeded. - On the second and third deliveries, you look up
evt_123, see it succeeded, and immediately return success without creating anything new.
Notice what did not matter: the webhook was duplicated, the accounting API might have been flaky, and your own worker may have retried. Idempotency made all of those conditions survivable.
If you want a further safety net, you can also make downstream calls idempotent. For example, pass the same idempotency key to the accounting system if it supports it, or include a unique “source reference” field so the remote system can de-duplicate.
Implementation checklist
Copy this list into your runbook or backlog. It is intentionally tool-agnostic.
- Pick the key: Identify a stable, unique identifier for the business event (not the attempt).
- Define the boundary: Decide which side effects must be “at most once” (invoice creation, charge creation, ticket creation).
- Create an outcome store: Choose a durable place to write records with an atomic uniqueness guarantee on the key.
- Record before side effects: Write an in-progress marker before calling external systems.
- Store external IDs: Persist identifiers from external systems so retries can reuse them.
- Handle in-progress: Decide your behavior when the key is locked (wait/back off, return 202, or queue).
- Set timeouts and TTL: Prevent stuck in-progress records from blocking forever. Include a safe reprocessing strategy.
- Log with the key: Every log line for the workflow should include the idempotency key for quick tracing.
- Test duplicates intentionally: Replay the same event payload multiple times and confirm only one side effect happens.
Common mistakes
Most idempotency bugs are not algorithmic. They come from choosing the wrong key or storing the wrong thing.
- Using a per-attempt UUID: This guarantees duplicates instead of preventing them.
- Only storing “done”: If you do not store in-progress, concurrent workers can both proceed before “done” exists.
- Not persisting external IDs: A retry that cannot reuse the previous result often repeats the side effect.
- Hashing unstable payloads: If you fingerprint a JSON body that contains timestamps or reordered fields, identical intents may not match.
- Ignoring partial success: If invoice creation succeeded but ticket creation failed, the retry might create a second invoice unless you store progress per step or store the invoice ID as soon as you have it.
A practical fix for partial success is to treat the workflow as a set of checkpoints: store each external ID as soon as it is created, and on retries skip steps that already have IDs recorded.
When not to use idempotency
Idempotency is powerful, but it is not always the right lever.
- When the action is inherently additive: “Append a log entry” or “increment a counter” may need a different design, such as deduplicating inputs or using exactly-once semantics in a streaming system.
- When you cannot define a stable key: If you do not have any trustworthy unique identifier and cannot derive one safely, you may need to introduce one earlier in the process (for example, generate a business transaction ID at the source).
- When duplicates are acceptable: Some analytics events or low-stakes notifications can tolerate repeats. Keep the complexity proportional to the risk.
Even in these cases, you can often apply a lighter version: de-duplicate within a time window, or make only the high-risk steps idempotent.
Conclusion
Retries, redeliveries, and ambiguous outcomes are not edge cases in automation. They are standard operating conditions. Idempotency is the simplest, most durable way to make your workflows safe under repeat execution.
If you implement only one improvement this quarter, make it this: pick a stable idempotency key, store outcomes, and ensure every retry reuses the original result instead of redoing the work.
FAQ
Where should I store idempotency records?
Use the most reliable store you already operate: a database table with a unique constraint on the key is common. For smaller automations, a durable key-value store works well. The key requirement is atomic “insert if absent” behavior.
How long should I keep idempotency records?
Keep them at least as long as duplicates may arrive or retries may be attempted. For webhooks, that might be days. For scheduled batch jobs, it might be weeks if humans can replay jobs. Use a TTL if volume is high, but be intentional about the risk of late duplicates.
What if the request payload changes between retries?
If the business intent is the same, retries should not change the intent. Store a request fingerprint and compare it on duplicate keys. If the same key shows up with a different fingerprint, treat it as suspicious: alert, reject, or route to manual review.
Do I still need idempotency if my vendor supports it?
Vendor idempotency helps, but it only covers a single API call. Your workflow may involve multiple calls and steps. Keeping your own outcome store protects the whole automation and makes your system resilient even when one vendor does not support idempotency.