The network will lie to you. The client will give up and retry. The message broker will deliver the same event twice. The load balancer will forward the request again after the upstream timed out, and by the time the second copy arrives, the first one is halfway through charging the customer’s card. Idempotency is the property that makes this survivable. Without it, every retry is a potential data corruption event.
In twenty years of building distributed systems that process real money, real orders, and real inventory counts, the conversations that haunt me most are the ones where I had to explain to a business owner that yes, some customers were charged twice, or that inventory went negative, or that audit logs have duplicates we cannot safely deduplicate. Every one of those incidents traces back to the same root cause: we forgot to make something idempotent.
This is not exotic distributed systems theory. It is the engineering discipline behind every reliable payment processor, every order management system, and every event-driven architecture that runs without corrupting its data over time.
What Idempotency Actually Means
An operation is idempotent if applying it multiple times produces the same result as applying it once. In mathematics, f(f(x)) = f(x). In distributed systems, the definition is more practical: retry all you want, you will not get a different outcome.
This is not the same as saying the operation is “safe.” An HTTP DELETE is idempotent: deleting a resource twice leaves you with no resource, same as deleting it once. But it is not safe; it changes state. HTTP GET is both safe and idempotent. HTTP POST is neither, which is why every payment API, webhook processor, and message queue consumer in the world needs to solve idempotency explicitly.
The reason idempotency matters in practice is that exactly-once delivery is a fiction. At the network level, you have two choices: at-most-once (lose messages sometimes) or at-least-once (deliver duplicates sometimes). You choose at-least-once for reliability, which means your consumers see duplicates, and your APIs get retried, and your message queues fire twice. The only way out is to make each consumer and handler idempotent.
The Kafka consumer group protocol guarantees at-least-once delivery during rebalances. SQS standard queues can deliver the same message multiple times if your consumer does not acknowledge before the visibility timeout expires. Neither of these are bugs. They are deliberate design choices that trade correctness guarantees for availability. Idempotency is your responsibility, not the broker’s.
The HTTP Method Baseline
Before getting into custom idempotency implementations, it helps to understand what HTTP already gives you.
GET, HEAD, OPTIONS, and TRACE are safe and idempotent by specification. PUT and DELETE are idempotent but not safe. POST is neither. PATCH is complicated: a PATCH that says “set quantity to 5” is idempotent; a PATCH that says “add 1 to quantity” is not.
In practice, you cannot rely on HTTP semantics alone because:
- Clients retry on timeouts, and your server may have already processed the request before timing out.
- Load balancers and reverse proxies sometimes retry internally on 5xx responses.
- Message queues that invoke webhooks retry on non-200 responses.
- Mobile clients have unpredictable connectivity and retry aggressively.
The HTTP specification says PUT is idempotent. It does not enforce this. Your database sees two INSERT statements. You need to handle it.
The Idempotency Key Pattern
Stripe invented (or at least popularized) the idempotency key pattern, and it is now the de facto standard for mutation APIs. The pattern is straightforward: the client generates a unique key (typically a UUID v4) and sends it in an Idempotency-Key header. The server stores the key along with the response from the first successful processing. On subsequent requests with the same key, the server returns the stored response without reprocessing.
POST /v1/charges
Idempotency-Key: 4e3a5c2a-b1a2-4f3e-9c7d-1a2b3c4d5e6f
Content-Type: application/json
{"amount": 2000, "currency": "usd", "source": "tok_visa"}
The server flow:
- Look up the idempotency key in a fast store (Redis, or a database table with a unique index on the key).
- If found and already completed: return the stored response.
- If found but still in-flight: return 409 Conflict (prevents concurrent duplicate processing).
- If not found: reserve the key, process the request, store the response, return it.
The reservation step is critical. Without it, two concurrent requests with the same key both see “not found,” both begin processing, and you have a race condition. Redis SET NX (set if not exists) is the right primitive here: it is atomic at the single-key level, so only one writer wins. Combining this with the rate limiting and distributed counters patterns in Redis gives you a complete request-level control plane.
Stripe retains idempotency keys for 24 hours. This is a reasonable window: long enough to cover any realistic retry storm or client-side recovery scenario, bounded enough that your Redis memory does not grow forever.
A few things trip people up consistently:
The key is scoped to the client and operation type. If a client sends the same key for two different operations, detect and reject this. Stripe returns a 400 error if the request parameters do not match the key’s original payload.
Keys must be unguessable. If an attacker can predict your idempotency key scheme, they can intentionally collide with another client’s in-flight request. UUID v4 is the standard.
Storing failures counts as idempotent too. If the first request fails with a 500, store that 500. The retry should return the same 500, not retry the operation. This prevents a subtle class of bug where a partially-applied operation fails, and retries apply part of it again.

Server-Side Implementation
The idempotency store needs two properties: atomicity and durability.
Redis SET NX satisfies atomicity for the reservation step. You SET the key with a short TTL as a “processing” marker, then UPDATE it with the final response once processing completes. If the server crashes mid-processing, the TTL expires and the next retry sees the key as absent, triggering a clean retry.
But Redis alone is not durable. A Redis crash or flush between when you store the response and when you commit your primary database means the client received a success response for an operation that was never persisted. The consequences range from financial discrepancies to ghost records.
The robust pattern stores the idempotency record in the same database transaction as the primary operation:
BEGIN;
INSERT INTO idempotency_keys (key, status, created_at)
VALUES ($1, 'processing', NOW())
ON CONFLICT (key) DO NOTHING;
-- If 0 rows affected, key already exists: read status and return stored response.
-- If 1 row affected: proceed with business logic.
INSERT INTO charges (id, amount, status)
VALUES ($charge_id, $amount, 'captured');
UPDATE idempotency_keys
SET status = 'completed',
response = $2,
completed_at = NOW()
WHERE key = $1;
COMMIT;
If this transaction rolls back for any reason, the idempotency key is never stored. The next retry finds no record and processes cleanly. This is exactly the behavior you want.
For high-throughput services, Redis as a pre-filter still makes sense: check Redis first, return immediately if you find a completed result, fall through to the database transaction only for fresh requests or those in an ambiguous state. This pattern reduces database pressure for the common case (repeated retries of already-completed operations) while keeping durability guarantees on the path that matters.
Database-Level Consumer Deduplication
For event consumers and background workers, the API layer idempotency key pattern does not apply. You need something simpler at the database level.
The pattern is a processed events table with a unique index on the event identifier:
CREATE TABLE processed_events (
event_id TEXT PRIMARY KEY,
processed_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
source_topic TEXT,
result JSONB
);
Your consumer:
INSERT INTO processed_events (event_id, source_topic)
VALUES ($1, $2)
ON CONFLICT (event_id) DO NOTHING
RETURNING event_id;
If the INSERT returns a row, this event has not been processed before. Process it. If it returns nothing, this is a duplicate. Skip it.
The constraint is that event processing and the INSERT must happen in the same transaction, and in the correct order:
- Begin transaction.
- INSERT into processed_events (claims this event).
- Apply business logic.
- Commit.
Reversing steps 2 and 3 creates a window where the business logic succeeds but the processed_events record is never committed (crash between 3 and the implicit COMMIT). The next retry sees no record and reruns the business logic. Keep the deduplication INSERT as the first write inside the transaction.
Saga Pattern and Idempotency at the Step Level
The Saga pattern, which models distributed transactions as a sequence of local transactions with compensating actions for rollback, depends on idempotency at every step. The saga orchestrator retries individual steps on transient failures. If each step is not idempotent, a retry after a partial success creates inconsistent state that is nearly impossible to repair without manual intervention.
This is one of the places where distributed SQL databases like CockroachDB or YugabyteDB have a practical production advantage: they can span a transaction across nodes, reducing the number of compensating transactions you need to write. But even with distributed SQL, remote calls to external services remain non-transactional and require explicit idempotency keys.
A critical principle: every external call should have an idempotency key derived from something stable about the current business operation, not generated fresh on each attempt. If you generate a new UUID every time your service retries calling Stripe, you defeat the purpose entirely. Each retry must send the same key.
UUID v5 gives you deterministic key derivation:
import uuid
NAMESPACE = uuid.UUID("6ba7b810-9dad-11d1-80b4-00c04fd430c8")
def derive_idempotency_key(order_id: str, operation: str) -> str:
"""Given the same inputs, always returns the same UUID."""
content = f"{order_id}:{operation}"
return str(uuid.uuid5(NAMESPACE, content))
# Regardless of how many times this code runs for order_id=42:
key = derive_idempotency_key("42", "stripe_charge")
# => always "2b1c4b0e-..." -- same key, same Stripe response
The transactional outbox pattern is the natural complement to this. Outbox ensures your events are reliably published; idempotent consumers ensure those events are reliably processed without duplication. Neither is sufficient alone.

Multi-Region Considerations
Multi-region architectures introduce a specific idempotency challenge: the same event or request may arrive at two different regional processing nodes before the idempotency store has replicated.
If you use a global Redis cluster (Redis Enterprise Global Active-Active, or ElastiCache Global Datastore), the conflict resolution for concurrent writes to the same key defaults to last-write-wins. This can produce wrong results for idempotency: if region A and region B both process the same payment within the replication window, one charge wins in LWW resolution, but the actual Stripe charge may already have gone through twice.
The safer pattern for multi-region idempotency is geographic key routing: hash the idempotency key to a home region and always route requests with that key to that region. This is deterministic and avoids replication races. The downside is that a region failure makes idempotency checks for keys hashed to that region unavailable, which you need to handle gracefully (fail closed: do not process if you cannot check).
The alternative is a globally consistent database as the authoritative idempotency store. Google Spanner supports conditional mutations. CockroachDB and YugabyteDB support serializable isolation across regions. The cross-region write latency is real (tens of milliseconds per the laws of physics) but is often acceptable for operations like payment processing where correctness matters more than speed.
The distributed caching architecture considerations around Redis replication topology apply directly here. The same tradeoffs between eventual consistency and strong consistency in a caching layer show up in idempotency stores.
Designing APIs for Idempotency
If you are building an API that external clients consume, idempotency is a first-class design concern, not something you bolt on later.
Declare which operations require idempotency keys. Stripe makes this explicit: POST /charges, POST /refunds, and similar mutation endpoints require keys. Idempotent-by-design operations like GET do not. Document this in your API specification.
Specify your retention window. Clients need to know how long their key is valid. If your retention window is 24 hours and the client holds the key for 48 hours before retrying (slow jobs, overnight batches), the key has expired and the retry is treated as a fresh operation. Match the window to your clients’ actual retry patterns.
Be explicit about what “same key, different body” means. Stripe rejects it with a 400 error. Different operations must always use different keys. This is the correct behavior and prevents subtle bugs where a client accidentally reuses a key for a different payment amount.
Return idempotency metadata in responses. Include a header like X-Idempotency-Replay: true on replayed responses so clients can distinguish between fresh processing and a cached result. This helps dramatically in debugging.
Handle the in-flight case properly. Return 409 Conflict with a Retry-After header when a duplicate arrives for a key that is currently being processed. Clients should wait and retry rather than treating this as a permanent failure.
At the API gateway layer, you can implement idempotency checking as a gateway-level policy, before requests reach your services. This works well if your gateway shares a Redis cluster with your services, and it keeps business logic services focused on their core responsibility.
Common Pitfalls
Forgetting side effects outside your database. Your database transaction is idempotent. The email you sent inside that transaction is not. The Slack notification is not. Extract all external calls from inside transactions and apply idempotency keys to each one independently.
Deduplication window too short. I have seen systems with 60-second idempotency key TTLs. Exponential backoff with jitter can easily span 60 seconds, especially in mobile applications with intermittent connectivity. Use 24 hours as your minimum. For batch processing jobs that run overnight, consider longer windows.
Treating operation-level idempotency as message-level idempotency. An operation that is idempotent when called in isolation may not produce correct behavior when called out of order. Kafka messages delivered out of sequence during rebalancing can produce wrong state even if each individual message is processed idempotently. Think about ordering constraints separately from duplication constraints.
Confusing idempotency with uniqueness constraints. Idempotency says “this operation, applied twice, has the same effect as applying it once.” Uniqueness says “this record should exist only once.” They overlap but are not the same. A system can enforce uniqueness (unique constraint on email address) without being idempotent in the retry sense. A unique constraint rejects the duplicate; an idempotency check returns the original result.
Not testing idempotency. The classic gap: the happy path is tested, the retry path is not. Write explicit tests that replay requests and verify the result is identical. Write tests that fire concurrent duplicate requests and verify only one takes effect. Run chaos tests that kill your service mid-operation and verify clean recovery. Idempotency bugs surface in production under load, not in development.

Putting It Together: A Production Checklist
When I review a new service or API for production readiness, idempotency gets its own checklist:
- Every POST and non-trivially-safe PATCH endpoint accepts an
Idempotency-Keyheader. - Idempotency keys are stored transactionally with the primary operation, not separately.
- Redis is used as a fast pre-check, not as the authoritative store.
- All external calls (Stripe, email, SMS, other services) use deterministic idempotency keys derived from stable business identifiers.
- Event consumers check for duplicate event IDs before processing.
- The deduplication check and business logic write happen inside the same database transaction.
- Idempotency key TTL is at least 24 hours.
- Concurrent duplicate requests during the in-flight window return 409, not 500.
- Integration tests explicitly replay requests and verify no duplicates.
The CQRS and event sourcing patterns that many teams adopt as they scale their event-driven architectures only function correctly when the underlying command handlers are idempotent. Event sourcing gives you a replayable log; if replaying that log produces duplicate state changes, the log is unreliable.
Conclusion
Idempotency is not a nice-to-have. In any system where retries happen, non-idempotent operations accumulate data corruption over time. Charges double. Events process twice. Inventory goes negative. The incidents that follow are expensive to investigate, often impossible to fully repair, and always avoidable.
The implementation is straightforward. An idempotency_keys table with a unique index, a Redis pre-check for performance, and deterministic key derivation from your business identifiers covers the vast majority of production scenarios. The hard part is discipline: auditing every mutation, every external call, every event consumer, and asking “what happens when this runs twice?”
The patterns here are what payment processors, order management systems, and reliable event-driven services run in production. The common thread is that retries are expected, not exceptional, and idempotency is the engineering response to that reality. Build it in from the start. Retrofitting it onto a live system that has already been accumulating subtle duplicates for months is not a project you want.
Get Cloud Architecture Insights
Practical deep dives on infrastructure, security, and scaling. No spam, no fluff.
By subscribing, you agree to receive emails. Unsubscribe anytime.
