Idempotency Keys Turn Retries into One Logical Operation
A client can lose the response to a successful request. The server may commit a charge, create an order, or enqueue a job, then the connection can fail before the response reaches the caller. From the client’s view, success and failure are now ambiguous.
Retrying is necessary for availability, but an ordinary retry can repeat the side effect. An idempotency key gives the client a way to say that several HTTP attempts represent one logical operation.
The server records the first accepted outcome under that key. A later attempt with the same key can receive the recorded outcome rather than execute the operation again.
The ambiguous outcome is the core failure mode
Consider a request that creates a payment:
client API database
| | |
| POST /payments | |
| key: p-481 | |
|-------------------->| INSERT payment |
| |-------------------->|
| | commit |
| |<--------------------|
| connection lost |
x<--------------------| |The client cannot infer from the broken connection whether the database committed. Sending the same business payload again with a new identity may create a second payment.
A stable idempotency key carries the identity of the logical operation across attempts:
attempt 1: key=p-481 -> execute -> payment 901
attempt 2: key=p-481 -> replay -> payment 901The key does not make the network reliable. It makes repeated delivery safe at the application boundary that enforces the key.
Key scope is part of the API contract
A key must be unique within a defined scope. That scope might be an account, tenant, API credential, endpoint, or operation type.
A raw key such as p-481 is rarely sufficient as the storage identity. A server can derive a composite identity:
tenant_id + operation_name + idempotency_keyScoping prevents unrelated clients from colliding on a common key and lets the service apply retention rules with a precise namespace.
The client should generate the key before the first attempt and reuse it for every retry of that operation. A fresh key on each retry defeats deduplication.
Reservation and side effect need one correctness boundary
A naive sequence can still race:
1. check that key is absent
2. perform side effect
3. store key and resultTwo concurrent requests can both pass step 1 before either stores the record. Both may then execute.
For a side effect stored in the same relational database, a unique constraint and transaction can make reservation and mutation atomic. One pattern inserts an idempotency record with a unique composite key, performs the business mutation, stores the response metadata, and commits them together.
CREATE UNIQUE INDEX idempotency_once
ON idempotency_records (tenant_id, operation, key);If the business effect lives outside that transaction, the design needs another protocol. A local idempotency row cannot atomically cover an arbitrary remote API call. Durable state machines, provider-side idempotency, transactional messaging, or reconciliation may be required depending on the boundary.
The same key must not authorize different requests
A client bug can accidentally reuse a key with a different payload. Returning the first result without checking request identity can hide the error and associate the caller with an unintended operation.
The server can store a canonical request fingerprint beside the key:
key: p-481
request_hash: sha256(canonical relevant fields)
status: completed
response_status: 201
resource_id: 901A later request with the same key and the same fingerprint is a retry. The same key with a different fingerprint should produce a conflict rather than execute or replay silently.
Canonicalization must be defined carefully. Hashing raw JSON bytes treats harmless formatting or object-key order changes as different input. A service can instead fingerprint the semantic fields that define the operation, using a deterministic representation.
In-progress attempts need an explicit state
Concurrent retries may arrive while the first attempt is still running. The idempotency record therefore often needs states such as in_progress, completed, and failed.
A second request that finds in_progress should not start another copy of the operation. Depending on the API, it can wait, return a conflict or retryable response, or expose operation status through a separate resource.
The state transition must also survive process crashes. If a worker dies after reserving a key, a permanent in_progress row can block the operation forever. Recovery may use a lease, attempt epoch, timeout plus reconciliation, or a queue whose ownership rules permit safe takeover.
The recovery rule must match the side effect. Expiring an idempotency row and blindly executing again is unsafe when the earlier attempt may already have committed elsewhere.
Replay semantics should be deliberate
A completed record can store enough information to reproduce the contract of the first successful attempt. That may be the HTTP status, selected response fields, and the identifier of the created resource.
Storing the entire response is simple but can retain sensitive or bulky data. Storing only a resource identifier reduces storage, but reconstructing the response later may reflect state that changed after the original call.
The API should choose which semantics it promises. A replay can mean “return the original response” or “return the current representation of the original resource,” but those are different contracts.
Failures need similar care. Validation errors that occur before any side effect may be safe to repeat without persistence. A failure recorded after partial external work may require durable treatment. A blanket rule that caches every error or no error at all is usually too coarse.
Retention defines the deduplication window
Idempotency records consume storage, so services commonly expire them. Expiration also limits the period during which a repeated key is recognized.
If records are retained for 24 hours, a retry after that window can be treated as a new operation. Clients need a contract that makes this boundary explicit when delayed retries are plausible.
Retention should be based on the longest retry and reconciliation horizon the system intends to support, plus operational margin. Deleting records sooner merely to reduce table size can reopen duplicate side effects.
Cleanup also needs to avoid removing active records. A background job can delete completed records past their retention deadline while preserving active or quarantined states until their recovery path finishes.
Idempotency is narrower than exactly-once execution
An idempotency key does not prove that code executed exactly once. A handler can run more than once while producing one accepted durable effect, or it can crash between external effects that do not share a transaction.
The useful guarantee is narrower: within the documented scope and retention window, repeated attempts carrying the same valid key are resolved as one logical operation at the protected boundary.
That distinction keeps the design concrete. The key identifies intent, durable state records its progress, atomic enforcement prevents concurrent duplication, request fingerprints reject conflicting reuse, and retention defines how long the guarantee remains available.