Proposed flow for this scenario
- Stable customer operation key
- Validate request intent
- Conditional job claim
- Produce durable result
- Record completion
- Acknowledge delivery
This flow describes a design to evaluate. The local experiment explores one stated mechanism. Its scope appears with the controls; the proposed services are not provisioned.
Decision checkpoints
| Choice | Fits when | Watch for |
|---|---|---|
| Stable operation key + conditional completion | A retried request represents the same intent and the durable store can enforce the condition. | Creating a fresh key per delivery, or sharing a key across unrelated requests. |
| Transactional outbox | A committed application update must eventually publish an event. | Assuming the publisher cannot repeat an event after an ambiguous acknowledgement. |
| Provider idempotency + reconciliation | An external action can complete before the worker records its outcome. | Blindly repeating an ambiguous charge or notification without checking the provider outcome. |
The duplicate starts at a crash boundary
Consider a hypothetical LetX export job, LetX-export-42. The worker writes an artifact, then stops before recording completion or deleting its queue message. A later delivery still describes the same customer request. If the next worker treats that delivery as a new request, it may create a second artifact or send another notification. These are synthetic scenarios, not measurements from a production LetX deployment.
Separate the customer operation identifier from the delivery identifier and attempt counter. Record the intended inputs with the operation key. A repeated key with different inputs should produce a conflict or another explicit decision, rather than silently reusing an unrelated result.
A claim is useful only when recovery is defined
Use a conditional write to distinguish an unfinished job from one already completed. A competing worker must inspect the existing record rather than overwrite it. For replaceable workers, a claim needs an expiry or another recovery mechanism; otherwise a crash can leave the job stuck forever.
Expiry also permits overlap. A slow original worker may still finish after a replacement starts. Where the target supports it, an attempt generation or fencing condition can prevent an old worker from committing over a newer attempt. A lock in worker memory cannot protect state after that process disappears.
Close the gap between effects and completion
For a LetX artifact, a stable result location and durable completion record help a retry recognize existing work. Specify which result is authoritative and how incomplete output is detected. Do not mark a job complete before its usable artifact exists.
For a third-party action, a crash after the provider succeeds but before the local record updates leaves an ambiguous outcome. Reuse the provider-supported idempotency key, inspect the provider result when available, and reconcile uncertain records. If the provider offers neither mechanism, describe the remaining duplicate risk honestly.
Publishing and retries still need boundaries
An outbox can store an application change and pending event in one transaction. A separate publisher sends committed events and records progress. That solves the local dual-write gap; consumers still need to handle repeated publication.
Use bounded retry attempts with backoff and jitter where appropriate. Classify failures: repeating an unchanged invalid request is different from retrying a temporary dependency failure. A parked dead-letter message needs diagnosis and a recovery decision before replay. More attempts can consume capacity while producing no additional useful results.
Test the failure you claim to handle
In a real implementation, inject interruption before the effect, after the effect, before recording completion and before acknowledging delivery. Check the durable result and external provider outcome, not merely the count of handler invocations. Repeat the same intent concurrently and verify the conflict path.
The experiment below models retry amplification and bounded capacity. It does not execute conditional database writes, provider reconciliation, distributed locks or actual queue consumers. In the console, the standard SQS model lets a visibility lease expire and exposes repeated receives; it does not emulate distributed duplicate injection.
Original practice question
QuantumSketch now sends a notification through a provider that accepts no idempotency key. The worker records “sending,” the provider accepts the request, and the worker crashes. Can a new worker prove that resending is safe from that local record alone?
No. The record captures an attempt, not the provider outcome. Look for an authoritative provider lookup or reconciliation path, then define what to do if uncertainty remains. Choosing FIFO delivery does not automatically remove a crash boundary around an external effect.
Predict it. Test it. Change one thing.
Local educational model. No account or cloud charges. Nothing is deployed to AWS.
Nested retries multiply attempts in the worst case. Backoff and jitter change timing; they do not reduce this configured upper bound. Real failures need an end-to-end retry budget.
sim.shahriarlabs.com · Free to explore
Sources and scope
Reviewed against these official references. The model’s supported scope appears alongside its controls.
- AWS Builders’ Library: making retries safe with idempotent APIs
- AWS: transactional outbox pattern
- AWS: SQS visibility timeout