Proposed flow for this scenario
- Client submits request
- Durable job record
- Queued job reference
- Worker processes idempotently
- Result stored
- Client reads job status
This flow describes a design to evaluate. The experiment below models queue balance; it does not provision or run these services.
Decision checkpoints
| Choice | Fits when | Watch for |
|---|---|---|
| Synchronous request | Work fits the entire request deadline and customer interaction benefits from an immediate answer. | Keeping a connection open for slow, variable work and allowing client retries to start it again. |
| Queued background job | Completion may happen later; arrivals vary; work can be represented durably and retried safely. | Promising immediate completion, ignoring backlog age, or treating a successful enqueue as a finished job. |
| Orchestrated workflow | Multiple dependent steps, waits, compensations or callbacks need recorded progress. | Using an orchestration tool as a substitute for defining idempotent side effects and recovery. |
Return acceptance, not a false success
In this synthetic QuantumSketch scenario, a customer asks for a render. The initial request validates its input and records a job. It returns an accepted response with a stable job identifier and a status location. The customer can leave and return without depending on an open socket.
The API must distinguish accepted, queued, running, succeeded and failed states. An enqueue acknowledgement is evidence of delivery to the buffer, not evidence that the final artifact exists. If the initial request fails halfway through, a retried request should not create a second unrelated customer job.
Make the job record and queue agree
Writing a job row and sending a message are two operations. If the process stops between them, a row can exist with no queued work, or a message can arrive before the expected row is usable. Define the repair path before calling this design reliable.
One candidate is a transactional outbox: record the job and pending event atomically in the chosen data store, then publish and retry from that record. Another is a clearly bounded reconciliation process for pending jobs. The right implementation depends on the data store. Neither pattern means a distributed side effect becomes exactly once.
Keep retries safe and acknowledge after success
A standard SQS queue can deliver a message more than once. Give every render a stable job key. Claim and complete work with conditions that distinguish already-completed jobs from unfinished attempts. Write the final artifact in a way that lets a retried worker recognize it.
Visibility timeout gives a consumer time to work; it is not a guarantee that two workers can never overlap. Configure it for the consumer’s expected processing and supported extension behavior. Delete the message only after the durable outcome is recorded. A worker that crashes after completion must be able to recognize that completion on redelivery.
Model burst demand and choose a recovery budget
Start the experiment with arrivals above completion. Backlog growth is the expected result. Then reduce arrivals or increase bounded worker capacity. Explain whether the queue eventually drains and how long a customer might wait.
The experiment uses fixed completion per step. Real job sizes differ, workers start at different times, and dependencies can fail. Track oldest-message age, job latency and useful successes in addition to queue length. A slowly draining queue may still violate the customer deadline even when no job is permanently lost.
A dead-letter queue needs an operator decision
A message that repeatedly fails should stop blocking normal work. A dead-letter path preserves evidence, but a parked message is not a repaired customer job. Retain the job identifier, error category and attempts needed to understand what happened.
Before redriving, identify whether the cause was transient, a bad input, a permission error or an incompatible deployment. Replaying the same malformed input through unchanged code repeats the failure. Record whether the customer job should be retried, corrected, cancelled or marked failed.
Choose the worker independently of the queue
A short independent transform can fit a standard Lambda consumer. A continuous render that exceeds that model may fit an ECS task or another compute option. SQS buffering does not lengthen a function timeout or provide missing memory, CPU or hardware.
Original practice question: every render succeeds individually, but the queue grows during a release-day burst. Would increasing the retry count fix throughput? No. More attempts consume capacity; they do not create it. First compare arrivals with useful completion, then inspect whether the worker ceiling or downstream dependency is the limiting factor.
Predict it. Test it. Change one thing.
Local educational model. No account or cloud charges. Nothing is deployed to AWS.
Constant rates; no polling, retries, batch effects or service quotas. This models work conservation.
sim.shahriarlabs.com · Free to explore
Sources and scope
Reviewed against these official references. The model’s supported scope appears alongside its controls.
- AWS: SQS at-least-once delivery
- AWS: SQS visibility timeout
- AWS: SQS dead-letter queues
- AWS: transactional outbox pattern
- AWS: using Lambda with SQS