Proposed flow for this scenario
- Accepted customer job
- Queue wait
- Worker attempt
- Dependency operation
- Durable result
- Useful completion signal
This flow describes a design to evaluate. The local experiment explores one stated mechanism. Its scope appears with the controls; the proposed services are not provisioned.
Decision checkpoints
| Choice | Fits when | Watch for |
|---|---|---|
| Job outcome and latency | Measure whether accepted work becomes a usable result within its deadline. | Counting a received message or a successful enqueue as a completed customer job. |
| Queue age and backlog | Detect waiting work and compare arrivals with useful completion. | Using queue length alone for jobs with very different sizes or ignoring dead-letter work. |
| Correlated logs and dependency evidence | Follow one job across attempts, permissions, waits and errors. | Logging full documents, credentials or every unique job ID as a metric dimension. |
Healthy CPU is one observation
In a hypothetical LetX incident, customers wait for exports while worker CPU remains low. Several explanations fit: workers could be blocked on a dependency, failing authorization quickly, receiving no jobs, waiting for rate limits or running below the required capacity. CPU alone cannot distinguish them.
Begin with the customer symptom and its time range. Check accepted jobs, useful completions, time to result and oldest unfinished work. Then inspect the particular stage where work stops. A worker that successfully runs but produces no usable result is not a successful customer outcome.
Follow one job and compare the rates
Assign the synthetic export a stable job ID and record stage changes with an attempt number and timestamps. Find evidence of acceptance, delivery, dependency access and durable completion. Keep error categories and request correlation identifiers available for diagnosis.
Compare arrivals with useful completions over the same interval. A backlog can grow because arrivals increased, completion capacity fell, or retries consumed the workers. Queue age adds a waiting-time signal, but its precise service definition matters. Inspect delayed, in-flight and dead-letter work rather than assuming one queue count represents every unfinished job.
Use evidence to choose the repair
If logs show access denied at an object read, inspect the workload identity and required resource action. Increasing worker count does not repair that permission. If attempts spend time waiting for a dependency, inspect its latency, timeout and capacity before raising concurrency. If no consumers receive messages, inspect their configuration and ability to reach the queue.
These are candidate diagnoses, not guarantees derived from low CPU. Change one justified control, then compare useful completion and waiting work. An apparent CPU improvement without customer recovery is insufficient evidence that the incident is fixed.
Alarm on the outcome and test missing data
Choose a threshold, statistic and evaluation window from the workload’s expected behavior and response budget. A single failed job, a sustained error rate and a growing wait are different signals. Also define how an alarm treats missing observations; absent data can mean inactivity or a broken telemetry path.
Correlate the alarm with logs and deployment changes. Test both a known failure and recovery, then inspect the transition and its input periods. A missing-data policy can make an alarm appear healthy despite an absent publisher, so test that case separately.
Keep observation cost and privacy under control
Use stable low-cardinality metric dimensions such as workload and environment. Put per-job identifiers into appropriately protected logs or traces instead of creating a new metric identity for every customer operation. Select retention and sampling deliberately.
Record the evidence needed to repair the job without retaining full uploaded documents or secrets. This platform’s local audit omits message payloads, while explicitly entered local log messages remain inspectable. Learners should use the synthetic project data rather than personal documents.
What you can practice here
The CloudWatch console contains a manual synthetic LetX diagnostic. Its failure and recovery inputs publish custom observations with the same low CPU value, illustrating why outcome metrics matter. These values are authored scenario inputs, not measured CPU or a running database workload. Create an ApplicationErrors Sum alarm, publish failure and recovery in separate completed simulated minutes, and inspect alarm history and logs.
The arithmetic experiment below models arrivals, bounded completion and backlog. It does not emit real SQS metrics or inspect operating-system CPU. Original question: arrivals remain unchanged, useful completion drops and dependency wait increases. Is adding unlimited workers justified? No: first establish whether more concurrent requests would exceed the dependency budget.
Predict it. Test it. Change one thing.
Local educational model. No account or cloud charges. Nothing is deployed to AWS.
Constant rates; no polling, retries, batch effects or service quotas. This models work conservation.
sim.shahriarlabs.com · Free to explore
Sources and scope
Reviewed against these official references. The model’s supported scope appears alongside its controls.
- AWS: available CloudWatch metrics for SQS
- AWS: handling SQS event source errors in Lambda
- AWS: alarm evaluation and missing data