Proposed flow for this scenario
- Detect affected zone
- Route to surviving healthy workers
- Reconnect to available data endpoint
- Recover accepted jobs
- Verify useful completion
- Restore intended redundancy
This flow describes a design to evaluate. The local experiment explores one stated mechanism. Its scope appears with the controls; the proposed services are not provisioned.
Decision checkpoints
| Choice | Fits when | Watch for |
|---|---|---|
| Active application capacity across zones | The remaining zones can serve the required workload after one zone fails. | Dividing capacity so thinly that every surviving zone overloads. |
| RDS Multi-AZ DB-instance standby | The selected deployment needs database failover support. | Treating its standby as a readable replica or a backup. |
| Regional recovery strategy | Loss of a whole Region is in the threat model. | Calling a same-Region two-zone design regional disaster recovery. |
Begin with the failure you intend to survive
For the synthetic LetX design, choose one Availability Zone as unavailable while customer export requests continue. List which workers, routes and data dependencies stop being useful. A diagram with duplicated boxes is not enough unless those boxes are placed and operated across the intended failure boundary.
Specify the minimum customer service during the event: new submissions, status reads, completion of accepted jobs or only read-only access. These can have different capacity and recovery needs. A degraded but explicit service target is more useful than an undefined promise that the application stays up.
Check surviving capacity before assuming scale-out saves you
If half the active workers disappear, the other half inherit demand unless admission or routing changes. Determine whether they can sustain it without saturating shared dependencies. Capacity that will be launched later does not serve the request during its startup interval.
The failure-domain experiment below lets you compare assumed remaining capacity with demand after a zone is removed. It does not inject an AWS outage, reserve replacement capacity or simulate every service’s zonal behavior. Choose assumptions deliberately and validate the real system with an appropriate recovery drill.
Data availability has service-specific meaning
For an RDS Multi-AZ DB-instance deployment, the standby exists for failover rather than customer read scaling. Applications still need suitable connection handling and retry behavior. A read replica addresses another requirement and is not automatically the same recovery path.
The local RDS console stores configuration and models selected failover or replica states. It contains no database pages or SQL. A local failover outcome cannot prove transaction preservation, replication timing, reconnection duration or the correctness of an application’s recovery.
Avoid a hidden dependency in the failed zone
Review network egress, credentials retrieval, deployment artifacts and operational access alongside the frontend workers. An application spread over two zones can still depend on a single unsuitable access path. Identify which dependencies are regional, zonal or outside the selected topology rather than relying on labels.
Accepted jobs should have durable state outside replaceable worker memory. After recovery, determine which jobs completed, which need a safe retry and which have an ambiguous external outcome. Recreating workers without resolving those states can turn an availability incident into duplicate customer work.
Write a recovery checklist for the scenario before toggling the failed zone: one status read, one new accepted job and one safe retry of an unfinished job. Assign each check to its required data and worker path. The checklist reveals when surviving capacity exists but a critical dependency still prevents the customer outcome.
The changed requirement and recovery evidence
Question: a bad release deletes an important row, and the deletion reaches every replica. Will moving traffic to another zone recover it? No. Redundancy can reproduce the incorrect change. Restore or reconstruct the intended data from a tested recovery source.
After a drill, verify useful requests, outstanding jobs and restored redundancy, then record observed recovery against the chosen objective. Include extra steady-state capacity, data availability and transfer in the cost model. Multi-AZ improves a defined failure response; it does not eliminate logical errors, backups or a separate regional recovery decision.
Practice the supported console workflow
Inspect target registration across selected zones
Uses simulated resources stored on this device. No website account or AWS credentials required.
ExploreInspect RDS failover versus reader metadata
Uses simulated resources stored on this device. No website account or AWS credentials required.
ExploreInspect recovery signals
Uses simulated resources stored on this device. No website account or AWS credentials required.
Explore
Predict it. Test it. Change one thing.
Local educational model. No account or cloud charges. Nothing is deployed to AWS.
Two independent zones, equal static capacity and immediate traffic redistribution are assumed. Real failover includes detection, connection recovery, data state and dependencies. This is a failure-domain capacity experiment, not a database implementation.
sim.shahriarlabs.com · Free to explore
Sources and scope
Reviewed against these official references. The model’s supported scope appears alongside its controls.
How we review explanations