Proposed flow for this scenario

  1. Identify availability or read-pressure problem
  2. Select deployment and engine
  3. Route writes and reads correctly
  4. Inspect failure or lag
  5. Verify application recovery

This flow describes a design to evaluate. The local experiment explores one stated mechanism. Its scope appears with the controls; the proposed services are not provisioned.

Decision checkpoints

ChoiceFits whenWatch for
Multi-AZ DB-instance standbyThe primary database needs the selected failover arrangement.Sending reporting queries to a standby that cannot serve them.
Read replicaSupported read work can tolerate its consistency and routing behavior.Assuming an asynchronous copy is always current or already a write endpoint.
Backup and restoreA past safe data state is required after a logical error.Expecting replicated deletion to preserve the deleted record.

Name the deployment, then the requirement

Suppose LetX needs its private database to recover from a primary-instance availability failure. Separately, QuantumSketch’s reports are loading the primary with reads. These requirements can coexist but are not interchangeable.

This comparison addresses an RDS Multi-AZ DB-instance standby and supported ordinary read replicas. Multi-AZ DB clusters and Aurora reader/writer topologies differ. Avoid transferring one deployment’s standby rule to every service carrying a Multi-AZ label.

Reads, writes and failover are different paths

The DB-instance standby is maintained for failover rather than a separate customer-read endpoint. A read replica provides a read path whose freshness depends on replication. Choose which queries may tolerate that difference and how the application selects its endpoint.

Promotion of a replica is a lifecycle decision, not a request that turns every replica into the primary automatically. Plan reconnection, endpoint selection and the consequences for its source relationship. Verify engine-specific behavior before relying on a recovery procedure.

Practice intent and configuration locally

Create a synthetic private database configuration and a selected replica. In the local connectivity diagnostic, distinguish an unreachable endpoint from a reachable read-only endpoint unsuitable for a write intent. Both can prevent a write, but they require different repairs.

Use the failover control to inspect the modeled state transition and retained configuration. This console contains no database data or SQL. It cannot measure lag, prove synchronous transaction durability, connect a real client or establish a production failover duration.

Original comparison worksheet

Create a worksheet with three desired outcomes: retain an available write path after a primary fault, offload suitable report reads and recover a deleted record. Assign each to its actual mechanism and verification evidence. If one label is used for all three, the design still hides an important assumption.

For the synthetic LetX drill, inspect a selected write intent against the replica, then change only the intent to an eligible read. Separately inspect the primary failover state and a configuration snapshot. Explain why these screens answer different questions. A connection-path result and a resource transition are teaching evidence, not proof of recovered customer transactions.

An application can need both a standby arrangement and a reporting replica. That combination should be intentional: define which caller selects which endpoint and how each path behaves when unavailable. Recheck cost and authorization for both rather than treating the second resource as an invisible consequence of the first.

A copy is not a historical recovery point

Question: an accidental application deletion is replicated. Will moving reads to the replica or failing over recover the row? No. Choose a tested backup or another authoritative reconstruction path. Availability and read scaling do not replace logical recovery.

Estimate additional database capacity, storage, recovery and transfer from the chosen engine and Region. The failure-domain experiment below examines assumed surviving capacity; it does not implement database replication or restore. Check the actual deployment before choosing a label as the answer.

Practice the supported console workflow

SimAWS · ShahriarLabs

Predict it. Test it. Change one thing.

Local educational model. No account or cloud charges. Nothing is deployed to AWS.

Available capacity: 160/s. Unserved demand: 0/s.

Two independent zones, equal static capacity and immediate traffic redistribution are assumed. Real failover includes detection, connection recovery, data state and dependencies. This is a failure-domain capacity experiment, not a database implementation.

sim.shahriarlabs.com · Free to explore

Sources and scope

Reviewed against these official references. The model’s supported scope appears alongside its controls.

How we review explanations