Managed relational databases · Chapter 1 of 2
How RDS keeps databases online
High availability is not performance, and understanding how RDS replicates data and manages DNS failover is critical for designing resilient database tiers.
FoundationsBuilds the idea from nothing. No prior AWS assumed.
Multi-AZ replication and standby database instances
How does RDS protect your data from a physical data centre failure?
An application resolves the primary database IP address via the RDS canonical DNS endpoint.
- 01
Availability is not scalability
A classic Multi-AZ DB instance deployment creates a primary database instance and a standby instance in a different availability zone. This standby provides high availability and does not serve reads. Multi-AZ DB clusters are a different architecture with readable instances.
- 02
Synchronous block level replication
Every write to the primary database is replicated synchronously to the standby instance at the storage layer. The database transaction is only committed after the data is successfully written in both locations.
- 03
Automatic failover detection
If the primary database loses power, fails a hardware check, or suffers network loss, RDS detects the failure. It automatically promotes the standby instance to become the new primary database.
Check yourself
What is the primary benefit of a classic Multi-AZ DB instance deployment?
Under the hoodThe same thing from underneath: limits, failure modes, numbers.
DNS endpoints and in flight connection failures
What happens to active database connections during an automatic failover?
An application resolves the primary database IP address via the RDS canonical DNS endpoint.
- 01
The endpoint DNS swap
RDS does not expose the physical IP addresses of your database instances. Instead, you connect to a stable DNS endpoint that RDS updates during a failover to point to the promoted standby instance.
- 02
Handling connection drops
During failover, existing connections can fail and clients must establish connections to the promoted primary. Connection timeouts or resets require suitable reconnection and retry behavior. The original instance is not necessarily permanently terminated.
- 03
The role of client DNS caching
Because failover relies on DNS updates, your application driver must resolve the database endpoint again. If your application or language runtime caches DNS results indefinitely, it will attempt to connect to the dead instance.
The numbers
- Typical failover time
- 60 to 120 secondsThe duration required to detect the failure, update DNS, and complete database recovery on the standby.
- JVM DNS cache default
- Depends on runtime configurationInspect the runtime DNS cache policy. Excessive caching can keep clients connecting to an old address after failover.
- Backup retention range
- 0 to 35 daysFor supported RDS DB instance configurations, 0 disables automated backups. Multi-AZ DB clusters have different retention requirements; verify the selected deployment type.
Check yourself
An RDS failover occurs, and your application fails to reconnect for ten minutes, even though the database is healthy. What is the most likely cause?
Official references & further reading
These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.
Report an error or suggest a clearer explanation →