Failure domains and blast radius · Chapter 1 of 1
How failure domains limit damage
Partitioning resources into isolated logical and physical boundaries prevents localised component failures from causing global system outages.
FoundationsBuilds the idea from nothing. No prior AWS assumed.
The certainty of component failure
Why must we design systems assuming that individual parts will fail?
Workloads can be structured as: a single shared database, AZ-isolated redundant databases, or cell-based partitions.
- 01
Hardware breaks constantly at scale
When running a system on ten servers, hardware failure is rare. When running on ten thousand servers, hard drives, power supplies, and network switches fail every day. Instead of attempting to build unbreakable hardware, cloud architecture focuses on containing the damage when parts fail.
- 02
Defining the blast radius
The blast radius is the maximum portion of a system that can fail when a single component malfunctions. A bug in a utility function that crashes one compute instance has a tiny blast radius. A bug in a shared database configuration that takes down the entire application has a global blast radius.
- 03
Designing physical walls
AWS builds its infrastructure with physical containment zones. A server rack shares a power distribution unit. An Availability Zone shares a physical building. A Region shares nothing but a global billing endpoint. Designing for high availability means aligning application components with these physical walls.
Check yourself
What is the main goal of reducing the blast radius in an application architecture?
Under the hoodThe same thing from underneath: limits, failure modes, numbers.
Control planes versus data planes
How does AWS isolate configuration logic from runtime operations?
Workloads can be structured as: a single shared database, AZ-isolated redundant databases, or cell-based partitions.
- 01
Separating configuration from execution
The control plane consists of the APIs and services that create, modify, and delete resources, such as launching an EC2 instance. The data plane consists of the active infrastructure running your workloads, such as routing packets or executing virtual machine instructions.
- 02
Control plane failures do not stop data planes
Control planes are complex and subject to API rate limits and configuration errors. Data planes are kept simple and isolated. If the AWS EC2 control plane API suffers an outage, you cannot launch new servers, but your running servers continue to process user traffic unaffected.
- 03
Accounts are logical fortresses
The AWS account is the ultimate administrative and security boundary. Accounts share no resources, quotas, or IAM policies by default. An API rate limit exhaustion or a destructive script execution in a staging account cannot exhaust resources or delete data in your production account.
- 04
Cells isolate user groups
Cell-based architecture divides an application into multiple identical, isolated instances called cells. By partitioning your users across these cells, a database corruption bug or a malicious attack triggered by one user only affects the single cell hosting that user.
The numbers
- Availability Zone separation
- Tens of kilometresZones are close enough for low-latency networking but distant enough to avoid shared physical hazards.
- Control plane rate limits
- Varies by API (often 100-200 RPS)AWS limits API call rates to prevent control plane exhaustion, while data planes handle millions of calls.
- Inter-AZ data transfer charge
- USD 0.01 per gigabyteData sent across Availability Zones in a VPC is billed in both directions, unlike free intra-AZ traffic.
- Multi-Region network latency
- 70 to 100 millisecondsPhysical distance prevents synchronous database writes across global Regions.
Check yourself
What occurs to running EC2 instances if the EC2 control plane API experiences a total outage?
In practiceThe judgement call you actually have to make.
Balancing isolation and complexity
When should you introduce multi-Region or cell-based architecture?
- 01
Multi-Region is a last resort
Many teams design multi-Region architectures to improve availability, but they end up increasing complexity and cost. Managing data consistency across global regions is difficult, and routing errors often create outages. Focus on multi-AZ resilience before attempting multi-Region.
- 02
Isolate environments with separate accounts
Never run production and non-production workloads in the same AWS account. A developer misinterpreting a command or deleting a resource in a test environment can accidentally destroy production resources. Use AWS Organizations to enforce account separation.
- 03
Scale out using cell partitioning
When database scaling limits prevent further growth, divide your workload into independent cells. Instead of buying a larger database server, create ten small cells, each with its own load balancer and database. This limits the blast radius of a database failure to ten percent.
Check yourself
Why is a cell-based architecture used in large-scale systems?
Official references & further reading
These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.
Report an error or suggest a clearer explanation →