SimAWSby ShahriarLabs
Search

Failure domains and blast radius · Chapter 1 of 1

How failure domains limit damage

Partitioning resources into isolated logical and physical boundaries prevents localised component failures from causing global system outages.

FoundationsBuilds the idea from nothing. No prior AWS assumed.

The certainty of component failure

Why must we design systems assuming that individual parts will fail?

1. Shared Account2. AZ Isolated3. Cell-BasedApp ServersSharedDatabaseShared DBApp ServersAZ IsolatedDatabaseAZ SplitApp ServersCell 1 (10%)DatabaseCell 1 DBCell-based architectures partition workloads into small, independent units, limiting the blast radius of failures compared to shared regional setups.

Workloads can be structured as: a single shared database, AZ-isolated redundant databases, or cell-based partitions.

1/5
  1. 01

    Hardware breaks constantly at scale

    When running a system on ten servers, hardware failure is rare. When running on ten thousand servers, hard drives, power supplies, and network switches fail every day. Instead of attempting to build unbreakable hardware, cloud architecture focuses on containing the damage when parts fail.

  2. 02

    Defining the blast radius

    The blast radius is the maximum portion of a system that can fail when a single component malfunctions. A bug in a utility function that crashes one compute instance has a tiny blast radius. A bug in a shared database configuration that takes down the entire application has a global blast radius.

  3. 03

    Designing physical walls

    AWS builds its infrastructure with physical containment zones. A server rack shares a power distribution unit. An Availability Zone shares a physical building. A Region shares nothing but a global billing endpoint. Designing for high availability means aligning application components with these physical walls.

Check yourself

What is the main goal of reducing the blast radius in an application architecture?

Under the hoodThe same thing from underneath: limits, failure modes, numbers.

Control planes versus data planes

How does AWS isolate configuration logic from runtime operations?

1. Shared Account2. AZ Isolated3. Cell-BasedApp ServersSharedDatabaseShared DBApp ServersAZ IsolatedDatabaseAZ SplitApp ServersCell 1 (10%)DatabaseCell 1 DBCell-based architectures partition workloads into small, independent units, limiting the blast radius of failures compared to shared regional setups.

Workloads can be structured as: a single shared database, AZ-isolated redundant databases, or cell-based partitions.

1/5
  1. 01

    Separating configuration from execution

    The control plane consists of the APIs and services that create, modify, and delete resources, such as launching an EC2 instance. The data plane consists of the active infrastructure running your workloads, such as routing packets or executing virtual machine instructions.

  2. 02

    Control plane failures do not stop data planes

    Control planes are complex and subject to API rate limits and configuration errors. Data planes are kept simple and isolated. If the AWS EC2 control plane API suffers an outage, you cannot launch new servers, but your running servers continue to process user traffic unaffected.

  3. 03

    Accounts are logical fortresses

    The AWS account is the ultimate administrative and security boundary. Accounts share no resources, quotas, or IAM policies by default. An API rate limit exhaustion or a destructive script execution in a staging account cannot exhaust resources or delete data in your production account.

  4. 04

    Cells isolate user groups

    Cell-based architecture divides an application into multiple identical, isolated instances called cells. By partitioning your users across these cells, a database corruption bug or a malicious attack triggered by one user only affects the single cell hosting that user.

The numbers

Availability Zone separation
Tens of kilometresZones are close enough for low-latency networking but distant enough to avoid shared physical hazards.
Control plane rate limits
Varies by API (often 100-200 RPS)AWS limits API call rates to prevent control plane exhaustion, while data planes handle millions of calls.
Inter-AZ data transfer charge
USD 0.01 per gigabyteData sent across Availability Zones in a VPC is billed in both directions, unlike free intra-AZ traffic.
Multi-Region network latency
70 to 100 millisecondsPhysical distance prevents synchronous database writes across global Regions.

Check yourself

What occurs to running EC2 instances if the EC2 control plane API experiences a total outage?

In practiceThe judgement call you actually have to make.

Balancing isolation and complexity

When should you introduce multi-Region or cell-based architecture?

  1. 01

    Multi-Region is a last resort

    Many teams design multi-Region architectures to improve availability, but they end up increasing complexity and cost. Managing data consistency across global regions is difficult, and routing errors often create outages. Focus on multi-AZ resilience before attempting multi-Region.

  2. 02

    Isolate environments with separate accounts

    Never run production and non-production workloads in the same AWS account. A developer misinterpreting a command or deleting a resource in a test environment can accidentally destroy production resources. Use AWS Organizations to enforce account separation.

  3. 03

    Scale out using cell partitioning

    When database scaling limits prevent further growth, divide your workload into independent cells. Instead of buying a larger database server, create ten small cells, each with its own load balancer and database. This limits the blast radius of a database failure to ten percent.

Check yourself

Why is a cell-based architecture used in large-scale systems?

Official references & further reading

These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.

Report an error or suggest a clearer explanation →