SimAWSby ShahriarLabs
Search

Load balancing and scaling · Chapter 3 of 3

What a health check actually proves

A health check determines whether a server is capable of processing work, but poorly designed checks can easily trigger self-inflicted outages.

FoundationsBuilds the idea from nothing. No prior AWS assumed.

The role of automated health monitoring

Why do load balancers and auto scaling groups monitor server health?

Target GroupApp ClientsLoad BalancerTarget 110.0.1.10 (Healthy)Target 210.0.1.20 (Healthy)When a target fails health checks, ELB routes new requests to healthy targets while allowing active in-flight requests to complete during the connection draining window.

The load balancer routes client traffic evenly between Target 1 and Target 2, while checking target health in the background.

1/5
  1. 01

    Avoiding black hole servers

    When a server process crashes or freezes, it stops responding to requests. Without health checks, a load balancer will continue sending traffic to the dead server, resulting in connection errors for users.

  2. 02

    Determining healthy states

    A health check is a periodic request sent by the load balancer to a specific endpoint on your server. If the server returns a successful response code within the timeout window, it is considered healthy.

  3. 03

    Automating replacement

    When an instance fails its health check repeatedly, the load balancer stops routing traffic to it. The Auto Scaling Group detects this unhealthy state and terminates the instance, launching a fresh one to replace it.

Check yourself

What is the primary action a load balancer takes when a server fails its health check?

Under the hoodThe same thing from underneath: limits, failure modes, numbers.

The danger of deep health checks

How can a database outage take down your entire compute fleet?

Target GroupApp ClientsLoad BalancerTarget 110.0.1.10 (Healthy)Target 210.0.1.20 (Healthy)When a target fails health checks, ELB routes new requests to healthy targets while allowing active in-flight requests to complete during the connection draining window.

The load balancer routes client traffic evenly between Target 1 and Target 2, while checking target health in the background.

1/5
  1. 01

    Shallow versus deep health checks

    A shallow health check only verifies that the web server process is running and responding. A deep health check verifies external dependencies, such as database connectivity or third-party APIs.

  2. 02

    The cascading failure risk

    If a deep health check queries a database, a temporary database slowdown will cause all health checks to fail. The load balancer will mark every server as unhealthy, terminating the entire fleet when the database was the actual problem.

  3. 03

    Connection draining and deregistration

    When a server is marked unhealthy, the load balancer enters a deregistration delay. During this window, it stops sending new requests but allows in-flight connections to finish processing before disconnecting.

The numbers

ALB default thresholds
5 checks to become healthy, 2 to become unhealthyAsymmetric on purpose: pull a bad target out fast, put a recovered one back slowly.
NLB default thresholds
3 and 3A network load balancer has no application-level signal to go on, so it is more conservative in both directions.
Default deregistration delay
300 secondsThe time the load balancer waits for active requests to complete before closing connections.
Minimum check interval
5 secondsThe frequency of health check requests, which determines detection speed.

Check yourself

Why are deep health checks that query database tables considered risky?

Official references & further reading

These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.

Report an error or suggest a clearer explanation →