Load balancing and scaling · Chapter 1 of 3
How load balancers route incoming traffic
Clients need a single entry point that survives server failures. A load balancer decouples incoming requests from individual backend instances to keep systems online.
FoundationsBuilds the idea from nothing. No prior AWS assumed.
Why applications need a single entry point
Why do we place a load balancer in front of multiple servers?
Client sends GET /api/v1/users request to ALB listener.
- 01
The entry point problem
Running a single server works until that server fails or receives too much traffic. If you point clients directly at a server IP address, you cannot add more servers or replace a broken one without updating client configuration.
- 02
Decoupling clients from servers
An Elastic Load Balancer provides a single, stable DNS name for clients to connect to. It receives incoming traffic and distributes it across a group of backend instances, allowing you to add or remove servers behind the scenes.
- 03
The necessity of multiple zones
To survive an availability zone outage, your load balancer must be configured in multiple zones. AWS requires you to specify subnets in at least two availability zones when creating an Application Load Balancer.
Check yourself
Why does AWS require you to select subnets in at least two availability zones when creating an Application Load Balancer?
Under the hoodThe same thing from underneath: limits, failure modes, numbers.
Comparing layer seven routing with layer four performance
How does AWS evaluate routing rules and handle connections under the hood?
Client sends GET /api/v1/users request to ALB listener.
- 01
Layer seven versus layer four routing
Application Load Balancers operate at layer seven, inspecting HTTP headers, cookies, and paths to make routing decisions. Network Load Balancers operate at layer four, routing raw TCP and UDP connections with extremely low latency.
- 02
Listeners and target groups
Listeners define the port and protocol the load balancer monitors for incoming connections. Rules on the listener evaluate the request and forward it to a target group, which contains the IP addresses or instance IDs of your servers.
- 03
The mechanics of cross-zone load balancing
By default, Application Load Balancers distribute traffic evenly across all target instances in all enabled zones. This cross-zone routing ensures even distribution but can introduce small latencies and cross-zone data transfer costs.
The numbers
- Default idle timeout
- 60 secondsThe load balancer closes connections if no data is sent or received within this window.
- Cross-zone load balancing
- Always on and free for ALB; off by default and chargeable for NLBAn NLB with cross-zone off sends traffic only to targets in the same AZ as the node that received it, so uneven targets per AZ means uneven load.
- Load Balancer Capacity Units (LCU)
- Based on 4 dimensionsALB pricing evaluates new connections, active connections, processed bytes, and rule evaluations.
Check yourself
An application requires routing incoming traffic based on the HTTP request path /api. Which load balancer type should you choose?
Official references & further reading
These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.
Report an error or suggest a clearer explanation →