Load balancing and scaling · Chapter 2 of 3
Why scaling out takes longer than expected
Adding servers automatically is not instantaneous, and understanding the delay between a metric spike and a running application is key to preventing system outages.
FoundationsBuilds the idea from nothing. No prior AWS assumed.
Understanding horizontal scaling policies
How does an Auto Scaling Group decide when to add or remove servers?
Sudden load spike pushes Instance 1 CPU utilization to 99%.
- 01
Scaling out instead of up
When traffic increases, scaling up by changing to a larger server requires a reboot and has a hard ceiling. Scaling out adds identical servers horizontally, allowing your system to handle larger workloads without downtime.
- 02
Triggering actions with metrics
Scaling policies monitor metrics like CPU utilisation or request count. Target tracking policies adjust the fleet size to maintain a specific target value, whilst scheduled scaling prepares capacity ahead of known events.
- 03
The risk of scaling on the wrong metric
Scaling on a metric that does not reflect actual work, such as memory usage for a memory-leak-prone application, will cause the group to scale unnecessarily. You must choose metrics that correlate directly with throughput or resource exhaustion.
Check yourself
Which scaling policy type is best suited for maintaining a constant average CPU utilisation of 70%?
Under the hoodThe same thing from underneath: limits, failure modes, numbers.
The timeline of a scale-out event
Why does it take several minutes for a new instance to start receiving traffic?
Sudden load spike pushes Instance 1 CPU utilization to 99%.
- 01
The detection delay
CloudWatch metrics are aggregated over time, meaning a sudden spike in traffic is not detected instantly. Depending on your alarm configuration, it can take several minutes of sustained high utilisation to trigger a scaling action.
- 02
Boot time and application warm-up
Once the scaling group requests a new instance, the virtual machine must boot, run initialization scripts, and start the application. During this phase, the instance is running but cannot accept traffic.
- 03
The role of the cooldown period
The cooldown period prevents the group from launching additional instances before the previous ones have started processing traffic. Without a cooldown, a single traffic spike could trigger multiple scaling actions in rapid succession.
The numbers
- Default cooldown period
- 300 secondsPrevents the group from launching or terminating more instances before previous changes take effect.
- Standard metric granularity
- 5 minutesBasic monitoring provides metrics at this interval, which can delay scaling decisions.
- Detailed monitoring granularity
- 1 minuteEnabling detailed monitoring on EC2 instances reduces the metric delay at an additional cost.
- Typical scale-out duration
- 3 to 10 minutesIncludes alarm evaluation, EC2 provisioning, instance booting, and application initialization.
Check yourself
You notice your Auto Scaling Group launches too many instances during a brief traffic spike. What setting should you adjust to prevent this behaviour?
Official references & further reading
These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.
Report an error or suggest a clearer explanation →