SimAWSby ShahriarLabs
Search

Load balancing and scaling · Chapter 2 of 3

Why scaling out takes longer than expected

Adding servers automatically is not instantaneous, and understanding the delay between a metric spike and a running application is key to preventing system outages.

FoundationsBuilds the idea from nothing. No prior AWS assumed.

Understanding horizontal scaling policies

How does an Auto Scaling Group decide when to add or remove servers?

Auto Scaling Group (Min: 1, Max: 3)CloudWatchCPU > 80% AlarmInstance 1CPU 99% BusyCPU 99%Instance 2Launching / InitAuto Scaling involves metric delay, alarm evaluation, EC2 provisioning, and user boot scripts before handling load.

Sudden load spike pushes Instance 1 CPU utilization to 99%.

1/5
  1. 01

    Scaling out instead of up

    When traffic increases, scaling up by changing to a larger server requires a reboot and has a hard ceiling. Scaling out adds identical servers horizontally, allowing your system to handle larger workloads without downtime.

  2. 02

    Triggering actions with metrics

    Scaling policies monitor metrics like CPU utilisation or request count. Target tracking policies adjust the fleet size to maintain a specific target value, whilst scheduled scaling prepares capacity ahead of known events.

  3. 03

    The risk of scaling on the wrong metric

    Scaling on a metric that does not reflect actual work, such as memory usage for a memory-leak-prone application, will cause the group to scale unnecessarily. You must choose metrics that correlate directly with throughput or resource exhaustion.

Check yourself

Which scaling policy type is best suited for maintaining a constant average CPU utilisation of 70%?

Under the hoodThe same thing from underneath: limits, failure modes, numbers.

The timeline of a scale-out event

Why does it take several minutes for a new instance to start receiving traffic?

Auto Scaling Group (Min: 1, Max: 3)CloudWatchCPU > 80% AlarmInstance 1CPU 99% BusyCPU 99%Instance 2Launching / InitAuto Scaling involves metric delay, alarm evaluation, EC2 provisioning, and user boot scripts before handling load.

Sudden load spike pushes Instance 1 CPU utilization to 99%.

1/5
  1. 01

    The detection delay

    CloudWatch metrics are aggregated over time, meaning a sudden spike in traffic is not detected instantly. Depending on your alarm configuration, it can take several minutes of sustained high utilisation to trigger a scaling action.

  2. 02

    Boot time and application warm-up

    Once the scaling group requests a new instance, the virtual machine must boot, run initialization scripts, and start the application. During this phase, the instance is running but cannot accept traffic.

  3. 03

    The role of the cooldown period

    The cooldown period prevents the group from launching additional instances before the previous ones have started processing traffic. Without a cooldown, a single traffic spike could trigger multiple scaling actions in rapid succession.

The numbers

Default cooldown period
300 secondsPrevents the group from launching or terminating more instances before previous changes take effect.
Standard metric granularity
5 minutesBasic monitoring provides metrics at this interval, which can delay scaling decisions.
Detailed monitoring granularity
1 minuteEnabling detailed monitoring on EC2 instances reduces the metric delay at an additional cost.
Typical scale-out duration
3 to 10 minutesIncludes alarm evaluation, EC2 provisioning, instance booting, and application initialization.

Check yourself

You notice your Auto Scaling Group launches too many instances during a brief traffic spike. What setting should you adjust to prevent this behaviour?

Official references & further reading

These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.

Report an error or suggest a clearer explanation →