Amazon EC2 Auto Scaling
EC2 Auto Scaling
You need your server fleet to shrink and grow automatically based on real-time CPU demand or traffic spikes.
Reach for it when
- Scaling a web application tier up during daytime peak usage and down at night to save costs.
- Automatically replacing unhealthy or crashed EC2 instances to maintain a target minimum fleet size.
- Processing a message queue where the number of instances scales based on the backlog size.
Do not reach for it when
- Running monolithic applications that require manual state synchronization and cannot handle sudden termination — use manual scaling instead.
- Hosting a primary relational database where instances cannot be dynamically duplicated and terminated — use RDS scaling instead.
- Scaling serverless workloads that scale natively without server management — use Lambda or ECS Fargate instead.
Alternatives, and how to choose
| Service | Pick it instead when |
|---|---|
| ECS Service Auto Scaling | Choose it when scaling container tasks inside an ECS cluster rather than full EC2 virtual machines. |
| Lambda | Choose it when you want automatic scaling from zero to thousands of concurrent requests without server fleets. |
How you pay
- The model
- Free service, but you pay for the underlying EC2 instances, EBS volumes, and CloudWatch alarms launched.
- The line item that surprises people
- A flapping scaling policy can trigger rapid instance launch and termination cycles, compounding billing costs.
What trips people up
- Cooldown periods must be configured; otherwise, a single spike launches multiple redundant instance groups before the first boot finishes.
- Instances are terminated by default in the AZ with the most instances; this can kill instances with active user sessions.
- Failing to configure a lifecycle hook can lead to instances terminating before they finish processing active jobs.
Verify the live service
This page is a concept reference. Cost models are qualitative; confirm the current offering, Region and pricing before deploying.