Knowing what is happening · Chapter 1 of 1
Why you should alarm on symptoms rather than causes
Alarming on every CPU spike leads to paging fatigue, whereas targeting end-user symptoms ensures you only wake engineers for real outages.
FoundationsBuilds the idea from nothing. No prior AWS assumed.
Designing symptom based alarms
Why does monitoring internal server utilization lead to poor alerts?
CloudWatch evaluates metrics against a threshold over consecutive periods to prevent transient alarms.
- 01
The cause versus symptom dilemma
A server running at 90% CPU is not necessarily broken; it may be processing queries efficiently. If you alarm on CPU usage, you will page engineers for normal operations, causing alert fatigue.
- 02
Focusing on customer impact
Symptom-based monitoring alerts on user-facing metrics, such as HTTP 5xx error rates, response latencies, or failed queue deliveries. These metrics indicate that users are actually experiencing issues.
- 03
The risk of ignoring the symptom
If you only alarm on CPU, a database deadlock that causes HTTP requests to fail immediately with 500 errors will go undetected because CPU usage remains low. Always start alerts at the customer boundary.
Check yourself
Which metric is the best candidate for a high-priority pager alarm?
Under the hoodThe same thing from underneath: limits, failure modes, numbers.
Evaluation periods and metric economics
How do alarm evaluation periods and custom metric charges affect system cost?
CloudWatch evaluates metrics against a threshold over consecutive periods to prevent transient alarms.
- 01
Smoothing spikes with evaluation periods
To prevent transient spikes from triggering alarms, you configure evaluation periods. An alarm should fire only when a metric exceeds the threshold for several consecutive data points.
- 02
The price of custom metrics
AWS charges for every metric path you publish to CloudWatch. If you publish per-customer latency metrics, the cost of custom metrics can easily exceed the cost of the compute resources.
- 03
Log ingestion and storage costs
CloudWatch Logs charges based on the volume of data ingested and stored. Storing verbose debug logs indefinitely in production creates a large recurring bill that yields little value.
The numbers
- Standard metric resolution
- 60 secondsThe default interval for standard CloudWatch metrics, which determines detection speed.
- High-resolution metric interval
- 1 secondCan be configured for custom metrics but increases cost and requires specialized alarms.
- Custom metric price
- $0.30 per metric/monthCharged per metric path, which multiplies quickly if using high cardinality dimensions.
- Log ingestion price
- $0.50 per GBThe cost to ingest log data, plus storage charges of $0.03 per GB per month.
Check yourself
You configure an alarm with a period of 1 minute and datapoints to alarm set to 3 out of 3. What does this configuration prevent?
Official references & further reading
These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.
Report an error or suggest a clearer explanation →