SimAWSby ShahriarLabs
Search

Knowing what is happening · Chapter 1 of 1

Why you should alarm on symptoms rather than causes

Alarming on every CPU spike leads to paging fatigue, whereas targeting end-user symptoms ensures you only wake engineers for real outages.

FoundationsBuilds the idea from nothing. No prior AWS assumed.

Designing symptom based alarms

Why does monitoring internal server utilization lead to poor alerts?

CloudWatchAlarm (3 Periods)CPU: 45%SNS Topicoperators-notifyOperatorEmailCloudWatch Alarms require a metric to breach a threshold for a consecutive number of evaluation periods before changing state and triggering SNS notifications.

CloudWatch evaluates metrics against a threshold over consecutive periods to prevent transient alarms.

1/7
  1. 01

    The cause versus symptom dilemma

    A server running at 90% CPU is not necessarily broken; it may be processing queries efficiently. If you alarm on CPU usage, you will page engineers for normal operations, causing alert fatigue.

  2. 02

    Focusing on customer impact

    Symptom-based monitoring alerts on user-facing metrics, such as HTTP 5xx error rates, response latencies, or failed queue deliveries. These metrics indicate that users are actually experiencing issues.

  3. 03

    The risk of ignoring the symptom

    If you only alarm on CPU, a database deadlock that causes HTTP requests to fail immediately with 500 errors will go undetected because CPU usage remains low. Always start alerts at the customer boundary.

Check yourself

Which metric is the best candidate for a high-priority pager alarm?

Under the hoodThe same thing from underneath: limits, failure modes, numbers.

Evaluation periods and metric economics

How do alarm evaluation periods and custom metric charges affect system cost?

CloudWatchAlarm (3 Periods)CPU: 45%SNS Topicoperators-notifyOperatorEmailCloudWatch Alarms require a metric to breach a threshold for a consecutive number of evaluation periods before changing state and triggering SNS notifications.

CloudWatch evaluates metrics against a threshold over consecutive periods to prevent transient alarms.

1/7
  1. 01

    Smoothing spikes with evaluation periods

    To prevent transient spikes from triggering alarms, you configure evaluation periods. An alarm should fire only when a metric exceeds the threshold for several consecutive data points.

  2. 02

    The price of custom metrics

    AWS charges for every metric path you publish to CloudWatch. If you publish per-customer latency metrics, the cost of custom metrics can easily exceed the cost of the compute resources.

  3. 03

    Log ingestion and storage costs

    CloudWatch Logs charges based on the volume of data ingested and stored. Storing verbose debug logs indefinitely in production creates a large recurring bill that yields little value.

The numbers

Standard metric resolution
60 secondsThe default interval for standard CloudWatch metrics, which determines detection speed.
High-resolution metric interval
1 secondCan be configured for custom metrics but increases cost and requires specialized alarms.
Custom metric price
$0.30 per metric/monthCharged per metric path, which multiplies quickly if using high cardinality dimensions.
Log ingestion price
$0.50 per GBThe cost to ingest log data, plus storage charges of $0.03 per GB per month.

Check yourself

You configure an alarm with a period of 1 minute and datapoints to alarm set to 3 out of 3. What does this configuration prevent?

Official references & further reading

These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.

Report an error or suggest a clearer explanation →