SimAWSby ShahriarLabs
Search

Retries, idempotency and backpressure · Chapter 1 of 2

How retries multiply system load

Uncoordinated client retries amplify downstream database and service failures, transforming minor network glitches into permanent system outages.

FoundationsBuilds the idea from nothing. No prior AWS assumed.

The danger of immediate retries

Why can retrying failed network calls bring down an unhealthy service?

App ClientsHigh concurrencyAPI GatewayImmediate RetriesDatabaseSlow DownstreamUncontrolled retries from clients can flood a slow downstream service, causing a retry storm, which is mitigated by exponential backoff and jitter.

Clients process requests through a gateway. The backend database suddenly slows down due to high resource usage.

1/5
  1. 01

    Network failures are normal

    In distributed systems, packets are dropped and connections time out. If an application aborts at the first network error, users see a failure. Retrying the call is a simple way to mask these transient glitches and improve reliability.

  2. 02

    Amplifying downstream stress

    If a database slows down, it takes longer to process requests, causing upstream queues to fill. If clients retry their timed-out requests, they send new queries before the database can finish the original ones. The database must now handle double its normal traffic.

  3. 03

    The anatomy of a retry storm

    A retry storm occurs when retries from multiple clients combine to create a self-sustaining overload. The downstream service becomes so busy handling retries and dropping connections that it cannot process any work. The system remains down until clients stop sending traffic.

Check yourself

How do client retries turn a temporary downstream slowdown into a complete outage?

Under the hoodThe same thing from underneath: limits, failure modes, numbers.

Backoff, jitter, and tokens

How do clients coordinate retries to prevent synchronised traffic spikes?

App ClientsHigh concurrencyAPI GatewayImmediate RetriesDatabaseSlow DownstreamUncontrolled retries from clients can flood a slow downstream service, causing a retry storm, which is mitigated by exponential backoff and jitter.

Clients process requests through a gateway. The backend database suddenly slows down due to high resource usage.

1/5
  1. 01

    Exponential backoff delays retries

    Instead of retrying immediately, clients wait longer after each consecutive failure. If the first retry waits 100 milliseconds, the second waits 200, then 400, and then 800. This delay gives the downstream service time to clear its backlog.

  2. 02

    Jitter breaks up synchronised waves

    If a server slows down, hundreds of clients might fail at the same millisecond. If they all use exponential backoff, they will retry together in 100 milliseconds, producing another spike of traffic. Adding random variation, or jitter, spreads these retries over time.

  3. 03

    Token buckets limit retry volume

    AWS SDKs do not retry indefinitely. They use a retry token bucket. Each successful request adds tokens, and each retry consumes them. If the failure rate rises, the bucket empties, and the client stops retrying, returning the error immediately.

  4. 04

    Circuit breakers stop traffic

    A circuit breaker monitors call success rates. If the failure rate crosses a threshold, the circuit trips open, and all subsequent calls fail fast without touching the network. This completely isolates the failing service, allowing it to recover.

The numbers

AWS SDK default retries
3 attempts (2 retries)The default behavior for most services like DynamoDB and S3.
AWS SDK backoff base
100 millisecondsThe initial delay before the first retry, which doubles exponentially on subsequent attempts.
Retry token bucket size
500 tokensA pool shared by clients to prevent a failing system from generating endless retries.
Retry token cost
5 tokens for a retried error, 10 for a timeoutA successful call returns 1 token. The asymmetry is the point: sustained failure drains the bucket and retries stop, so the client backs off from the whole service rather than from one request.

Check yourself

Why is adding jitter (randomness) to backoff logic critical in distributed systems?

In practiceThe judgement call you actually have to make.

Configuring safe client behaviours

How do you configure timeouts and retries to protect your services?

  1. 01

    Limit retry attempts aggressively

    Do not configure infinite retries. Most user-facing applications should retry at most twice before returning an error to the user. If a service is down for 10 seconds, retrying 10 times will not save the request, but it will make it harder for the service to recover.

  2. 02

    Set tight timeouts to fail fast

    Configure short timeouts. If a database call normally takes 30 milliseconds, set the client timeout to 150 milliseconds, not 5 seconds. Waiting 5 seconds holds database connections and memory open, which eventually exhausts resources at your API gateway.

  3. 03

    Identify storms in CloudWatch

    A retry storm is visible in metrics. You will see a sudden rise in 5xx error rates, accompanied by an API request volume that exceeds your peak traffic. The request volume increases because clients are multiplying their calls, even though user traffic remains constant.

Check yourself

What is the primary danger of setting client timeouts to a high value (like 10 seconds) for a fast database call?

Official references & further reading

These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.

Report an error or suggest a clearer explanation →