SimAWSby ShahriarLabs
Search

Caching, and what it really costs · Chapter 1 of 1

How caching hit rates dictate origin load

A cache trades data consistency for read throughput, but a slight drop in the hit rate can double the traffic hitting your database.

FoundationsBuilds the idea from nothing. No prior AWS assumed.

The non-linear math of hit rates

Why does a small decrease in cache hit rate cause a large increase in database load?

MissApp Traffic1,000 rpsElastiCacheIn-Memory DBDatabaseDisk StorageHigh cache hit rates protect databases from traffic spikes, whereas minor decreases in hit rate disproportionately multiply database load.

App traffic is routed to ElastiCache first, with cache misses falling through to the persistent backend database.

1/5
  1. 01

    Caching is a leverage game

    Caches store computed results in fast memory so that subsequent requests do not need to query the database or recompute logic. Because memory reads take less than a millisecond, caching is the standard tool for increasing read throughput.

  2. 02

    The non-linear load math

    The impact of a cache is determined by its hit rate. If your hit rate is 95%, only 5% of requests reach your database. If the hit rate drops to 90%, the database receives 10% of requests. This small 5% drop in hit rate doubles the load on your database.

  3. 03

    The correctness trade-off

    By placing a cache in front of a database, you accept that your application will serve stale data. A cache turns a database capacity problem into a data consistency problem. You must decide how long old data can be served before it becomes a business error.

Check yourself

If an application’s cache hit rate drops from 99% to 98%, what is the relative increase in traffic hitting the origin database?

Under the hoodThe same thing from underneath: limits, failure modes, numbers.

Invalidation, expiry, and placement

How do caches keep data fresh without overloading the origin database?

MissApp Traffic1,000 rpsElastiCacheIn-Memory DBDatabaseDisk StorageHigh cache hit rates protect databases from traffic spikes, whereas minor decreases in hit rate disproportionately multiply database load.

App traffic is routed to ElastiCache first, with cache misses falling through to the persistent backend database.

1/5
  1. 01

    Time to Live is the simple default

    Time to Live (TTL) is the duration a cache holds an item before marking it as expired. TTL is simple because it requires no coordination: the cache automatically discards old entries. The trade-off is that data is guaranteed to be stale for up to the TTL duration.

  2. 02

    Active invalidation is immediate but complex

    Active invalidation deletes cached entries whenever the underlying database changes. This keeps the cache fresh, but it requires your application code to know every cache key associated with a write. Missing a key leads to permanent cache inconsistency.

  3. 03

    Thundering herd on expiry

    When a popular cache key expires, hundreds of concurrent requests will see a cache miss at the same millisecond. They will all query the database and attempt to write the result back to the cache. This thundering herd can crash your database.

  4. 04

    CloudFront vs ElastiCache placement

    CloudFront is an edge cache located close to users, which saves network transit time. ElastiCache is a database cache located inside your VPC, which saves database query time. An optimal architecture uses both layers with different TTL settings.

The numbers

Origin load multiplier formula
1 / (1 - hit_rate)Calculates the division of labor between the cache and the origin database.
CloudFront edge location count
600+ points of presenceDistributes content globally to minimise latency to the client.
ElastiCache Redis read latency
Under 1 millisecondProvides extremely fast local memory access within a VPC.
CloudFront default TTL
24 hours (86,400 seconds)The default expiry if no Cache-Control headers are provided by the origin.

Check yourself

Which pattern prevents a thundering herd (cache stampede) when a popular key expires?

In practiceThe judgement call you actually have to make.

Designing caching strategies

When is caching the wrong solution for a slow system?

  1. 01

    Do not use caches to mask slow queries

    Adding a cache to fix a slow database query is a common anti-pattern. If the query takes 10 seconds, the first user always waits 10 seconds. When the cache expires, or under a cache stampede, your system will still crash. Fix the index first.

  2. 02

    The cache-aside pattern

    In the cache-aside pattern, the application queries the cache first. If the key is missing, it reads from the database, writes the data to the cache, and returns it. This ensures that only requested data is cached, saving memory.

  3. 03

    Write-through for immediate read consistency

    If your application needs immediate consistency for reads, use a write-through pattern. The application writes to the cache and the database in the same transaction. This prevents stale reads but adds latency to your write operations.

Check yourself

Why is caching a slow, unindexed database query dangerous?

Official references & further reading

These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.

Report an error or suggest a clearer explanation →