Caching, and what it really costs · Chapter 1 of 1
How caching hit rates dictate origin load
A cache trades data consistency for read throughput, but a slight drop in the hit rate can double the traffic hitting your database.
FoundationsBuilds the idea from nothing. No prior AWS assumed.
The non-linear math of hit rates
Why does a small decrease in cache hit rate cause a large increase in database load?
App traffic is routed to ElastiCache first, with cache misses falling through to the persistent backend database.
- 01
Caching is a leverage game
Caches store computed results in fast memory so that subsequent requests do not need to query the database or recompute logic. Because memory reads take less than a millisecond, caching is the standard tool for increasing read throughput.
- 02
The non-linear load math
The impact of a cache is determined by its hit rate. If your hit rate is 95%, only 5% of requests reach your database. If the hit rate drops to 90%, the database receives 10% of requests. This small 5% drop in hit rate doubles the load on your database.
- 03
The correctness trade-off
By placing a cache in front of a database, you accept that your application will serve stale data. A cache turns a database capacity problem into a data consistency problem. You must decide how long old data can be served before it becomes a business error.
Check yourself
If an application’s cache hit rate drops from 99% to 98%, what is the relative increase in traffic hitting the origin database?
Under the hoodThe same thing from underneath: limits, failure modes, numbers.
Invalidation, expiry, and placement
How do caches keep data fresh without overloading the origin database?
App traffic is routed to ElastiCache first, with cache misses falling through to the persistent backend database.
- 01
Time to Live is the simple default
Time to Live (TTL) is the duration a cache holds an item before marking it as expired. TTL is simple because it requires no coordination: the cache automatically discards old entries. The trade-off is that data is guaranteed to be stale for up to the TTL duration.
- 02
Active invalidation is immediate but complex
Active invalidation deletes cached entries whenever the underlying database changes. This keeps the cache fresh, but it requires your application code to know every cache key associated with a write. Missing a key leads to permanent cache inconsistency.
- 03
Thundering herd on expiry
When a popular cache key expires, hundreds of concurrent requests will see a cache miss at the same millisecond. They will all query the database and attempt to write the result back to the cache. This thundering herd can crash your database.
- 04
CloudFront vs ElastiCache placement
CloudFront is an edge cache located close to users, which saves network transit time. ElastiCache is a database cache located inside your VPC, which saves database query time. An optimal architecture uses both layers with different TTL settings.
The numbers
- Origin load multiplier formula
- 1 / (1 - hit_rate)Calculates the division of labor between the cache and the origin database.
- CloudFront edge location count
- 600+ points of presenceDistributes content globally to minimise latency to the client.
- ElastiCache Redis read latency
- Under 1 millisecondProvides extremely fast local memory access within a VPC.
- CloudFront default TTL
- 24 hours (86,400 seconds)The default expiry if no Cache-Control headers are provided by the origin.
Check yourself
Which pattern prevents a thundering herd (cache stampede) when a popular key expires?
In practiceThe judgement call you actually have to make.
Designing caching strategies
When is caching the wrong solution for a slow system?
- 01
Do not use caches to mask slow queries
Adding a cache to fix a slow database query is a common anti-pattern. If the query takes 10 seconds, the first user always waits 10 seconds. When the cache expires, or under a cache stampede, your system will still crash. Fix the index first.
- 02
The cache-aside pattern
In the cache-aside pattern, the application queries the cache first. If the key is missing, it reads from the database, writes the data to the cache, and returns it. This ensures that only requested data is cached, saving memory.
- 03
Write-through for immediate read consistency
If your application needs immediate consistency for reads, use a write-through pattern. The application writes to the cache and the database in the same transaction. This prevents stale reads but adds latency to your write operations.
Check yourself
Why is caching a slow, unindexed database query dangerous?
Official references & further reading
These lessons simplify selected behaviors for learning. Verify current service limits, Region support and production requirements with the official references. Experiments describe their own assumptions.
Report an error or suggest a clearer explanation →