The idea: same cache, four ways to use it
A cache (here Redis) keeps a copy of hot rows in memory, so a read costs about 1 ms instead of a 10 ms database round trip. Putting it in front of the database raises two questions, and the four strategies answer them differently:
- Who loads the cache on a miss? The app (cache-aside) or the cache itself (read-through).
- When is the database written? By the app, around the cache (cache-aside); by the cache, before it answers (write-through); or by the cache, later (write-behind).
The page runs the same script in four lanes at once. Every lane has an app, a cache and a PostgreSQL database with three keys A, B, C. Values are versions (v1, v2, ...), so you can see at a glance when the cache and the database disagree. Under each lane, counters show latency, database load, stale reads (a read that returned something older than the last acknowledged write) and lost writes (acknowledged, but never reached the database). The strip shows one cell per request.
The model: cache round trip 1 ms, database round trip 10 ms, TTL 30 s on every entry, write-behind delay 5 s, refresh-ahead factor 0.5. A crash takes the cache down for 5 s, and it restarts empty.
Cache-aside (lazy loading)
The app does all the work, and the cache is a plain key-value store:
- Read:
GET A. On a hit, done (1 ms). - On a miss (
nil), the app runsSELECTon the database, thenSET A v1 EX 30. That is three trips: 12 ms on the page. - Write:
UPDATEthe database, thenDEL A. It deletes the key instead of updating it, so a race between two writers cannot leave the older value in the cache. The next read reloads it.
Only data that is actually read gets cached, and the cache never talks to the database. So when Redis is down, the app can go around it and read the database directly: slower, but still up (Demo: crash: write-behind loses queued writes, the read during the outage). The costs: every miss is slow, the cache logic is repeated in every app that uses the data, and data can go stale (see TTL). This is the usual pattern with Memcached and Redis.
Write-through (with read-through)
The app talks to the cache only, and the cache is the data access layer. A miss is loaded by the cache itself (read-through, 11 ms on the page). A write goes to the cache, which writes the database synchronously and answers only after the database commits (11 ms).
After a write the cache already holds the new value, so reads right after writes are hits and are never stale for writes that go through the cache (Demo: write, then read). The costs: every write pays the database round trip; data that is written but never read still takes cache memory; and the cache is on the data path. When it is down, both reads and writes fail. A new or restarted cache node is empty until reads warm it up again.
Write-behind (write-back)
Like write-through, but the cache answers a write at once (1 ms), marks the entry dirty and queues it. After the write-delay (5 s on the page), it writes the whole queue to the database as one batch, one row per key. Several writes to the same key in that window coalesce into one row (Demo: write-behind coalesces writes: 4 writes, 2 rows). The database sees fewer, larger writes, and write bursts are absorbed.
The price is durability. Until the flush, the only copy of an acknowledged write is in cache memory. Demo: crash: write-behind loses queued writes writes A and B and then crashes Redis: both writes are gone, and after the restart the lane reads the old v1 from the database. Real write-behind caches reduce the risk by replicating the queue (a backup copy on another node) or by writing it to a persistent log first. The CPU does the same with its write-back caches (see CPU Cache), and there a dirty line is only lost if power fails.
Refresh-ahead
The cache predicts which entries will be read again and reloads them before they expire. The page uses the rule from Oracle Coherence's refresh-ahead-factor: when a read hits an entry older than 0.5 × TTL (15 s), the cache answers from memory at once and reloads the entry from the database in the background, which resets its TTL. A key that keeps being read never misses (Demo: hot key crosses the TTL: at t = 35 s the other lanes miss, this one hits).
When the prediction is wrong, the reload was wasted database work. Demo: refresh-ahead guesses wrong reads A, B and C at t = 20 s, which triggers three reloads, but only A is read again. B and C expire unread, and the lane counts 2 wasted reloads. With many keys and a poor guess, refresh-ahead can put more load on the database than no refresh at all. The page's refresh-ahead lane writes through, like the write-through lane.
TTL, staleness and invalidation
Every strategy is only as fresh as its invalidation. A write that goes around the cache, such as another service, a migration or a manual UPDATE, is invisible to all four (Demo: DB updated behind the cache). Nothing tells the cache, so it serves the old value until the entry expires. The TTL is the bound on that staleness: a short TTL means fresher data and more misses, a long one fewer misses and older data. Refresh-ahead shortens the stale window for hot keys, because it reloads them early.
In write-behind a direct database write is worse. If the key is dirty, the next flush writes the cached, older value over it. Every other writer must go through the same cache, or the database must reject old versions.
Systems that need tighter freshness invalidate from the source: they read the database's change stream (CDC from the MySQL binlog or PostgreSQL logical replication) and delete the affected keys, or they publish invalidation messages.
When the cache dies
- Cold start. A restarted cache is empty, so every key misses again and the database takes the full load until the cache is warm (Demo: crash + restart: cold cache). With many clients this is a thundering herd (cache stampede): thousands of misses for the same hot key at once. Common fixes are request coalescing (one loader per key, the others wait), a lock or a "being loaded" marker in the cache, TTL jitter so keys don't all expire together, and warming the cache before it takes traffic.
- Cache on the data path. With read-through and write-through, an unavailable cache means unavailable data. Cache-aside only loses speed.
- Lost writes. Only write-behind loses acknowledged data. The other three keep only copies in the cache.
For keeping the cache itself available, see Redis Sentinel (failover to a replica) and Redis Cluster (sharding).
Side by side
| Cache-aside | Write-through | Write-behind | Refresh-ahead | |
|---|---|---|---|---|
| read miss | 3 trips, app loads (12 ms) | cache loads (11 ms) | cache loads (11 ms) | cache loads (11 ms), rare for hot keys |
| write latency | DB + DEL (11 ms) | cache + DB (11 ms) | cache only (1 ms) | as write-through |
| DB writes | one per write | one per write | one per key per batch (coalesced) | one per write |
| read right after a write | miss | hit | hit | hit |
| cache down | still works, from the DB | errors | errors | errors |
| cache crash loses data | no | no | yes, the unflushed queue | no |
| what is cached | only what is read | everything written, plus reads | everything written, plus reads | reads, kept warm |
| good for | read-heavy data, a general-purpose default | data read soon after it is written; simple app code | write-heavy counters, metrics, bursts; data you can afford to lose or that is replicated | a small set of hot keys with a known access pattern |
They combine. Cache-aside for reads plus write-through for writes is common, and it also covers new nodes that start empty. Write-behind is usually paired with read-through.
See also LRU Cache, which shows what the cache evicts when memory is full, and CPU Cache, which shows the same write-through vs write-back choice in hardware.
What the page leaves out
Eviction when the cache is full (every key fits here); concurrent clients and the races they cause (two cache-aside readers filling the cache around a writer, stale sets), and the fixes for them (leases, versioned sets); request coalescing and stampede protection; Redis replication and persistence (RDB/AOF), which would let a restarted cache keep its data; write-behind retries when the database is down; negative caching of missing keys; multi-level caches (in-process + Redis); and network failures that are neither "up" nor "down".