The idea: one copy is easy, three copies disagree for a while
A system keeps copies of its data on several replicas, so that it survives a machine failing and serves readers close to them. Copying takes time. For a while after a write, some replicas have it and some do not. A consistency model is the promise the system makes about what a reader can see during that while.
The page runs the same client operations on the same three replicas under six levels. R1, R2 and R3 hold three keys: x (a profile), q (a question) and a (an answer). Alice, Bob and Carol normally use their home replica (R1, R2, R3). An operation can also go to another replica, as when a load balancer sends a request elsewhere. A message between a client and a replica takes 1 tick. R1–R2 takes 2 ticks, R2–R3 takes 1 tick, and the slow link R1–R3 takes 8 ticks. The slow link is where all the trouble comes from.
Every reply lands in the history: one lane per client, and a bar from request to reply. A checker tests each reply against five guarantees and turns the badge red at the first reply that breaks one. "(promised)" marks the guarantees the current level is supposed to keep. Changing the level starts the store over.
Weak consistency
After a write, a read may or may not see it, and nothing promises it ever will. The replica that takes the write applies it, answers OK and sends it once to the others. If that message is lost, the copy stays wrong for good. That is fine for data that is soon overwritten or not worth repairing: a cache (memcached), a live video or voice call, the position of a player in a game. (Demo: lost message on the Weak tab: R3 keeps x = v0 for good, and the Converged badge turns red.)
Eventual consistency
If writes stop, every replica ends up with the same value. Reads in the meantime can be stale. The write path is the same as with weak consistency. What the level adds is a repair path. Here that path is anti-entropy: every 10 ticks R2 and R3 exchange a digest of the writes they know, and R1 and R2 do the same 5 ticks later. The side that is missing something pulls it. Real systems use Merkle trees for the digest (Cassandra, Riak, DynamoDB's ancestor Dynamo), plus read repair and hinted handoff. DNS, email and asynchronous database replicas are eventually consistent too.
When two replicas get different writes to the same key, they need a rule to agree on a winner. This page uses last write wins (LWW) on the timestamp (tick, replica).
Demo: stale read (eventual): Alice's write is acknowledged at t2. Carol reads R3 at t2 and gets v0 at t4. The value only reaches R3 at t9. Demo: lost message repaired: anti-entropy between R2 and R3 at t10 repairs R3 at t13.
Strong consistency (linearizability)
Once a write is acknowledged, every read that starts later sees it (or something newer), whichever replica it asks. The system behaves as if there were one copy. The usual way to build it is a single leader and a replicated log (Raft, Paxos). Here the leader is R2.
- Writes: the leader appends the write and sends
AppendEntriesto R1 and R3. It commits the entry once 2 of 3 replicas have it, then answers. - Reads: the leader first confirms, with a heartbeat to a majority, that it is still the leader (ReadIndex), then answers from its own state. A follower that gets a request forwards it.
The price is latency. Alice's write takes 8 ticks instead of 2, and Carol's read takes 6 instead of 2. (Demo: same ops on the Strong tab.) Examples are etcd, ZooKeeper, Consul and Spanner, and also a single-primary SQL database read from the primary.
Session guarantees: read-your-writes and monotonic reads
Eventual consistency is hard to use: a user saves their profile, reloads the page, and the change is gone. The session guarantees (Terry et al., 1994) remove one anomaly each. They don't make the whole system strong, only the view of one client.
- Read-your-writes: a client always sees its own writes. The OK carries the write's version. The client sends that version as a session token with its next read, and a replica that is behind holds the read until it catches up. Another option is to send the read to the replica that took the write. (Demos: my update vanished (eventual) → Alice reads
v0at t4; read-your-writes → R3 holds her read until t9 and answersv1at t10.) - Monotonic reads: a client never sees an older value after a newer one. The token is the newest version the client has read. Sticking each user to one replica also gives it. (Demos: time goes backwards (eventual) → Bob reads
v1from R2, thenv0from R3; monotonic reads → R3 holds the read.) - The two are independent. Demo: not monotonic on the Read-your-writes tab runs the monotonic-reads script under read-your-writes. Bob never wrote
x, so his token is empty and time still goes backwards. - Two more guarantees, text only: monotonic writes (a client's writes are applied in the order it made them) and writes-follow-reads (a write is ordered after the writes its client had read).
Causal consistency
If Bob reads Alice's question and then writes an answer, the answer depends on the question. Causal consistency promises that nobody sees the answer without the question. It includes all four session guarantees, and it can still be served by any single replica, even during a network partition. That is why it is the strongest model an always-available system can offer.
Each write carries its dependencies: everything its client had read or written. Real systems use vector clocks or dependency lists (COPS, MongoDB causal sessions, Cosmos DB "session", Bolt-on causal consistency). A replica that gets a write before its dependencies keeps it in a pending buffer, where reads can't see it.
Demo: answer before question (eventual): Alice writes q = v1 at R1. Bob reads it at R2 and writes a = v2. The answer takes the fast path to R3 (t6), while the question is still on the slow link until t9. Carol reads R3 at t7 and gets a = v2, q = ∅.
Demo: causal: R3 keeps a pending, and Carol's first read returns neither value. That is consistent, but it isn't linearizable: the question had already been acknowledged, so the Linearizable badge still turns red. Her second read gets both.
The levels side by side
| level | converges | read-your-writes | monotonic reads | causal | linearizable | Alice's write / Carol's remote read | examples |
|---|---|---|---|---|---|---|---|
| weak | ✗ | ✗ | ✗ | ✗ | ✗ | 2 / 2 ticks | memcached, VoIP, game state |
| eventual | ✓ | ✗ | ✗ | ✗ | ✗ | 2 / 2 | DNS, Cassandra ONE, DynamoDB default reads, async replicas |
| read-your-writes | ✓ | ✓ | ✗ | ✗ | ✗ | 2 / 2, or until the replica catches up | "read from the primary after a write", LSN / GTID tokens |
| monotonic reads | ✓ | ✗ | ✓ | ✗ | ✗ | 2 / 2, or until caught up | sticky sessions, Cosmos DB "consistent prefix" (a related guarantee) |
| causal | ✓ | ✓ | ✓ | ✓ | ✗ | 2 / 2, writes may wait in pending | MongoDB causal sessions, COPS, Cosmos DB "session" |
| strong | ✓ | ✓ | ✓ | ✓ | ✓ | 8 / 6 | etcd, ZooKeeper, Spanner, DynamoDB ConsistentRead |
Azure Cosmos DB exposes five of these as a setting: strong, bounded staleness, session, consistent prefix and eventual. Cassandra makes the choice per query with ONE / QUORUM / ALL. With R + W > N, a read quorum overlaps the last write quorum, which gives read-your-writes and more, but on its own still not linearizability.
Concurrent writes
Below strong, two replicas can take writes to the same key at the same time. Demo: concurrent writes, LWW (eventual): Alice writes v1 at R1 and Carol writes v2 at R3, both at t0, and both get OK at t2. Both timestamps are t1, so the tie goes to the higher replica id and v2 wins everywhere. Alice's acknowledged write is gone, and nobody is told. Causal consistency doesn't help here, because neither write depends on the other.
The fixes are to keep both versions as siblings (vector clocks, as in Riak and Dynamo) and let the application merge them, or to use data types that merge by themselves (CRDTs). The other option is to go strong. Demo: concurrent writes on the Strong tab: the leader orders the two writes (Carol's arrives first, at t2), both are committed, and the final value v1 is the last one in the order everyone saw.
The price
Every step up costs waiting. Strong needs a round trip to the leader and a majority. The session levels sometimes hold a read until a replica catches up. PACELC names this: even when there is no Partition, you trade Latency against Consistency. During a partition, the trade is availability against consistency (the CAP theorem). A strong system refuses requests on the minority side. The weaker levels, up to causal, keep answering. See Primary–Replica Replication for lag, failover and lost writes on a single primary.
Consistency is not isolation
These models are about copies: what a reader sees on a replica. Isolation levels (read committed, snapshot, serializable) are about transactions: what one transaction sees of another's unfinished work, even on a single machine. See MVCC and Isolation Levels. Strict serializability is both: serializable transactions that also respect real time.
What the page leaves out
Network partitions and leader failover (see Raft). Sequential consistency and bounded staleness as separate levels. Quorum reads and writes (R + W > N). Clock skew: the page's timestamps come from one perfect clock, which makes LWW look better than it is. Vector clocks: dependencies are drawn as a list of write names. Merkle trees: the digest is the list itself.
Simplifications:
- Each replica sends a write straight to the other two, once.
- Anti-entropy runs only between neighbours (R2–R3 at t = 10, 20, …; R1–R2 at t = 5, 15, …), and only when their digests differ.
- Session tokens name write versions. Real systems use LSNs, GTIDs or cluster times.
- The strong leader never fails, and its term is always 1.
- An anomaly is reported once per badge: the first one.