The idea: split the keys, let the client route

A single Redis server keeps all its data in the memory of one machine and runs commands on one thread. Redis Cluster spreads the keys over several masters, each with one or more replicas, and fails over by itself when a master dies. There is no proxy and no central coordinator: the nodes talk to each other on a separate port (the cluster bus, client port + 10000), and the clients know which node has which key.

The page runs the smallest recommended layout: three masters and one replica each.

MasterSlotsReplica
A0 – 5460A1
B5461 – 10922B1
C10923 – 16383C1

Hash slots and hash tags

The key space is cut into 16384 hash slots. The slot of a key is CRC16(key) mod 16384, and every slot belongs to exactly one master. The slots, not the keys, are what the cluster moves around: adding a node means moving some slots to it. 16384 is small enough that the bitmap of a node's slots (2 KB) fits in every heartbeat packet, and large enough for clusters of up to about 1000 masters. For how fixed slots compare with a consistent-hashing ring and with hash % N, see Consistent Hashing.

A bar of slots 0 to 16383 split into master A (0 to 5460), B (5461 to 10922) and C (10923 to 16383), each master with a replica A1, B1, C1. The key user:1 hashes with CRC16 mod 16384 to slot 10778, which belongs to B; keys with the hash tag {user:1} land in the same slot
Every key maps to one of 16384 slots and every slot to one master, so user:1 (slot 10778) always lives on B.

If a key contains {...}, only the part between the first { and the next } is hashed. This hash tag puts related keys in the same slot on purpose: user:1, {user:1}:cart and {user:1}:orders all hash user:1 and land in slot 10778. The slots in the key table on the page are the real CRC16 values. Too many keys behind one tag make a hot slot that cannot be split, so tag only what must be used together.

Smart clients, MOVED and ASK

A cluster client asks any node for the slot map (CLUSTER SHARDS, or CLUSTER SLOTS in older clients), caches it, and sends each command straight to the master that owns the key's slot. Nodes never forward commands. A node that gets a command for a slot it does not serve answers with a redirect:

-MOVED 10778 C-ASK 10778 C
Meaningslot 10778 belongs to C nowslot 10778 is being migrated; this key is already on C
The clientupdates its slot map (usually refreshes all of it) and resends to Csends ASKING and then the command to C, once
Slot map changed?yesno: the next command for the slot goes to the old owner again

A client that starts with only a seed node learns the map this way (Demo: MOVED, learning the slot map). A client whose node went away gets a connection error instead and fetches a fresh map from another node.

Two sequence diagrams. MOVED: the client sends GET user:1 to A, A answers -MOVED 10778 B, the client updates its slot map and resends to B, which returns the value. ASK: during a migration the client sends GET user:1 to B, B answers -ASK 10778 C because the key already moved, the client sends ASKING and then GET to C and gets the value, without changing its map
MOVED means the slot lives elsewhere for good, so the client updates its map; ASK is a one-time detour while a slot is half migrated.

Multi-key commands

MSET, MGET, SUNION, transactions and Lua scripts can only touch keys of one slot; otherwise the node answers -CROSSSLOT. This holds even when the two slots are on the same node, as user:1 (10778) and user:2 (6777) are on B: a slot can move to another node at any time, so the rule is about slots, not nodes. With hash tags, MSET {user:1}:cart … {user:1}:orders … works. While a slot is half migrated, a multi-key command whose keys are split between the two nodes gets -TRYAGAIN.

Resharding: moving a slot while it is used

  1. CLUSTER SETSLOT 10778 IMPORTING B on the target C, then CLUSTER SETSLOT 10778 MIGRATING C on the source B.
  2. CLUSTER GETKEYSINSLOT 10778 on B, then MIGRATE for each key: B sends the key to C, C stores it, B deletes it, atomically. While this goes on, B serves the keys it still has and answers -ASK for the others (and for new keys); C serves the slot only after ASKING.
  3. CLUSTER SETSLOT 10778 NODE C on C and on B. C raises its configEpoch without a vote, so its claim wins, and the other nodes learn it from gossip. From now on B answers -MOVED.

redis-cli --cluster reshard and --cluster rebalance run exactly these commands.

Failure detection: PFAIL and FAIL

Every node PINGs other nodes on the cluster bus all the time, and each PING/PONG carries the sender's epochs, its slots and a gossip section about a few other nodes. A node that gets no PONG from another for cluster-node-timeout flags it PFAIL ("possibly failing"): one node's opinion, which could be its own network problem. The flag travels in gossip. When a node has PFAIL/FAIL reports about the same node from a majority of masters within a time window, it flags it FAIL and broadcasts a FAIL message, which every node accepts at once.

Replica election and epochs

A replica whose master is FAIL starts an election after a delay of 500 ms + a random part + 1 s for each replica of the same master that has more data (its rank), so the most up-to-date replica usually goes first. It increments currentEpoch and sends FAILOVER_AUTH_REQUEST to all masters. A master votes (FAILOVER_AUTH_ACK) at most once per epoch and only if it too sees the master as FAIL. With votes from a majority of masters, the replica becomes a master, sets its configEpoch to the new epoch and claims the slots.

Epochs are the cluster's version numbers. When two nodes claim the same slot, the claim with the higher configEpoch wins, everywhere. That is how every node, including an old master that comes back, converges to the same map: the old master sees a higher configEpoch on its slots (an UPDATE message), turns itself into a replica of the winner and resyncs.

A replica that has been disconnected from its master for too long does not try (cluster-replica-validity-factor): its data would be too old. If a master and all its replicas are down, nobody can take over.

Lost writes and the minority side

Replication is asynchronous: a master answers +OK and sends the write to its replicas afterwards. If it dies in between, the promoted replica does not have the write, and when the old master returns it throws the write away in a full resync (Demo: acked write lost). WAIT numreplicas timeout makes one command wait for replicas, but a failover can still pick a replica without the write, so Redis Cluster does not promise that acknowledged writes survive.

A partition makes this larger. A master cut off with some clients keeps accepting writes until it has missed the majority of masters for cluster-node-timeout; then it sets its cluster state to fail and answers -CLUSTERDOWN. Meanwhile the majority side promotes its replica. Every write the old master accepted in that window is lost when the partition heals (Demo: master in a minority partition). A shorter node timeout shrinks the window but causes failovers for short network hiccups.

Full coverage

With cluster-require-full-coverage yes (the default), a node that sees any slot without a working owner stops serving all commands, even for keys whose master is fine (Demo: master and replica down): the cluster prefers to fail visibly rather than serve only part of the data. Set it to no to keep serving the other slots; cluster-allow-reads-when-down yes keeps reads working while the state is fail.

Cluster vs Sentinel

Redis Sentinel (page)Redis Cluster
Dataone master holds everything; replicas are copieskeys split over several masters by hash slot
Scales writes / memorynoyes, by adding masters and moving slots
Who detects failure and votesseparate Sentinel processes (quorum, then a majority of Sentinels)the data nodes themselves (a majority of masters)
How clients find the masterask a Sentinel (get-master-addr-by-name)slot map + MOVED / ASK redirects
Multi-key commands, several databasesany keys, SELECT worksone slot only, database 0 only
Lost writes on failoverpossible (async replication)possible (async replication)

What the page leaves out

Replica migration (a spare replica moving to an orphaned master), manual failover (CLUSTER FAILOVER, which loses no writes), reads from replicas (READONLY), pub/sub and sharded pub/sub, how new nodes join (CLUSTER MEET and the handshake), configEpoch collision handling, the exact failure-report window and the random parts of the delays, RDB transfer during a full resync, and ports and addresses in the redirects.