The idea: split the keys, let the client route
A single Redis server keeps all its data in the memory of one machine and runs commands on one thread. Redis Cluster spreads the keys over several masters, each with one or more replicas, and fails over by itself when a master dies. There is no proxy and no central coordinator: the nodes talk to each other on a separate port (the cluster bus, client port + 10000), and the clients know which node has which key.
The page runs the smallest recommended layout: three masters and one replica each.
| Master | Slots | Replica |
|---|---|---|
| A | 0 – 5460 | A1 |
| B | 5461 – 10922 | B1 |
| C | 10923 – 16383 | C1 |
Hash slots and hash tags
The key space is cut into 16384 hash slots. The slot of a key is CRC16(key) mod 16384, and every slot belongs to exactly one master. The slots, not the keys, are what the cluster moves around: adding a node means moving some slots to it. 16384 is small enough that the bitmap of a node's slots (2 KB) fits in every heartbeat packet, and large enough for clusters of up to about 1000 masters. For how fixed slots compare with a consistent-hashing ring and with hash % N, see Consistent Hashing.
If a key contains {...}, only the part between the first { and the next } is hashed. This hash tag puts related keys in the same slot on purpose: user:1, {user:1}:cart and {user:1}:orders all hash user:1 and land in slot 10778. The slots in the key table on the page are the real CRC16 values. Too many keys behind one tag make a hot slot that cannot be split, so tag only what must be used together.
Smart clients, MOVED and ASK
A cluster client asks any node for the slot map (CLUSTER SHARDS, or CLUSTER SLOTS in older clients), caches it, and sends each command straight to the master that owns the key's slot. Nodes never forward commands. A node that gets a command for a slot it does not serve answers with a redirect:
-MOVED 10778 C | -ASK 10778 C | |
|---|---|---|
| Meaning | slot 10778 belongs to C now | slot 10778 is being migrated; this key is already on C |
| The client | updates its slot map (usually refreshes all of it) and resends to C | sends ASKING and then the command to C, once |
| Slot map changed? | yes | no: the next command for the slot goes to the old owner again |
A client that starts with only a seed node learns the map this way (Demo: MOVED, learning the slot map). A client whose node went away gets a connection error instead and fetches a fresh map from another node.
Multi-key commands
MSET, MGET, SUNION, transactions and Lua scripts can only touch keys of one slot; otherwise the node answers -CROSSSLOT. This holds even when the two slots are on the same node, as user:1 (10778) and user:2 (6777) are on B: a slot can move to another node at any time, so the rule is about slots, not nodes. With hash tags, MSET {user:1}:cart … {user:1}:orders … works. While a slot is half migrated, a multi-key command whose keys are split between the two nodes gets -TRYAGAIN.
Resharding: moving a slot while it is used
CLUSTER SETSLOT 10778 IMPORTING Bon the target C, thenCLUSTER SETSLOT 10778 MIGRATING Con the source B.CLUSTER GETKEYSINSLOT 10778on B, thenMIGRATEfor each key: B sends the key to C, C stores it, B deletes it, atomically. While this goes on, B serves the keys it still has and answers-ASKfor the others (and for new keys); C serves the slot only afterASKING.CLUSTER SETSLOT 10778 NODE Con C and on B. C raises itsconfigEpochwithout a vote, so its claim wins, and the other nodes learn it from gossip. From now on B answers-MOVED.
redis-cli --cluster reshard and --cluster rebalance run exactly these commands.
Failure detection: PFAIL and FAIL
Every node PINGs other nodes on the cluster bus all the time, and each PING/PONG carries the sender's epochs, its slots and a gossip section about a few other nodes. A node that gets no PONG from another for cluster-node-timeout flags it PFAIL ("possibly failing"): one node's opinion, which could be its own network problem. The flag travels in gossip. When a node has PFAIL/FAIL reports about the same node from a majority of masters within a time window, it flags it FAIL and broadcasts a FAIL message, which every node accepts at once.
Replica election and epochs
A replica whose master is FAIL starts an election after a delay of 500 ms + a random part + 1 s for each replica of the same master that has more data (its rank), so the most up-to-date replica usually goes first. It increments currentEpoch and sends FAILOVER_AUTH_REQUEST to all masters. A master votes (FAILOVER_AUTH_ACK) at most once per epoch and only if it too sees the master as FAIL. With votes from a majority of masters, the replica becomes a master, sets its configEpoch to the new epoch and claims the slots.
Epochs are the cluster's version numbers. When two nodes claim the same slot, the claim with the higher configEpoch wins, everywhere. That is how every node, including an old master that comes back, converges to the same map: the old master sees a higher configEpoch on its slots (an UPDATE message), turns itself into a replica of the winner and resyncs.
A replica that has been disconnected from its master for too long does not try (cluster-replica-validity-factor): its data would be too old. If a master and all its replicas are down, nobody can take over.
Lost writes and the minority side
Replication is asynchronous: a master answers +OK and sends the write to its replicas afterwards. If it dies in between, the promoted replica does not have the write, and when the old master returns it throws the write away in a full resync (Demo: acked write lost). WAIT numreplicas timeout makes one command wait for replicas, but a failover can still pick a replica without the write, so Redis Cluster does not promise that acknowledged writes survive.
A partition makes this larger. A master cut off with some clients keeps accepting writes until it has missed the majority of masters for cluster-node-timeout; then it sets its cluster state to fail and answers -CLUSTERDOWN. Meanwhile the majority side promotes its replica. Every write the old master accepted in that window is lost when the partition heals (Demo: master in a minority partition). A shorter node timeout shrinks the window but causes failovers for short network hiccups.
Full coverage
With cluster-require-full-coverage yes (the default), a node that sees any slot without a working owner stops serving all commands, even for keys whose master is fine (Demo: master and replica down): the cluster prefers to fail visibly rather than serve only part of the data. Set it to no to keep serving the other slots; cluster-allow-reads-when-down yes keeps reads working while the state is fail.
Cluster vs Sentinel
| Redis Sentinel (page) | Redis Cluster | |
|---|---|---|
| Data | one master holds everything; replicas are copies | keys split over several masters by hash slot |
| Scales writes / memory | no | yes, by adding masters and moving slots |
| Who detects failure and votes | separate Sentinel processes (quorum, then a majority of Sentinels) | the data nodes themselves (a majority of masters) |
| How clients find the master | ask a Sentinel (get-master-addr-by-name) | slot map + MOVED / ASK redirects |
| Multi-key commands, several databases | any keys, SELECT works | one slot only, database 0 only |
| Lost writes on failover | possible (async replication) | possible (async replication) |
What the page leaves out
Replica migration (a spare replica moving to an orphaned master), manual failover (CLUSTER FAILOVER, which loses no writes), reads from replicas (READONLY), pub/sub and sharded pub/sub, how new nodes join (CLUSTER MEET and the handshake), configEpoch collision handling, the exact failure-report window and the random parts of the delays, RDB transfer during a full resync, and ports and addresses in the redirects.