The idea: replicas give copies, Sentinel gives failover
A single Redis server is a single point of failure. Replication keeps copies of its data on other servers (replicas), but by itself it changes nothing when the master dies: the replicas refuse writes and the application keeps talking to a dead server. Redis Sentinel is a separate process (redis-sentinel, or redis-server --sentinel) that watches the master, decides together with the other Sentinels that it is really down, promotes a replica, reconfigures the rest and tells the clients where the new master is. Sentinels store no data; they are a small distributed system for one decision: who is the master now?
The page runs the smallest setup that makes sense:
| Process | Role at start | On the page |
|---|---|---|
| redis-1 | master | yellow title |
| redis-2, redis-3 | replicas of redis-1 | white title |
| S1, S2, S3 | Sentinels, sentinel monitor mymaster redis-1 6379 2 (quorum 2), down-after-milliseconds 3 ticks | purple boxes at the top |
| app | client that asks a Sentinel for the master and writes to it | top left |
Asynchronous replication, offsets and PSYNC
The master executes a write, answers the client and, in the same event-loop turn, appends the command to the replication stream it sends to every replica. It does not wait for the replicas. Each byte of the stream has a position, the replication offset (on the page, one write = one step), and the history is named by a replication ID. A replica reports its offset back with REPLCONF ACK every second, so the master knows how far behind each replica is.
The master also keeps the recent stream in a ring buffer, the replication backlog (repl-backlog-size, 1 MB by default; 6 writes on the page). A replica that reconnects sends PSYNC <replid> <offset>:
+CONTINUE(partial resync): the master knows that history and still has everything after the offset in its backlog. It sends just the missing part.+FULLRESYNC <replid> <offset>: otherwise. The master makes an RDB snapshot, sends it, and the replica throws away its data set and loads the snapshot. This is how writes a replica had but the master never had disappear.
When a replica is promoted it gets a new replication ID but remembers the old one as replid2 together with the offset where it switched (PSYNC2, Redis 4.0+). The other replicas of the old master share that history, so they can continue partially from the new master. Run Demo: replica link break to see both kinds of resync.
SDOWN and ODOWN: "I think" and "we agree"
Every Sentinel sends PING to the master, the replicas and the other Sentinels once per second. If the master has not answered validly for down-after-milliseconds, that Sentinel marks it SDOWN, subjectively down. One Sentinel's opinion is not enough: it may itself be cut off from the network. So it asks the others, with SENTINEL is-master-down-by-addr, whether they see the master as down too. When at least quorum Sentinels (itself included) say yes, the master is ODOWN, objectively down. Only ODOWN starts a failover. Replicas and Sentinels can be SDOWN, but only a master can be ODOWN.
The leader vote: why quorum is not enough
Only one Sentinel may carry out the failover, or two could promote two different replicas. Before it starts, a Sentinel increases its current epoch and asks the others for their vote in that epoch, sending is-master-down-by-addr again, this time with its own run ID. Each Sentinel votes once per epoch, for the first candidate that asks, and remembers the vote (a Raft-style election). A candidate needs max(quorum, majority of all Sentinels) votes. With 3 Sentinels that is 2.
So the quorum only decides when the master counts as down; the majority decides whether a failover may happen. With 5 Sentinels and quorum 2, two Sentinels can declare ODOWN, but the failover still needs 3 votes. A Sentinel that is in a minority of the network can therefore never fail over on its own (Demo: Sentinels in a minority cannot fail over). If nobody wins an epoch, the candidates wait (failover-timeout) and try again in a later epoch. Real Sentinels randomize their timers so they rarely ask at the same moment; the page staggers them instead (S1 first, then S2, then S3).
Choosing the replica
The leader first drops the replicas it cannot use: not reachable, disconnected from the master for too long, or configured with replica-priority 0. It sorts the rest by
replica-priority: lower is better (default 100);- replication offset: the replica that received the most data wins, so the fewest writes are lost;
- run ID: the lexicographically smaller one, just to make the choice deterministic.
It sends the winner REPLICAOF NO ONE, waits until the replica's INFO reports role:master, then sends REPLICAOF <new master> to the other replicas (parallel-syncs of them at a time).
Config epochs and hello messages
The new configuration ("master of mymaster is redis-2") is stamped with the epoch of the election: the config epoch. Every 2 seconds each Sentinel publishes a hello message with its configuration on the __sentinel__:hello Pub/Sub channel of every instance it monitors; this is also how Sentinels discover each other. A Sentinel that receives a configuration with a higher config epoch adopts it. Because an epoch has one leader, two different configurations can never have the same epoch, and after a partition heals every Sentinel ends up with the newest one.
A master that returns after a failover still believes it is the master. A Sentinel notices from its INFO that it reports role:master while the configuration says replica, and sends it REPLICAOF. Its history is not the new master's, so it does a full resync.
How clients find the master
A Sentinel-aware client is configured with the Sentinel addresses and the name mymaster, not with the master's address. It asks any Sentinel SENTINEL get-master-addr-by-name mymaster, connects to the answer (and should check it with ROLE), and subscribes to the Sentinel's +switch-master event. After a failover it learns the new address from that event, or from a connection error and a new question. An instance that is turned from master into replica closes its client connections, so no client keeps writing to it.
Lost writes and split brain
Asynchronous replication means that an OK is not a promise that the write survives a failover:
- Crash before replication: the master answers
OKand dies before the write left in the replication stream. The replica that is promoted never had it (Demo: acked write lost). - Split brain: a partition leaves the old master on one side with some clients, and a majority of Sentinels with the replicas on the other. The majority promotes a replica; the old master keeps accepting writes from the clients on its side. For a while there are two masters. When the partition heals, the old master becomes a replica and a full resync throws away everything it accepted meanwhile (Demo: split brain).
min-replicas-to-write N with min-replicas-max-lag S makes a master refuse writes (-NOREPLICAS) when fewer than N replicas have acknowledged within S seconds. An isolated master then stops accepting writes after at most S seconds, so at most that window of writes is lost (Demo: split brain with min-replicas-to-write 1). The price: if all replicas are down, the master refuses writes even though it is fine. WAIT numreplicas timeout lets a client wait until its write has reached N replicas, but it does not make Redis strongly consistent: a failover can still pick a replica that did not get it.
Deploying Sentinels
Use at least three Sentinels, on machines that fail independently (not all on the master's host), and an odd number so a majority exists. Put Sentinels where the clients are: a Sentinel should see the network the way the application does. Sentinel rewrites its sentinel.conf with every change (current epoch, configuration, known Sentinels and replicas), so a restarted Sentinel remembers where it was.
What the page leaves out
How the RDB is transferred (disk-based or diskless sync), parallel-syncs, PINGs to replicas and Sentinels (only drawn on Tick) and replica SDOWN, the INFO and hello periods, TILT mode, notification scripts and client reconfiguration scripts, Sentinel auto-discovery, ACLs and TLS, and the announce options needed behind NAT or Docker. A restarted instance keeps its data here (as with AOF); without persistence it would come back empty, which is dangerous for a master that Sentinel has not yet failed over.