The problem: one order, three databases
An order needs three things to happen: Payment charges the customer, Inventory reserves the item and Shipping books a courier. Each service has its own database, so no single local transaction can cover all three. Either all three must happen or none of them, even when a service says no or a machine crashes halfway. The page runs the same order through the two classic answers at the same time, on one clock: two-phase commit on the left, a saga on the right.
Every message takes one tick. The timeline under each side has one cell per tick for each database, showing what it holds for the order at the end of that tick. The stats line counts messages, lock-ticks (ticks a participant holds a lock for the order, summed over participants) and ticks in which a half-done order was visible: one database already shows the order's change while another shows the state without it, and nothing stops a reader from seeing both.
Two-phase commit
A coordinator (transaction manager) drives the participants (resource managers):
- Prepare. The coordinator logs BEGIN and sends
PREPAREto every participant. A participant does the work, keeps its locks, force-writes a PREPARED record and votesYES. After a YES it has promised to commit if told to, even after a crash, so it may no longer decide alone. A participant that cannot do its part votesNOand keeps nothing. - Commit. With every vote in, the coordinator force-writes the decision: COMMIT if all voted YES, otherwise ABORT. Writing that record is the moment the transaction commits. It sends the decision; the participants commit or undo, release their locks and send
ACK. With every ACK in, the coordinator logs END and forgets the transaction.
In the happy path that is 4 message delays and 12 messages. The order is all or nothing, and it is isolated: from PREPARE to COMMIT every changed row is locked, so no one can read a half-done order (see the orange band in the left timeline and the zero in the stats).
The weak spot: blocking
If the coordinator crashes after the participants voted YES but before they hear the decision, they are in doubt. They cannot commit (someone may have voted NO) and cannot abort (the coordinator may have logged COMMIT). They ask the coordinator again and again, and keep their locks until it comes back. Every other transaction that needs those rows waits too; in Demo: coordinator crash that is order B. This is why 2PC is called a blocking protocol.
Recovery works from the logs. The page uses presumed abort: a coordinator that restarts and finds BEGIN but no decision aborts, and answers "abort" for any transaction it has no record of. A participant that restarts with PREPARED and no outcome takes its locks again and asks (Demo: participant crash); the coordinator keeps resending its decision until every participant has acknowledged it.
Real systems: XA transactions (JTA in Java, XA START / XA PREPARE / XA COMMIT in MySQL, PREPARE TRANSACTION in PostgreSQL), and inside distributed databases such as Spanner and CockroachDB, where each participant is a replicated Paxos/Raft group, so a "crash" of the coordinator is survived by its replicas. Three-phase commit tries to remove the blocking, but only under assumptions (bounded delays, no network partition) that real networks do not give.
Saga
A saga (Garcia-Molina and Salem, 1987) gives up the single atomic step. It is a sequence of local transactions T1, T2, T3, each of which commits on its own at once, and for each one a compensating transaction C1, C2, C3 that undoes its effect in business terms. If Tk fails, the saga runs C(kâ1), ..., C1. The page uses an orchestrator: one service that calls the steps one by one and writes each step to a saga log before and after it.
- No distributed locks. Every lock lives only for one local transaction, so a slow or dead service never blocks the others' data. The right timeline has no orange cells.
- No isolation. Between T1 and the end, the charge and the reservation are committed and visible. When a later step fails, the customer sees a charge and then a refund, not "nothing happened".
- A compensation is not a rollback. It is a new transaction that runs after others may have read or used the earlier state: a refund, a release, a cancellation e-mail. Some actions cannot be compensated at all (an e-mail that was sent, money that was paid out); put those last, after the step that can still fail.
- Recovery is forward. A restarted orchestrator replays its saga log and carries on (Demo: coordinator crash), so nothing waits on it except the order itself.
- Retries need idempotency. When a reply is lost, the orchestrator cannot tell whether the step ran, so it sends it again. Each request carries a key (saga id, step) and a participant stores the answer for each key, so a repeated request returns the stored answer instead of charging twice (Demo: participant crash). Compensations must be idempotent too, and must not fail for good: they are retried until they succeed.
It usually costs more latency (the steps run one after another: 6 ticks against 4 on the page) and about half the messages, and every service must be designed with its compensation.
The isolation anomaly: the last copy
Demo: last copy gives order A the last copy in stock and makes Shipping refuse. Order B, from another customer, wants the same copy right after A reserved it.
| 2PC | Saga | |
|---|---|---|
| when B arrives | the copy is locked by A's prepared transaction: B waits | A's reservation is committed: stock is 0 |
| B's result | A aborts, the lock is released, B buys the copy | "out of stock" |
| in the end | B has the copy, A is cancelled | A is cancelled and the copy is back on the shelf, unsold |
B made a decision on data that later "never happened", the saga version of a dirty read. Sagas counter it with semantic locks (mark the row "reserved, pending" and let others treat it specially), commutative updates, re-reading values before acting, and by ordering the steps so that the ones likely to fail come first (the pivot transaction: after it the saga only moves forward).
Orchestration and choreography
The page draws an orchestrated saga. In a choreographed saga there is no central orchestrator: each service reacts to the events of the previous one (PaymentCharged â Inventory reserves â StockReserved â Shipping books ...), usually through a message broker, and the failure events trigger the compensations. It avoids a central service but makes the flow harder to see and to change. Either way a service must update its database and publish its event or reply atomically, which is what the transactional outbox pattern is for.
Which one?
| Two-phase commit | Saga | |
|---|---|---|
| atomicity | yes, one decision | eventually, through compensations |
| isolation | yes, locks until commit | no: half-done state is visible |
| locks across services | held for the whole protocol | none |
| coordinator failure | participants block, in doubt | the order pauses; nothing else is blocked |
| needs from participants | support for prepare (XA) | a compensation per step, idempotent handlers |
| good for | few, reliable, close resources (databases in one data centre, or inside a distributed database) | long-running business processes across services owned by different teams |
See also Optimistic vs Pessimistic Locking: the same lock-or-check trade-off inside one database. And Hexagonal Architecture: one service's side of the story, where a failed save is compensated by a refund through a payment port.
What the page leaves out
Participant-to-participant termination (cooperative termination protocol), presumed commit, read-only participants that skip phase 2, heuristic decisions (an administrator forcing an in-doubt transaction), replication of the coordinator, parallel saga steps, choreography, the outbox, and timeouts that give up on a saga step.