The problem: one order, three databases

An order needs three things to happen: Payment charges the customer, Inventory reserves the item and Shipping books a courier. Each service has its own database, so no single local transaction can cover all three. Either all three must happen or none of them, even when a service says no or a machine crashes halfway. The page runs the same order through the two classic answers at the same time, on one clock: two-phase commit on the left, a saga on the right.

Every message takes one tick. The timeline under each side has one cell per tick for each database, showing what it holds for the order at the end of that tick. The stats line counts messages, lock-ticks (ticks a participant holds a lock for the order, summed over participants) and ticks in which a half-done order was visible: one database already shows the order's change while another shows the state without it, and nothing stops a reader from seeing both.

Two-phase commit

A coordinator (transaction manager) drives the participants (resource managers):

  1. Prepare. The coordinator logs BEGIN and sends PREPARE to every participant. A participant does the work, keeps its locks, force-writes a PREPARED record and votes YES. After a YES it has promised to commit if told to, even after a crash, so it may no longer decide alone. A participant that cannot do its part votes NO and keeps nothing.
  2. Commit. With every vote in, the coordinator force-writes the decision: COMMIT if all voted YES, otherwise ABORT. Writing that record is the moment the transaction commits. It sends the decision; the participants commit or undo, release their locks and send ACK. With every ACK in, the coordinator logs END and forgets the transaction.

In the happy path that is 4 message delays and 12 messages. The order is all or nothing, and it is isolated: from PREPARE to COMMIT every changed row is locked, so no one can read a half-done order (see the orange band in the left timeline and the zero in the stats).

Sequence diagram, time flowing down: the Coordinator sends PREPARE to Payment, Inventory and Shipping, which lock their rows and vote YES; the Coordinator logs COMMIT, sends COMMIT, and the participants release their locks and send ACK. If the coordinator crashes after the votes, participants wait in doubt with locks held
2PC holds every participant's locks from PREPARE to COMMIT, so the order is all-or-nothing and never half visible, but a coordinator crash leaves everyone waiting.

The weak spot: blocking

If the coordinator crashes after the participants voted YES but before they hear the decision, they are in doubt. They cannot commit (someone may have voted NO) and cannot abort (the coordinator may have logged COMMIT). They ask the coordinator again and again, and keep their locks until it comes back. Every other transaction that needs those rows waits too; in Demo: coordinator crash that is order B. This is why 2PC is called a blocking protocol.

Recovery works from the logs. The page uses presumed abort: a coordinator that restarts and finds BEGIN but no decision aborts, and answers "abort" for any transaction it has no record of. A participant that restarts with PREPARED and no outcome takes its locks again and asks (Demo: participant crash); the coordinator keeps resending its decision until every participant has acknowledged it.

Real systems: XA transactions (JTA in Java, XA START / XA PREPARE / XA COMMIT in MySQL, PREPARE TRANSACTION in PostgreSQL), and inside distributed databases such as Spanner and CockroachDB, where each participant is a replicated Paxos/Raft group, so a "crash" of the coordinator is survived by its replicas. Three-phase commit tries to remove the blocking, but only under assumptions (bounded delays, no network partition) that real networks do not give.

Saga

A saga (Garcia-Molina and Salem, 1987) gives up the single atomic step. It is a sequence of local transactions T1, T2, T3, each of which commits on its own at once, and for each one a compensating transaction C1, C2, C3 that undoes its effect in business terms. If Tk fails, the saga runs C(k−1), ..., C1. The page uses an orchestrator: one service that calls the steps one by one and writes each step to a saga log before and after it.

Sequence diagram: the Orchestrator runs T1 charge 100 on Payment and T2 reserve 1 on Inventory, each committing at once; T3 book courier fails on Shipping; the orchestrator then runs the compensations C2 release 1 and C1 refund 100 in reverse order. The charge is visible to readers from T1 until C1
A saga commits each step on its own and undoes completed steps with compensations, so nothing blocks, but readers can see the charge before the refund.
  • No distributed locks. Every lock lives only for one local transaction, so a slow or dead service never blocks the others' data. The right timeline has no orange cells.
  • No isolation. Between T1 and the end, the charge and the reservation are committed and visible. When a later step fails, the customer sees a charge and then a refund, not "nothing happened".
  • A compensation is not a rollback. It is a new transaction that runs after others may have read or used the earlier state: a refund, a release, a cancellation e-mail. Some actions cannot be compensated at all (an e-mail that was sent, money that was paid out); put those last, after the step that can still fail.
  • Recovery is forward. A restarted orchestrator replays its saga log and carries on (Demo: coordinator crash), so nothing waits on it except the order itself.
  • Retries need idempotency. When a reply is lost, the orchestrator cannot tell whether the step ran, so it sends it again. Each request carries a key (saga id, step) and a participant stores the answer for each key, so a repeated request returns the stored answer instead of charging twice (Demo: participant crash). Compensations must be idempotent too, and must not fail for good: they are retried until they succeed.

It usually costs more latency (the steps run one after another: 6 ticks against 4 on the page) and about half the messages, and every service must be designed with its compensation.

The isolation anomaly: the last copy

Demo: last copy gives order A the last copy in stock and makes Shipping refuse. Order B, from another customer, wants the same copy right after A reserved it.

2PCSaga
when B arrivesthe copy is locked by A's prepared transaction: B waitsA's reservation is committed: stock is 0
B's resultA aborts, the lock is released, B buys the copy"out of stock"
in the endB has the copy, A is cancelledA is cancelled and the copy is back on the shelf, unsold

B made a decision on data that later "never happened", the saga version of a dirty read. Sagas counter it with semantic locks (mark the row "reserved, pending" and let others treat it specially), commutative updates, re-reading values before acting, and by ordering the steps so that the ones likely to fail come first (the pivot transaction: after it the saga only moves forward).

Orchestration and choreography

The page draws an orchestrated saga. In a choreographed saga there is no central orchestrator: each service reacts to the events of the previous one (PaymentCharged → Inventory reserves → StockReserved → Shipping books ...), usually through a message broker, and the failure events trigger the compensations. It avoids a central service but makes the flow harder to see and to change. Either way a service must update its database and publish its event or reply atomically, which is what the transactional outbox pattern is for.

Which one?

Two-phase commitSaga
atomicityyes, one decisioneventually, through compensations
isolationyes, locks until commitno: half-done state is visible
locks across servicesheld for the whole protocolnone
coordinator failureparticipants block, in doubtthe order pauses; nothing else is blocked
needs from participantssupport for prepare (XA)a compensation per step, idempotent handlers
good forfew, reliable, close resources (databases in one data centre, or inside a distributed database)long-running business processes across services owned by different teams

See also Optimistic vs Pessimistic Locking: the same lock-or-check trade-off inside one database. And Hexagonal Architecture: one service's side of the story, where a failed save is compensated by a refund through a payment port.

What the page leaves out

Participant-to-participant termination (cooperative termination protocol), presumed commit, read-only participants that skip phase 2, heuristic decisions (an administrator forcing an in-doubt transaction), replication of the coordinator, parallel saga steps, choreography, the outbox, and timeouts that give up on a saga step.