Cloud Native Patternsadvanced10 min

Saga

Coordinate a business operation across many services by chaining local transactions and undoing them with compensations when a step fails.

Imagine placing an online order. Behind that one button click, several independent services have to cooperate: one records the order, another charges your card, a third reserves the item from stock, and a fourth schedules the shipment. Each of those services owns its own database. So what happens when the payment succeeds but inventory then reports the item is out of stock? You've taken the customer's money for something you can't deliver, and there's no single "undo" button across four separate systems.

The problem

Inside a single database, this is a solved problem: wrap everything in one ACID transaction and the database guarantees it all commits together or not at all. But a business operation that spans place order → charge payment → reserve stock → ship touches four services, each with a separate database, often on different machines.

There is no shared transaction manager that can lock all four databases at once and roll them back as a unit. Holding a global lock across services for the duration of the whole operation would also be disastrous for availability — you'd be freezing inventory and payment systems while you wait on the network. We need a way to keep the operation consistent without a single all-or-nothing transaction.

How it works

A saga reframes the operation as a sequence of local transactions, one per service. Each service does its own small piece of work, commits it to its own database, and then triggers the next step. Place the order, then charge the payment, then reserve the stock, then ship — each commit is real and durable on its own.

The trick is the failure path. If a later step fails, the saga walks backward through the steps that already succeeded and runs a compensating transaction for each one: refund the payment, then cancel the order. These compensations don't magically restore the old database state — they're new transactions that semantically undo the effect of the earlier ones. It's a logical rollback, not a true database rollback.

Step through one order below. When Inventory comes back out of stock, predict which steps get undone, and in what order, before you look. Keep an eye on what the customer sees along the way. Then flip to No saga to watch the same failure with nobody to clean it up.

Note

Compensation isn't time travel. A database rollback erases a transaction as if it never happened. A compensation is a brand-new action that counteracts a committed one — a refund offsets a charge, it doesn't delete the original charge. That's why your domain has to define a sensible "undo" for every step; some effects (an email already sent, a missile already launched) can't truly be taken back.

Check yourself

An order saga runs four steps: create order, charge card, reserve stock, book a courier. Booking the courier fails. What does the saga do?

Two ways to coordinate: choreography vs. orchestration

Sagas come in two coordination styles. In choreography, there is no central brain: each service publishes an event when it finishes its local transaction, and the next service listens for that event and reacts. The order service emits OrderPlaced, payment reacts and emits PaymentCharged, inventory reacts to that, and so on. This is naturally pub/sub-driven and pairs well with event sourcing, keeping services loosely coupled — but the overall flow is implicit, smeared across many services and hard to see in one place.

In orchestration, a central coordinator (the orchestrator) owns the workflow. It calls each service in turn, waits for the result, and decides what to do next — including which compensations to fire on failure. The logic lives in one place, which makes the saga far easier to understand, monitor, and modify, at the cost of a component that every step now depends on.

Below, the order saga is paused at the moment it fails. Choose who should drive the compensations and watch it play out, then flip the switch to see the other style at the same step. Notice where the rule "if stock runs out, refund and cancel" ends up living.

Consistency, isolation, and idempotency

Because each step commits independently, a saga gives you eventual consistency, not the instant all-or-nothing consistency of a single ACID transaction. For a window of time the system sits in a partial state — the order exists and the payment is taken, but the stock isn't reserved and the shipment hasn't been arranged yet.

That partial state is also visible to other readers: a saga has no isolation, so the customer can see a charge for an order that's about to be cancelled, and another request can see a half-finished order as if it were final. And because messages get retried and steps can re-run after a crash, every step and every compensation must be idempotent — running it twice must produce the same result as running it once, or you'll double-charge a card or release the same stock twice.

Watch out

The intermediate states are real, and others can see them. With no isolation, a customer might briefly see an order marked "placed" that's about to be cancelled by a compensation, and a concurrent operation can act on stock that's only tentatively reserved. Design every step to tolerate these in-between states — use pending/confirmed status flags rather than assuming each step is the final word.

Check yourself

A customer sees "Order placed" and an $80 charge on their card, then a minute later gets a refund and a cancellation email. What happened?

The trade-offs

Sagas buy you cross-service consistency, but they aren't free:

  • Complexity — you've replaced one transaction with a distributed workflow, plus a mirror-image set of compensation logic and the messaging or coordination plumbing to drive it. That's a lot more moving parts to build, test, and reason about.
  • No isolation — intermediate states leak, so every part of the system has to be written with partial progress in mind.
  • Designing compensations is hard — for each step you must define a meaningful undo, handle the case where the compensation itself fails (and must be retried), and accept that some real-world effects simply can't be reversed and need human or manual handling instead.

When to use it

Reach for a saga when a single business operation must span multiple services or databases and you genuinely cannot wrap it in one ACID transaction — the classic microservices situation of distributed data you still need to keep consistent. It's the standard answer for long-running, multi-step workflows like order fulfillment, booking and travel itineraries, or any "do these five things across five teams, and undo them if one fails" process.

If, on the other hand, your operation lives inside a single service and database, don't reach for a saga at all — use a plain local transaction and let ACID do the work. Sagas are the price you pay for distribution, so only pay it when distribution is what forced your hand.

Key takeaways

  • A saga breaks one business operation into a sequence of local transactions, one per service, instead of a single transaction spanning all of them.
  • When a step fails, the saga runs compensating transactions that semantically undo the earlier steps — a logical rollback, not a database rollback.
  • Choreography lets each service react to events with no central brain; orchestration uses a coordinator to drive the steps explicitly.
  • Sagas trade strict consistency for availability: the system is eventually consistent and intermediate states are visible to others.
  • Every step and every compensation must be idempotent, because failures and retries mean they can run more than once.

Keep going