You book a holiday: flight, hotel, rental car — three separate bookings made one after another. The flight and hotel go through, but the car company has nothing available. You can't "rewind time" to before you booked the flight; that booking really happened. What you can do is cancel the flight and the hotel — new actions that undo the earlier ones. That's a compensating transaction.
The problem
Inside a single database, undo is easy: an ACID transaction either commits everything or rolls it all back, and a failed step vanishes as if it never ran. But a business operation that spans multiple services or databases can't be wrapped in one transaction. Each step commits to its own store the moment it succeeds. So when step three fails, steps one and two are already durably committed — there's no shared transaction manager to roll them back as a unit. You're left holding partial, real, irreversible-by-default progress.
Step through the trip below. When the car fails, predict what happens to the flight and hotel before you look. Then flip to 1 database to replay the same failure inside a single transaction, where the undo comes for free.
How it works
For every step that changes state, you define a paired compensating step that semantically reverses it: charge payment pairs with refund payment, reserve stock pairs with release stock, book flight pairs with cancel flight. When a later step fails, you walk backward through the steps that already succeeded and run each one's compensation in turn. Crucially, this is not a database rollback — a refund is a brand-new transaction that offsets the original charge; the charge still exists in history. It's a logical undo built from forward actions.
That also means a compensation follows the business rules of the moment, not the arithmetic of an undo. Cancelling a hotel may be free; cancelling a flight may cost a fee. Each compensation is a real request to a real system, with its own price and its own ways to fail.
Step through the unwinding below. First, predict where the traveler ends up once everything is cancelled. Then the airline's cancel call times out: predict what the planner should do before you see it.
Your app buys a $60 concert ticket, then fails to reserve parking, so it compensates by cancelling the ticket. The seller keeps a $5 fee. What should the user's card statement show afterwards?
When the undo fails
Compensations run over the same unreliable network as everything else, so they fail too: a timeout, a crashed worker, a service that's down for an hour. Unlike a forward step, you usually can't answer that by giving up. A half-unwound trip, with the hotel cancelled and the flight still booked, is worse than either a finished booking or a clean cancellation. So a compensation is retried until it succeeds, and that only works if it follows a few rules:
- Idempotent. Running the same compensation twice must have the same effect as running it once. Send an idempotency key with each request, like A7 in the scene, so the receiver can spot a repeat and return its original answer instead of refunding again.
- Retryable. Keep a durable record of which compensations are still owed, and retry them with a backoff. A crash halfway through the unwinding should resume where it stopped, not forget the flight.
- Escalate in the end. Some compensations will keep failing. After a bounded number of attempts, park the work in a queue for a human, with everything they need to finish it by hand.
Flip the scene's cancel call to Not idempotent to see what a retry without a key costs.
Record the undo before you need it. Each time a forward step commits, write down its compensation ("cancel seat 14C, key A7") in durable storage. If the process crashes mid-unwind, a fresh worker can read the list and carry on. That bookkeeping is exactly what a saga orchestrator or a Scheduler Agent Supervisor does for you.
A compensation gets you somewhere consistent, not back where you started. Fees, side effects and time don't reverse: the hotel room you released may be resold a minute later, and the confirmation email was already read. For each step, write down what its compensation really restores, and tell users the truth ("refunded $350: $400 minus the $50 fare fee") instead of pretending nothing happened.
When to use it
Compensating transactions are the undo half of a saga: use them whenever a multi-step operation spans services that can't share one ACID transaction and you still need a way to back out of partial progress. They fit long-running workflows — order fulfillment, travel booking, provisioning — where each step is independently committed and any step might fail.
The hard part is that not everything is cleanly reversible. An email already sent or a physical package already dispatched can't simply be cancelled, so some compensations must be approximate (a follow-up correction), a fallback, or a hand-off to a human. If your operation lives in a single database, don't bother — let a plain ACID rollback do the work for free.
A "refund payment" compensation times out, and your workflow retries it. What makes that retry safe?