Stage 12: Microservices, lesson 5 of 7

Distributed data and the saga pattern

Advanced3 min readall versions
Explain it forThe essentials plus production detail and pitfalls.

With a database per service you can't wrap several services in one ACID transaction, and two-phase commit (XA) doesn't scale well. A saga is a sequence of local transactions; if a step fails, compensating transactions undo the steps that already succeeded.

Two styles:

  • Choreography: each service reacts to events and emits new ones. Simple for two to four steps, harder to follow as it grows.
  • Orchestration: a central orchestrator tells each service what to do next and tracks the saga's state. Clearer for complex flows (tools include Temporal and Camunda, or a state machine you write).

Related patterns: the transactional outbox for reliable event publishing, CQRS for separate read models built from events, and event sourcing, which stores events as the source of truth.

Example

Plain text
Course purchase saga (orchestrated)

1. order-service      create order           PENDING
2. payment-service    charge ₹1,999          ok: next step   failed: cancel order
3. user-service       grant course access    ok: next step   failed: refund, cancel order
4. notification-svc   send receipt           retried until it works, never compensated
5. order-service      mark order COMPLETED

Compensations run in reverse order: refund payment, then cancel order
Orchestrator compensation handlers
public void on(PaymentFailed e) {
    orders.cancel(e.orderId(), "PAYMENT_FAILED");           // undo step 1
}

public void on(AccessGrantFailed e) {
    payments.refund(e.paymentId());                         // undo step 2
    orders.cancel(e.orderId(), "ACCESS_FAILED");            // undo step 1
}

Common mistake

Treating a compensating action like a database rollback. It's a new business operation, such as a refund or a cancellation email, that the customer may see.

Under the hood

Sagas give eventual consistency, not isolation: other requests can see in-between states, such as an order that is still PENDING. Design for it with status fields (semantic locks), updates that can be applied in any order, and read-your-own-writes where users expect it. Compensations must be idempotent and must themselves be retried until they succeed.

Check yourself

What undoes a completed saga step?

How this connects

Where this leads

You've reached the end of this thread. Try a learning path for what's next.

Part of Microservices and production.

Was this lesson helpful?

Finished reading? Mark it complete to track your progress.