Any service that writes to its own database and then publishes an event about that write has a consistency problem hiding in plain sight: those are two separate operations against two separate systems, and one can succeed while the other fails. The outbox pattern makes them effectively atomic without a distributed transaction.
The Dual-Write Problem
Consider an order service that saves a new order and publishes an OrderCreated event to a message broker so downstream services — billing, shipping, notifications — can react.
async function createOrder(order: Order) {
await db.orders.insert(order); // succeeds
await broker.publish('OrderCreated', order); // ...then the process crashes
}
If the process crashes, the network blips, or the broker is briefly unavailable between those two calls, you get an order that exists in the database with no event ever published — downstream services never hear about it. Swap the order and you get the opposite failure: an event published for an order that never actually committed, if the database write then fails or rolls back. There’s no version of “just call both” that’s safe, because a database transaction and a broker publish aren’t part of the same atomic unit.
The Outbox Table
The fix: don’t publish directly. Instead, write the event into an outbox table in the same database, in the same transaction as the actual business write. Since it’s the same transaction, it’s genuinely atomic — either both rows commit or neither does, guaranteed by the database itself.
CREATE TABLE outbox (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
aggregate_type TEXT NOT NULL,
aggregate_id TEXT NOT NULL,
event_type TEXT NOT NULL,
payload JSONB NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
published_at TIMESTAMPTZ
);
BEGIN;
INSERT INTO orders (id, customer_id, total_cents, status)
VALUES ('ord_9f2a', 'cus_11', 4899, 'created');
INSERT INTO outbox (aggregate_type, aggregate_id, event_type, payload)
VALUES ('order', 'ord_9f2a', 'OrderCreated',
'{"orderId": "ord_9f2a", "totalCents": 4899}');
COMMIT;
Now the order and the fact that it needs to be announced either both exist or neither does. What’s left is getting rows out of the outbox table and onto the broker, which is a separate, retriable problem, not an atomicity problem.
The Relay Process
A separate process — a poller or a change-data-capture stream — reads unpublished rows from the outbox, publishes them to the broker, and marks them published.
async function relayOutbox() {
const rows = await db.query(
`SELECT * FROM outbox WHERE published_at IS NULL
ORDER BY created_at LIMIT 100`,
);
for (const row of rows) {
await broker.publish(row.event_type, row.payload);
await db.query(`UPDATE outbox SET published_at = now() WHERE id = $1`, [row.id]);
}
}
The simplest version of the relay is a polling loop like this. A more efficient version reads the database’s write-ahead log directly via change data capture — tools like Debezium do this — pushing new outbox rows to the broker within milliseconds of the commit, without polling overhead at all.
This Only Gives You At-Least-Once
The relay can crash between publishing to the broker and marking the row published. When it restarts, it will publish that row again. The outbox pattern gives you at-least-once delivery, not exactly-once — and that’s a feature, not a shortfall, because exactly-once delivery across two independent systems isn’t achievable without cooperation from the consumer anyway.
Outbox vs Two-Phase Commit vs CDC-Only
| Approach | Atomicity guarantee | Broker coupling | Operational cost |
|---|---|---|---|
| Direct dual write | None — can silently diverge | Tight, synchronous | Low, until it breaks |
| Two-phase commit (XA) | Strong, if broker supports XA | Very tight | High, most brokers/queues don’t support it well |
| Transactional outbox + polling relay | Strong (write side), at-least-once delivery | Loose | Moderate, relay process to run |
| Transactional outbox + CDC | Strong (write side), at-least-once, low latency | Loose | Higher setup cost (Debezium/CDC pipeline), lower ongoing polling load |
Common Pitfalls
- Forgetting to prune the outbox table — published rows accumulate forever unless something archives or deletes them, and a huge outbox table slows the polling query.
- Publishing events out of order because the relay parallelizes across rows instead of respecting
created_atorder for events on the same aggregate. - Letting the relay’s polling interval become the de facto latency floor for every event in the system — a 30-second poll interval means every downstream reaction is delayed up to 30 seconds, which is often a surprise to whoever built assuming near-real-time delivery.
- Treating outbox delivery as exactly-once and building consumers that aren’t idempotent, which silently corrupts data the first time a redelivery happens.
Takeaway
The outbox pattern avoids the dual-write problem by never actually doing two writes to two systems — it writes the business change and the event into the same database, in the same transaction, and lets a separate relay process handle the (retriable, at-least-once) job of getting that event onto a broker. It buys you strong consistency on the write side in exchange for requiring idempotent consumers on the read side, which is a trade almost every distributed system ends up needing to make somewhere.