All notebook entries

When saving a request does not queue its work

Introducing a broker left a gap between committing application state and publishing a task. A transactional outbox made the intent to dispatch durable.

Published

Moving delivery into a queue let me control how much work reached ZanLIS at once. It also gave the processor another responsibility: an accepted request needed to reach the queue reliably.

Saving a request in the processor's PostgreSQL database and publishing a task to RabbitMQ are two separate operations. Completing the first does not guarantee the second will happen.

The database can enforce uniqueness and commit its own changes atomically. Those guarantees do not extend to the separate broker publication. This problem remains even after duplicate database writes are prevented.

The gap between two successful steps

Consider a producer that saves its request and then publishes a task. If the database commit succeeds but the process stops before publication, the request exists without the background work needed to deliver it.

Changing the order does not make the two operations atomic. A task published before the database commit could run before its state is ready, or the database transaction could subsequently fail.

There is also uncertainty at publication: the broker may accept a message while the producer fails to receive confirmation. Retrying publication can then publish the same logical work again.

The architecture needed a durable record of what still had to be dispatched.

Save the intent with the request

I introduced a transactional outbox. In my implementation, a producer writes a TaskDispatch row inside the same database transaction as the associated application state.

Either both commit, or both roll back. Once committed, the database holds both the request state and the intent to publish its task.

A dispatcher then claims due outbox rows and publishes them to RabbitMQ. Confirmed publications remove their outbox rows. Failed publications are backed off; exhausted publication attempts remain available for operator recovery.

The producer no longer needs RabbitMQ to be available at the instant it commits the request. Delivery can catch up after publication recovers, subject to those retry and recovery paths.

Separate publication failure from delivery failure

A failed broker publication is different from a task that ran and failed to deliver an order to ZanLIS.

I separated those responsibilities:

  • The request's recovery state decides whether another delivery attempt is needed.
  • The outbox owns publishing that authorized attempt.
  • The worker executes the attempt when it receives the task.

A broker outage therefore does not consume a domain delivery attempt as though ZanLIS had rejected the request. Long delays can also be represented by a due time in the database until publication is appropriate.

The outbox is not the permanent task history. Successfully published rows are removed. Operational history and the request's recovery state have their own records.

Durable intent still allows duplicate publication

The outbox does not eliminate the uncertain acknowledgement problem. If RabbitMQ accepts a message and the dispatcher stops before recording success, a later dispatcher can publish that row again.

Stable operation keys help avoid creating multiple active intents for the same logical attempt. They do not prove that a worker will only receive the task once.

Workers still need state checks and safe repeat handling. For sample creation in ZanLIS, that includes checking whether an uncertain earlier create already produced a sample.

This distinction matters when describing what the architecture guarantees. The outbox makes committed dispatch intent recoverable. Completing the downstream operation correctly still depends on the worker and its recovery rules.

Each boundary needs an owner

The queue controls how work reaches workers. The outbox records work that still needs publication. The request's recovery model decides whether another attempt is justified.

Keeping those responsibilities explicit helped avoid one generic retry mechanism making decisions for every failure. At each boundary, the useful questions are what has been confirmed, what remains uncertain, and which component owns the next step.

Share this entry

Back to the notebook