All notebook entries

When a deployment left two databases running

A failed deployment left old and replacement database and broker services running together. Recovery meant reconciling their state, replaying accepted orders, and changing how stateful services update.

Published

In July 2026, the ZanLIS Processor was accepting laboratory orders, but processing them became unreliable in a particularly confusing way. Requests appeared stuck or missing. Within a batch, their states could alternate between pending and processing. A worker could receive work and immediately report “LabRequest not found.”

The problem was underneath the request-processing code. A failed deployment had left both the old and replacement database and message-broker tasks running.

One service name hid two different realities

The deployment had encountered image-pull failures. Replacement tasks eventually started while the previous tasks were still running, and the update policy permitted that overlap.

Database traffic could reach diverging views of the data behind the same service address. The broker side had separate queues. Accepting a request through one path did not mean a healthy worker on another path could find either its message or its database record.

This explained why checking whether a container was running was insufficient. The containers were running. I needed to establish which database and broker the application and workers were actually using, and whether those services represented the same processing state.

The update policy was part of the data architecture

The stateful services were singletons: one database, one broker, one Redis service. Their deployment policy needed to preserve that arrangement during an update as well as after it.

Docker's start-first update order starts the replacement before stopping the old task, allowing them to overlap. That can keep an application available during a rolling update. It was the wrong policy for these single-instance stateful services.

I changed their update order to stop-first: stop the old task before starting its replacement. The API could retain its overlapping update strategy; the services holding the processing state needed a different lifecycle.

There was a cost. Updating a database or broker could now cause a brief interruption while its replacement started. For this deployment, accepting that interruption was preferable to allowing two instances to serve incompatible state. A genuinely redundant database or broker would require an architecture designed for that redundancy, beyond changing an update setting.

Restarting was only the beginning of recovery

Removing the duplicate services did not reconcile what had happened while they overlapped.

I recovered from the newer database backup and brought across records missing from the older view. That reconciliation included laboratory requests, integration requests, and an already delivered outcome. I also reset 76 stuck requests for controlled replay.

The processor had already accepted those orders. Asking the source systems to send everything again would have moved my recovery problem back to them and introduced another opportunity for duplicates.

Instead, recovery used the saved original requests. Administrative actions could rebuild and replay the work while retaining the existing duplicate-prevention checks. The uncertain-create handling still mattered: a stranded processor task did not prove that the laboratory had done nothing.

I checked restored record counts and worker and scheduler behavior after redeployment, alongside the automated tests for the recovery changes.

Check the transition, not just the desired state

A configuration saying “one replica” had not been enough to protect the system during replacement. The deployment transition could temporarily violate the assumption the application depended on.

That incident made deployment order an explicit engineering decision for me. Post-deployment checks now include looking for duplicate stateful tasks, alongside verifying that services are healthy and accepted work can actually progress.

Share this entry

Back to the notebook