Accepting batches without overwhelming ZanLIS
The integrations sent batches, but ZanLIS accepted one order per API call and struggled with concurrent writes. Celery and RabbitMQ let me control the handoff.
One of the early problems in the ZanLIS integration was a mismatch between how requests arrived and how I could deliver them.
The systems sending laboratory orders could submit a batch. The ZanLIS endpoint I was integrating with accepted individual requests. Splitting the batch was straightforward. Deciding how quickly to send those individual requests became the harder problem.
Sending everything at once made things worse
It was tempting to split a batch into individual calls and send them concurrently. That would avoid making the source system wait for each request to finish before starting the next.
In this deployment, however, the write path struggled under that pressure. I remember runs where sending about 20 concurrent requests could leave roughly half failing. That is an approximate recollection of the behavior I encountered, rather than a benchmark for every deployment.
ZanLIS uses ZODB. Its transaction model uses optimistic concurrency: conflicting changes can require a transaction to be aborted and retried. The workload was running into concurrent-write problems. Simply opening more connections did not make ZanLIS able to accept more useful work.
I needed somewhere for the rest of the batch to wait.
Put a queue between arrival and delivery
I introduced asynchronous processing with Celery and RabbitMQ. The processor could split an incoming batch into individual tasks, while workers delivered those tasks to ZanLIS at a controlled level of concurrency.
The responsibilities became clearer:
- The processor accepted and prepared the incoming orders.
- RabbitMQ held messages waiting for workers.
- Celery workers performed the individual handoffs to ZanLIS.
The queue let the processor accept a burst without turning the whole burst into simultaneous writes downstream. A request could wait for a worker slot instead of competing with every other request in its batch.
Why I settled on four concurrent handoffs
For that deployment, four concurrent handoffs worked well. The arrangement gave each of four execution slots one delivery task at a time. When a task finished, the slot could take the next one.
With a batch of 20, that meant up to four tasks executing while the remaining work waited. The tasks did not have to finish together in groups of four; a free slot could immediately take another task.
The limit belonged to the worker configuration. A message broker does not, by itself, make four worker instances equivalent to four concurrent requests. Celery workers can each run multiple tasks. Its concurrency and prefetch settings govern different things: executing tasks and reserving work are not the same operation.
What mattered was the total number of active delivery slots feeding that ZanLIS instance. Adding worker replicas or increasing their concurrency would need to respect the same downstream limit. Other traffic reaching ZanLIS also remained outside this queue's control.
The tradeoff was waiting
This did not make ZanLIS process an individual request faster. It gave me a way to keep the volume of concurrent work within what the deployment could handle.
The cost was queueing time. Accepting an order into the integration could no longer mean that its sample already existed in ZanLIS. I had to track delivery separately and make failures recoverable after the original HTTP call had ended.
A queue also cannot absorb an indefinitely growing workload. If arrivals consistently exceed completions, the backlog grows. I still need to understand how long work waits and whether workers are making progress.
Four was a useful operating choice for the workload I had, not a permanent property of ZanLIS or ZODB. The lasting decision was to control delivery independently of arrival.
That separation led to another problem: what should happen when a delivery times out after ZanLIS may already have created the request?