All notebook entries

Retrying a timeout without creating another sample

A timeout left me unsure whether ZanLIS had created a request. Blind retries created duplicates, so recovery had to begin with a searchable external order ID.

Published

Retries helped the integration recover from temporary failures, but they also introduced a failure of their own: duplicate samples in ZanLIS.

The processor would send a create request to ZanLIS and wait. Sometimes the call timed out. I treated that as an error and retried the stage.

But a timeout only told me that the processor had not received a definitive response. ZanLIS might already have created the sample.

One order, more than one sample

When the first create had succeeded and its response was lost or delayed, replaying the create could produce another sample. Both samples carried the same external order ID, but each had a different internal sample ID in ZanLIS.

At that point, there was no uniqueness check preventing that external order from being created again. The retry mechanism could report progress while quietly making the downstream state more ambiguous.

The recovery question had to change from “Did the HTTP call succeed?” to “Does the requested sample already exist?”

First, make the order searchable

I could not reliably answer that question with the API lookup I had at the time. The external order ID was the identifier available to the processor, and it was not searchable through the path I needed.

The fix therefore crossed the integration boundary. I added catalogue indexing in ZanLIS so the external order could be looked up. The current sample catalogue includes a ClientOrderNumber field index.

The index made order IDs searchable, but did not enforce uniqueness. ZODB offered no relational-style UNIQUE constraint for the field; that would have required a custom guard designed and tested for concurrent writes. I chose verification to stop blindly retrying uncertain creates without taking on that larger change.

The processor also needed the source and destination context, so it could look for the right order in the right ZanLIS instance.

An uncertain create needs verification

After a create whose outcome is uncertain, the processor now performs a confirmation lookup before deciding whether another create is appropriate.

What verification establishes What happens next
A matching request exists Treat delivery as completed; do not create another sample.
A trusted lookup reports no match Allow the guarded create path to retry.
The lookup itself fails Keep the create uncertain and retry verification.

That third case is essential. A failed search is not evidence that the sample does not exist.

The current implementation has a separate verification task and a verification_pending state. The verification task searches; it does not create samples. A confirmed absence sends work back through the create path, which performs its own checks.

Multiple matches require care. Before a new create, existing multiple matches stop automatic creation for investigation. After an uncertain create, finding multiple matches prevents another create and records a warning. Treating delivery as completed in that branch does not mean the duplicates have been repaired.

Retry the part that still needs work

There is a related distinction after a successful create. If creating the sample succeeds but a subsequent state transition fails, replaying the whole operation is unnecessary. The retry should target that state transition.

I use structured retry hints to preserve these differences: retry the request, retry a state change, verify an uncertain create, or stop for investigation. The next action depends on what the processor can confirm happened, rather than just the presence of an error.

What verification guarantees

A lookup before a write still leaves a race if independent writers both observe no match and then create. Search visibility also matters when deciding that a record is absent. Verification reduces duplicate risk without making the lookup and creation atomic.

It addressed the concrete failure I had: repeatedly creating samples because an earlier response was uncertain. It gave the processor a way to ask ZanLIS what had happened before trying to make it happen again.

A related uncertainty appears when saving a request and publishing its background task: committing the processor database does not confirm publication to the broker.

Share this entry

Back to the notebook