Todd Watts

What an API Timeout Does and Does Not Tell You

Todd WattsUpdated 9 min read

A reproducible synthetic lending experiment shows why a lost reply needs stable request identity, separate notification state and a recovery plan before another retry.

A lime record is stored in an archive box while its return path breaks before a waiting clock; a separate tray holds notifications.
A missing reply does not undo a saved operation.Illustration · Todd Watts
03 / Knowledge ≠ state

A timeout changes what you know.

  1. Stored factR-001 existsOne reservation.
    One notification intent.
  2. Caller knowsOutcome uncertainThe commit survived. Its reply did not.
  3. NotificationPending send0 send attempts.
    Delivery is not established.

Next decision: Repeat the original key and payload to recover the stored result. A new key means a different request.

Computed from the same deterministic TypeScript model as the lending lab. No network, database or actual notification is involved.

A caller can stop waiting while the requested change has already happened. If the interface turns that uncertainty into “nothing was saved,” its next suggestion may make the situation worse.

The telescope lending lab makes this problem small enough to inspect. One fictional member wants the last telescope. The model can lose the reservation reply or interrupt a notification acknowledgement. Each action exposes the caller's knowledge alongside the stored result.

This is an original, in-memory demonstration with invented data. Its commands run synchronously in a prescribed order. There is no HTTP server, database or notification provider behind the experiment, and it does not reconstruct an employer or client system.

First, identify which timeout you mean

A client deadline expiring is an observation at the caller. It may occur without any HTTP response. A received HTTP error has more specific semantics: 408 concerns a server waiting for a complete request, while 504 concerns a gateway waiting for an upstream response. Those are different events, even if the interface labels all three “timeout.” HTTP semantics, sections 15.5.9 and 15.6.5.

For a request whose reply is missing, I want to distinguish three questions:

  1. What did the caller observe?
  2. What change, if any, was committed?
  3. Which operation is safe to repeat?

The lab deliberately shows a reply disappearing after its commit. This is one possible failure sequence. It does not simulate every place a real connection might fail.

Reproduce the lost reply

Open the lending lab, choose Duplicate request under Try a failure, and press Replay if a trace is already in progress. Press Start scenario, then Next event to advance one step at a time.

Scroll horizontally if needed to read every column.

StepEventWhat to inspect
1A member asks for the telescopeThe caller keeps request key loan-017. No reservation exists yet.
2One commit records the reservationR-001 and notification intent N-001 appear together. The caller is still awaiting a reply.
3The response is lostThe caller shows Reply unknown. R-001 still exists.
4The caller retries the same requestThe request-attempt counter becomes 2. The original key and payload are reused.
5The saved result is reusedThe store shows Existing result reused, with 0 available and 1 reservation.
6The caller receives R-001The caller can now show Reservation confirmed.
7–8The notification is attempted and acknowledgedOne intent has one send attempt. The reservation is unchanged.

The important ordering is inside commit: check for the saved request before rejecting a new reservation for lack of stock. Otherwise, the successful original request would make its own retry appear unavailable.

The trace exposes the entire model for learning. A real caller would need an authorized response or lookup to learn the authoritative result; it could not inspect server memory as the diagram does.

A key identifies the same intent

In this model, request identity includes a key, member and kit. Repeating loan-017 with the same member and kit returns R-001. Reusing that key with a different member or kit returns key-conflict and leaves the original reservation and notification intent intact.

Those mismatch branches are executable in the reducer, but the three-option lab interface does not expose them. The recorded checks exercise them directly. An additional check uses a new key after losing the original reply. That attempt returned unavailable: the single telescope was already reserved. It did not recover the original confirmation.

That last result illustrates why generating a fresh key is a poor recovery mechanism. In this particular model, zero stock prevents a second reservation. It says nothing about a service with enough stock to accept both requests.

For an HTTP API, a header named Idempotency-Key only helps if the server implements an appropriate contract. Define the key's caller or tenant scope, which payload fields must match, how long results are retained, and what concurrent or expired requests do. The lab compares two payload fields and has no authentication, expiry or concurrent execution.

HTTP method semantics also matter. The standard discourages automatic retries of non-idempotent requests unless the client knows the operation is safe to repeat or knows the original was not applied. A POST should not acquire an unconditional retry loop merely because its response went missing. HTTP retry semantics.

Give the notification its own outcome

The model creates three linked facts in one transition: the reservation, its request receipt and a notification intent. The receipt is represented by the request saved on the reservation. The notification is sent in a later transition.

That models the boundary of a transactional outbox: persist the business change and its intent to notify together, then have separate work send the notification. It avoids relying on a successful database write followed by an unrelated send that could be forgotten after a crash. A real implementation needs an actual durable transaction and a recoverable worker. Duplicate sends still need handling. AWS transactional outbox guidance.

In the lab, “atomic store” is an assumption implemented as one JavaScript state transition. It is not evidence that a database has provided isolation, crash safety or durable storage.

Now select Notification timeout. The selection resets the trace; press Start scenario and continue with Next event:

Scroll horizontally if needed to read every column.

StepCallerNotification
3Reservation confirmedN-001 is queued, with 0 send attempts.
4Reservation confirmedThe worker awaits acknowledgement after attempt 1.
5Reservation confirmedAcceptance unknown. The reservation is still R-001.
6Reservation confirmedThe same N-001 is attempted again. The send-attempt counter becomes 2.
7Reservation confirmedProvider accepted. No additional reservation request was made.

The simulated acknowledgement says nothing about arrival in an inbox. The model does not count real deliveries or emulate provider deduplication. If a real provider accepted the first attempt and its acknowledgement was lost, repeating the send could produce a duplicate message.

The interface should preserve that distinction: reservation confirmed, notification uncertain. A notification failure should not silently roll back a reservation that has already been confirmed.

Put a boundary around retries

The notification scenario contains exactly one scripted retry. The reducer itself has no retry limit, timer, backoff or jitter. Calling send and timeout repeatedly would keep increasing its attempt count. The trace's finite length is not a production retry policy.

For a real integration, I would write that policy before implementing the retry button:

  • Retry only outcomes the operation's contract classifies as retryable. Correct a payload conflict rather than resubmitting it unchanged.
  • Keep the original operation identity. Retrying a notification uses its notification identity, not a new reservation request.
  • Bound both the total attempts and the total waiting time. Account for retries already performed by an SDK or another layer.
  • Schedule eligible retries with backoff and jitter, and honor applicable server delay guidance. Stop when the remaining deadline cannot accommodate another attempt.
  • Leave an explicit unresolved state when the budget ends, with a route to reconciliation or human review.

Exact limits depend on the operation, latency requirements and dependency contract. AWS SDK documentation provides a concrete example of capped attempts, retry classification and jittered backoff; it is not a universal policy to copy into every application. AWS SDK retry behavior.

Reconciliation needs a trustworthy result

Reconciliation means resolving uncertainty against authoritative state. For the lost reservation reply, this lab demonstrates one such path: repeat the same keyed operation and obtain the stored result. A real API might instead expose an authenticated status lookup by operation ID. The lab does not implement that endpoint.

For a notification, a provider's status lookup or idempotency support may help, if its documented guarantees cover the operation. If neither exists, the product needs an explicit choice about duplicate risk, delayed communication or manual review. Re-running the reservation is not a substitute for finding out what happened to the message.

I would also define the meaning of “not found.” It could be a terminal result, a lookup made before processing finished, or a record that has expired. A recovery flow cannot treat all three as permission to repeat a business action.

An executable check and the recorded observations

The observations for this article were generated on October 2, 2026, by executing every step of the three authored scenarios and the three extra key probes through the same TypeScript reducer used by the browser lab. The recorded JSON trace includes commands, intermediate states, runtime and a hash of the model source. It contains synthetic state, not request logs or timing measurements.

Scroll horizontally if needed to read every column.

Completed scenarioReservation requestsReservationsNotification intentsSend attempts
Normal request1111
Duplicate request2111
Notification timeout1112

To run the check yourself, download the standalone model source and timeout check into the same folder, keeping those filenames. With Bun installed, run bun timeout-check.ts from that folder. No repository checkout or application setup is required. The model download is an exact copy of the independently written source used by the browser lab; a test checks that the copies stay identical.

The downloadable check contains this code:

import assert from 'node:assert/strict';
import {
  applyLendingCommand,
  LENDING_REQUEST,
  lendingSnapshot,
} from './lending-lab-model';

const lost = lendingSnapshot('duplicate', 3).state;
assert.equal(lost.client, 'uncertain');
assert.equal(lost.reservation?.id, 'R-001');

const replayed = lendingSnapshot('duplicate', 6).state;
assert.equal(replayed.client, 'confirmed');
assert.deepEqual(replayed.reservation, lost.reservation);
assert.deepEqual(replayed.outbox, lost.outbox);

const changed = applyLendingCommand(lost, {
  type: 'receive',
  request: { ...LENDING_REQUEST, member: 'member-18' },
});
const rejected = applyLendingCommand(changed, { type: 'commit' });
assert.deepEqual(rejected.reply, { kind: 'key-conflict' });
assert.deepEqual(rejected.reservation, lost.reservation);

const timeout = lendingSnapshot('timeout', 5).state;
assert.equal(timeout.client, 'confirmed');
assert.equal(timeout.outbox?.status, 'uncertain');
console.log('Replay, payload conflict and notification checks passed.');

These checks establish behavior of this deterministic model. Before relying on the design in a service, I would still need tests for simultaneous requests, real transaction rollback, a crash after commit, worker recovery, caller authorization, key retention and provider acknowledgement loss. None of those are measured by this article's trace.

The useful design artifact is the relationship between a fact, the caller's knowledge of that fact and the next permitted action. You can step through the lab, inspect other engineering work, or read about fractional CTO work if this kind of decision is part of a system your business needs to build.

Todd Watts

Software engineer and fractional CTO through Shell Command, LLC. Technical direction, architecture and hands-on development.

← All posts