Skip to main content

Retry, DLQ, and replay

Retry is a durable policy decision, not an unbounded loop around producer.send(). Every new destination attempt must be eligible, admitted, and recoverable inside the message's pinned policy and lifetime.

Retryable failure​

A definitely-not-published transient result may schedule another attempt when retry budget, delivery lifetime, Lane state, and capacity allow it. Backoff becomes durable scheduling state so a restart does not forget how the next attempt was chosen.

Permanent errors, exhausted budgets, or expired admission windows do not keep retrying forever. They move the generation toward its registered terminal path.

Uncertain is not a normal retry signal​

UNCERTAIN means the target may have accepted the side effect. Retrying immediately can create a duplicate, so the system preserves the attempt and its evidence boundary. Resolution may come from capability-specific evidence recovery, an explicit operator decision, or a later workflow that knowingly accepts duplicate risk.

The uncertainty state is not relabeled as a destination failure merely to simplify dashboards.

Lane-level failure​

Repeated target failures can open the affected Lane's circuit and stop new Claims or admissions for that Lane. The Shard continues applying Commands and serving healthy Lanes. When prerequisites recover, the Lane re-enters readiness through its capability and resource checks rather than a blind timer flip.

Dead-letter state​

DLQ is a durable internal message outcome with export evidence. It is not defined only by whether a best-effort write reached an external topic. The Shard records why the message entered DLQ, the generation and attempt involved, and the result of any configured export.

This keeps the authoritative state queryable even if the external DLQ destination is temporarily unavailable.

Replay is explicit​

Replay creates a new controlled delivery generation or operation under a registered policy. It does not erase the old terminal result, mutate the original attempt, or quietly requeue every record from an external DLQ topic.

Administrative replay and uncertainty resolution use authenticated, idempotent Control Operations. Their registration, per-Shard application, and overall completion are separate stages, all ordered through the relevant Shard Log.

Operator questions​

Before retrying or replaying, answer:

  • Is the prior outcome definitely not published, or merely uncertain?
  • Is the original generation still inside its allowed policy and lifetime?
  • Is the Destination Lane ready with the same capability and resource identity?
  • Does the action create a new generation, and can the consumer tolerate a duplicate?
  • Is the action durably registered and auditable?

See delivery guarantees and uncertainty and query and read barrier.