Scheduler, fairness, and backpressure
The scheduler does more than find timestamps. It must discover durable work, prevent one target from monopolizing a Worker, and keep overload from consuming the reserves needed for correctness and control.
Persistent Timeline
Each Delay Shard stores its due index in RocksDB. The Timeline is therefore recoverable and supports long delays, cancellation, rescheduling, and bounded scans without keeping every future message in memory.
Trusted time decides eligibility. A message becomes due only when the safe time boundary permits its action. Being due does not itself create a publish attempt; Claim and durable Publish Admission remain separate steps.
Destination Lane
A Destination Lane groups work that shares destination, tenancy, ordering, capability, retry, and capacity boundaries. Each Lane owns its own readiness, inflight limits, circuit state, and adapter admission.
If a destination is unavailable or a Lane is saturated:
- that Lane can stop Claim or Publish Admission;
- healthy Lanes on the same Shard can continue;
- Command application continues unless a real source-safety gate fails.
A Lane is not a Shard, an owner, or a checkpoint unit.
Two-level bounded DRR
V1 uses bounded deficit round robin at two levels:
- choose among ready Shards hosted by the Worker;
- choose among ready Destination Lanes inside the selected Shard.
Weights express relative service share, while per-turn bounds prevent one large backlog from occupying the execution loop indefinitely. Under the certified healthy-load envelope, a continuously ready healthy Lane must receive service within registered discovery and service-gap bounds.
This is a fairness contract, not a promise that every message has the same latency. Ordering heads, target throttling, retry delay, clock gating, and capability drift remain visible reasons for later admission.
Work-class and capacity isolation
Source apply, due discovery, Publish Outcome application, checkpointing, GC, query, control, and destination I/O do not share one unbounded queue. They have bounded work classes and separate reserves.
Admission accounts for more than message count. Relevant dimensions include bytes, inflight attempts, open channels, file descriptors, local disk, RocksDB pressure, Object Store operations, and durable SLO evidence capacity. Control and outcome reserves remain available even when ordinary scheduling is full, so the system can drain, record outcomes, or close unsafe work.
What to monitor
- due-but-not-admitted counts by closed reason;
- ready-Lane discovery delay and DRR service gap;
- per-Lane pending and inflight messages/bytes;
- circuit state, retry age, and oldest uncertainty;
- Worker disk, file descriptors, compaction, cache, and adapter channel limits;
- source apply lag separately from destination scheduling lag.
See Destination Lane isolation and observability and release gates.