Skip to main content

Frequently asked questions

Why does Nereus Delay need a Command Topic?​

The Command Topic is the Shard's durable ordered log, not only an intake queue. It orders user Commands and internal events such as Publish Admission, Publish Outcome, expiration, DLQ results, and control markers so the state machine can replay them deterministically.

Why keep RocksDB when the Command Topic is durable?​

The log stores history; RocksDB stores current message state and persistent indexes. Without the materialized database, query, Cancel, Reschedule, due scanning, retry, and restart would require replaying the entire retained log.

Why does each Shard use a separate RocksDB database?​

It aligns the database with the ownership, checkpoint, restore, migration, local-deletion, and failure boundary. The tradeoff is that a Worker must govern shared cache, file descriptors, compaction, and disk across multiple databases.

Why can a producer callback not update state directly?​

A callback races with Cancel, Reschedule, timeout, expiration, and operator actions. Turning it into a Shard Log mutation gives all of those events one Source Position order and keeps replay deterministic.

How are duplicate Commands handled?​

Physical retries reuse the same Prepared Command. The Shard compares commandId + commandHash: the same pair is idempotent, while the same ID with different semantics is rejected as a conflict.

How does Reschedule reject an old callback?​

Reschedule creates a new message generation. Claims, Attempts, and outcomes name their generation, so an old callback cannot overwrite the current schedule.

What if a Worker crashes during delivery?​

Producer I/O starts only after durable Publish Admission records PUBLISHING. Recovery therefore finds the Attempt and either recovers evidence or preserves UNCERTAIN; it does not assume the request never left the process.

How does Nereus Delay prevent two active owners?​

Activation requires both the source assignment and the matching Oxia Owner Lease. The owner restores, replays, and crosses an activation barrier before applying new Commands. The owner epoch fences local state, while remote side effects still require Attempt identity and evidence handling.

Can one failed destination block a whole Shard?​

It should not. Target work is isolated by Destination Lane, with lane-local readiness, buffers, inflight limits, retry, and circuit state. Command application and healthy Lanes continue unless a genuine source-safety gate fails.

Why use both checkpoints and Shard Log replay?​

A checkpoint provides fast physical restoration. Replay applies the ordered facts after its Source Position. A checkpoint alone misses later events; log-only recovery becomes increasingly expensive.

Why is Object Store not the live state database?​

Current delayed-message state needs frequent point updates and ordered range scans. Object Store is used for immutable large payloads and complete checkpoint objects, while RocksDB serves mutable local state.

Does Nereus Delay guarantee exactly-once delivery?​

Not universally. V1 provides deterministic command application, idempotent identities, recoverable Attempts, and bounded at-least-once delivery. Stronger effects require a certified Kafka/Pulsar capability or consumer-side idempotency and remain scoped to that boundary.

Why not use a hierarchical timing wheel as the source of truth?​

An in-memory timing wheel is useful for near-term acceleration but not for long-lived durable state, point cancellation, rescheduling, and disk-loss recovery. V1 keeps the RocksDB Timeline authoritative and allows caches to remain replaceable.

Does submit success mean the destination message was delivered?​

No. A managed queued receipt means the ingress Broker durably accepted a Command. The application must distinguish queued, applied, published, and uncertain outcomes.

What does Pulsar delayed handoff optimize?​

It moves network and persistence work before deliverAt and leaves the final bounded wait to certified Pulsar delayed delivery. It does not weaken the not-before boundary or turn native handoff into a managed lifecycle.

Is V1 complete?​

Yes. V1 has completed its protocol, integration, correctness, no-early, capacity, soak, upgrade, operations, and distribution gates. Subsequent work is optimization, refactoring, review, and further evolution. Environment-specific deployment approval remains a separate operational decision.