NEW live workshops: Leaving TIBCO or Solace for NATS Adding an enterprise backbone above MQTT Active/active multi-cloud architectures
All posts

A common community question is how to randomize the order of jobs stored in a JetStream stream. The short answer is: do not try to make JetStream globally shuffle a large stream. If you only need to break strict sequence and improve distribution, randomize over a bounded window instead.

Why global randomization is hard

A JetStream stream is an ordered log. Consumers normally see messages according to stream and consumer delivery semantics, with redelivery behavior added when messages are not acknowledged.

Picking a truly random message from a large, constantly changing backlog is a different problem. To do it fairly across the whole stream, a system needs global state: which messages exist, which have already been delivered, which are pending, how long each message has waited, and how to weight new arrivals against old messages. Maintaining and querying that state can become very expensive for large backlogs.

This is especially true when:

  • messages arrive continuously,
  • producers enqueue faster than consumers process,
  • the backlog is large,
  • the desired randomization spans the entire backlog rather than a small window.

For most job-processing workloads, the practical goal is not mathematically fair random selection. It is usually to avoid processing jobs in the exact order they were published. That lower-quality randomization is much cheaper.

Option 1: randomize into subject buckets

One straightforward pattern is to split jobs across a fixed set of subject buckets at publish time:

jobs.bucket.0
jobs.bucket.1
jobs.bucket.2
...
jobs.bucket.N

The publisher chooses a bucket randomly, and the stream captures those subjects.

For the buckets to actually affect processing order, the consuming side has to read them in a way that interleaves buckets. A single consumer with one wildcard filter such as jobs.bucket.* still delivers every message in strict stream sequence, so bucketing alone does nothing for it. The disorder comes from running a separate subject-filtered consumer per bucket (or per group of buckets) and processing those consumers concurrently at independent rates, or from pulling from per-subject filtered consumers in a randomized order. Each individual bucket is still consumed first-in, first-out; the randomness exists only across buckets.

This approach can work well when you can change the publisher and when the number of buckets is bounded.

Tradeoffs:

  • Randomization happens at publish time, not at consume time.
  • Randomness is over buckets first, not over every individual message in the stream.
  • A very large number of subjects increases memory and operational overhead.
  • More buckets can reduce visible ordering, but millions of tiny buckets are not a free shuffle.

There is not a simple universal subject-count number that is the right limit for every stream. Practically, JetStream keeps per-subject indexing state, and community guidance has estimated this at roughly hundreds of bytes per subject, around 300 bytes per subject, before considering the rest of the system. Treat the bucket count as a capacity and operations decision, not as an unbounded randomization mechanism.

Option 2: use delayed NAK redelivery to randomize a window

If you want publisher-independent randomization, a useful pattern is to consume normally but randomly delay some first deliveries. Many NATS clients expose this as a negative acknowledgement with a delay, often called delayed NAK or NAK with delay.

The idea is to randomize within a window of messages, not across the entire stream.

A typical algorithm looks like this:

  1. Create a durable JetStream consumer with explicit acknowledgements.
  2. Consume messages normally.
  3. For each message, inspect the delivery count (num_delivered) from the JetStream message metadata.
  4. If this is the first delivery:
    • with probability P, process the message immediately and ACK it;
    • otherwise, send a delayed NAK using a random delay interval.
  5. If this is a redelivery, process the message immediately and ACK it.
  6. Monitor the number of unacknowledged or pending messages and tune the delay interval.

In pseudocode:

for each message:
metadata = message.metadata
if metadata.num_delivered == 1:
if random() < process_immediately_probability:
process(message)
ack(message)
else:
delay = random_delay(min_delay, max_delay)
nak_with_delay(message, delay)
else:
process(message)
ack(message)

This breaks the strict delivery sequence because some first deliveries are deferred and later messages can be processed before them. When the delayed message returns, the consumer recognizes it as a redelivery and processes it instead of delaying it again.

Keep the randomization window bounded

The delayed NAK approach works best when the number of delayed, unacknowledged messages stays bounded. A practical rule of thumb is to keep the consumer’s unacknowledged message count below roughly 100,000. This is not a hard protocol limit, and some systems may operate above it, but very large pending sets can slow the consumer and increase overhead.

There is also a concrete consumer setting to account for: MaxAckPending (max_ack_pending). It caps the number of delivered-but-unacknowledged messages a consumer may have outstanding at once. A message that has been NAK’d — including a delayed NAK — is not acknowledged, so it stays outstanding and counts toward this limit until it is redelivered and finally ACKed. When the limit is reached, the server stops delivering new messages until some are acknowledged, which is the main reason an over-delayed consumer effectively slows down.

For explicit-ack consumers MaxAckPending defaults to 1000, which is far below the window sizes discussed here. If you want a large randomization window, you must raise MaxAckPending so it comfortably exceeds your intended window (or set it to -1 to disable the cap, accepting that you then give up that backpressure safeguard). Also review AckWait, which defaults to 30 seconds: it should be longer than your normal processing time so messages are not redelivered while still being processed, and note that an AckWait expiry — not only your delayed NAK — can also trigger a redelivery.

The expected window is related to processing rate and delay interval. If consumers process R messages per second and you delay messages for an average of D seconds, the delayed set can grow toward approximately R * D, depending on the probability of delaying messages and the workload shape.

Use this relationship to choose conservative values:

  • lower the maximum delay when processing throughput is high,
  • reduce the probability of delaying messages when pending counts grow,
  • increase the delay only when the randomized window is too small,
  • set MaxAckPending above the largest window you expect, so the cap acts as a safety net rather than an unexpected throttle.

The goal is to create enough disorder to improve distribution without turning the consumer into a huge redelivery scheduler.

Why not shuffle batches and republish?

Another possible design is a separate service that consumes a batch, shuffles it, and republishes to another stream or subject.

That can be valid when you need a separate randomized output stream, but it adds moving parts:

  • extra storage and publish work,
  • new acknowledgement and failure-handling paths,
  • duplicate or idempotency concerns,
  • decisions about preserving headers, metadata, and ordering expectations,
  • operational responsibility for the shuffler service.

It also only randomizes within the batches it reads unless the service itself keeps a large global buffer. Once the buffer becomes large enough to represent the entire backlog, the design runs into the same global-state problem.

Choosing between the approaches

Use subject buckets when:

  • you can modify publishers,
  • you want simple, durable distribution across partitions,
  • the bucket count can remain modest and predictable,
  • bucket-level randomization is good enough.

Use delayed NAK redelivery when:

  • you cannot or do not want to change publishers,
  • low-quality randomization is enough,
  • you only need to break sequence within a bounded window,
  • you can monitor and control pending message counts.

Use a shuffler service when:

  • downstream systems require a separate randomized stream,
  • you need explicit control over the shuffle buffer,
  • you are prepared to handle the extra reliability and operational complexity.

Practical recommendation

For a high-backlog job queue where the goal is simply to break strict sequence, start with delayed NAK redelivery over a bounded window. Process redelivered messages immediately, monitor pending counts, and keep the delay interval small enough that the unacknowledged set stays manageable.

If you already control publishing and want a simpler model, use a fixed number of random subject buckets. Avoid treating a huge number of subjects as a substitute for a global shuffle; it moves the cost into stream indexing and consumer behavior.

The key design point is to decide what kind of randomness you actually need. Fair random selection across an entire dynamic stream is expensive. Bounded disorder is usually much cheaper and often sufficient for job distribution.


Want help from the NATS experts? Meet with our architects to get help tailored to your use case and environment.

Get the NATS Newsletter

News and content from across the community


Cancel