A common community question is whether end-user clients should connect directly to NATS, or whether every deployment should put a custom gateway in front of the NATS servers.
The short answer: direct client connections can be a reasonable design, especially for persistent, low-message-rate, low-latency delivery. A gateway is not automatically required just because the clients are end users or because the connection count is large. However, a direct model makes authentication, authorization, reconnect behavior, and subject design part of your core system architecture.
Consider an APNS-like service where each user keeps a persistent connection open and receives occasional messages, perhaps around ten per day. Delivery should happen as soon as possible, with a small amount of processing or batching delay acceptable.
For that kind of workload, keeping clients connected is usually the right direction if you need near-immediate delivery. NATS is designed for many concurrent connections and low-latency message delivery, but the exact capacity of a deployment depends on hardware, topology, TLS and authentication settings, number of subscriptions, message size, message rate, and operational limits. Treat 100,000 concurrent clients as a capacity-planning exercise, not as a reason by itself to introduce a gateway.
Direct connections are often a good fit when:
If your backend already tracks each client’s last delivered message and can republish missed messages after reconnect, you may not need JetStream persistence for every user-facing payload. Core NATS can handle online delivery, while your application-level state handles gap detection and replay.
That said, a gateway may still be valuable for reasons unrelated to NATS scalability.
A gateway can be the right choice when you need to:
The tradeoff is that a gateway becomes another distributed system component: it must be deployed, scaled, monitored, upgraded, tested, and made highly available. If the gateway only forwards messages one-for-one between clients and NATS, it may add operational complexity without solving the hardest parts of the design.
The biggest issue in this kind of system is often not normal-state delivery. It is the thundering herd after a restart, network outage, deploy, DNS issue, regional failover, or load balancer event.
NATS clients and servers can generally handle reconnect behavior, subject to normal capacity planning and configuration. But NATS cannot automatically make your application-level reconnect handshake cheap. If 100,000 clients reconnect and each immediately asks your backend to compute missed messages, your backend has to absorb that work somehow.
Practical mitigations include:
A plain core NATS queue subscription can distribute handshake requests across backend workers, but it does not by itself provide durable buffering. If no worker is subscribed when a request is published, the message is simply not delivered, and if a worker receives a request and then fails before completing it, there is no automatic redelivery. If you need burst absorption and retry semantics, consider placing handshake requests into a JetStream stream and having workers consume them as a work queue. With a work-queue stream, each request is retained until a consumer acknowledges it, and a request whose worker fails before acknowledging can be redelivered. That lets reconnect work drain at a rate your backend can sustain.
One possible design looks like this:
This keeps the latency path simple while allowing the expensive reconnect path to be controlled.
Per-session subjects can be useful, but there are two common pitfalls.
First, subject names are not security boundaries by themselves. Use NATS authorization to restrict what each client can publish and subscribe to. Do not rely on an unguessable subject token as the only protection.
Second, queue groups do not guarantee that all messages for a session go to the same worker unless you design for that. If several workers share the same queue subscription, NATS will load-balance messages across them. That is excellent for distributing independent work, but it is not session affinity.
If one specific worker must handle all messages for a session, you need an explicit routing strategy. For example, after the handshake, the selected worker could instruct the client to publish session traffic to a worker-specific subject that only that worker consumes. The downside is that you then need a plan for worker failure, session migration, and cleanup. In many systems, it is simpler to keep workers stateless and store session state in a shared backend, so any worker can process the next message.
Also consider using separate subjects for each direction, such as client-to-worker and worker-to-client, rather than using one bidirectional subject. Separate subjects make permissions clearer and reduce accidental message loops or self-delivery surprises.
If end-user clients connect directly, security and tenancy should be designed up front:
For multi-tenant products, NATS accounts and carefully scoped permissions are important tools for isolating subjects and capabilities. The exact credential model depends on how you provision users and sessions.
Not necessarily for every user-facing message.
If clients stay connected and your application already has durable knowledge of what each user has received, core NATS may be enough for online delivery. On reconnect, the client can report the last message it received, and your backend can republish any missing data.
JetStream becomes more attractive when you need NATS itself to provide durability, replay, backpressure, or buffering. In this design space, a common compromise is:
That avoids creating a durable consumer for every client while still giving the backend a buffer during reconnect storms.
Direct end-user connections to NATS are not inherently a bad idea. For a persistent, low-message-rate, latency-sensitive system, they can be simpler than building a gateway, as long as clients are tightly authorized and reconnect behavior is engineered deliberately.
Start with the direct model if your clients can speak NATS and your security model supports scoped credentials. Add a gateway when you need protocol translation, stronger edge controls, or a product boundary that NATS clients should not cross. In either model, design the reconnect path carefully: buffer handshakes, drain work through workers, make replay idempotent, and test outage recovery as a first-class scenario.
Want help from the NATS experts? Meet with our architects to get help tailored to your use case and environment.
News and content from across the community