DZ

Writing

Scaling Long-Running Operations: Distributed Event Routing, Protocol Constraints, and Resilient UI Architecture

HTTP/2, a pub/sub broker, and an authoritative job snapshot are what make Server-Sent Events safe for a long-running workflow. Polling stays the stable choice when those are not in place.

Moving a long-running workflow (such as heavy data recalculations) from client-driven adaptive polling to server-driven updates is frequently treated as a minor frontend quality-of-life refactor. In reality, shifting the delivery mechanism from polling to Server-Sent Events (SSE) changes the core characteristics of the system's network utilization, memory layout, and state topology.

When evaluated at scale, this architectural shift requires navigating severe protocol bottlenecks, managing a pivot from high request velocity to high connection concurrency, and ensuring state consistency across transient transport layers.

1. The protocol bottleneck: HTTP/1.1 vs. HTTP/2 multiplexing

From a UI and application architecture perspective, the single biggest risk of deploying SSE is the browser's underlying transport layer.

Under the HTTP/1.1 protocol, browsers enforce a strict limit of 6 concurrent open connections per domain. Because an SSE connection is a sustained, indefinite HTTP request, opening a single application tab locks one of these 6 channels.

If a user utilizes tabbed browsing or opens multiple application modules simultaneously:

  • The risk. Opening 6 tabs completely starves the browser's network pool for that domain. All subsequent REST API fetches, asset downloads (.js, .css), and analytics beacons will be queued indefinitely, rendering the UI unresponsive.
  • The requirement. SSE should never be deployed to production without guaranteeing HTTP/2 or HTTP/3 end-to-end (from the browser through to the application origin). HTTP/2 uses multiplexing, allowing a single TCP connection to concurrently handle the persistent SSE stream alongside thousands of standard REST requests.

2. Infrastructure scale math: request velocity vs. connection concurrency

Switching to server-driven updates does not magically erase infrastructure load; it shifts the bottleneck from the CPU/network bandwidth layer to the memory/OS file descriptor layer.

Adaptive polling

High request-per-second spikes

CPU / stateless bandwidth bound

Server-Sent Events

High persistent connections

Memory / file descriptor bound

Adaptive polling load profile

The load is stateless. If 50,000 active users poll every 10 seconds, the load balancers and API layers must process a high volume of requests per second. However, these connections are closed immediately upon delivering the payload, allowing infrastructure to scale horizontally using standard stateless auto-scaling metrics.

SSE load profile

The load is stateful and persistent. With 50,000 active users, the infrastructure must maintain 50,000 concurrent, open TCP sockets.

  • OS constraints. Each open connection consumes an operating system file descriptor and requires a persistent memory heap allocation on the target node.
  • Load balancer tunnels. Reverse proxies and load balancers (like Nginx or AWS ALB) must be specifically tuned. Aggressive write/idle timeouts will prematurely sever the stream, causing reconnection storms, while too-lenient settings can lead to memory exhaustion across the ingress layer.

3. Distributed event routing in horizontally scaled topologies

In a modern cloud architecture, stateless API endpoints sit behind an ingress controller or load balancer. This introduces a structural mismatch: the user's SSE socket is anchored to a specific API instance, but the calculation job executes asynchronously across a separate worker pool.

A load balancer fans out to three API nodes. Node 1 holds the UI stream and node 3 executes the job. Both meet at a pub/sub broker.

Load balancer

API node 1

Holds the UI stream

API node 2

API node 3

Executes the job

Pub/sub broker

Redis, RabbitMQ, or Kafka

To bridge this gap without breaking stateless design principles, introduce a decoupled pub/sub broker (Redis, RabbitMQ, or Kafka):

  1. Connection anchorage. The client establishes a persistent SSE connection with a random node (API node 1). That node immediately subscribes to a message channel mapped to the user or job ID on the pub/sub broker.
  2. Asynchronous execution. The heavy calculation executes isolated on a dedicated background worker or a different node (API node 3).
  3. Event propagation. As API node 3 mutates state, it publishes progress frames to the pub/sub broker. The broker broadcasts the message to all active subscribers, so API node 1 receives the frame and pipes it down the open HTTP stream to the browser.

4. State topology: volatile streams vs. authoritative recovery

An event stream is an excellent delivery mechanism, but it is a highly volatile source of truth. Network micro-outages, corporate proxies, VPN switches, and browser sleep states frequently drop TCP connections.

If the application relies solely on the stream to maintain UI state, a disconnected client will miss mid-flight progress packets and end up locked in an inaccurate visual state upon automatic reconnection.

The architecture must separate volatile transport from durable state:

Volatile event stream

SSE as the delivery layer: real-time deltas.

Persistent database or cache

The authoritative layer: the latest snapshot.

The resilient reconnection lifecycle

When the client detects an SSE disconnect, it must execute a multi-phase reconciliation loop:

  1. State fetch. The UI immediately falls back to a standard stateless REST endpoint (GET /api/v1/jobs/:id) to pull the complete, authoritative current snapshot of the operation from the database or cache.
  2. UI synchronization. The frontend state tree is reset using this fresh baseline, so no events were missed during the dead zone.
  3. Stream reconnection. The UI re-initiates the SSE handshake, using the nativeLast-Event-ID header if tracking historical offsets, or re-subscribing cleanly to live events from that point forward.

5. UI thread stability: throttling the event loop

When decoupling the server's execution speed from the client, architects often overlook browser rendering performance. If a backend calculation updates thousands of times per second (granular data processing, for example), streaming every update down the wire can crash the frontend.

Flooding the UI state with high-frequency updates forces constant layout recalculations, causing heavy frame drops, blocked main-thread execution, and a frozen user interface.

Mitigating UI jank

  • Server-side batching. The execution engine or API layer should debounce progress events, emitting updates only on logical steps (every 1% change) or at a fixed time interval (max once every 250ms).
  • Client-side throttling. If the server cannot be restricted, the UI architecture must introduce an ingestion buffer. Incoming stream events should be collected in a queue and batched via arequestAnimationFrame loop or a debounced state hook, preventing the rendering engine from strangling the application.

Direct decision matrix: when to pivot

To choose between these paradigms, use a strict architectural scoring system rather than intuition:

MetricPollingServer-Sent Events
Primary network costHigh request/second overhead; repetitive TLS handshakes.High persistent TCP connection overhead; memory allocation per user.
Proxy and firewall compatibilityTransports over standard short-lived HTTP; virtually zero proxy friction.Corporate firewalls and proxies often sever long-lived streaming connections.
Mobile client battery impactHigh. Constant radio wake-ups for repeated polling cycles.Low. The radio enters a low-power state once the single TCP pipe is established.
Tab proliferationScales linearly (N × tabs requests). Optimized via shared Web Workers.Suffers connection exhaustion under HTTP/1.1; requires HTTP/2.
Real-time latencyGoverned by the poll interval (T/2 average lag).Near zero. Instant push when the backend generates the event.

Summary for the team

Don't move to SSE because it feels like a cleaner paradigm. Move to SSE if the workload demands near-real-time updates, if end-to-end HTTP/2 network routing is guaranteed, and if you are ready to provision the pub/sub backend infrastructure needed to route volatile events across a stateless cluster. Otherwise, optimizing polling intervals or implementing an exponential backoff engine remains the most stable, cost-effective architectural choice.