Web Services

WebSocket Reverse Proxy: Nginx Configuration and Timeout Traps

WebSocket connections keep dropping behind Nginx, handshakes return 400, or messages vanish in a cluster? Start from the minimal Upgrade header setup, then fix silent disconnects caused by the 60s proxy_read_timeout default, sticky sessions, and a copy-paste debugging checklist.

By LaoHand Team·9 min read·Updated 2026-09-30

How a WebSocket differs from a plain HTTP request

A normal HTTP proxy is one request, one response: Nginx receives, forwards, waits, returns, and the connection may close. A WebSocket changes the semantics of the connection permanently — after the `Upgrade: websocket` handshake the TCP connection becomes a full-duplex channel and neither side may close it on its own.

That semantic change is the root of every WebSocket proxy problem. `proxy_pass` talks HTTP/1.0 upstream by default, rewrites the `Connection` header to `close`, and has no notion of the Upgrade hop, so the handshake either gets rejected or succeeds and then the connection is immediately treated as idle and closable.

One more detail is easy to miss: after the handshake Nginx no longer interprets the bytes on that connection, it merely shuttles them in both directions. So every default about timeouts and buffering turns into an unexplained user-visible disconnect rather than a clear error in the log.

The right debugging order is therefore: confirm the handshake stage first (the response code must be 101), then measure how long the established connection survives, and only then look at sticky-session problems in your application layer.

curl -i -N -H "Connection: Upgrade" -H "Upgrade: websocket" \
     -H "Sec-WebSocket-Version: 13" -H "Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==" \
     http://127.0.0.1:8080/ws
# 期望看到:HTTP/1.1 101 Switching Protocols

Forwarding the Upgrade headers: the minimal working setup

The documented approach adds two directives inside the `location`: `proxy_set_header Upgrade $http_upgrade;` and `proxy_set_header Connection "upgrade";`. The first forwards the client header verbatim, the second forces the Connection header to upgrade.

The documentation also offers a refinement: derive `Connection` from a `map` so it falls back to `close` when the client sent no Upgrade header. That lets one `location` serve both ordinary HTTP requests and WebSocket handshakes instead of maintaining two configurations.

The most frequent mistake is assuming `proxy_pass` plus `proxy_http_version 1.1` is enough. The version directive alone does not add `Connection: upgrade`, so the backend sees a plain GET, answers 200 instead of 101, and the handshake fails. A second common omission is `proxy_set_header Host $host`, which breaks virtual-host matching on the backend.

There is a subtler trap: `$http_upgrade` is the raw incoming header, and a client sending `WebSocket` with a capital W can trip strict case-sensitive matching in some frameworks. Using a `map` with an explicit lowercase default makes the configuration deterministic.

# http 级:集中管理 Upgrade 映射,避免每个 location 重复写
map $http_upgrade $connection_upgrade {
    default upgrade;
    ''      close;
}

server {
    listen 80;
    location /ws {
        proxy_pass http://127.0.0.1:3000;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection $connection_upgrade;
        proxy_set_header Host $host;
        proxy_read_timeout 600s;
        proxy_send_timeout 600s;
    }
}

The 60s proxy_read_timeout default: the most hidden disconnect

Nginx `proxy_read_timeout` defaults to **60 seconds**, meaning the upstream connection is closed if the upstream sends Nginx no data at all for that long. For ordinary HTTP this is irrelevant because responses end quickly. For WebSocket it is a timer bomb: any connection idle for 60 seconds gets closed unilaterally.

The classic symptom is "everything works, then it drops — very regularly". The log usually shows `upstream timed out (110: Connection timed out) while reading upstream`, and clients report a disconnect about once a minute. Users blame the application; it is only the timeout.

The fix has two layers. First, raise `proxy_read_timeout` explicitly (300s to 3600s is typical) and set `proxy_send_timeout` for the opposite direction. Second and more fundamentally, have the application send heartbeat pings. The Nginx timer resets on any byte received, so a ping every 30 seconds keeps the connection alive indefinitely regardless of the configured timeout.

Also consider `proxy_buffering`. With buffering on, Nginx accumulates upstream data before forwarding it downstream, which adds latency for small low-latency WebSocket frames and interacts with timeout evaluation in ways that are hard to reason about. WebSocket proxying normally sets `proxy_buffering off` explicitly.

location /ws {
    proxy_pass http://127.0.0.1:3000;
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection $connection_upgrade;
    proxy_set_header Host $host;
    proxy_buffering off;        # WebSocket 关闭缓冲,避免额外延迟
    proxy_read_timeout 3600s;    # 默认 60s 必须调大
    proxy_send_timeout 3600s;
}

Multi-instance deployments and sticky sessions

With a single instance everything works. Scale out and new problems appear immediately: a connection receives pushes that do not match what was published, or subscription state vanishes. The cause is that a WebSocket is stateful state — established against instance A, all subsequent messages must come from A, but the load balancer has no idea which client sits on which instance.

There are three remedies, in order of preference. **Sticky routing** via consistent hashing (`ip_hash` or `hash $remote_addr consistent`) always maps the same client to the same instance: simplest to implement and semantically correct. **Application-level shared state** stores subscriptions in Redis so any instance can answer for any client, removing the dependency entirely. **A message bus** broadcasts between instances via Redis Pub/Sub or Kafka.

Note that open-source Nginx has no sticky-cookie build (`nginx-plus` only). If a copied `sticky` directive produces `"sticky" directive is not allowed here`, the configuration is not wrong — the build simply does not support it. Use the `hash` directive for an equivalent effect.

Finally, do not confuse this with the **stream** module used for L4 TCP proxying. Stream configuration differs (`proxy_pass` with `proxy_timeout`, no Upgrade headers needed) and stream naturally allocates per connection, so stickiness is a non-issue there. This article is about application-layer proxying in the HTTP module.

upstream ws_backend {
    hash $remote_addr consistent;   # 开源版实现粘性会话
    server 10.0.0.11:3000;
    server 10.0.0.12:3000;
    server 10.0.0.13:3000;
}

server {
    listen 80;
    location /ws {
        proxy_pass http://ws_backend;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection $connection_upgrade;
        proxy_set_header Host $host;
        proxy_read_timeout 3600s;
    }
}

A debugging checklist with verification commands

The first step is always the handshake response code. `curl` with Upgrade headers should return `101 Switching Protocols`, proving the Nginx-to-backend leg is fine and the problem lies after establishment. A 200 or 400 means the handshake itself failed — check the `Upgrade`, `Connection` and `Host` headers.

Second, read the Nginx error log. `upstream timed out` is a timeout, `upstream prematurely closed connection` means the backend closed it (restart, heartbeat expiry, load shedding), and `recv() failed (104: Connection reset by peer)` means the backend reset it. Different codes mean different root causes; do not lump them together as "a proxy problem".

Third, measure how long the connection actually lives. With `wscat` or the browser DevTools Network panel, if the disconnect interval is a tidy 60 or 120 seconds the cause is almost certainly the timeout setting. Fifth, stress the multi-instance case: open many connections in a row, watch for lost subscriptions, and compare against single-instance behaviour.

Finally, lock the configuration in. Validate with `nginx -t`, reload gracefully with `nginx -s reload`, then re-test the handshake with `wscat`. There are only a handful of relevant knobs — Upgrade forwarding, the Host header, timeouts, buffering, stickiness — so freeze them into a template and copy it for every new service.

nginx -t                                   # 校验语法
nginx -s reload                              # 平滑重载
tail -f /var/log/nginx/error.log | grep -i "ws\|upstream"
# 用 wscat 复测:npx wscat -c ws://127.0.0.1:8080/ws

Official References

Each command links to its official documentation below, so you can verify the latest usage and read deeper.