How a WebSocket differs from a plain HTTP request
A normal HTTP proxy is one request, one response: Nginx receives, forwards, waits, returns, and the connection may close. A WebSocket changes the semantics of the connection permanently — after the `Upgrade: websocket` handshake the TCP connection becomes a full-duplex channel and neither side may close it on its own.
That semantic change is the root of every WebSocket proxy problem. `proxy_pass` talks HTTP/1.0 upstream by default, rewrites the `Connection` header to `close`, and has no notion of the Upgrade hop, so the handshake either gets rejected or succeeds and then the connection is immediately treated as idle and closable.
One more detail is easy to miss: after the handshake Nginx no longer interprets the bytes on that connection, it merely shuttles them in both directions. So every default about timeouts and buffering turns into an unexplained user-visible disconnect rather than a clear error in the log.
The right debugging order is therefore: confirm the handshake stage first (the response code must be 101), then measure how long the established connection survives, and only then look at sticky-session problems in your application layer.
curl -i -N -H "Connection: Upgrade" -H "Upgrade: websocket" \
-H "Sec-WebSocket-Version: 13" -H "Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==" \
http://127.0.0.1:8080/ws
# 期望看到:HTTP/1.1 101 Switching ProtocolsForwarding the Upgrade headers: the minimal working setup
The documented approach adds two directives inside the `location`: `proxy_set_header Upgrade $http_upgrade;` and `proxy_set_header Connection "upgrade";`. The first forwards the client header verbatim, the second forces the Connection header to upgrade.
The documentation also offers a refinement: derive `Connection` from a `map` so it falls back to `close` when the client sent no Upgrade header. That lets one `location` serve both ordinary HTTP requests and WebSocket handshakes instead of maintaining two configurations.
The most frequent mistake is assuming `proxy_pass` plus `proxy_http_version 1.1` is enough. The version directive alone does not add `Connection: upgrade`, so the backend sees a plain GET, answers 200 instead of 101, and the handshake fails. A second common omission is `proxy_set_header Host $host`, which breaks virtual-host matching on the backend.
There is a subtler trap: `$http_upgrade` is the raw incoming header, and a client sending `WebSocket` with a capital W can trip strict case-sensitive matching in some frameworks. Using a `map` with an explicit lowercase default makes the configuration deterministic.
# http 级:集中管理 Upgrade 映射,避免每个 location 重复写
map $http_upgrade $connection_upgrade {
default upgrade;
'' close;
}
server {
listen 80;
location /ws {
proxy_pass http://127.0.0.1:3000;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $connection_upgrade;
proxy_set_header Host $host;
proxy_read_timeout 600s;
proxy_send_timeout 600s;
}
}Multi-instance deployments and sticky sessions
With a single instance everything works. Scale out and new problems appear immediately: a connection receives pushes that do not match what was published, or subscription state vanishes. The cause is that a WebSocket is stateful state — established against instance A, all subsequent messages must come from A, but the load balancer has no idea which client sits on which instance.
There are three remedies, in order of preference. **Sticky routing** via consistent hashing (`ip_hash` or `hash $remote_addr consistent`) always maps the same client to the same instance: simplest to implement and semantically correct. **Application-level shared state** stores subscriptions in Redis so any instance can answer for any client, removing the dependency entirely. **A message bus** broadcasts between instances via Redis Pub/Sub or Kafka.
Note that open-source Nginx has no sticky-cookie build (`nginx-plus` only). If a copied `sticky` directive produces `"sticky" directive is not allowed here`, the configuration is not wrong — the build simply does not support it. Use the `hash` directive for an equivalent effect.
Finally, do not confuse this with the **stream** module used for L4 TCP proxying. Stream configuration differs (`proxy_pass` with `proxy_timeout`, no Upgrade headers needed) and stream naturally allocates per connection, so stickiness is a non-issue there. This article is about application-layer proxying in the HTTP module.
upstream ws_backend {
hash $remote_addr consistent; # 开源版实现粘性会话
server 10.0.0.11:3000;
server 10.0.0.12:3000;
server 10.0.0.13:3000;
}
server {
listen 80;
location /ws {
proxy_pass http://ws_backend;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $connection_upgrade;
proxy_set_header Host $host;
proxy_read_timeout 3600s;
}
}A debugging checklist with verification commands
The first step is always the handshake response code. `curl` with Upgrade headers should return `101 Switching Protocols`, proving the Nginx-to-backend leg is fine and the problem lies after establishment. A 200 or 400 means the handshake itself failed — check the `Upgrade`, `Connection` and `Host` headers.
Second, read the Nginx error log. `upstream timed out` is a timeout, `upstream prematurely closed connection` means the backend closed it (restart, heartbeat expiry, load shedding), and `recv() failed (104: Connection reset by peer)` means the backend reset it. Different codes mean different root causes; do not lump them together as "a proxy problem".
Third, measure how long the connection actually lives. With `wscat` or the browser DevTools Network panel, if the disconnect interval is a tidy 60 or 120 seconds the cause is almost certainly the timeout setting. Fifth, stress the multi-instance case: open many connections in a row, watch for lost subscriptions, and compare against single-instance behaviour.
Finally, lock the configuration in. Validate with `nginx -t`, reload gracefully with `nginx -s reload`, then re-test the handshake with `wscat`. There are only a handful of relevant knobs — Upgrade forwarding, the Host header, timeouts, buffering, stickiness — so freeze them into a template and copy it for every new service.
nginx -t # 校验语法
nginx -s reload # 平滑重载
tail -f /var/log/nginx/error.log | grep -i "ws\|upstream"
# 用 wscat 复测:npx wscat -c ws://127.0.0.1:8080/ws