The upstream block: where distribution actually happens
Writing just `proxy_pass http://backend;` and calling it load balancing is a common misconception: `backend` must first be declared as an `upstream` inside the `http` block. Only an upstream name can carry weight, failure parameters, connection reuse, and health checks. A raw IP:port makes each `proxy_pass` an isolated unit, and Nginx will not even reuse the backend connection.
An `upstream` block belongs only at `http` level — not inside `server`, not inside `location`, and not at main level. This placement rule is one of the most frequent beginner syntax errors, surfacing as `"upstream" directive is not allowed here` from `nginx -t`.
Servers inside an upstream can be hostnames rather than only IPs. Nginx resolves them once at startup and caches the result, so the config does not follow DNS changes. Supporting dynamic service discovery means either Nginx Plus `resolve`, or community-mode `resolver` plus a variable-form proxy_pass (which costs you most upstream capabilities), or simply pointing at a Service VIP.
Another underappreciated fact: Nginx proxies at the HTTP level, and `proxy_pass` names a logical backend — Nginx is not doing L4 balancing here. The default is `proxy_http_version 1.0` with no keepalive, meaning one connection per request, which is exactly why the keepalive directives below matter.
http {
upstream backend {
server 10.0.1.11:8080 weight=3 max_fails=3 fail_timeout=30s;
server 10.0.1.12:8080 weight=3 max_fails=3 fail_timeout=30s;
# 备份节点:只有所有主节点都不可用时才启用
server 10.0.1.20:8080 backup;
keepalive 64;
}
server {
listen 80;
location / {
proxy_pass http://backend;
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}
}The three default balancing methods and their distribution traps
The default is round robin, handing each request to the next backend. It weights every request equally — whether it is a 100-byte health check or an 8MB file download. Under keepalive connections this breaks badly: a client reusing one connection sends all its requests to a single backend, and real load skews hard.
Weighted round robin (`weight=N`) gives a proportional share of requests and is the most common, most predictable choice. Note that weight affects assignment probability only, not service time: a weight-5 backend that is five times slower per request may easily be far more loaded than a weight-1 healthy node. Weights double as traffic-splitting levers for canary rollouts.
`ip_hash` pins a client IP to one backend, which suits local sessions or in-memory caches. It has two clear costs: an entire cohort sharing one NAT egress lands on a single machine, and adding or removing backends reshuffles the hash so a large share of sessions jump to new nodes and hit cold caches.
Add-on modules provide `least_conn` and `hash $key`. least_conn balances better when request durations vary wildly; `hash` ships in Nginx Plus and needs ngx_http_upstream_hash on community builds. Whichever you choose, validate the real distribution under load — much apparent imbalance comes from keepalive connections left in place or long-tail backend latency.
# 三种策略对比:同一组配置只改一个关键字
upstream by_roundrobin {
server 10.0.1.11:8080;
server 10.0.1.12:8080;
server 10.0.1.13:8080;
}
upstream by_weight {
server 10.0.1.11:8080 weight=5; # 承接 50% 流量
server 10.0.1.12:8080 weight=3;
server 10.0.1.13:8080 weight=2;
}
upstream by_ip {
ip_hash;
server 10.0.1.11:8080;
server 10.0.1.12:8080;
}
# 校验配置后再 reload
nginx -t && nginx -s reloadFailover and retries: max_fails, fail_timeout, proxy_next_upstream
`max_fails=N fail_timeout=T` defines failure detection: within a T-second window, if a backend accumulates more than N failures, Nginx stops sending it new requests (the docs call it "considered unavailable"), and it becomes eligible again once the window expires. This is purely passive — Nginx judges from real forwarding failures and needs no probe.
The prerequisite is `proxy_next_upstream`, which decides "when should this request be retried against the next backend". Its values include `error`, `timeout`, `invalid_header`, `http_500`, `http_502`, `http_503`, `http_504`, `http_403`, `http_404`, and `non_idempotent`; the default is only `error timeout`. Without it, quotas keep hitting a broken node — because never retrying means exactly one failure per request and no path to recovery.
One trap you must know: retrying non-idempotent requests (POST, PUT, DELETE) can duplicate writes. To include those status codes you must explicitly add `non_idempotent`, otherwise Nginx will not retry them. If your backend cannot make its endpoints idempotent, split write paths into their own `location` blocks using an instruction without `non_idempotent`, so retries only ever happen on read paths.
The failure counter tracks connection-level failures only, not business-level semantics. Worse, Nginx judges health against the address in its config, which may diverge from what clients experience — a backend process that is alive but no longer able to serve still looks healthy. Such "zombie alive" nodes need active health checks, which on open-source Nginx means Nginx Plus `health_check`, a third-party module, or a component with its own active prober in front of the upstream.
upstream backend {
server 10.0.1.11:8080 max_fails=3 fail_timeout=30s;
server 10.0.1.12:8080 max_fails=3 fail_timeout=30s;
}
server {
listen 80;
# 只读路径:允许对多种错误重试到下一台
location /api/ {
proxy_pass http://backend;
proxy_next_upstream error timeout http_502 http_503 http_504;
proxy_next_upstream_tries 2;
proxy_next_upstream_timeout 10s;
}
# 写路径:明确不重试,避免重复写入
location ~ ^/api/(orders|payments)/ {
proxy_pass http://backend;
proxy_next_upstream off;
proxy_connect_timeout 3s;
proxy_read_timeout 30s;
}
}Upstream keepalive: the payoff, the cost, and the correct setup
Upstream keepalive lets Nginx reuse TCP connections to backends, skipping the handshake and slow start on every request. For high-concurrency short-request services that shows up immediately in P99. But three conditions must all hold, and missing any one of them makes the directive a no-op.
First, the `upstream` block needs `keepalive N;` to size the pool. Second, the location needs `proxy_http_version 1.1;` — the 1.0 default can only close the connection after the response, so there is nothing to reuse. Third, `proxy_set_header Connection "";` — you must clear the default `Connection: close` header, and an empty value tells the upstream this connection may persist.
The cost lands on tail latency. While a pooled connection is busy serving a 30-second request it cannot serve anything else until the keepalive timeout (60s by default) elapses, so one slow request removes capacity from the pool. Keep the keepalive window coordinated with the backend `keepalive_timeout` and ensure single-request durations stay well inside it.
Another widespread myth is that a larger keepalive pool is always better. Pool size should match how many requests the backend can process concurrently; exceeding that only piles up connections, grows memory footprint, and can trip the backend's own connection limits. A practical starting point is that `keepalive 64` multiplied by actual worker concurrency must stay under the backend's idle-connection ceiling.
http {
upstream backend {
server 10.0.1.11:8080;
server 10.0.1.12:8080;
keepalive 64; # 每个 worker 的空闲连接数
keepalive_timeout 60s; # 空闲连接保留时长(upstream 级,1.15.3+)
}
server {
location / {
proxy_pass http://backend;
# 以下三行缺一不可,否则 keepalive 不生效
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_set_header Host $host;
}
}
}Rate limiting and WebSocket: two parameter sets that are easy to get wrong
When WebSocket runs through an Nginx proxy, the connection upgrades from HTTP to a long-lived stream after the handshake. From then on the default `proxy_read_timeout` of 60 seconds kills it while idle, and clients see "reconnects every minute". Fix it by raising `proxy_read_timeout` substantially (3600s, for example), and set `proxy_send_timeout` alongside it in the WebSocket location.
The upgrade headers must also be correct. `proxy_set_header Upgrade $http_upgrade;` passes the client header through, and `proxy_set_header Connection "upgrade";` tells the upstream this is an upgrade. The second header has two hard rules: hardcoding `"upgrade"` stamps upgrade semantics on every request including plain HTTP, so in a location serving both protocols use a `map` directive keyed on `$http_upgrade` instead.
Rate limiting uses `limit_req_zone`, which must be declared in the `http` block. The key is typically `$binary_remote_addr` (binary form saves memory) or, for authenticated users, something derived from the token. Avoid `=` and similar characters in the zone name, size the zone by expected key count, apply `limit_req` inside the location, and use `burst` for short spikes with `nodelay` to serve burst requests immediately instead of queueing them.
WebSocket adds one more constraint: `limit_conn` counts **connections**, not requests, and a long-lived connection holds its slot the whole time. Sizing `limit_conn` for a WebSocket location must be based on concurrent users rather than QPS, or legitimate online users get 503s for exceeding a limit that was never meant for them.
http {
# WebSocket / 普通请求分开处理 Connection 头
map $http_upgrade $connection_upgrade {
default upgrade;
'' close;
}
limit_req_zone $binary_remote_addr zone=api_zone:10m rate=20r/s;
limit_conn_zone $binary_remote_addr zone=conn_zone:10m;
server {
listen 80;
location /ws {
proxy_pass http://backend;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection $connection_upgrade;
proxy_set_header Host $host;
proxy_read_timeout 3600s;
proxy_send_timeout 3600s;
# 长连接按「同时在线数」限制,而不是 QPS
limit_conn conn_zone 200;
}
location /api/ {
limit_req zone=api_zone burst=40 nodelay;
proxy_pass http://backend;
proxy_http_version 1.1;
proxy_set_header Connection "";
}
}
}