Two different kinds of protection
limit_conn caps the number of concurrent connections. Because one client may hold many connections, it protects against a single source tying up workers or a backend connection pool. The quota is released only when the connection closes.
limit_req caps request rate, counting requests per unit time regardless of connection reuse. It is the right tool for APIs and search endpoints facing bursty peaks. The two are complementary: connections govern resource occupancy, requests govern flow rate.
Both directives require a shared memory zone. The size after the zone name determines how many states fit; roughly 1MB holds 4000 to 8000 entries. When the table fills, old entries are evicted, which can let traffic through slightly early. That is a precision loss rather than a failure.
http {
limit_req_zone $binary_remote_addr zone=api_perip:10m rate=10r/s;
limit_conn_zone $binary_remote_addr zone=api_conn:10m;
server {
location /api/ {
limit_req zone=api_perip burst=20 nodelay;
limit_conn api_conn 20;
}
}
}Choosing between burst and nodelay
burst sets the queue length for excess traffic: requests above the rate wait in the queue instead of being rejected outright. Without burst, anything above the rate is rejected, and even a modest spike can lock out legitimate users.
By default queued requests wait at the configured rate, so clients feel latency climb. With nodelay, requests admitted into the burst queue are forwarded immediately without waiting, at the cost of transient concurrency exceeding what the rate implies.
A workable rule: use nodelay for short-cycle endpoints so bursts pass quickly, and keep queueing for long-running work to protect the backend. Size burst as an allowed multiple of the rate; with rate=10r/s and burst=20 you tolerate 2x for a moment, and anything beyond 30 in that instant returns 429.
# 突发 2 倍,立即放行:适合短请求 API
limit_req zone=api_perip burst=20 nodelay;
# 突发排队等待:适合慢接口/长任务,保护后端
limit_req zone=slow_perip burst=10;
# 按接口 key 分别计数,而不是只按 IP
limit_req_zone $server_name$request_uri zone=api_route:10m rate=5r/s;Why a wrong zone silently disables throttling
limit_req_zone and limit_conn_zone must live in the http block because they maintain state shared across locations and workers. Defining them in server or location blocks makes Nginx fail at startup, so this particular mistake is loud rather than silent.
The real trap is the opposite case: a mistyped zone name in the location, or omitting the zone argument entirely. Older releases allowed no zone, keeping state inside a single worker, so with worker_processes greater than 1 each worker counted separately and the effective threshold was multiplied by the worker count.
Modern Nginx (since 1.17.6) no longer supports that form without a shared zone. When debugging, first confirm the effective worker count, then observe decisions through the $limit_req_status variable: PASSED for admitted, REJECTED for throttled, DELAYED for requests waiting in the queue.
log_format limits "$remote_addr req=$request_uri status=$status "
"limit=$limit_req_status conn=$limit_conn_status rt=$request_time";
access_log /var/log/nginx/limits.log limits;
worker_processes 4;
nginx -t && nginx -s reloadDebugging a 429 spike
First establish who is rejecting. A CDN or an upstream WAF may be returning 429 before Nginx ever sees the request. Confirm via the response body and $limit_req_status in the log format: only REJECTED attributes the rejection to Nginx rate limiting.
Then decide whether it is configuration or real traffic growth. Cross-tabulate $limit_req_status against $remote_addr: rejections concentrated on a few addresses point to a crawler or a single client, which can be handled with per-API-key or keyed limits; uniform rejections mean the rate is genuinely too tight.
Finally verify that the pass policy is sensible. Common improvements are giving login and search endpoints their own looser zone, exempting static assets with a separate location, and setting limit_req_status and limit_conn_status explicitly so clients receive a deliberate 429 rather than whatever the default happens to be.
limit_req_status 429;
limit_conn_status 429;
location /static/ { # 静态资源不限流
limit_req off;
limit_conn off;
}
location /login/ { # 登录接口单独放宽
limit_req zone=login_perip burst=10 nodelay;
}Rate limiting behind a reverse proxy
When Nginx reverse proxies, $remote_addr is the address of the upstream load balancer or CDN, so every client collapses into one source and the whole site shares a single quota. This is the number one reason limiting behind a proxy appears broken or hits the wrong users.
Fix it with the realip module: set_real_ip_from for the trusted internal ranges, real_ip_header set to CF-Connecting-IP or X-Forwarded-For, and real_ip_recursive on to walk the trusted chain. Keep the trusted range narrow. Writing 0.0.0.0/0 lets anyone spoof a source address and bypass the limit entirely.
After editing, verify with `nginx -t` and reload, then log $remote_addr next to $http_x_forwarded_for to confirm you are seeing the real client. If a CDN sits in front, remember its own rate limiting is a separate layer, and the two together usually mean the CDN absorbs the coarse-grained rejections.
http {
set_real_ip_from 10.0.0.0/8;
set_real_ip_from 172.16.0.0/12;
real_ip_header X-Forwarded-For;
real_ip_recursive on;
limit_req_zone $binary_remote_addr zone=api_perip:10m rate=10r/s;
}
log_format realip "$remote_addr fwd=$http_x_forwarded_for rl=$limit_req_status";