Step 1: read the state distribution with ss
The first command in any TCP investigation is `ss -s`, which reports counts per state: ESTAB, TIME_WAIT, CLOSE_WAIT, and so on. A TIME_WAIT pile-up and a CLOSE_WAIT pile-up are entirely different failures with opposite causes, and the state counts tell them apart in one second.
`ss -tan state time-wait` lists every TIME_WAIT connection. The flags matter: `-t` TCP, `-a` all states, and `-n` skip service-name resolution — omitting `-n` triggers a reverse DNS lookup per connection, which can add tens of seconds when there are tens of thousands of them. Pipe through `wc -l` to count.
Establishing whether you are the client or the server is essential: TIME_WAIT accumulates on the side that closed first, CLOSE_WAIT on the side that received the close. TIME_WAIT concentrated in outbound connections means your host is a chatty client; TIME_WAIT on thousands of inbound connections is a server-side problem. Same symptom, completely different location.
Aggregating by local port finds the port hog directly: `ss -tan | awk '{print $4}' | sed 's/.*://' | sort | uniq -c | sort -rn | head`. Thousands of connections on a service port, combined with the ESTAB count, tells you whether you are hitting the application backlog or the file-descriptor ceiling.
ss -s
ss -tan state time-wait | wc -l
ss -tan state close-wait | wc -l
ss -tan state time-wait dst 10.0.0.1 | wc -l # 判断方向Port exhaustion: range, reservations, and real causes
Every TCP connection consumes a four-tuple, made unique by source IP, source port, destination IP, and destination port. For a server the ceiling is therefore the product of destination ports times source IPs; for a client opening connections it is source IPs times source ports. These two ceilings differ enormously, and conflating them is why triage goes in the wrong direction.
The default ephemeral range is usually 32768–60999, roughly 28000 ports. That is small for any workload making many outbound connections — crawlers, pooled HTTP clients, message-queue consumers, database connection pools. Raise it with `net.ipv4.ip_local_port_range`, persist it in `/etc/sysctl.d/99-network.conf`, and apply with `sysctl --system`.
When you see `Cannot assign requested address` (EADDRNOTAVAIL), check that parameter first — not disk, not DNS. Verify with `sysctl net.ipv4.ip_local_port_range` and compare against `ss -s`: if the TIME_WAIT count is of the same order as the port range, the diagnosis is essentially confirmed.
One more often-missed knob: `net.ipv4.ip_local_reserved_ports` reserves a range for privileged services so a restart does not collide with ports handed out moments earlier. It does not add capacity, it only keeps the kernel from allocating them. Also note cloud-specific outbound port restrictions: some environments only permit a handful of destination ports, which shows up as a connection that resets right after the handshake.
sysctl net.ipv4.ip_local_port_range
sysctl -w net.ipv4.ip_local_port_range="10000 65000"
echo 'net.ipv4.ip_local_port_range = 10000 65000' > /etc/sysctl.d/99-network.conf
ss -s # 看 TIME_WAIT 是否逼近范围上限What TIME_WAIT really is, and whether to act on it
TIME_WAIT is part of the TCP design, not an error. After sending the final FIN and receiving the ACK, the side that closed actively must linger for 2MSL (60 seconds on Linux by default). The state exists so the last ACK can be retransmitted and stale packets dissipate before a new connection reuses the same four-tuple.
So a large TIME_WAIT count is not itself a fault. Two things matter: whether it costs throughput or memory (the cost on modern kernels is small), and whether it exhausts source ports. If ports are plentiful, hundreds of thousands of TIME_WAIT entries mainly slow down `ss` queries and consume a little kernel memory — usually not worth acting on.
What is worth fixing is the pattern "we are the side that closes first". Database connection pools and HTTP clients that close eagerly pile up TIME_WAIT on the client host. The remedy is to change application behaviour — let the peer close first, or use long-lived pooled connections — not to flip a kernel knob.
Shrinking `net.ipv4.tcp_fin_timeout` is a widely repeated piece of advice that is simply wrong here: it does **not** affect TIME_WAIT (only orphaned FIN_WAIT2 sockets), so lowering it will not reduce your TIME_WAIT count. Knowing this saves an afternoon every time it comes up.
ss -tan state time-wait | wc -l
cat /proc/sys/net/ipv4/tcp_fin_timeout # 与 TIME_WAIT 数量无关
# 排查应用侧:确认是哪个组件频繁主动关闭连接backlog and somaxconn: three queues routinely confused
Three layers of connection queueing must be kept distinct. First, the application accept queue, set by the listen backlog argument. Second, the `net.core.somaxconn` ceiling for that listen socket — the effective value is `min(backlog, somaxconn)`. Third, the SYN queue, governed by `tcp_max_syn_backlog` (in modern kernels the SYN queue is organized by connection hash buckets and rarely the real bottleneck).
The classic misreading is that writing `listen(fd, 4096)` means the queue is 4096 deep, when `somaxconn` may still cap it far lower. When a queue overflows, clients typically see a timeout rather than an explicit refusal, which is why this misconfiguration is hard to diagnose from the client side.
Raising `somaxconn` is trivial, but it only takes effect if the application sets backlog correctly. Modern Linux also allows deferring handshake completion with `TCP_DEFER_ACCEPT`, where the server accepts on SYN without waiting for the full handshake — measurably cheaper under short-lived high-concurrency connections.
One more layer is easy to miss: even with room in the queue, connections can fail under SYN flood or half-open accumulation. `net.ipv4.tcp_syncookies=1` substitutes cookie-encoded state for server-side queue storage when the SYN queue overflows. It is the standard defence against SYN floods, at the cost of some TCP options.
sysctl net.core.somaxconn net.ipv4.tcp_max_syn_backlog net.ipv4.tcp_syncookies
ss -ltn # Send-Q 列即当前 backlog 上限,Recv-Q 为当前排队数
ss -ltn "sport = :8080"The correct boundary for tcp_tw_reuse
`net.ipv4.tcp_tw_reuse=1` lets the kernel reuse TIME_WAIT four-tuples for **outbound** connections only. It has no effect on inbound server connections, since a server never "creates" a connection in that sense. Once enabled, the client skips the 2MSL wait and source-port exhaustion on short-lived connections usually eases.
The boundary matters: reuse requires the new four-tuple to be older than the TIME_WAIT linger, and safety depends on four-tuples not being recycled too quickly, otherwise a stale packet could be mistaken for part of a new connection. In practice, pairing `tcp_tw_reuse` with a shortened `tcp_fin_timeout` is a dangerous combination — the latter does not shorten the TIME_WAIT linger, it only creates the illusion that it does.
There is also a version difference: since Linux 4.12 the parameter only reuses when the peer supports TCP timestamps, which makes it safer in practice because the timestamp acts as a generation marker. It is therefore reasonable on modern kernels, but you should still confirm the symptom is genuine source-port exhaustion rather than a CLOSE_WAIT leak or descriptor exhaustion.
Never confuse it with `tcp_tw_recycle`, which was removed in Linux 4.12 because it broke correct connection tracking by middleboxes behind NAT. The names are similar but the mechanisms are unrelated, and no modern system should have any recycle-related parameter set.
sysctl net.ipv4.tcp_tw_reuse net.ipv4.tcp_tw_recycle
sysctl -w net.ipv4.tcp_tw_reuse=1
uname -r # 4.12+ 语义已变,recycle 已被移除
ss -tan state time-wait | wc -l # 生效后应逐步下降