SysOps

Diagnosing Disk IO Bottlenecks: from iostat to Per-Process Attribution

Slow system at 30% CPU usually means the disk queue is the bottleneck. This guide covers what %util and %await actually mean, why %util 100% is not a verdict, how to read iostat -x, how PSI pressure reveals who is being slowed, and how to attribute wait time to specific processes.

By LaoHand Team·10 min read·Updated 2026-09-30

First: CPU-bound or IO-bound

A quick discriminator: compare load average against CPU utilization in `top`. If load is high (say 8) while `%Cpu(s)` us+sy is far below 100% (say 40%), many processes sit in uninterruptible sleep (D state) waiting on IO. The D-state processes in `top` are the direct evidence.

Better evidence is `/proc/pressure/io`. Read the `some avg10` line: a value near or above 10 means that over the last 10 seconds, an average of 10% of wall time had at least one task stalled on IO. That number tracks perceived sluggishness far better than any single device metric.

High load can also come from non-D artifacts — zombie processes, a fork storm inside containers. So order the checks: count D-state processes, then PSI, then iostat. Only when all three agree should you conclude IO-bound.

If PSI io `some` is low while load stays high, find the load source: `ps -eo state,pid,comm | grep -c "^D"` for an exact count, then `vmstat 1` and watch the `b` column (tasks blocked on IO) alongside `wa`. A persistently positive `b` with high `wa` confirms IO wait rather than CPU contention.

ps -eo state,pid,comm | awk '$1=="D"' | head
grep -E "some|full" /proc/pressure/io
vmstat 1 5      # 关注 wa(iowait)与 b(阻塞数)

Reading iostat -x column by column

%util only means "fraction of the sampling window during which the device was not idle". For block devices it is the fraction of time with at least one request in flight; it says nothing about how deep the queue is. An NVMe can sit at %util 100% with perfectly healthy throughput and latency — normal behavior for a parallel device, not a fault signal.

The columns that matter are await and the queue-related ones. Modern kernels split `r_await` / `w_await` in `iostat -x`. `%aqu-sz` and `%arq-sz` are the average queue wait and average queue size for read requests — rising values mean requests are piling up in front of the device.

Diagnose by combination. High await with low %arq-sz means the device itself is slow: replace the disk or change storage class. High await with high %arq-sz means queue congestion, typically too much concurrency or random IO saturating queue depth. Low await with high %util means many small requests overwhelming parallelism — the classic small-file storm.

Get the sampling interval right. With `iostat -x 1` the first row is the since-boot average and is useless — skip it and read from the second row onward. Use `iostat -x 1 10` for ten samples, or `iostat -x 0.5 20` when you suspect momentary spikes.

iostat -x 1 10
# 关注列:r_await/w_await(平均延迟,含排队)
#        %arq-sz / %aqu-sz(请求平均队列深度 / 排队时长)
#        rkB/s、wkB/s(吞吐)、%util(只是非空占比)

PSI: turning "the device is slow" into "who is being slowed"

PSI (Pressure Stall Information) measures the fraction of time tasks lose scheduling opportunities because a resource class is unavailable. It is the only metric that maps kernel-side waiting onto "how many tasks are affected", and its overhead is low enough for continuous collection.

`/proc/pressure/io` has two lines. `some` is the fraction of time at least one task was stalled on IO; `full` is the fraction of time every non-idle task was stalled on IO — meaning the box is fully serialized on disk. `some` reflects individual experience, `full` reflects systemic stall; both high is a catastrophe.

Combine PSI with cgroups. systemd services expose `IOReadPressureIOPressure` and friends via `systemctl show <unit> -p IOReadPressureIOPressure`, or you can read the cgroup `io.pressure` file directly. That answers "which service is dragging the machine down" rather than just "the machine is slow".

For monitoring, apply `rate()` to `node_pressure_io_wait_seconds_total`, normalize by time, and alert on thresholds (for example above 40% for 5 minutes). Track `avg10/avg60/avg300` together to separate short spikes from sustained degradation.

cat /proc/pressure/io
cat /sys/fs/cgroup/system.slice/<unit>/io.pressure
systemctl show nginx.service -p IOReadPressureIOPressure

Attributing to processes: who is actually reading and writing

iostat tells you the device is busy, not who is making it busy. Attribution needs `pidstat -d 1`, which reports `kB_rd`, `kB_wr`, `kB_ccwr` (cancelled writes — a high value means requests were issued then aborted, wasting bandwidth) and `iodelay` per process or thread.

pidstat output rolls: each refresh shows only the current interval, so watch for rows where `kB_wr` stays high rather than reading totals. Piping through `sort` does not work because output is streaming; in practice, watch a few screens and note PIDs that keep recurring.

`fiotop` is better for interactive triage: it lists processes ordered by IO rate with readable output, far more intuitive than pidstat. Its downside is that it redraws on a timer, so it is not scriptable and keeps no history.

When iotop is unavailable there is a dependency-free fallback: `/proc/<pid>/io` exposes per-process `read_bytes` / `write_bytes` counters (actual bytes to storage, excluding page-cache hits). Sample twice and difference them to get a real rate — handy inside minimal containers with nothing installed.

pidstat -d 1
# 关注 kB_wr 高、kB_ccwr 也高的行:写风暴 + 请求被取消
cat /proc/$(pgrep -n nginx)/io | grep -E "read_bytes|write_bytes"
fiotop -o -d 1      # -o 按写入速率排序

Classify the root cause, then act

Small-file storm (high %util, low await, high %arq-sz): the classic inode and fsync-heavy workload where metadata operations saturate the queue. Fix by merging small files into a few large ones, or by switching to a storage format that supports direct IO. More RAM only defers the IO, it does not remove it.

Large sequential throughput (low await, high %arq-sz): almost always backup, sync, or export jobs. Verify whether the work is necessary, move it to a quiet window, or throttle it with `ionice -c3` (idle priority) or cgroup IO weights. This class has the most tuning headroom and the largest quick win.

Heavy random writes (high await with a long latency tail): usually a database checkpoint or log flush. Check whether log/backup configuration exceeds realistic performance expectations, and whether a volume was mistakenly configured for synchronous commits. This is especially visible on network storage — review the sync mount options in `/etc/fstab` together with the storage-side cache policy.

Do not forget the kernel layer: confirm the scheduler at `/sys/block/*/queue/scheduler` (`none` suits high-concurrency NVMe, `bfq` suits desktop and low-concurrency workloads), tune depth at `/sys/block/*/queue/nr_requests`, and identify rotational disks at `/sys/block/*/queue/rotational`. Persist any change via `/etc/sysctl.d/` or udev rules so the change survives reboot.

cat /sys/block/nvme0n1/queue/scheduler
# 高并发 NVMe 常用 none;低并发/桌面可试 bfq
cat /sys/block/nvme0n1/queue/nr_requests
cat /sys/block/nvme0n1/queue/rotational   # 0 = 非机械盘

Official References

Each command links to its official documentation below, so you can verify the latest usage and read deeper.