Two independent ledgers: data blocks and inodes
Every POSIX-permission Linux filesystem maintains two resource classes: **data blocks**, which store file contents, and **inodes**, which store metadata — permissions, ownership, timestamps, size, and pointers to data blocks. The inodes themselves live on disk but their total count is **fixed at format time**.
Because the inode total is fixed, every inode can be consumed while not a single data block is used. That is the entire explanation behind "No space left on device at 20% usage": the capacity ledger is fine, the file-count ledger is exhausted.
A further trap: inode exhaustion reports the **exact same** errno as space exhaustion (ENOSPC). The kernel does not distinguish the cause, so the only reliable habit is to always read `df -h` and `df -i` together. Looking at one alone will mislead you.
Symlinks, hard links, empty files, and directories each consume one inode, including zero-byte files. That is why a 0-byte file can be the culprit: it occupies no data blocks yet still burns an inode. Enough of them will exhaust an entire filesystem.
df -h # 数据块水位
df -i # inode 水位,IUse% 达到 100% 即耗尽
# 两条一起看,缺一不可Which workloads are most dangerous
① Massive small-file counts: static asset directories, build artifacts (`node_modules`, `.venv`, `target/`), backup extraction dirs. Each file burns an inode while using very few blocks — the archetypal "capacity empty, inodes full". With a 4 KB block and 1 KB of content, three quarters of the block is wasted per inode.
② Session and cache directories: PHP/Java session dirs, Redis persistence shards, application local caches (`~/.cache`), Nginx `proxy_cache`. Short file lifetimes plus fast growth mean a single batch job can write hundreds of thousands of tiny files overnight.
③ Mail and queue spools: `/var/spool/postfix`, mail queues, unarchived log segments. One file per message, several per message with attachments — steep growth under bursts.
④ Container overlay layers: every image layer and every container writable layer counts against the overlay inode budget, so a laptop running dozens of local images can hit this easily. It is easy to miss because containers are intuitively assumed to live elsewhere.
df -i /
# 典型输出:iused 100%,avail 为 0,而 df -h 仍有大量剩余容量Locating it: counting files layer by layer
The approach mirrors big-directory hunting, replacing `du` with `find | wc -l`. Keep the "partition first, drill down later" strategy — a full-disk find on an inode-starved machine can take tens of minutes of IO.
Start with a one-level count inside the full mount. Expect a small discrepancy against the used count from `df -i` because directories themselves consume inodes; do not be confused when the numbers do not line up exactly.
Step into the largest directory and repeat; two or three rounds usually pin it down. The fastest signal is a large ratio of "file count versus apparent data size": a directory with 500,000 files occupying 200 MB must be dominated by its inode usage.
Often you do not need precise location — running `ls | wc -l` on a known suspect (session dir, cache dir, `node_modules`) confirms it immediately. Before doing this on production, make sure `find` will not trigger heavy directory reads; add `-xdev` to stay within the mount.
find /var -xdev -type f -o -type d 2>/dev/null | wc -l
# 逐层下钻,对比每个子目录的文件数
for d in /var/*; do printf "%s " "$(find "$d" -xdev 2>/dev/null | wc -l)"; echo "$d"; done | sort -rn | headCleaning up safely and expanding
The first rule of cleanup is "delete regenerable small files, never touch data". Caches, sessions, temp files, and build artifacts are regenerable: removing them lets the app rebuild at the cost of a brief slowdown. Database files, user uploads, and archived logs are not — even if they are small files, deleting them is a data incident.
Prefer the application cleanup command over `rm -rf`. Redis has `SCAN` for incremental non-blocking deletion, Nginx has `proxy_cache_purge`, and package managers have `apt clean` / `yum clean all`. A raw `rm -rf` over hundreds of thousands of files generates heavy directory operations that can hang or fail outright on an already inode-starved filesystem.
Confirm up front that cleanup actually frees resources: inodes are reclaimed immediately on unlink, with no restart or sync required — the same as disk blocks. So you can clean incrementally and watch `df -i` as you go.
If cleanup cannot free enough, or the file count is genuinely required by the business, the only real options are expansion or a different filesystem. ext4 fixes its inode count at format time and cannot change it online; XFS can grow online with `xfs_growfs` but its inode ceiling is likewise fixed at format time. Note that filesystems like `tmpfs` take their size from the mount option, so remounting with a larger `size=` is a fast mitigation.
du -sh /var/cache/* 2>/dev/null | sort -rh | head
apt clean
rm -rf /var/lib/app/sessions/* # 可再生内容,可安全清理
df -i / # 清理后立即复查
mount -o remount,size=4G /dev/sdb1 # tmpfs 类可直接放大Prevention: monitoring and limits
Make inode usage a first-class alert item rather than folding it into disk usage. Typical thresholds are 80% warning and 90% critical. Note that inode consumption is **step-wise**: once you are alerted, you are often close to 100%, which leaves far less reaction time than capacity exhaustion does.
Give small-file-heavy directories their own quota or mount point. Putting sessions, caches, and log spools on a separate ext4/xfs mount means exhausting them cannot take down the root partition. systemd also offers `TemporaryFileSystem=/var/cache/app`, which keeps such files in an in-memory filesystem that vanishes when the process exits.
The application owns part of the fix: change the "one file per record" pattern into batched files (merged log segments, a database instead of a file tree) or enforce a minimum file size — many exports can simply ship as a tarball. Such changes remove inode pressure structurally rather than marginally.
Finally, discipline: after every "No space left on device" incident, leave behind a one-line conclusion — capacity or inodes? The next engineer on call should be able to answer in ten seconds instead of spending an hour re-deriving it.
# 告警表达式思路:inode 使用率独立成一条规则
# node_filesystem_files_free / node_filesystem_files < 0.1 即已耗尽
systemctl show app.service -p TemporaryFileSystem