Approach: decide whether the problem is writing, rotating, or querying
Almost every logging failure lands in one of three stages. Writing: the service never emitted the log, the level is wrong, or the disk filled up. Rotating: files got cut but never compressed, old files were never deleted, or a rename-based rotation left the open write handle pointing at an unlinked inode so the space is never reclaimed. Querying: the data is still there and you filtered it out.
Separate them by looking at space first, content second. Use `df -h` and `du` to find which directory ate the disk. If there is free space but nothing shows up in your query, it is a query problem. If the disk is full, it is a write or rotate problem. This order saves a lot of futile filter debugging.
journald and plain file logs have different entry points. journald is a binary, indexed format with precise filtering on structured fields such as `_SYSTEMD_UNIT=`, `_PID=`, and `PRIORITY=`, so no regex is needed; `/var/log/*.log` is text and grep is all you have. When unsure, answer the most basic question first with journalctl: did this unit emit anything at all in this window?
# 空间侧:谁吃掉了磁盘
df -h
du -xh --max-depth=2 /var/log 2>/dev/null | sort -rh | head -20
du -xh --max-depth=1 /var/lib/docker 2>/dev/null | sort -rh | head
# journal 自身占用
journalctl --disk-usage
# 目录清单:有没有被轮转后没删掉的旧文件
ls -lhtr /var/log/ | tail -30Common journalctl queries: get the time units and filter fields right
journalctl defaults to "since this boot", an implicit `-b`. To see earlier boots you must say so explicitly with `journalctl -b -1`, or check `journalctl --list-boots` first for the list of boots and their time ranges. This is one of the top causes of "my logs vanished": the service restarted two days ago and the entries you want live in the previous boot.
Time filtering uses `--since` and `--until`, accepting relative expressions (`-1h`, `-30min`, `today`, `yesterday`) as well as absolute timestamps. `-p` filters by priority, where `-p err` means err and worse because smaller priority numbers are more severe. Field filters take the form `FIELD=value`, for example `_SYSTEMD_UNIT=nginx.service`, and multiple conditions are ANDed.
For output shape, `-o short-precise` keeps millisecond timestamps, which is invaluable for millisecond-scale races. `-o json` emits structured records you can pipe into `jq` to aggregate, for instance counting errors per hour for a unit. `-f` follows live, the equivalent of `tail -f` on the journal. Combining `-n`, `-p`, and `-u` covers most day-to-day needs.
journald also captures console output from the boot itself under `-b -k` for kernel messages — essential when investigating networking, mounts, or OOM kills. A powerful combination is a two-field filter such as `_SYSTEMD_UNIT=cron.service _PID=1234`, which narrows to one specific run instead of pulling the entire unit output.
# 先看有哪些次启动,再逐次排查
journalctl --list-boots | tail -10
journalctl -b -1 -p err --no-pager | tail -50
# 时间范围 + 优先级 + unit 三条件组合
journalctl -u nginx.service --since "2 hours ago" -p warning --no-pager
# 精确到某次任务
journalctl _SYSTEMD_UNIT=cron.service _PID=23117 --since "today" --no-pager
# 聚合统计:每小时 ERROR 数
journalctl -u api.service --since today -p err -o json --no-pager | jq -r '_TIME
| (.[11:13])' | sort | uniq -c
# 实时跟随 / 内核 / 导出归档
journalctl -f -u api.service
journalctl -b -k --since "-30min" --no-pager
journalctl --since "-7d" -o export > /tmp/journal-export.txtDisk full from logs: the correct order of emergency reclamation
At 100% disk, services fail in bizarre ways, so resist deleting files blindly. The highest-yield first step is letting journald shrink itself: `journalctl --vacuum-size=500M` removes the oldest entries until the total is under 500M, while `--vacuum-time=3d` retains by age. Both flags combine; start with `--disk-usage` to see the current figure.
Second, handle plain file logs. Use `du` ordered by modification time to find the biggest files. If they belong to a service that is still writing, `rm` leaves a deleted-but-open inode and the space is not reclaimed — `df` still shows full. Truncate instead of delete: `truncate -s 0 /var/log/app.log` returns the space immediately while keeping the write handle valid.
Third, clean archives and container logs. Historical `/var/log/*.gz` files can be removed outright; Docker and containerd logs can also blow up the disk (one JSON log per container under `/var/lib/docker/containers`), so check whether the daemon config caps log size. Do not touch files other than `*_journal` for journald, and leave alone any directory an audit tool is actively reading.
Fourth, restore rotation and confirm. Force one pass with `logrotate -f /etc/logrotate.conf`, or `systemctl restart systemd-journald`. Then verify with `df -h` — if space has not returned, someone still holds a deleted file, and `lsof +L1` lists every such open-but-unlinked handle.
Actions to avoid throughout: do not `kill -9` the service doing the logging, do not `rm -rf` the log directory, and do not delete only the `.gz` files while leaving a huge current file. Disk exhaustion is always a missing rotation policy or an oversized cap; after the emergency, set `maxsize` and a retention count or it will recur unchanged in three days.
# 1) 现状
df -h /
journalctl --disk-usage
du -xh --max-depth=1 /var/log | sort -rh | head
# 2) 收缩 journal
journalctl --vacuum-time=3d --vacuum-size=500M
# 3) truncate 而非 delete(写句柄仍有效,空间立即回收)
truncate -s 0 /var/log/app.log
# 4) 空间没回来?查被持有的已删除文件
lsof +L1 2>/dev/null | head -20
# 5) 恢复轮转并复核
logrotate -f /etc/logrotate.conf
df -h /
# 6) 防止复发:给容器日志加上限
# /etc/docker/daemon.json -> {"log-driver":"json-file","log-opts":{"max-size":"100m","max-file":"3"}}"Yesterday's logs are missing": logrotate copytruncate and file naming
Traditional rotation comes in two flavours. `create` renames first and then creates a new file, so the application's write handle still points at the renamed file and content splits across two files. `copytruncate` copies first and then zeroes the original, leaving the handle untouched. For programs that open logs with O_APPEND — which is most logging libraries — rename mode makes the process keep writing to the old file while the new one stays empty until restart. That is the root cause of "rotation is configured but logs never appear".
copytruncate costs you the writes that land in the window between copy and truncate, plus a short IO spike on large files. But it is the only scheme that does not require the application to reopen its log, so it remains the correct choice for third-party libraries without reopen support — unless you add SIGHUP-based reopening to the application.
What actually deletes history is a missing or too-small `rotate N`. Only rotations older than the retained count are removed; omit `rotate` or set it too low and logs from yesterday or the day before are already gone. `dateext` controls whether names carry a date suffix (`app.log-20261002`); without it you get plain numeric suffixes, and across midnight `app.log.1` ambiguously means both "yesterday" and "the most recent prior rotation".
Diagnosis order: run `ls -lhtr /var/log/` to see which files and suffix conventions actually exist, then find the matching rule under `/etc/logrotate.d/` and read `rotate`, the `daily/weekly` schedule, `compress`, and `delaycompress`. With `compress` on, the uncompressed one is `.1` and `.1.gz` is the previous rotation — read it with `zcat`. `delaycompress` keeps the newest rotation uncompressed to cut IO.
A related trap when journald is in the mix: once files under `/var/log/journal` have been removed by `journalctl --vacuum-*`, the entries already recorded in the journal still exist, so people see "the file is gone but journal has the record" — and in that situation the journal is the correct source of truth. The mirror image bites collectors: Filebeat and Fluent Bit track files by inode, and copytruncate versus rename demand different collection strategies, otherwise you duplicate or silently drop events.
# /etc/logrotate.d/app
/var/log/app/*.log {
daily
rotate 14 # 保留 14 份;省略或写 1 就会丢历史
missingok
notifempty
compress # 历史轮次压缩
delaycompress # 最近一份不压缩,减少 IO
copytruncate # 不要求应用 reopen 文件
dateext
dateformat -%Y%m%d
su root adm
create 0640 root adm
}
# 排查三连
ls -lhtr /var/log/app/
logrotate -d /etc/logrotate.d/app # -d 仅模拟,不实际执行
logrotate -v /etc/logrotate.d/app # -v 输出处理了哪些文件Multi-line logs: three different joining rules
Stack traces are inherently multi-line. Line one is `java.lang.NullPointerException` and a dozen `at com.example...` lines follow. If a collector processes line by line, the trace is shredded into a dozen independent records and alert deduplication, error counts, and traceId correlation all degrade. The join rule must be decided where the log is produced and kept consistent.
Inside a systemd unit, `StandardOutput=` accepts `journal` (the default) and `journal+console`, but what really governs line parsing is `SyslogIdentifier=` plus the log format itself; in practice the cleanest solution is to make the application emit single lines with newlines escaped as `\n`. If that is impossible, do the joining either in the application or at the collector with a multiline pattern.
At the collector, Filebeat and Fluent Bit both ship multiline support, usually deciding continuation by "this line starts with an exception class name or timestamp" or the inverse, "lines that do not match the start pattern belong to the previous record". Filebeat uses a `multiline.pattern` block, Fluent Bit a `[MULTILINE] parser`. Their regex dialects differ, so configurations do not port directly and must be retested line by line.
On the shell side, grep cannot join lines. Practical substitutes are a multi-line-capable matcher such as `pcregrep -M`, or normalising the file to one-line JSON with `awk` before processing. journald needs no extra configuration — if the application already emits a single record, journald stores it verbatim. That is the argument for single-line application logs: one change fixes journald, logrotate, and the collector at once.
# Filebeat:把堆栈合并成一条记录
filebeat.inputs:
- type: filestream
id: app-log
paths:
- /var/log/app/*.log
parsers:
- multiline:
type: pattern
pattern: '(^[0-9]{4}-[0-9]{2}-[0-9]{2}.*|Exception in thread|Caused by:|^\\s+at .*)$'
negate: false
match: after
timeout: 5s
# 读压缩的历史轮次
zcat -f /var/log/app/app.log.* 2>/dev/null | grep -n "ERROR" | tail -50
# 侧栏:合并后先看堆栈头部数量,确认有没有被切开
grep -c "Exception in thread" /var/log/app/app.log