How the two mechanisms actually differ
RDB is a point-in-time snapshot: the whole dataset is compressed into a binary `dump.rdb`. It is compact and fast to load, but nothing is recorded between two snapshots, so the loss window equals the snapshot interval.
AOF is an operation log: every write command is appended to `appendonly.aof` in the Redis wire format and replayed at startup. The loss window shrinks to "at most the writes not yet flushed", but the file grows without bound and must be rewritten with `BGREWRITEAOF` to compact it.
Redis 7 introduced the multi-part AOF: a base file (usually an RDB-format snapshot) plus a set of incremental files. Rewriting no longer needs the "write a temp file then atomically rename" dance — new increments during a rewrite land in their own file, and recovery loads the base then replays each increment. Rewrite cost and recovery time both drop sharply.
One commonly underappreciated fact: writes during an RDB snapshot are not lost. BGSAVE records the current dictionary version first; concurrent writes enter the replication buffer and are saved again if needed once the snapshot completes. Data consistency across a snapshot is guaranteed — "writes are lost during BGSAVE" is a widespread myth.
# 状态查看
INFO persistence
# rdb_changes_since_last_save : 上次快照后的写入次数(越大丢失风险越高)
# rdb_last_bgsave_status : ok / err
# aof_rewrite_in_progress : 是否正在重写
# aof_last_bgrewrite_status : ok / err
CONFIG GET appendonly
CONFIG GET appendfsync
CONFIG GET dir
CONFIG GET dbfilenameThe three appendfsync policies and their trade-offs
`always`: fsync on every write before replying to the client. The loss window is effectively zero (at most the last unacknowledged write, and OS page-cache loss remains theoretically possible), but throughput becomes disk-bound — typically a few thousand to ten thousand writes per second. Use it for orders and ledgers where losing money is not an option.
`everysec`: fsync once per second. This is the officially recommended balance — worst case you lose the last second of writes, and throughput usually reaches tens of thousands per second. On Redis 7+ the write path was optimised so the main thread no longer blocks directly on fsync.
`no`: leave it to the operating system, usually every 30 seconds, sometimes longer, and possibly all of it on power loss. Highest throughput, but your durability promise does not exist. Only use it when Redis is provably a cache with another rebuildable source of truth.
How to choose: ask whether data can be rebuilt if Redis dies. If yes, disable persistence or use `no`. If no, use `everysec`. Only escalate to `always` when a single second of loss is unacceptable — for example a sole-source queue or lock service. Most teams should start at `everysec` and tighten only after measuring real fsync latency.
# 纯缓存(可重建)
appendonly no
# save 置空以关闭 RDB 快照
# 通用推荐
appendonly yes
appendfsync everysec
# 强一致要求
appendonly yes
appendfsync alwaysFailover and data integrity
Redis replication is asynchronous by default. The primary may have already replied "write accepted" while the write has not yet reached a replica; if the primary dies at that moment the data is lost after failover. The size of that window depends on network latency and the replica `repl_ack` frequency.
The pair `min-replicas-to-write` and `min-replicas-max-lag` makes the risk explicit: with `min-replicas-to-write=1` and `min-replicas-max-lag=2`, the primary refuses writes whenever fewer than one replica is present or the lag exceeds two seconds. You are trading "possible data loss" for "possible unavailability" as an explicit decision.
A further integrity detail is that expired keys are not strictly identical between primary and replicas. From Redis 7 `REPLICAOF` propagates expire times through the replication stream but deletion still requires each replica to judge it, so the two can briefly disagree during network jitter. Critical logic should never rely on key expiry as its only consistency guarantee.
When you truly need synchronous semantics, use the `WAIT` command: after the write, call `WAIT 1 1000` to require one replica acknowledgement with a one-second timeout. This gives per-write confirmation on top of asynchronous replication, keeping most of the throughput while protecting critical writes.
REPLICAOF NO ONE
# 查看复制状态
INFO replication
# master_repl_offset / slave0:offset=<n>,lag=<n seconds>
SET balance:1001 9900
WAIT 1 1000 # 等待 1 个从节点确认,最多等 1000ms
# 拒绝写入模式
min-replicas-to-write 1
min-replicas-max-lag 2A configuration template and how to verify it
A good generic starting point: AOF enabled, `appendfsync everysec`, `auto-aof-rewrite-percentage 100`, `auto-aof-rewrite-min-size 64mb`. The latter two make Redis rewrite automatically once the file passes 64MB and has doubled since the last rewrite, with no human intervention.
Verifying persistence is more reliable by behaviour than by reading the config file. The direct test is to write a key, kill the Redis process, restart and check whether the key survived. In production avoid `DEBUG RELOAD` (it blocks); in test environments compare `SHUTDOWN NOSAVE` against `SHUTDOWN SAVE`.
For monitoring, watch two signals: a continuously rising `rdb_changes_since_last_save` means the snapshot interval is too long or snapshots keep failing, and `aof_rewrite_in_progress` stuck at 1 means a rewrite is hung. `aof_last_write_status: err` in `INFO persistence` is a hard metric that deserves an alert.
One migration trap to remember: the AOF format changed between Redis 6 and 7 (multi-part AOF), while older versions used a single file. Check the current version before upgrading and confirm the AOF loads afterwards. If loading fails, Redis starts with an empty dataset — the worst kind of silent data loss.
# 推荐的通用配置
appendonly yes
appendfsync everysec
auto-aof-rewrite-percentage 100
auto-aof-rewrite-min-size 64mb
# 验证
redis-cli INFO persistence
redis-cli CONFIG GET appendonly
redis-cli LASTSAVE