First, separate the three problems: different causes, different fixes
A cache breakdown (often called a hot key expiring) happens when a single extremely popular key happens to expire at one instant. If rebuilding that key is expensive — aggregating several tables or calling an external service — every concurrent request discovers the empty cache at the same moment and hundreds of them hit the database together. The database is not genuinely broken, but it momentarily absorbs read traffic the cache was supposed to serve, the connection pool saturates, and latency jumps from milliseconds to seconds.
Cache penetration is different: the query targets a key that simply does not exist. Since the key was never cached, the cache can never hit, the database is never warmed, and every such request walks the full path of cache lookup, miss, database query. An attacker can generate a sequence of random IDs against your endpoint and hammer the database without ever producing a cache hit. It has nothing to do with concurrency or overall traffic volume — only with the fact that the key is absent.
A cache avalanche is a mass invalidation event. If you bulk-load a batch of entries at 02:00 with a TTL of 3600 seconds, they all expire together at 03:00 and a flood of traffic reaches the database in the same second. The other form is the Redis instance itself going down or losing network connectivity: all traffic has nowhere to go, which looks identical to an avalanche even though the cause is storage unavailability rather than colliding expiry times.
Learn the triage signature of each one, because it saves enormous time during an incident. A breakdown shows a narrow, tall spike in database QPS while the hot key hit rate momentarily drops to zero. A penetration shows steadily rising database QPS without a corresponding rise in total request volume, many empty results, and a Redis keyspace that barely grows. An avalanche shows hit rate falling off a cliff, with the timing lining up with a bulk write or an instance event.
# ① 击穿:单点瞬时高 QPS + 该 key 命中率归零
redis-cli --no-raw OBJECT FREQ "hot:config:global" # LFU 计数,用于识别热点键
redis-cli --no-raw INFO stats | grep -E "keyspace_(hits|misses)"
redis-cli --no-raw TTL "hot:config:global" # 剩余生存时间
# ② 穿透:命中率正常/偏高,但库里有大量「查不到」的查询
redis-cli --no-raw DBSIZE # 观察 keyspace 是否随流量增长
redis-cli --no-raw SLOWLOG GET 10 # 定位被穿透打中的慢查询
# ③ 雪崩:命中率整体掉头,DB 侧出现成批同构慢查询
redis-cli --no-raw INFO memory | grep -E "used_memory_human|maxmemory_policy"
redis-cli --no-raw INFO persistence | grep -E "rdb_last_bgsave_status|aof_last_write_status"
redis-cli --no-raw INFO replication | grep -E "role|master_link_status"Taming breakdowns: mutex rebuilds and logical expiry
The most direct fix is a mutex around the rebuild: on a cache miss, first try `SET lock:key token NX EX 10`. The winner rebuilds, and everyone else either retries the cache after a short wait or returns a degraded value. However many requests arrive, exactly one reaches the database. The lock value must be a unique identifier such as a UUID, and unlocking must go through a Lua script that compares the token before deleting, otherwise a holder whose lock already expired will delete the lock acquired by a later owner.
Waiters should not sleep for a fixed duration. A common implementation spins a few times with short intervals, re-reading the cache each time and returning immediately on a hit. After a bounded number of retries with no hit, decide based on business tolerance whether to return a stale value, a default, or pass through to the database. Keep the retry count and interval configurable, because sensible values depend entirely on rebuild cost and traffic volume — constants copied from another project rarely fit.
Logical expiry, often called soft expire, takes a different route and sidesteps waiting altogether: the cache entry stores both the value and a logical expiry timestamp. On read, if the timestamp has passed, return the possibly stale value immediately and kick off a rebuild in the background. Readers are never blocked; the cost is that callers may observe data that is at most a bounded interval old.
Logical expiry is naturally immune to breakdowns because expiry detection and rebuild initiation are decoupled: no matter how many readers simultaneously see the entry as expired, only one rebuild really runs. The price is that you must define a business-level tolerance for staleness and require callers to accept occasional inconsistency. Product pages and home-page recommendations — read-heavy, rarely written, fine with a few seconds of lag — are ideal fits. Balances and stock levels are not.
# 互斥锁重建:抢锁 -> 重建 -> 放锁(Node 风格伪代码,语义与 redis-cli 一致)
# 1) 抢锁,NX 保证只有一个赢家,EX 保证持有者崩溃后锁会自动释放
redis-cli SET "lock:user:1001" "b1f2c0de-uuid" NX EX 10
# 返回 OK 表示抢到,nil 表示已有别人在重建
# 2) 等待方:短间隔自旋,先看缓存是否已被别人填好
redis-cli GET "cache:user:1001"
redis-cli EXISTS "lock:user:1001"
# 3) 释放锁必须用 Lua,保证「比对 + 删除」原子,
# 否则持有者超时后会误删新持有者的锁
redis-cli EVAL "if redis.call(\"get\",KEYS[1])==ARGV[1] then return redis.call(\"del\",KEYS[1]) else return 0 end" 1 "lock:user:1001" "b1f2c0de-uuid"
# 4) 逻辑过期写法:值里带时间戳,过期也照常返回,由后台重建
redis-cli SET "cache:user:1001" \n "{\"data\":{\"name\":\"Alice\"},\"expireAt\":1758000000000}" EX 300Stopping penetration: null-value caching and Bloom filters
Null-value caching is the cheapest fix with the fastest payoff. When the database lookup returns nothing, do not report a miss and let the next request fall through again — store the fact that the key is absent, with a very short TTL such as 30 to 60 seconds. A scanner hitting random IDs then penetrates the database at most once per key, and even a sustained attack is throttled by that TTL.
The TTL for negative entries should be short, which is the opposite of the instinct used for positive entries. Real cached data may stay valid for minutes or hours, whereas a judgement of nonexistence ages quickly: a stale negative entry means genuinely existing data stays invisible far too long. Prefer leaking a few lookups over caching emptiness for long.
When negative caching is not enough — attackers using identifiers that are structurally valid but semantically dead, or a key space so large that the negative entries themselves cost real memory — reach for a Bloom filter. It is a bit vector plus several hash functions, into which every known-existing key is mapped ahead of time. A query consults the filter first: if the relevant bits are zero the key definitely does not exist, and you can reply without even writing a negative entry.
Two properties define the trade-off. A Bloom filter has no false negatives, only false positives: saying "absent" is always safe, saying "present" may just be a collision, so the real cache or database lookup still happens. Second, the classic structure cannot delete entries; deletion needs a counting variant such as a Cuckoo Filter, which costs more memory. Populating the filter is tightly coupled to the write path, and a missed insertion produces the catastrophic "exists but reported absent" failure, so populate on creation rather than backfilling later.
# 空值缓存(负缓存):给不存在的键一个短 TTL
redis-cli SET "cache:article:999999" "__nil__" EX 45
redis-cli GET "cache:article:999999"
# 布隆过滤器:用 EXISTS 代替逐个试探,标准 BF 支持 FEXISTS 扩展
redis-cli BF.RESERVE "bf:article_id" 0.01 1000000
redis-cli BF.ADD "bf:article_id" "1001"
redis-cli BF.EXISTS "bf:article_id" "1001" # 返回 1:可能存在,需继续查缓存/库
redis-cli BF.EXISTS "bf:article_id" "999999" # 返回 0:一定不存在,直接返回空结果
redis-cli BF.INFO "bf:article_id" # 查看容量、错误率、已插入元素数
# 判定顺序:Bloom 过滤器 -> 缓存 -> 负缓存 -> 数据库Preventing avalanches: TTL jitter and graceful degradation
The cheapest and most effective defence against an avalanche is TTL jitter. Add a random offset to the base expiry, for example 600 plus a random 0 to 300 seconds. A batch of ten thousand entries then expires across a ten-minute window instead of detonating in the same second. A jitter ratio of 10 to 20 percent of the base is normally plenty; pushing it further weakens the hit rate for no benefit.
The second defence is not betting everything on one instance. Set a memory ceiling and an eviction policy so Redis discards old keys under pressure. If the cache regularly runs near its memory limit, a restart does far less damage, because the node has already been practising eviction in steady state. Note that allkeys-lru puts hot keys into the eviction pool as well, trading hit rate for survival, whereas volatile-lru only evicts keys that carry a TTL — the better choice when every cache entry is supposed to have one.
The third defence is degradation inside the application. When Redis is genuinely unavailable, the app must hold on its own: cap the concurrency of fallthrough requests with a semaphore so only a few dozen reach the database, briefly cache failed lookups as negative results, and if necessary return the last successful response or a static fallback page. The critical property is that a cache outage must not surface as an exception; a 500 caused by a Redis timeout is far worse than simply reading through to the database.
Finally, preventing avalanches means preventing their creation. Offset the TTLs of a bulk warm-up so entries do not align; run a keyspace-growth estimate in a staging environment before a release to see how many cache keys a dataset will create, and treat a large number during peak traffic as a risk to be reviewed. An easily overlooked equivalent is a maintenance script that flushes the cache for testing — that has exactly the same blast radius as an avalanche, so move it to an off-peak slot and run it in batches.
# TTL 抖动:为基准值叠加随机量,避免批量键同时到期
# shell 伪代码:base=600, jitter=120
for key in $(cat .ai-temp/reports/hot_keys.txt); do
ttl=$(( base + RANDOM % (base / 5) ))
redis-cli SETEX "$key" "$ttl" "$(cat .ai-temp/downloads/payload.json)" > /dev/null
done
# 内存上限与淘汰策略:让实例常态下就工作在驱逐状态
redis-cli CONFIG SET maxmemory 2gb
redis-cli CONFIG SET maxmemory-policy volatile-lru
redis-cli CONFIG GET maxmemory-policy
redis-cli --no-raw INFO memory | grep -E "used_memory_human|maxmemory_human|mem_fragmentation_ratio"
redis-cli --no-raw INFO stats | grep evicted_keys
# 观察命中情况:命中率长期低于 90% 就该重新评估 TTL 与 key 设计
redis-cli CONFIG RESETSTAT
redis-cli --no-raw INFO stats | grep -E "keyspace_hits|keyspace_misses"Hot keys: a second-level local cache and request spreading
When traffic on a single key is high enough that one instance CPU-bound or network-bound, the cache itself becomes the problem. Mutexes do not help here, because the bottleneck sits on the Redis side rather than on the fallthrough path. The first tool is a second-level local cache inside the application process: a very small in-memory cache with a short expiry of a few to a few dozen seconds holding a copy of the hot value. Hot-key responses are usually not millisecond-sensitive, so seconds of staleness buy you the elimination of most read traffic at the process level.
Be deliberate about consistency and memory when adding a local tier. The standard pattern is double-checked access: look locally, and on a miss read Redis and repopulate the local entry. On write, actively invalidate local copies, or rely on a very short TTL to expire naturally, so instances do not diverge for long. The local cache must also have a hard capacity ceiling — typically an LRU with a maximum entry count — otherwise a workload of mostly unique keys turns it into a memory leak.
The second tool is spreading. Split a hot key into N replicas, such as cache:config:1 through cache:config:8, and let the caller pass a random offset when reading. On write, either update every replica (simple and reliable, at the cost of write amplification) or write one and let short TTLs refresh the rest (cheap writes, at the cost of brief inconsistency). For read-heavy, write-rare data, writing all replicas is almost always the better trade.
Finally, nothing gets fixed that is not first identified. OBJECT FREQ with an LFU policy reports per-key access frequency and is the most direct way to find hot keys; in production the OBJECT family can be unavailable behind proxies or restricted permissions, in which case fall back to a short MONITOR sample — never on a peak window — or to client-side instrumentation. Making "top N hot keys this week" a routine report beats debugging by hand at 3 AM.
# 识别热点键:LFU 策略下用 OBJECT FREQ 看访问频次
redis-cli CONFIG SET maxmemory-policy allkeys-lfu
redis-cli --no-raw OBJECT FREQ "hot:config:global"
redis-cli --no-raw SCAN 0 COUNT 500 | head -50 # 粗略抽样 key 空间
# 命中率与逐出观察
redis-cli --no-raw INFO stats | grep -E "keyspace_hits|keyspace_misses|evicted_keys"
redis-cli --no-raw INFO commandstats | grep -E "cmdstat_(get|set|evict)"
# 排障时抓一段慢查询(SCAN 不会阻塞,但生产慎用)
redis-cli SLOWLOG GET 20An observability checklist before going live
Cache incidents rarely announce themselves as outages; they quietly get slow or quietly get wrong. Instrument three dimensions before launch, not just availability. Hit rate, tracked as keyspace_hits over the sum of hits and misses, should be read as a time series and judged on percentiles rather than averages — a 99 percent average can hide a stretch where it sat at 50 percent. evicted_keys tells you whether the working set genuinely exceeds the memory ceiling, in which case you have merely moved the pressure onto the database.
Collect latency per command: look at the p99 of GET, SET, and EVAL separately. A noticeably higher p99 on EVAL than on ordinary commands usually points to intense lock contention and long lock hold times in the rebuild path. SLOWLOG GET is a tool for post-mortem attribution once something has already gone wrong, not for continuous monitoring.
Prepare a degradation switch. The cache path in code should sit behind a configuration flag so it can be disabled without a code change or a release, degrading to direct database reads: slow, but not a 500. In the same spirit, the fallthrough concurrency ceiling, the single-flight retry limit, and the negative-cache TTL should all be dynamically adjustable configuration rather than hard-coded constants — during an incident you want to tune parameters, not ship a release.
Finally, plan the regression checks. After every change to caching logic, run targeted cases: concurrent load against a hot key, a penetration simulation using random IDs, observation of the expiry-time distribution after a bulk warm-up, and the recovery behaviour after an instance restart. Script these and put them in CI or a scheduled job — their value far exceeds any one-off manual verification.
# 命中率(分时段采样,避免只看累计平均值)
redis-cli --no-raw INFO stats | grep -E "keyspace_hits|keyspace_misses"
# 逐出速率:持续增长 = 内存上限已不足
redis-cli --no-raw INFO stats | grep evicted_keys
# 延迟与命令统计
redis-cli --no-raw LATENCY LATEST
redis-cli --no-raw INFO commandstats | grep cmdstat_get
# 内存与碎片
redis-cli --no-raw INFO memory | grep -E "used_memory_human|maxmemory_human|allocator_frag_ratio"