한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L12 Database

Cache stampede Cache stampede / thundering herd

Cause ID db-cache-stampede · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

Open the interactive card with figures and simulations →

When cache entries for popular data expire at the same time, thousands of requests hit the DB all at once.

Why Popular data stored in Redis or a similar cache expires at the same time → Effect Requests trying to rebuild the same data rush to the DB all at once → On screen DB overload makes one feature after another slow down or freeze

Symptoms
Input lag, Freeze, Can’t connect / infinite loading
Factors
Stall, Latency
Who’s affected
Whole server
When
At regular intervals, When crowds gather
Owner
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Randomize expiry times, let a single request refresh the value while the rest keep using the old value.
Infra team action items
Set up Redis replicas and automatic failover so the cache doesn’t empty entirely on a restart or failure, confirm the DB has enough headroom to survive an empty cache.
On the graph
Periodic spikes · Cache hit ratio, DB queries per second
Where to look
Overlay keyspace_hits and keyspace_misses (hit ratio), expired_keys, and restarts (uptime_in_seconds) from Redis INFO on the DB’s queries per second, and count how many copies of the same query run at once on the DB at that moment (MySQL SHOW PROCESSLIST, PostgreSQL pg_stat_activity)
Confirmed if
At the moment cache misses shoot up, the DB query count spikes with them, and most concurrent queries are the same query reading the same data. Lines up with the expiry cycle of popular keys or a Redis restart
Ruled out if
Cache misses normal but only DB queries rise: a login storm (db-login-storm) or a batch job (db-batch)
Check with
Infra tools (no game code needed)
Learn more
The same thing happens when a Redis restart or failure empties the whole cache. The more an architecture relies on the cache and keeps the DB small, the greater the risk.

Sources

  1. Scaling Memcache at Facebook (NSDI '13) USENIX
    When a hot key is invalidated, many reads rush to the DB (thundering herd); prevented with leases (only one client refreshes) and by returning stale values; clusters with an empty cache are warmed up separately
  2. Optimal Probabilistic Cache Stampede Prevention VLDB Endowment
    When a popular item expires, many requests regenerate it at once (cache stampede); prevented by probabilistically refreshing early, before expiry
  3. High availability with Redis Sentinel Redis
    Automatic failover that promotes a replica when the primary dies
  4. INFO Redis
    keyspace_hits and keyspace_misses (successful and failed key lookups), expired_keys (number of expired keys), uptime_in_seconds (time since startup)
  5. SHOW PROCESSLIST Statement MySQL
    The statement each session is running (Info) and the time spent in its current state (Time, seconds)
  6. The Cumulative Statistics System (PostgreSQL Documentation) PostgreSQL
    pg_stat_activity: the query each session is running now (query)

See also

Same layer: L12 Database

Same symptom (Input lag), other layers

View the interactive card with figures and simulations