Redis processes commands one at a time, so a single slow command blocks every request behind it.
Why A full key search with KEYS in production, or reading or deleting a ranking or list with millions of elements in one go → Effect Every other request waits until that command finishes (tens of ms to several seconds) → On screen Features that use sessions, rankings, or the cache all hitch at once, slow logins
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Replace KEYS with SCAN, split large keys, delete with UNLINK (background deletion), spread out expiry times that cluster in the same second.
Infra team action items
Watch the slow command log (SLOWLOG), block dangerous commands such as KEYS on production servers, check for large keys regularly, turn off THP and keep enough spare memory for fork, run RDB and AOF persistence on replicas.
Ballpark numbers
A typical command takes under 1 ms. Handling millions of elements at once can take anywhere from hundreds of ms to several seconds.
On the graph
Random spikes · Redis response latency, slow command count
Where to look
SLOWLOG GET for commands over slowlog-log-slower-than; turn on the latency monitor (off by default) with CONFIG SET latency-monitor-threshold, then check per-event latency such as fork and expire-cycle with LATENCY LATEST and LATENCY DOCTOR. Also check fork time and large keys with latest_fork_usec in INFO and redis-cli --bigkeys
Confirmed if
At the time of the stall, SLOWLOG shows KEYS or commands handling a whole large key, or LATENCY records fork or expire-cycle events of tens of ms or more at the same time
Ruled out if
SLOWLOG and LATENCY are empty but it’s slow only as seen from the game server: the network or waiting inside the game server (SLOWLOG measures only command execution time, excluding time spent talking to the client)
Check with
Infra tools (no game code needed)
Learn more
Redis also stalls the moment it forks the process to create a save file (RDB snapshot) or rewrite the AOF. On modern servers this takes about 10 ms per GB of memory, so about 300 ms for 30 GB. With transparent huge pages (THP) turned on, every write after a fork copies an entire huge page (copy-on-write), sharply increasing stalls and memory use, so THP is usually turned off and plenty of spare memory is kept. Redis also stalls briefly to delete keys when a very large number of them expire in the same second.
Sources
Diagnosing latency issuesRedis One thread processes requests in turn, so a slow command blocks everything behind it; use SCAN in place of KEYS; fork measured at about 9–13 ms per GB on physical servers and modern VMs; THP causes latency and memory spikes from copying after fork; mass expiry in the same second causes stalls
KEYSRedis Use with extreme care in production; can ruin performance on large databases (40 ms for 1 million keys on an entry-level laptop)
UNLINKRedis Asynchronous deletion: unlinks the key immediately and reclaims its memory in another thread
SLOWLOGRedis Slow command log that records commands exceeding slowlog-log-slower-than; execution time excludes I/O with the client
Redis latency monitoringRedis latency-monitor-threshold defaults to 0 (off); LATENCY LATEST and LATENCY DOCTOR; records latency per event such as fork and expire-cycle
INFORedis latest_fork_usec: time taken by the last fork (microseconds)
Redis CLIRedis --bigkeys: scans the keyspace to find large keys