When memory runs short and the OS moves part of it out to disk, every access to that memory waits on a disk more than 1,000 times slower.
Why Memory in use exceeds physical RAM → Effect The OS moves part of it to disk and reads it back when needed → On screen Ticks balloon to hundreds of ms, and every player on the server sees slow motion and freezes
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Put a cap on process memory use (heap size and so on), check for leaks.
Infra team action items
Configure game servers not to use swap, act on memory alerts, provision RAM well above peak usage, since with swap off the process is killed (OOM) the moment memory runs out.
Ballpark numbers
A RAM read takes about 100 ns; reading back from an SSD takes about 100 µs (1,000×), a cloud disk over the network about 1 ms (10,000×), and an HDD 10 ms (100,000×).
On the graph
Slow climb · Swap usage, swap-in/out
Where to look
Overlay on tick time: the si and so columns of vmstat 1 (amount swapped in and out per second), some and full in /proc/pressure/memory (share of time stalled waiting for memory), and pidstat -r majflt/s for the game server process (page faults that had to read from disk)
Confirmed if
si above 0 at the time of the lag, with the game server’s majflt/s and the memory full value rising together
Ruled out if
si and so at 0 and memory pressure (PSI) near 0: swap isn’t the cause. No swap but majflt/s and PSI rising: memory is running out and code pages are being reread, so free up memory first
Check with
Infra tools (no game code needed)
Learn more
A server with GC reads all over the heap when it collects, so if even part of the heap is swapped out, a single GC can stretch to seconds or tens of seconds. With swap off, there is no slow swapping stage and the process goes straight to being killed (OOM), so secure spare memory first. Even without swap, when memory is nearly exhausted the OS may drop the executable’s code pages from memory and read them back again, so the whole server can slow down badly for a while before the OOM kill.
Documentation for /proc/sys/vm/Linux kernel swappiness: relative cost of swapping vs. reclaiming file pages; swap is random I/O and therefore expensive
Concepts overviewLinux kernel Reclaims page cache backed by files on disk and swappable pages; if that’s still not enough, the OOM killer kills a process
Solidigm™ D7-P5520 and D7-P5620 Product BriefSolidigm 99.99th percentile latency (four-nines latency) of 130 µs for server NVMe SSDs: basis for a single SSD read taking around 100 µs
PSI - Pressure Stall InformationLinux kernel some (share of time some tasks were stalled) and full (share of time all tasks were stalled at once) in /proc/pressure/memory