When the server crashes, writing several GB of memory to disk can delay the restart by several minutes.
Why A server crash writes all of memory to a file → Effect No restart until several GB have been written → On screen After the server dies and players disconnect, they can’t connect again for a long time
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Consider small dumps holding only the needed memory (minidumps), fix the cause of the crash.
Infra team action items
Limit dump size (OS core dump settings), use fast disks, decouple the restart from the dump (compress and upload dumps separately after the restart).
On the graph
Mass disconnect · Connection count, server restart times
Where to look
Line up the crash time, core file size (coredumpctl list and info, or the file where core_pattern points), the time the file finished writing, and the time the service came back up, and check wkB/s from iostat -x during that window
Confirmed if
After the crash, disk writes stay near the limit while a multi-GB core file is written, and the restart begins only after the write finishes
Ruled out if
Restart is still slow with core dumps off or finished small: points to the server startup process, such as map loading or a DB cold cache (db-cold-cache)
Check with
Infra tools (no game code needed)
Sources
core(5) — Linux manual pageLinux man-pages RLIMIT_CORE caps core file size, coredump_filter selects which memory regions to include, and core dumps can be piped to a program for separate handling
Minidump FilesMicrosoft A minidump holds only a useful subset of crash dump information, so it is fast and small
coredumpctl(1) — Linux manual pagesystemd list: core dumps recorded in the journal (TIME is the crash time reported by the kernel); info: details for each dump and the size written to disk