When the server process dies from an unhandled error, everyone on that server disconnects at the same time.
Why A fatal error such as a reference to something that doesn’t exist (null reference), bad data, or running out of memory → Effect The server (or zone) process exits → On screen Everyone disconnects at once, and progress since the last save may be rolled back
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Analyze crash dumps and fix the root cause, save often.
Infra team action items
Restart the process automatically, set up crash dump collection and retention, alert the moment a server goes down.
On the graph
Mass disconnect · Connection count, process restarts
Where to look
Core dump records in coredumpctl list (time, PID, terminating signal) and the service manager’s (systemd) records of abnormal exits and restarts. On Windows servers, dump files saved by WER
Confirmed if
At the moment the connection count dropped to near 0, the game server process exited abnormally and left a core dump
Ruled out if
Process stayed alive but connections dropped: points to network equipment or an idle timeout. A watchdog restart after a long hang: points to an infinite loop or deadlock
Check with
Infra tools (no game code needed)
Sources
Collecting User-Mode DumpsMicrosoft Configure Windows Error Reporting (WER) to collect full or mini dumps locally when a user-mode program crashes
systemd.service(5) — Linux manual pagesystemd Restart=on-failure automatically restarts the service after an abnormal exit, a kill by signal (including core dumps), or a watchdog timeout; recommended for long-running services
coredumpctl(1) — Linux manual pagesystemd list queries core dumps saved by systemd-coredump, showing crash time, PID, and the signal that caused the crash