한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L9 Server game process

Server crash Server process crash

Cause ID sp-crash · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Open the interactive card with figures and simulations →

When the server process dies from an unhandled error, everyone on that server disconnects at the same time.

Why A fatal error such as a reference to something that doesn’t exist (null reference), bad data, or running out of memory → Effect The server (or zone) process exits → On screen Everyone disconnects at once, and progress since the last save may be rolled back

Symptoms
Disconnect, Dropped action / rollback
Factors
Stall
Who’s affected
Specific zone/channel, Whole server
When
Randomly, During specific actions
Owner
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Analyze crash dumps and fix the root cause, save often.
Infra team action items
Restart the process automatically, set up crash dump collection and retention, alert the moment a server goes down.
On the graph
Mass disconnect · Connection count, process restarts
Where to look
Core dump records in coredumpctl list (time, PID, terminating signal) and the service manager’s (systemd) records of abnormal exits and restarts. On Windows servers, dump files saved by WER
Confirmed if
At the moment the connection count dropped to near 0, the game server process exited abnormally and left a core dump
Ruled out if
Process stayed alive but connections dropped: points to network equipment or an idle timeout. A watchdog restart after a long hang: points to an infinite loop or deadlock
Check with
Infra tools (no game code needed)

Sources

  1. Collecting User-Mode Dumps Microsoft
    Configure Windows Error Reporting (WER) to collect full or mini dumps locally when a user-mode program crashes
  2. systemd.service(5) — Linux manual page systemd
    Restart=on-failure automatically restarts the service after an abnormal exit, a kill by signal (including core dumps), or a watchdog timeout; recommended for long-running services
  3. coredumpctl(1) — Linux manual page systemd
    list queries core dumps saved by systemd-coredump, showing crash time, PID, and the signal that caused the crash

See also

Same layer: L9 Server game process

Same symptom (Disconnect), other layers

View the interactive card with figures and simulations