When a bug keeps a tick from ever finishing, the server stops, and the watchdog forces a restart.
Why A loop that never ends because of a wrong condition, or runaway recursion → Effect The tick never finishes, and the server stops → On screen A freeze, then everyone disconnects
Cap loop iterations, add a watchdog, write tests that reproduce the problem input.
On the graph
Mass disconnect · Connection count, per-thread CPU
Where to look
Per-thread CPU from pidstat -t 1 during the hang, and which function the thread spinning at 100% is in, from perf top -t (thread ID) or gdb. If it has already restarted, the watchdog timeout record (WatchdogSec in systemd, a failed Kubernetes liveness probe)
Confirmed if
While the server is stopped, one game thread sits at 100% CPU, and its stack keeps looping inside the same function or loop
Ruled out if
CPU near 0 during the hang: points to a deadlock or waiting on an external response
Check with
Infra tools (no game code needed)
Sources
systemd.service(5) — Linux manual pagesystemd WatchdogSec=: if the service doesn’t send a keep-alive signal (WATCHDOG=1) within the set time, it’s treated as failed and stopped, then restarted automatically depending on the Restart= setting
Liveness, Readiness, and Startup ProbesKubernetes A liveness probe catches a state where the app is running but can’t make progress and restarts it; by default it checks every 10 seconds and restarts after 3 consecutive failures
pidstat(1) — Linux manual pagesysstat -t shows per-thread statistics (CPU utilization and so on) for the threads in a process