한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L13 Server architecture and operations

Deploys and restarts Deploy / rolling restart

Cause ID in-deploy · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Open the interactive card with figures and simulations →

If you restart a server for an update without moving its connections, everyone on it disconnects, and the final saves before shutdown and the reconnects all hit at once.

Why Servers restart one after another to roll out a hotfix → Effect Each server shuts down without moving its connections, and saves for every player on it hit the DB at once → On screen Disconnects without notice, a surge of reconnects

Symptoms
Disconnect, Can’t connect / infinite loading, Input lag
Factors
Stall
Who’s affected
Whole server, Specific zone/channel
When
Randomly, Right after login or maintenance
Owner
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Add draining (block new connections only and wait for current players to leave), move characters to another server, spread out saves before shutdown, report ready after a restart only once cache loading and JIT warm-up are done, for hot reload load the new data ahead of time on a separate thread and swap it in all at once between ticks.
Infra team action items
Have the deploy tool wait for each server to drain before restarting it, route traffic to a restarted server only after confirming it’s ready (warm-up done), announce deploy times.
Ballpark numbers
With 5,000 players on one server, 5,000 saves hit the DB within a few seconds before shutdown.
On the graph
Mass disconnect · Connections per server, DB write count
Where to look
Deploy tool job history (restart time per server) overlaid as vertical lines (annotations) on graphs of connections, disconnects, DB writes, and login requests
Confirmed if
Connections per server drop sharply one server at a time at each restart, with DB writes spiking just before and login requests right after
Ruled out if
Disconnect times don’t line up with the deploy or restart history: points to a server crash or network equipment
Check with
Infra tools (no game code needed)
Learn more
Servers are also slow for a few minutes right after they come back. The cache is empty so DB queries pile up, and Java and C# servers haven’t yet finished optimizing code as it runs (JIT warm-up), so the same work takes longer. Reloading scripts or data tables without shutting down (hot reload) also stops the tick while it reads, causing a brief pause.

Sources

  1. Site Reliability Engineering, Chapter 20: Load Balancing in the Datacenter Google
    A server that receives SIGTERM goes into lame duck state, sending new requests to other servers and finishing only the ones in progress; for the first few minutes after a restart, before JIT optimization, it uses more resources, so it takes traffic only after warming up
  2. Liveness, Readiness, and Startup Probes Kubernetes
    A readiness check holds back traffic until connections are established, files are loaded, and caches are warmed
  3. Edit target group attributes for your Network Load Balancer AWS
    Deregistering a target stops new connections to it and drains existing ones (300 s by default)

See also

Same layer: L13 Server architecture and operations

Same symptom (Disconnect), other layers

View the interactive card with figures and simulations