한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L5 Data center network equipment

Network equipment failover Network device failover

Cause ID dc-failover · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), Game team (Client development)

Open the interactive card with figures and simulations →

When a router or firewall fails and traffic switches to the standby unit (failover), everyone freezes for a few seconds.

Why Switchover to standby equipment because of a failure or maintenance → Effect The switchover takes a few seconds, and connections reset if session state isn’t synced → On screen Every player on the server freezes at once; mass disconnects

Symptoms
Freeze, Disconnect
Factors
Packet loss
Who’s affected
Whole server
When
Randomly
Owner
Primary owner Infra team (Network infrastructure) · Also Game team (Server development), Game team (Client development)
Game team action items
Server: use timeouts that survive brief outages (a few seconds), let players resume their session with a session token when they reconnect after a disconnect. Client: reconnect automatically on disconnect (randomize retry intervals so everyone doesn’t pile in at once).
Infra team action items
Use redundancy that shares connection state, detect failures within 1 second with BFD, test failover regularly.
Ballpark numbers
About 1–3 seconds if the equipment detects the failure immediately. Without fast failure detection (BFD), relying only on default BGP timers, the route can be down for 90–180 seconds before neighboring equipment notices.
On the graph
Mass disconnect · Connections, total server traffic in/out
Where to look
Router and firewall event logs (VRRP role changes, BFD and BGP sessions going down, failover records) next to total server connections and traffic at the same time
Confirmed if
At the failover time in the device log, traffic for every server behind that device drops to 0 for a few seconds, or connection counts fall together
Ruled out if
Only one server’s connections drop: “Server crash” or “NIC driver and firmware problems.” Device logs clean and the frozen server is a single cloud VM: “Cloud host maintenance and live migration”
Check with
Infra tools (no game code needed)

Sources

  1. RFC 5880: Bidirectional Forwarding Detection (BFD) IETF
    Routing protocols’ Hello mechanisms take 1 second or more to detect a failure, so BFD was created to detect failures faster
  2. RFC 7938: Use of BGP for Routing in Large-Scale Data Centers IETF
    Relying only on BGP keepalives makes convergence slow; tearing down the session as soon as the link goes down detects failures within ms and reconverges
  3. RFC 5798: Virtual Router Redundancy Protocol (VRRP) Version 3 for IPv4 and IPv6 IETF
    VRRP advertisements default to 1 second, and the backup takes over the role when advertisements stop for more than about 3 intervals (just over 3 seconds with default settings)

See also

Same layer: L5 Data center network equipment

Same symptom (Freeze), other layers

View the interactive card with figures and simulations