When a router or firewall fails and traffic switches to the standby unit (failover), everyone freezes for a few seconds.
Why Switchover to standby equipment because of a failure or maintenance → Effect The switchover takes a few seconds, and connections reset if session state isn’t synced → On screen Every player on the server freezes at once; mass disconnects
Primary owner Infra team (Network infrastructure) · Also Game team (Server development), Game team (Client development)
Game team action items
Server: use timeouts that survive brief outages (a few seconds), let players resume their session with a session token when they reconnect after a disconnect. Client: reconnect automatically on disconnect (randomize retry intervals so everyone doesn’t pile in at once).
Infra team action items
Use redundancy that shares connection state, detect failures within 1 second with BFD, test failover regularly.
Ballpark numbers
About 1–3 seconds if the equipment detects the failure immediately. Without fast failure detection (BFD), relying only on default BGP timers, the route can be down for 90–180 seconds before neighboring equipment notices.
On the graph
Mass disconnect · Connections, total server traffic in/out
Where to look
Router and firewall event logs (VRRP role changes, BFD and BGP sessions going down, failover records) next to total server connections and traffic at the same time
Confirmed if
At the failover time in the device log, traffic for every server behind that device drops to 0 for a few seconds, or connection counts fall together
Ruled out if
Only one server’s connections drop: “Server crash” or “NIC driver and firmware problems.” Device logs clean and the frozen server is a single cloud VM: “Cloud host maintenance and live migration”