Connections pile onto one server, or players keep getting sent to a server that’s already dead.
Why The distribution rule is a poor fit, or the health check can’t see the real state → Effect One server alone is overloaded, or players try to connect to a dead server → On screen Only some channels or some players get slow motion, can’t connect, or get infinite loading
Right after login or maintenance, When crowds gather
Owner
Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
Game team action items
Implement a health check that answers the load balancer’s probes based on the real game state (tick progress, DB connections), report server load along with it.
Infra team action items
Switch to health checks that verify real game responses, distribute by server load, monitor differences in connection counts between servers.
On the graph
Outliers only · Connections/CPU utilization per server
Where to look
Overlay connection counts (ss -s) and CPU utilization of each server behind the load balancer on one graph, and compare the load balancer’s target health status (HealthyHostCount and UnHealthyHostCount in CloudWatch on AWS) with the game servers’ actual state
Confirmed if
Only one or two servers have far higher connections and CPU than the rest, or a server whose tick has stopped stays “healthy” and keeps taking new connections
Ruled out if
Connection counts even across servers but one channel is slow: load inside that channel (“Single-threaded zone overload (hotspot)”)
Load Balancing in the DatacenterGoogle Plain round robin lets CPU usage differ by up to 2× between tasks; weighted distribution where backends report their load in responses and health checks; a lame duck state in which a backend asks not to be sent new requests
Health checks for Network Load Balancer target groupsAWS Default health check every 30 seconds, target removed after 2 failures; UDP services are checked with TCP or HTTP health checks, so configuring them to reflect the real service state is recommended