한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L5 Data center network equipment

Load balancer skew and misjudged health checks LB imbalance, bad health checks

Cause ID dc-lb-imbalance · Primary owner Infra team (Network infrastructure) · Also Game team (Server development)

Open the interactive card with figures and simulations →

Connections pile onto one server, or players keep getting sent to a server that’s already dead.

Why The distribution rule is a poor fit, or the health check can’t see the real state → Effect One server alone is overloaded, or players try to connect to a dead server → On screen Only some channels or some players get slow motion, can’t connect, or get infinite loading

Symptoms
Slow motion, Can’t connect / infinite loading
Factors
Stall, Packet loss
Who’s affected
Specific zone/channel
When
Right after login or maintenance, When crowds gather
Owner
Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
Game team action items
Implement a health check that answers the load balancer’s probes based on the real game state (tick progress, DB connections), report server load along with it.
Infra team action items
Switch to health checks that verify real game responses, distribute by server load, monitor differences in connection counts between servers.
On the graph
Outliers only · Connections/CPU utilization per server
Where to look
Overlay connection counts (ss -s) and CPU utilization of each server behind the load balancer on one graph, and compare the load balancer’s target health status (HealthyHostCount and UnHealthyHostCount in CloudWatch on AWS) with the game servers’ actual state
Confirmed if
Only one or two servers have far higher connections and CPU than the rest, or a server whose tick has stopped stays “healthy” and keeps taking new connections
Ruled out if
Connection counts even across servers but one channel is slow: load inside that channel (“Single-threaded zone overload (hotspot)”)
Check with
Infra tools (no game code needed)
Real incidents
AWS 2025: AWS us-east-1 DynamoDB DNS outage and long recovery

Sources

  1. Load Balancing in the Datacenter Google
    Plain round robin lets CPU usage differ by up to 2× between tasks; weighted distribution where backends report their load in responses and health checks; a lame duck state in which a backend asks not to be sent new requests
  2. Health checks for Network Load Balancer target groups AWS
    Default health check every 30 seconds, target removed after 2 failures; UDP services are checked with TCP or HTTP health checks, so configuring them to reflect the real service state is recommended
  3. CloudWatch metrics for your Network Load Balancer AWS
    HealthyHostCount and UnHealthyHostCount: number of targets judged healthy and unhealthy

See also

Same layer: L5 Data center network equipment

Same symptom (Slow motion), other layers

View the interactive card with figures and simulations