한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L13 Server architecture and operations

Autoscaling delay Autoscaling lag

Cause ID in-autoscale · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

Open the interactive card with figures and simulations →

When players flood in, servers are added automatically, but getting them ready takes several minutes, and the existing servers are overloaded in the meantime.

Why Connections spike when an event starts → Effect Several minutes pass before new servers boot and are ready → On screen Slow motion, and players who can’t connect, for the first few minutes after the event starts

Symptoms
Slow motion, Can’t connect / infinite loading
Factors
Stall
Who’s affected
Whole server
When
When crowds gather, Right after login or maintenance
Owner
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Spread players across channels (players already in a crowded channel can’t be moved to a new server), cut startup and data loading time for new servers.
Infra team action items
Scale out ahead of events, keep warmed-up spare servers, when scaling in shut a server down only after its remaining players leave.
Ballpark numbers
1 to a few minutes to detect the load (because metrics are averaged over several minutes), then several more minutes to boot a new server, read game data, and fill caches.
On the graph
Surge after opening · Instance count, CPU utilization, queued logins
Where to look
Autoscaling activity history (when scale-out was decided, when new instances went into service) overlaid on CPU utilization and connection graphs. On AWS, the Auto Scaling group metrics (must be enabled to appear) GroupDesiredCapacity (target count), GroupPendingInstances (starting up), and GroupInServiceInstances (in service)
Confirmed if
For several minutes after a connection spike, only the desired count and pending instances rise while existing servers sit at their CPU limit, and things ease from the moment in-service instances increase
Ruled out if
Still slow after new instances come in: points to a cause other than server count (a shared resource such as the DB, a cascading failure)
Check with
Infra tools (no game code needed)
Learn more
Autoscaling is mostly used where a new server can simply take new players, such as login, gateway, and dungeon servers. Scaling in causes trouble too. If you scale in during the early-morning lull and shut servers down without waiting for the remaining players to leave, those players disconnect.
Real incidents
AWS 2021: AWS us-east-1 internal network congestion
AWS 2025: AWS us-east-1 DynamoDB DNS outage and long recovery

Sources

  1. Target tracking scaling policies for Amazon EC2 Auto Scaling AWS
    EC2 basic metrics come at 5-minute intervals (1 minute with detailed monitoring), so metrics at 1-minute or finer intervals are recommended for a fast reaction
  2. Amazon CloudWatch metrics for Amazon EC2 Auto Scaling AWS
    Group metrics are published every minute only when enabled; GroupDesiredCapacity (the count the group tries to maintain), GroupPendingInstances (instances not yet in service), GroupInServiceInstances (instances in service)
  3. Scheduled scaling for Amazon EC2 Auto Scaling AWS
    Adds and removes capacity ahead of time, at set times, to match predictable load changes
  4. Decrease latency for applications with long boot times using warm pools AWS
    For apps that take a long time to boot, a pool of pre-initialized instances (warm pool) cuts scale-out delay

See also

Same layer: L13 Server architecture and operations

Same symptom (Slow motion), other layers

View the interactive card with figures and simulations