When players flood in, servers are added automatically, but getting them ready takes several minutes, and the existing servers are overloaded in the meantime.
Why Connections spike when an event starts → Effect Several minutes pass before new servers boot and are ready → On screen Slow motion, and players who can’t connect, for the first few minutes after the event starts
When crowds gather, Right after login or maintenance
Owner
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Spread players across channels (players already in a crowded channel can’t be moved to a new server), cut startup and data loading time for new servers.
Infra team action items
Scale out ahead of events, keep warmed-up spare servers, when scaling in shut a server down only after its remaining players leave.
Ballpark numbers
1 to a few minutes to detect the load (because metrics are averaged over several minutes), then several more minutes to boot a new server, read game data, and fill caches.
On the graph
Surge after opening · Instance count, CPU utilization, queued logins
Where to look
Autoscaling activity history (when scale-out was decided, when new instances went into service) overlaid on CPU utilization and connection graphs. On AWS, the Auto Scaling group metrics (must be enabled to appear) GroupDesiredCapacity (target count), GroupPendingInstances (starting up), and GroupInServiceInstances (in service)
Confirmed if
For several minutes after a connection spike, only the desired count and pending instances rise while existing servers sit at their CPU limit, and things ease from the moment in-service instances increase
Ruled out if
Still slow after new instances come in: points to a cause other than server count (a shared resource such as the DB, a cascading failure)
Check with
Infra tools (no game code needed)
Learn more
Autoscaling is mostly used where a new server can simply take new players, such as login, gateway, and dungeon servers. Scaling in causes trouble too. If you scale in during the early-morning lull and shut servers down without waiting for the remaining players to leave, those players disconnect.
Amazon CloudWatch metrics for Amazon EC2 Auto ScalingAWS Group metrics are published every minute only when enabled; GroupDesiredCapacity (the count the group tries to maintain), GroupPendingInstances (instances not yet in service), GroupInServiceInstances (instances in service)