한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L13 Server architecture and operations

Auxiliary server outage Auxiliary service outage

Cause ID in-subservice · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Open the interactive card with figures and simulations →

When a server that runs separately from the game server, such as chat, party, or auction house, fails, only that feature stops working.

Why A server dedicated to one feature slows down or dies → Effect Only requests for that feature get no response → On screen Chat doesn’t work, party invites do nothing, the marketplace loads forever (combat is fine)

Symptoms
Dropped action / rollback, Can’t connect / infinite loading
Factors
Stall, Packet loss
Who’s affected
One feature only
When
Randomly, When crowds gather
Owner
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Design the game to keep running when a feature fails, show status per feature, don’t pile many features onto one central server.
Infra team action items
Set up health checks and alerts for each auxiliary server, add redundancy and automatic restarts.
On the graph
Mass disconnect · Request success rate per feature, auxiliary server connection count and health checks
Where to look
Health checks, process state, and connection count for each auxiliary server (chat, party, auction house), plus request success rate and response time per feature. Behind a load balancer, UnHealthyHostCount for the target group
Confirmed if
Only the server behind the reported feature fails health checks or shows a sharp drop in connections, while game server ticks and combat are normal
Ruled out if
Several features stopped at once: points to a central server that relays them all, or a cascading failure
Check with
Infra tools (no game code needed)
Learn more
If one central server (a world or manager server) relays parties, guilds, whispers, and cross-server moves, several features stop at once when that one server slows down.

Sources

  1. Bulkhead Pattern Microsoft Azure
    Isolating components into pools lets the rest keep working when one fails and keeps the failure from spreading
  2. REL05-BP01 Implement graceful degradation to transform applicable hard dependencies into soft dependencies AWS
    AWS Well-Architected. Design core functions to keep working on slightly stale or substitute data when a dependency fails
  3. CloudWatch metrics for your Network Load Balancer AWS
    UnHealthyHostCount: number of targets that health checks judged unhealthy

See also

Same layer: L13 Server architecture and operations

Same symptom (Dropped action / rollback), other layers

View the interactive card with figures and simulations