한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L13 Server architecture and operations

Login queue cap and insufficient reconnect grace Login queue cap / no reconnect grace

Cause ID in-login-queue · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)

Open the interactive card with figures and simulations →

When players flood in right after launch or maintenance, the login queue hits its cap and turns away new arrivals, and players who were already waiting lose their place during a brief disconnect and go back to the end of the line.

Why More people try to connect than the login server can take at once, so it keeps a queue, and when the queue gets too long it refuses new entries to protect the server → Effect The longer the queue, the longer the wait, and a brief Wi-Fi or mobile network drop during that time costs the player their place → On screen Can’t connect / infinite loading, the game quits with an error while waiting, the player starts over at the back of the line

Symptoms
Can’t connect / infinite loading, Disconnect
Factors
Stall
Who’s affected
Whole server, Just me
When
Right after login or maintenance, Evening peak hours
Owner
Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)
Game team action items
Server: set the queue cap to what the login server can actually handle, hold a disconnected player’s place for a set time (reconnect grace period), show queue position and estimated wait, record queue length, refusals, and disconnects while waiting as metrics. Client: when disconnected while waiting, reconnect automatically to the same place without quitting the game, spread out retries with exponential backoff and jitter.
Infra team action items
Servers/OS: measure login and lobby server capacity with load tests before launch, prepare spare machines that can be added quickly at launch, graph queue metrics alongside connection attempts.
Ballpark numbers
At the 2021 FINAL FANTASY XIV expansion launch, new entries were refused once the queue passed 17,000 players per logical data center (Error 2002). If a player disconnected while waiting, the lobby server waited tens of seconds to 1 minute, and a player who reconnected within that time kept their place in the queue.
On the graph
Hits a ceiling · Login queue length, refusals at the cap, disconnects while waiting
Where to look
Queue length, average wait time, refusals at the cap, and disconnects while waiting, as recorded by the login and lobby servers, on the same graph as connection attempts
Confirmed if
Right after launch or maintenance, refusals rise while queue length flattens at the cap, and disconnects while waiting concentrate on Wi-Fi and mobile network players
Ruled out if
The queue is short but login is slow: points to the DB (db-login-storm) or the OS connection queue (so-backlog)
Check with
Game server or client logs and metrics
Learn more
A login rush that slows down the DB is covered in “Login storm and N+1 queries,” and an OS connection queue overflow in “Connection queue (listen backlog) overflow.” This entry is about the design of the login queue the game keeps on purpose. The queue cap is a safeguard that protects the login server, so you can’t remove it: refusing excess requests early is what keeps the server processing the requests it can handle. What matters is reducing what refusals and disconnects cost players, and the longer the queue gets, the more the errors fall on players with unstable connections such as Wi-Fi and mobile networks.
Real incidents
Square Enix 2021: FINAL FANTASY XIV expansion launch congestion and login queue errors

Sources

  1. Response to Congestion (as of Dec. 11) Square Enix
    Once the queue passed 17,000 players per logical data center, new entries were refused so the login servers wouldn’t go down (Error 2002); a player disconnected while waiting got tens of seconds to 1 minute from the lobby server to reconnect and resume mid-queue, and went to the back of the line after that
  2. Using load shedding to avoid overload Amazon Builders' Library
    Load shedding: refusing excess requests early so the server keeps processing the requests it can handle

See also

Same layer: L13 Server architecture and operations

Same symptom (Can’t connect / infinite loading), other layers

View the interactive card with figures and simulations