한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L12 Database

DB failover Database failover

Cause ID db-failover · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)

Open the interactive card with figures and simulations →

When the primary DB dies, writes stop while it fails over to a standby, and the last data that hadn’t been replicated yet can be lost.

Why The primary DB fails and a standby is promoted → Effect No writes for seconds to minutes during the switch; with asynchronous replication, unreplicated data may be lost → On screen Every save fails for a moment, items and XP roll back

Symptoms
Dropped action / rollback, Freeze, Disconnect, Can’t connect / infinite loading
Factors
Stall, Packet loss
Who’s affected
Whole server
When
Randomly
Owner
Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Game team action items
Make saves retryable, configure the connection pool and DNS cache to drop broken connections quickly and reconnect to the new address, verify reconnection during failover drills.
Infra team action items
Use synchronous or semi-synchronous replication (at the cost of write latency), run failover drills, monitor failover time and replication lag.
Ballpark numbers
Automatic failover on a managed DB usually takes tens of seconds to 2 minutes. With asynchronous replication, you can lose recent saves equal to the replication lag (under 1 second to a few seconds).
On the graph
Mass disconnect · DB connection count, write error count
Where to look
Put the DB’s failover records (on RDS, events RDS-EVENT-0013 failover started and RDS-EVENT-0049 failover completed; on self-managed DBs, promotion logs) and the game server’s DB connection count and connection error count on one graph. With asynchronous replication, also check replication lag just before the failure (RDS ReplicaLag, replay_lag in PostgreSQL pg_stat_replication)
Confirmed if
Save failures cluster in one window that lines up with the span between failover start and completion. The amount rolled back is close to the replication lag just before the failure. A game server whose errors continue after the failover finishes is still using connections opened to the old address
Ruled out if
Disconnects at times with no failover record: points to the network or DB overload
Check with
Infra tools (no game code needed)
Real incidents
Riot Games 2021: League of Legends EUW 5-hour outage: one auxiliary DB halted the whole server

Sources

  1. Failing over a Multi-AZ DB instance for Amazon RDS AWS
    Multi-AZ failover usually takes 60–120 seconds; connections must be reestablished afterward, and a JVM DNS cache TTL of 60 seconds or less is recommended
  2. High availability for Amazon Aurora AWS
    Reads and writes fail during the outage, and recovery usually takes under 60 seconds (often under 30)
  3. Semisynchronous Replication MySQL
    With asynchronous replication, committed transactions may be missing from the replica when the primary dies; semi-synchronous replication narrows this by waiting for one replica to acknowledge receipt, at the cost of higher latency
  4. Log-Shipping Standby Servers (PostgreSQL Documentation) PostgreSQL
    Log shipping is asynchronous, so transactions not yet sent are lost if the primary dies; streaming replication lag is usually under 1 second
  5. Amazon RDS event categories and event messages AWS
    RDS-EVENT-0013: Multi-AZ failover started, RDS-EVENT-0049: Multi-AZ failover completed
  6. Amazon CloudWatch metrics for Amazon RDS AWS
    ReplicaLag: how far a read replica lags behind its source (seconds)
  7. The Cumulative Statistics System (PostgreSQL Documentation) PostgreSQL
    replay_lag in pg_stat_replication: time from the primary writing WAL until the replica reports having applied it

See also

Same layer: L12 Database

Same symptom (Dropped action / rollback), other layers

View the interactive card with figures and simulations