When the primary DB dies, writes stop while it fails over to a standby, and the last data that hadn’t been replicated yet can be lost.
Why The primary DB fails and a standby is promoted → Effect No writes for seconds to minutes during the switch; with asynchronous replication, unreplicated data may be lost → On screen Every save fails for a moment, items and XP roll back
Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Game team action items
Make saves retryable, configure the connection pool and DNS cache to drop broken connections quickly and reconnect to the new address, verify reconnection during failover drills.
Infra team action items
Use synchronous or semi-synchronous replication (at the cost of write latency), run failover drills, monitor failover time and replication lag.
Ballpark numbers
Automatic failover on a managed DB usually takes tens of seconds to 2 minutes. With asynchronous replication, you can lose recent saves equal to the replication lag (under 1 second to a few seconds).
On the graph
Mass disconnect · DB connection count, write error count
Where to look
Put the DB’s failover records (on RDS, events RDS-EVENT-0013 failover started and RDS-EVENT-0049 failover completed; on self-managed DBs, promotion logs) and the game server’s DB connection count and connection error count on one graph. With asynchronous replication, also check replication lag just before the failure (RDS ReplicaLag, replay_lag in PostgreSQL pg_stat_replication)
Confirmed if
Save failures cluster in one window that lines up with the span between failover start and completion. The amount rolled back is close to the replication lag just before the failure. A game server whose errors continue after the failover finishes is still using connections opened to the old address
Ruled out if
Disconnects at times with no failover record: points to the network or DB overload
Failing over a Multi-AZ DB instance for Amazon RDSAWS Multi-AZ failover usually takes 60–120 seconds; connections must be reestablished afterward, and a JVM DNS cache TTL of 60 seconds or less is recommended
High availability for Amazon AuroraAWS Reads and writes fail during the outage, and recovery usually takes under 60 seconds (often under 30)
Semisynchronous ReplicationMySQL With asynchronous replication, committed transactions may be missing from the replica when the primary dies; semi-synchronous replication narrows this by waiting for one replica to acknowledge receipt, at the cost of higher latency