한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L12 Database

Replication lag Replication lag

Cause ID db-replica-lag · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)

Open the interactive card with figures and simulations →

Writes go to the primary and reads come from replicas, so when a replica falls behind, data that was just written isn’t visible yet.

Why A burst of writes on the primary puts replicas several seconds behind → Effect Reading just-saved data from a replica finds it missing → On screen An item you just bought doesn’t show up, marketplace prices are stale, duplicate-reward bugs

Symptoms
Dropped action / rollback
Factors
Latency
Who’s affected
One feature only
When
When crowds gather, Evening peak hours
Owner
Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Game team action items
Read just-written data from the primary, check and grant rewards in one transaction on the primary (block duplicates with a unique key or a conditional UPDATE).
Infra team action items
Alert on replication lag, give replicas specs equal to or better than the primary and enable parallel replication, run bulk deletes in small chunks, manage long-running aggregate queries on replicas.
On the graph
Rises with load · Replication lag (seconds)
Where to look
MySQL: Seconds_Behind_Source from SHOW REPLICA STATUS on the replica (SHOW SLAVE STATUS on versions before 8.0.22). PostgreSQL: write_lag, flush_lag, and replay_lag in pg_stat_replication on the primary. RDS: ReplicaLag
Confirmed if
Lag is several seconds or more at the times of “it’s not showing up” reports, and things look normal once the lag clears. Lag grows during write bursts, bulk deletes, or long aggregate queries on the replica
Ruled out if
Lag near 0 but data still not showing up: points to the game server’s cache or sync
Check with
Infra tools (no game code needed)
Learn more
Replicas can fall behind even without heavy writes. A single bulk delete that took 10 minutes on the primary puts a replica that far behind while it replays, and long-running aggregate queries on a replica also slow its catch-up.

Sources

  1. SHOW REPLICA STATUS Statement MySQL
    Seconds_Behind_Source: time elapsed since the event the replica is currently applying was written on the primary (replication lag)
  2. Replica Server Options and Variables MySQL
    replica_parallel_workers lets multiple threads apply transactions in parallel (default 4; 0 means a single thread applies them in order)
  3. Log-Shipping Standby Servers (PostgreSQL Documentation) PostgreSQL
    Streaming replication is asynchronous by default, so there is a delay between commit and the replica applying it (usually under 1 second if the replica can keep up)
  4. MySQL 8.0 Reference Manual: SHOW REPLICA STATUS Statement MySQL
    From 8.0.22, SHOW REPLICA STATUS replaces SHOW SLAVE STATUS; earlier versions use SHOW SLAVE STATUS
  5. The Cumulative Statistics System (PostgreSQL Documentation) PostgreSQL
    write_lag, flush_lag, and replay_lag in pg_stat_replication: time from the primary writing WAL until the replica reports it has written, flushed to disk, and applied it
  6. Amazon CloudWatch metrics for Amazon RDS AWS
    ReplicaLag: how far a read replica lags behind its source (seconds)

See also

Same layer: L12 Database

Same symptom (Dropped action / rollback), other layers

View the interactive card with figures and simulations