한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L12 Database

Checkpoint / log flush Checkpoint / log flush stalls

Cause ID db-checkpoint · Primary owner Infra team (DB infrastructure)

Open the interactive card with figures and simulations →

Queries slow down at the moments the DB periodically writes its accumulated in-memory changes to disk in bulk.

Why Changes pile up and are periodically written to disk → Effect The disk gets busy at that moment and queries slow down → On screen Saves and loading slow down periodically

Symptoms
Input lag, Stutter
Factors
Latency
Who’s affected
Whole server, One feature only
When
At regular intervals
Owner
Primary owner Infra team (DB infrastructure)
Infra team action items
Spread checkpoints out in small, even steps, size the transaction log (redo log, WAL) generously, use fast disks.
On the graph
Periodic spikes · DB query latency, disk write volume
Where to look
PostgreSQL: checkpoint times and buffers written from the log_checkpoints log (on by default in recent versions), checkpoint counts (num_timed and num_requested in pg_stat_checkpointer on 17 and later, checkpoints_timed and checkpoints_req in pg_stat_bgwriter on 16 and earlier), and checkpoint_warning messages. MySQL: the gap between Log sequence number and Last checkpoint at in the LOG section of SHOW ENGINE INNODB STATUS. Overlay the server’s disk write volume and write latency
Confirmed if
Query latency spikes line up with checkpoint times, with disk write volume and write latency spiking at the same moments. In PostgreSQL, far more requested checkpoints (num_requested) than timed ones (num_timed) means WAL keeps hitting max_wal_size and checkpoints come early
Ruled out if
Spikes on a cycle unrelated to checkpoint times: backups or batch jobs (dk-backup, db-batch)
Check with
Infra tools (no game code needed)
Learn more
If the transaction log that records changes (the redo log in MySQL, WAL in PostgreSQL) is too small, the DB has to rush a checkpoint every time the log fills, and write throughput drops sharply for short periods.

Sources

  1. WAL Configuration (PostgreSQL Documentation) PostgreSQL
    By default, a checkpoint runs every 5 minutes or every 1 GB of WAL (max_wal_size) and is expensive because it writes all dirty pages. checkpoint_completion_target spreads the writes out to avoid I/O bursts. If checkpoints come closer together than checkpoint_warning, the log suggests raising max_wal_size
  2. Configuring Buffer Pool Flushing MySQL
    When the redo log fills, a sharp checkpoint briefly drops throughput; adaptive flushing spreads the writes out evenly
  3. Error Reporting and Logging (PostgreSQL Documentation) PostgreSQL
    log_checkpoints: logs the buffers written and time taken for each checkpoint; on by default
  4. The Cumulative Statistics System (PostgreSQL Documentation) PostgreSQL
    num_timed (checkpoints run because their scheduled time came) and num_requested (requested checkpoints) in pg_stat_checkpointer
  5. PostgreSQL 17 Release Notes PostgreSQL
    pg_stat_checkpointer added; checkpoint-related columns moved out of pg_stat_bgwriter
  6. The Cumulative Statistics System (PostgreSQL 16 Documentation) PostgreSQL
    Up to 16: checkpoints_timed and checkpoints_req in pg_stat_bgwriter
  7. InnoDB Standard Monitor and Lock Monitor Output MySQL
    LOG section: current log sequence number and last checkpoint position

See also

Same layer: L12 Database

Same symptom (Input lag), other layers

View the interactive card with figures and simulations