한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L12 Database

Cold cache (right after a restart) Cold buffer pool after restart

Cause ID db-cold-cache · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)

Open the interactive card with figures and simulations →

After a DB restart, the memory cache is empty, so for a while every lookup reads from disk.

Why The DB restarts for maintenance → Effect Frequently used data isn’t in memory, so it’s read from disk → On screen Logins and loading are slow for a while right after maintenance

Symptoms
Can’t connect / infinite loading, Input lag
Factors
Latency
Who’s affected
Whole server
When
Right after login or maintenance
Owner
Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Game team action items
Open gradually (use a login queue to ramp up the number of players logging in step by step).
Infra team action items
Warm the cache after a restart (check the buffer pool save/restore settings), pre-read the disk too on a DB restored from a snapshot.
On the graph
Surge after opening · Disk reads, buffer cache hit ratio
Where to look
MySQL: the ratio of Innodb_buffer_pool_reads (reads that missed the buffer pool and went to disk) to Innodb_buffer_pool_read_requests, and warm-up progress in Innodb_buffer_pool_load_status. PostgreSQL: blks_read and blks_hit in pg_stat_database. Also check disk reads on the DB server
Confirmed if
Right after the restart, disk reads spike and the hit ratio is low, recovering over time, and logins and loading are slow during that window
Ruled out if
Hit ratio normal but slow right after maintenance: a login storm and N+1 queries (db-login-storm) or the connection pool (db-pool)
Check with
Infra tools (no game code needed)
Learn more
MySQL saves the list of buffer pool pages at shutdown and reloads them in the background at startup, but filling the pool takes time. If a cloud DB was restored from a snapshot (a copy of a disk), the disk itself is also slow for every block read for the first time, so the slowdown lasts even longer.
Real incidents
Roblox 2021: Roblox 73-hour outage: contention in the service discovery (Consul) cluster

Sources

  1. Saving and Restoring the Buffer Pool State MySQL
    Saves the list of recently used pages (25% by default) at shutdown and reads them back at startup to shorten warm-up after a restart; both are on by default
  2. pg_prewarm — preload relation data into buffer caches PostgreSQL
    Periodically records shared buffer contents and loads them back after a restart (autoprewarm)
  3. Initialize Amazon EBS volumes AWS
    Volumes created from snapshots have higher latency and lower performance until all blocks have been fetched
  4. Server Status Variables MySQL
    Innodb_buffer_pool_reads (logical reads that missed the buffer pool and read straight from disk), Innodb_buffer_pool_read_requests, Innodb_buffer_pool_load_status (warm-up progress)
  5. The Cumulative Statistics System (PostgreSQL Documentation) PostgreSQL
    blks_read (blocks read from disk) and blks_hit (blocks found in the buffer cache) in pg_stat_database

See also

Same layer: L12 Database

Same symptom (Can’t connect / infinite loading), other layers

View the interactive card with figures and simulations