한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L11 Disk

Backup / compression / scan jobs Backup / compression / scans

Cause ID dk-backup · Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)

Open the interactive card with figures and simulations →

When early-morning backups, log compression, or security scans monopolize the disk, the game server’s reads and writes get held up.

Why A scheduled backup or compression job starts → Effect It takes most of the disk bandwidth and IOPS → On screen Lag at the same time every day

Symptoms
Stutter, Input lag
Factors
Stall, Latency
Who’s affected
Whole server
When
At regular intervals
Owner
Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)
Infra team action items
Servers/OS: lower the I/O priority of backup, compression, and scan jobs, stagger their times. DB hosts: take backups from a replica.
On the graph
Periodic spikes · Disk utilization, disk wait time
Where to look
Overlay %util, await, and aqu-sz by day from the past few days of sar -d history (daily files in /var/log/sa; sadc must collect disk data with -S DISK), find the processes with the highest kB_rd/s and kB_wr/s at that time with pidstat -d 1, and match them against cron and systemd timer schedules
Confirmed if
await and %util spike at the same time every day, and backup, compression, or scan processes account for most disk reads and writes at that time
Ruled out if
Spikes at a different time each day: a scheduled job is unlikely. The game server itself does most of the I/O at that time: points to saves or logging (dk-fsync, dk-sync-log)
Check with
Infra tools (no game code needed)

Sources

  1. ionice(1) — Linux manual page util-linux
    A job run in the idle class gets disk time only when no other program is using the disk
  2. Using Replication for Backups MySQL
    Stopping a replica to take a backup doesn’t affect the primary
  3. sar(1) — Linux manual page sysstat
    -d: per-device await, aqu-sz, and %util from daily history files (default /var/log/sa); disk data must be collected with sadc’s -S DISK option
  4. pidstat(1) — Linux manual page sysstat
    -d: per-process kB_rd/s and kB_wr/s (disk read and write volume per second)

See also

Same layer: L11 Disk

Same symptom (Stutter), other layers

View the interactive card with figures and simulations