Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)
Infra team action items
Servers/OS: lower the I/O priority of backup, compression, and scan jobs, stagger their times. DB hosts: take backups from a replica.
On the graph
Periodic spikes · Disk utilization, disk wait time
Where to look
Overlay %util, await, and aqu-sz by day from the past few days of sar -d history (daily files in /var/log/sa; sadc must collect disk data with -S DISK), find the processes with the highest kB_rd/s and kB_wr/s at that time with pidstat -d 1, and match them against cron and systemd timer schedules
Confirmed if
await and %util spike at the same time every day, and backup, compression, or scan processes account for most disk reads and writes at that time
Ruled out if
Spikes at a different time each day: a scheduled job is unlikely. The game server itself does most of the I/O at that time: points to saves or logging (dk-fsync, dk-sync-log)
Check with
Infra tools (no game code needed)
Sources
ionice(1) — Linux manual pageutil-linux A job run in the idle class gets disk time only when no other program is using the disk
sar(1) — Linux manual pagesysstat -d: per-device await, aqu-sz, and %util from daily history files (default /var/log/sa); disk data must be collected with sadc’s -S DISK option