한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L11 Disk

fsync surge fsync storms

Cause ID dk-fsync · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (DB infrastructure)

Open the interactive card with figures and simulations →

Asking for data to be written to disk “for sure” takes 0.1 ms to tens of ms per request depending on the disk, and when requests pile up, the queue grows.

Why Scheduled saves and logout rushes send a flood of durable write requests → Effect The disk queue grows → On screen Lag at every save time, slow logouts and channel changes

Symptoms
Stutter, Input lag
Factors
Stall, Latency
Who’s affected
Whole server
When
At regular intervals, When crowds gather
Owner
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (DB infrastructure)
Game team action items
Batch saves (many saves in one fsync), spread out save times.
Infra team action items
Servers/OS: use server SSDs with power-loss protection, monitor disk queue length and fsync latency. DB hosts: if saves go to the DB, put the DB log disk on the same kind of SSD, monitor commit latency.
Ballpark numbers
The time per call varies by hardware, but roughly: server SSD (with power-loss protection) 0.1 ms, regular SSD 1 to a few ms, cloud disk 1–2 ms, HDD 10 ms or more. With one thread waiting on each call in turn, an HDD can’t manage even 100 per second.
On the graph
Periodic spikes · Disk queue length, flush and write latency
Where to look
Overlay f/s and f_await (flushes the disk handled and how long they took), plus w/s, aqu-sz, and w_await, from iostat -x 1 on scheduled save and logout times. Older sysstat shows aqu-sz as avgqu-sz. On cloud disks, check EBS VolumeQueueLength and VolumeAvgWriteLatency
Confirmed if
At every save time and logout rush, flush count and queue length spike together, and w_await and f_await reach several times normal. Saves and channel changes slow down at those moments
Ruled out if
Queue spikes at times unrelated to saves or logouts: backup or compression (dk-backup) or the IOPS limit (dk-iops). Slower at the same flush count: points to burst credit depletion (dk-burst)
Check with
Infra tools (no game code needed)

Sources

  1. fsync(2) — Linux manual page Linux man-pages
    fsync empties even the disk cache and blocks until the device reports completion
  2. Reliability (PostgreSQL Documentation) PostgreSQL
    Ordinary SATA disks and many SSDs have write caches that are lost on power failure; durable writes need a cache with battery or power-loss protection
  3. Amazon EBS General Purpose SSD volumes AWS
    Default cloud disk (gp3) latency is single-digit ms; io2 Block Express averages under 500 µs for 16 KiB I/O
  4. Designs, Lessons and Advice from Building Large Distributed Systems (LADIS 2009 keynote) Google
    One HDD seek 10 ms
  5. iostat(1) — Linux manual page sysstat
    -x: f/s and f_await (flush requests the disk handled and their average time), w/s, w_await, aqu-sz (formerly avgqu-sz)
  6. Amazon CloudWatch metrics for Amazon EBS AWS
    VolumeQueueLength (requests waiting to complete), VolumeAvgWriteLatency (1-minute average write latency, Nitro instances)

See also

Same layer: L11 Disk

Same symptom (Stutter), other layers

View the interactive card with figures and simulations