한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L11 Disk

Synchronous log writes Synchronous logging

Cause ID dk-sync-log · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Open the interactive card with figures and simulations →

If the game thread waits for the disk to finish every log line, the game stalls too whenever the disk is busy.

Why Combat and trade logs are written straight to a file from the game thread → Effect When a durable write (fsync) is required or the OS write buffer (page cache) hits its limit, a single write takes tens of ms while the disk is busy → On screen Hitches in log-heavy fights

Symptoms
Stutter, Freeze
Factors
Stall
Who’s affected
Specific zone/channel, Whole server
When
When crowds gather
Owner
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Log asynchronously (memory buffer + separate thread), reduce log volume, don’t fsync on the game thread.
Infra team action items
Run log rotation and compression at low I/O priority, keep logs on a different disk from data, monitor disk write latency.
On the graph
Random spikes · Server tick time, disk write latency
Where to look
Overlay w_await and aqu-sz from iostat -x 1 on tick time, and use perf trace -p PID --duration 10 to find write and fsync calls in the game server that took over 10 ms, along with their threads
Confirmed if
At the tick spikes, the game thread’s write and fsync calls take tens of ms, and disk write latency spikes at the same moment. Often lines up with log rotation or compression
Ruled out if
Ticks spike with no slow system calls on the game thread: another cause such as GC, locks, or tick overrun. Only the dedicated logging thread is slow: no effect on gameplay
Check with
Infra tools (no game code needed)
Learn more
Normally the OS accepts writes into memory (the page cache) first and flushes them to disk later, so a log line usually completes right away. Stalls happen when fsync demands a durable write, when backed-up writes exceed the limit and the OS blocks the write call, and when log files are rotated or compressed. That’s why it’s fine most of the time and spikes only at moments when the disk is busy.

Sources

  1. fsync(2) — Linux manual page Linux man-pages
    fsync flushes modified data all the way to the disk (including the disk cache) and blocks until the device reports completion
  2. Documentation for /proc/sys/vm/ Linux kernel
    When backed-up (dirty) writes reach dirty_ratio, the writing process has to do the disk writeback itself
  3. ionice(1) — Linux manual page util-linux
    A job at idle I/O priority gets disk time only when no other program is using the disk
  4. iostat(1) — Linux manual page sysstat
    -x: w_await (average time per write request, including time waiting in the queue), aqu-sz (average queue length, formerly avgqu-sz)
  5. perf-trace(1) — Linux manual page perf
    -p traces the system calls of a running process; --duration shows only calls that took longer than the given ms

See also

Same layer: L11 Disk

Same symptom (Stutter), other layers

View the interactive card with figures and simulations