If the game thread waits for the disk to finish every log line, the game stalls too whenever the disk is busy.
Why Combat and trade logs are written straight to a file from the game thread → Effect When a durable write (fsync) is required or the OS write buffer (page cache) hits its limit, a single write takes tens of ms while the disk is busy → On screen Hitches in log-heavy fights
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Log asynchronously (memory buffer + separate thread), reduce log volume, don’t fsync on the game thread.
Infra team action items
Run log rotation and compression at low I/O priority, keep logs on a different disk from data, monitor disk write latency.
On the graph
Random spikes · Server tick time, disk write latency
Where to look
Overlay w_await and aqu-sz from iostat -x 1 on tick time, and use perf trace -p PID --duration 10 to find write and fsync calls in the game server that took over 10 ms, along with their threads
Confirmed if
At the tick spikes, the game thread’s write and fsync calls take tens of ms, and disk write latency spikes at the same moment. Often lines up with log rotation or compression
Ruled out if
Ticks spike with no slow system calls on the game thread: another cause such as GC, locks, or tick overrun. Only the dedicated logging thread is slow: no effect on gameplay
Check with
Infra tools (no game code needed)
Learn more
Normally the OS accepts writes into memory (the page cache) first and flushes them to disk later, so a log line usually completes right away. Stalls happen when fsync demands a durable write, when backed-up writes exceed the limit and the OS blocks the write call, and when log files are rotated or compressed. That’s why it’s fine most of the time and spikes only at moments when the disk is busy.
Sources
fsync(2) — Linux manual pageLinux man-pages fsync flushes modified data all the way to the disk (including the disk cache) and blocks until the device reports completion
Documentation for /proc/sys/vm/Linux kernel When backed-up (dirty) writes reach dirty_ratio, the writing process has to do the disk writeback itself
ionice(1) — Linux manual pageutil-linux A job at idle I/O priority gets disk time only when no other program is using the disk
iostat(1) — Linux manual pagesysstat -x: w_await (average time per write request, including time waiting in the queue), aqu-sz (average queue length, formerly avgqu-sz)
perf-trace(1) — Linux manual pageperf -p traces the system calls of a running process; --duration shows only calls that took longer than the given ms