Asking for data to be written to disk “for sure” takes 0.1 ms to tens of ms per request depending on the disk, and when requests pile up, the queue grows.
Why Scheduled saves and logout rushes send a flood of durable write requests → Effect The disk queue grows → On screen Lag at every save time, slow logouts and channel changes
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (DB infrastructure)
Game team action items
Batch saves (many saves in one fsync), spread out save times.
Infra team action items
Servers/OS: use server SSDs with power-loss protection, monitor disk queue length and fsync latency. DB hosts: if saves go to the DB, put the DB log disk on the same kind of SSD, monitor commit latency.
Ballpark numbers
The time per call varies by hardware, but roughly: server SSD (with power-loss protection) 0.1 ms, regular SSD 1 to a few ms, cloud disk 1–2 ms, HDD 10 ms or more. With one thread waiting on each call in turn, an HDD can’t manage even 100 per second.
On the graph
Periodic spikes · Disk queue length, flush and write latency
Where to look
Overlay f/s and f_await (flushes the disk handled and how long they took), plus w/s, aqu-sz, and w_await, from iostat -x 1 on scheduled save and logout times. Older sysstat shows aqu-sz as avgqu-sz. On cloud disks, check EBS VolumeQueueLength and VolumeAvgWriteLatency
Confirmed if
At every save time and logout rush, flush count and queue length spike together, and w_await and f_await reach several times normal. Saves and channel changes slow down at those moments
Ruled out if
Queue spikes at times unrelated to saves or logouts: backup or compression (dk-backup) or the IOPS limit (dk-iops). Slower at the same flush count: points to burst credit depletion (dk-burst)
Check with
Infra tools (no game code needed)
Sources
fsync(2) — Linux manual pageLinux man-pages fsync empties even the disk cache and blocks until the device reports completion
Reliability (PostgreSQL Documentation)PostgreSQL Ordinary SATA disks and many SSDs have write caches that are lost on power failure; durable writes need a cache with battery or power-loss protection
iostat(1) — Linux manual pagesysstat -x: f/s and f_await (flush requests the disk handled and their average time), w/s, w_await, aqu-sz (formerly avgqu-sz)
Amazon CloudWatch metrics for Amazon EBSAWS VolumeQueueLength (requests waiting to complete), VolumeAvgWriteLatency (1-minute average write latency, Nitro instances)