When requests exceed what the disk can handle per second, the queue grows and latency explodes.
Why Read and write requests approach the disk’s capacity → Effect The queue grows (usually exploding above 90% utilization) → On screen Slow saves and loading; freezes if the calls are blocking
Primary owner Infra team (Server infrastructure) · Also Game team (Server development), Infra team (DB infrastructure)
Game team action items
Merge requests, cache frequently read data, make disk access asynchronous so the game thread never waits on the disk.
Infra team action items
Servers/OS: use faster disks, check disk bandwidth and IOPS limits per instance type, alert on disk utilization and queue, run large file copies during quiet hours. DB hosts: alert on IOPS and throughput utilization for DB disks too, check the disk limits of the DB instance size.
Ballpark numbers
HDD about 150 IOPS, SATA SSD tens of thousands, NVMe hundreds of thousands. The default cloud disk (AWS gp3) gets 3,000 IOPS and 125 MiB per second. The per-second throughput limit is separate from IOPS, and when a large file copy fills it, even small writes get stuck behind it.
On the graph
Hits a ceiling · IOPS, disk queue length
Where to look
r/s and w/s, rkB/s and wkB/s, aqu-sz, and r_await and w_await from iostat -x 1. In the cloud, EBS VolumeReadOps, VolumeWriteOps, and VolumeQueueLength, the limit-exceeded checks VolumeIOPSExceededCheck and VolumeThroughputExceededCheck, and on the instance side InstanceEBSIOPSExceededCheck and InstanceEBSThroughputExceededCheck
Confirmed if
Requests per second or throughput flatten at the limit while aqu-sz and await shoot up together. In the cloud, the exceeded-check metric reads 1
Ruled out if
Even at 100% %util, low await means there may still be headroom. On SSDs and RAID that process requests in parallel, %util doesn’t mean the limit has been reached. High await without hitting the limit: latency of the disk itself (dk-hdd) or fsync (dk-fsync)
Check with
Infra tools (no game code needed)
Learn more
In the cloud, each server size (instance type) has its own disk bandwidth and IOPS limits, separate from the disk’s limits. Even with an expensive disk attached, a small server gets capped at the instance limit.
Sources
Exos X18 Data SheetSeagate 4K random reads on a 7,200 rpm server HDD: 170 IOPS (QD16)
D3-S4520 SSDSolidigm Server SATA SSD 4 KB random read/write up to 92K/48K IOPS
iostat(1) — Linux manual pagesysstat -x: r/s and w/s, rkB/s and wkB/s, aqu-sz (formerly avgqu-sz), r_await and w_await, %util. On RAID and modern SSDs that process requests in parallel, %util doesn’t indicate the performance limit
Amazon CloudWatch metrics for Amazon EBSAWS VolumeIOPSExceededCheck and VolumeThroughputExceededCheck: 1 if the volume tried to exceed its IOPS or throughput limit (Nitro instances); VolumeQueueLength