Some cloud disks and small server sizes have burst credits that let them run faster than baseline for a while, so when a busy period drags on and the credits run out, speed drops suddenly.
Why Sustained use above baseline performance → Effect Burst credits run out and performance drops sharply to baseline → On screen Lag starts a few hours into every evening
Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)
Infra team action items
Servers/OS: use disks with provisioned performance (gp3, provisioned IOPS), alert on credit balance, also check the instance’s disk bandwidth burst limit and CPU credits. DB hosts: move DB disks, including managed databases, to provisioned performance too, alert on credit balance.
Ballpark numbers
An AWS gp2 100 GB disk normally gets 300 IOPS, bursts to 3,000, and lasts about 30 minutes on a full credit balance. gp3 has no credits and always gets 3,000. Small Azure Premium SSDs also burst on credits for up to 30 minutes.
On the graph
Hits a ceiling · IOPS, burst credit balance
Where to look
In CloudWatch, EBS BurstBalance (gp2, st1, sc1), the instance’s EBSIOBalance% and EBSByteBalance% (some instances that burst), and CPUCreditBalance on burstable instances. On Azure, burst credit usage metrics such as Data Disk Used Burst IO Credits Percentage
Confirmed if
From the moment the balance drops near 0, IOPS (VolumeReadOps, VolumeWriteOps) flattens at baseline, and VolumeQueueLength and lag rise together. Starts after the peak has lasted a few hours
Ruled out if
All balances are healthy but IOPS is flat: a fixed volume or instance limit (dk-iops)
Check with
Infra tools (no game code needed)
Learn more
Even with a healthy disk, small virtual servers have a burst limit on the instance’s own disk bandwidth (for example, at least 30 minutes a day), which produces the same pattern. Low-cost servers that run on CPU credits also slow down to baseline performance once the credits run out.
Sources
Amazon EBS General Purpose SSD volumesAWS gp2 baseline is 3 IOPS per GiB (minimum 100), bursting to 3,000 IOPS on I/O credits; 5.4 million credits last at least 30 minutes. gp3 always gets 3,000 IOPS with no burst
Managed disk burstingMicrosoft Azure Premium SSD P20 and smaller use credit-based bursting; a full credit balance gives 30 minutes at max burst speed
Amazon EBS-optimized instance typesAWS Some instances sustain maximum EBS performance for only 30 minutes once every 24 hours, then return to baseline
Standard mode for burstable performance instancesAWS In standard mode, a burstable instance that runs out of CPU credits lowers CPU utilization to the baseline level (gradually, without a sudden drop)
Amazon CloudWatch metrics for Amazon EBSAWS BurstBalance: remaining I/O credits for gp2 and throughput credits for st1 and sc1 (%); VolumeReadOps, VolumeWriteOps, VolumeQueueLength
CloudWatch metrics that are available for your instancesAWS EBSIOBalance% and EBSByteBalance%: remaining EBS credits on some instances that burst for 30 minutes once every 24 hours; CPUCreditBalance: remaining CPU credits on burstable instances
Disk metricsMicrosoft Azure Disk and VM burst credit usage (5-minute intervals), such as Data Disk Used Burst IO Credits Percentage