한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L11 Disk

Cloud disk out of burst credits Burst credit depletion

Cause ID dk-burst · Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)

Open the interactive card with figures and simulations →

Some cloud disks and small server sizes have burst credits that let them run faster than baseline for a while, so when a busy period drags on and the credits run out, speed drops suddenly.

Why Sustained use above baseline performance → Effect Burst credits run out and performance drops sharply to baseline → On screen Lag starts a few hours into every evening

Symptoms
Stutter, Slow motion, Input lag
Factors
Stall, Latency
Who’s affected
Whole server
When
Evening peak hours, The longer it runs
Owner
Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)
Infra team action items
Servers/OS: use disks with provisioned performance (gp3, provisioned IOPS), alert on credit balance, also check the instance’s disk bandwidth burst limit and CPU credits. DB hosts: move DB disks, including managed databases, to provisioned performance too, alert on credit balance.
Ballpark numbers
An AWS gp2 100 GB disk normally gets 300 IOPS, bursts to 3,000, and lasts about 30 minutes on a full credit balance. gp3 has no credits and always gets 3,000. Small Azure Premium SSDs also burst on credits for up to 30 minutes.
On the graph
Hits a ceiling · IOPS, burst credit balance
Where to look
In CloudWatch, EBS BurstBalance (gp2, st1, sc1), the instance’s EBSIOBalance% and EBSByteBalance% (some instances that burst), and CPUCreditBalance on burstable instances. On Azure, burst credit usage metrics such as Data Disk Used Burst IO Credits Percentage
Confirmed if
From the moment the balance drops near 0, IOPS (VolumeReadOps, VolumeWriteOps) flattens at baseline, and VolumeQueueLength and lag rise together. Starts after the peak has lasted a few hours
Ruled out if
All balances are healthy but IOPS is flat: a fixed volume or instance limit (dk-iops)
Check with
Infra tools (no game code needed)
Learn more
Even with a healthy disk, small virtual servers have a burst limit on the instance’s own disk bandwidth (for example, at least 30 minutes a day), which produces the same pattern. Low-cost servers that run on CPU credits also slow down to baseline performance once the credits run out.

Sources

  1. Amazon EBS General Purpose SSD volumes AWS
    gp2 baseline is 3 IOPS per GiB (minimum 100), bursting to 3,000 IOPS on I/O credits; 5.4 million credits last at least 30 minutes. gp3 always gets 3,000 IOPS with no burst
  2. Managed disk bursting Microsoft Azure
    Premium SSD P20 and smaller use credit-based bursting; a full credit balance gives 30 minutes at max burst speed
  3. Amazon EBS-optimized instance types AWS
    Some instances sustain maximum EBS performance for only 30 minutes once every 24 hours, then return to baseline
  4. Standard mode for burstable performance instances AWS
    In standard mode, a burstable instance that runs out of CPU credits lowers CPU utilization to the baseline level (gradually, without a sudden drop)
  5. Amazon CloudWatch metrics for Amazon EBS AWS
    BurstBalance: remaining I/O credits for gp2 and throughput credits for st1 and sc1 (%); VolumeReadOps, VolumeWriteOps, VolumeQueueLength
  6. CloudWatch metrics that are available for your instances AWS
    EBSIOBalance% and EBSByteBalance%: remaining EBS credits on some instances that burst for 30 minutes once every 24 hours; CPUCreditBalance: remaining CPU credits on burstable instances
  7. Disk metrics Microsoft Azure
    Disk and VM burst credit usage (5-minute intervals), such as Data Disk Used Burst IO Credits Percentage

See also

Same layer: L11 Disk

Same symptom (Stutter), other layers

View the interactive card with figures and simulations