한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L11 Disk

Disk full Disk full

Cause ID dk-full · Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure), Game team (Server development)

Open the interactive card with figures and simulations →

When logs and dumps pile up and fill the disk, writes fail, and without safeguards the server crashes.

Why Logs, dumps, and temp files pile up to 100% → Effect Writes fail. Crash if there’s no error handling, failed saves if there is → On screen Disconnects, rolled-back progress

Symptoms
Disconnect, Dropped action / rollback
Factors
Stall
Who’s affected
Whole server
When
The longer it runs
Owner
Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure), Game team (Server development)
Game team action items
Handle write failures (retry the save and raise an alert so the server doesn’t crash), cut unnecessary logs and dumps.
Infra team action items
Servers/OS: log rotation, capacity alerts, separate log and data disks. DB hosts: watch that DB transaction logs (WAL, binlog) don’t pile up because replication stopped or a log backup was missed.
On the graph
Slow climb · Disk usage
Where to look
Usage from df -h and inode usage from df -i, plus ENOSPC errors in server and DB logs. For the DB: slots whose active is false in PostgreSQL pg_replication_slots, file count and size from MySQL SHOW BINARY LOGS, log_reuse_wait_desc in SQL Server sys.databases, and FreeStorageSpace on RDS
Confirmed if
Usage climbs steadily over several days, the moment it hits 100% lines up with crashes or failed saves, and ENOSPC shows up in the logs
Ruled out if
Writes fail with plenty of space left: another cause such as permissions or a file size limit
Check with
Infra tools (no game code needed)
Learn more
A DB’s transaction logs (WAL, binlog, and so on) are never deleted and keep piling up if a replica stops or a log backup is missed. When that disk fills, every write on the DB stops, and saves and trades fail all at once.

Sources

  1. write(2) — Linux manual page Linux man-pages
    Writes fail with ENOSPC when the device has no space left
  2. Monitoring Disk Usage (PostgreSQL Documentation) PostgreSQL
    If the WAL disk fills up, the DB server can panic and shut down
  3. Log-Shipping Standby Servers (PostgreSQL Documentation) PostgreSQL
    Replication slots keep WAL until the replica receives it, so they can fill pg_wal (limit with max_slot_wal_keep_size)
  4. Troubleshoot a full transaction log (SQL Server Error 9002) Microsoft SQL Server
    When the log is full, the DB is read-only and can’t be modified; missed log backups, replication lag, and long transactions are common causes that block log truncation; see what is blocking it in log_reuse_wait_desc of sys.databases
  5. df(1) — Linux manual page coreutils
    Usage per file system; -i shows inode usage in place of blocks
  6. pg_replication_slots (PostgreSQL Documentation) PostgreSQL
    active: whether the slot is currently streaming; wal_status: whether the WAL the slot retains has exceeded max_wal_size
  7. SHOW BINARY LOGS Statement MySQL
    List of the server’s binary log files and their sizes (File_size)
  8. Amazon CloudWatch metrics for Amazon RDS AWS
    FreeStorageSpace: storage space left on the DB instance

See also

Same layer: L11 Disk

Same symptom (Disconnect), other layers

View the interactive card with figures and simulations