한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L13 Server architecture and operations

Logging and monitoring overload Logging / monitoring overhead

Cause ID in-monitoring · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Open the interactive card with figures and simulations →

During an outage, log volume explodes, and servers that ship logs synchronously get even slower because of the logging.

Why Errors make log and metric volume explode → Effect The log collector falls behind, and servers that send synchronously wait on it → On screen Stutter and freezes during an outage get worse because of logging

Symptoms
Stutter, Freeze
Factors
Stall
Who’s affected
Whole server
When
When crowds gather, Randomly
Owner
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Send asynchronously, sample, drop when the buffer overflows, batch repeated error logs into one.
Infra team action items
Size log collector capacity for the burst volume seen during outages, alert on collector backlog.
On the graph
Random spikes · Log volume, log collector queue
Where to look
Log lines and bytes per second on the server, plus the log collection agent’s queue and dropped count, alongside tick time. If a thread is stalled, use bcc offcputime -p to see whether it’s waiting on log writes or shipping
Confirmed if
When ticks spike, log volume jumps to tens of times normal, and the game thread’s wait time concentrates in log write/send call stacks
Ruled out if
Log volume is normal or the game thread isn’t waiting on logging: the log burst is only a result of the outage, so look separately for whatever caused the first errors
Check with
Infra tools (no game code needed)

Sources

  1. Logging in C# Microsoft
    .NET logging methods are synchronous, so with slow storage it recommends writing to fast storage first and moving the logs later
  2. Asynchronous loggers Apache Software Foundation
    Asynchronous logging absorbs short bursts in a queue, but if output stays slow, the queue fills and logging drops to the speed of the slowest output, or logs are dropped (Discard) depending on policy
  3. Demonstrations of offcputime, the Linux eBPF/bcc version IO Visor
    Sums the time threads spent blocked off the CPU (off-CPU) per call stack; -p selects the process

See also

Same layer: L13 Server architecture and operations

Same symptom (Stutter), other layers

View the interactive card with figures and simulations