한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › Root causes of TCP retransmission

Send bursts overflow shallow buffers Sender bursts overflow shallow buffers

Cause ID rt-burst · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure)

Open the interactive card with figures and simulations →

When a server sends a whole tick of updates for thousands of players in one instant, a switch’s small buffer or a cloud instance’s short-term limit overflows in under 1 ms and some packets are dropped.

Why At the start of each tick, the server sends everyone’s packets all at once → Effect A switch port buffer where traffic from many servers converges (hundreds of KB to a few MB per port) or a cloud instance limit overflows for an instant (average utilization stays low) → On screen Many players teleport or hitch at the same time; averaged metrics don’t reveal the cause

Symptoms
Teleporting, Freeze, Fast-forward
Factors
Packet loss
Who’s affected
Specific zone/channel, Whole server
When
When crowds gather
Owner
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure)
Game team action items
Spread each tick’s sends across the tick (per-connection pacing does little for thousands of connections that all send at tick start), stagger tick start times across servers, cap the rate of connections that send large data with SO_MAX_PACING_RATE.
Infra team action items
Servers/OS: cap the whole server’s send rate (a shaper in the server OS, Linux tc), smooth out a single connection’s bursts with pacing (Linux fq qdisc, BBR). Network: use deep-buffer switches, check switch port output drop counters at short intervals (average utilization won’t show them).
Ballpark numbers
A 10 Gbps port can send about 1.25 MB in 1 ms. When several servers’ ticks line up and converge on one port, the buffer fills almost instantly.
On the graph
Rises with load · Switch port output drops, retransmission rate
Where to look
Output discards (ifOutDiscards) collected every few seconds on the switch port the server connects to and on the port above it; in the cloud, bw_out_allowance_exceeded and pps_allowance_exceeded from ethtool -S. Line these up against retransmissions collected at the same moments with bcc tcpretrans
Confirmed if
Output discards or allowance overruns grow while per-minute average utilization stays low, and they scale with concurrent users and crowding in one spot. Retransmissions hit many connections on that server at the same moment, with no concentration on particular player IP ranges (ISP or region)
Ruled out if
CRC and input errors rising on the same port point to “Physical errors (bad cable, optics, connectors).” NIC drop counters or softnet dropped rising on the receiving server point to “Packet drops on the receiving host”
Check with
Infra tools (no game code needed)
Learn more
Pacing works per connection. When thousands of connections each send one or two packets at tick start, per-connection pacing does little to spread them out, so the game server has to split up its send timing itself. By contrast, when one connection sends a large amount of data, the NIC cuts tens of KB into packet-sized pieces and sends them back to back (TSO), and pacing spreads that kind of burst out well.

Sources

  1. High-Resolution Measurement of Data Center Microbursts Meta
    Over 70% of bursts at data center rack switches end within tens of µs, and per-minute average utilization correlates only weakly with drops (IMC 2017)
  2. tc-fq(8) — Linux manual page iproute2
    The fq qdisc paces each socket (connection), and SO_MAX_PACING_RATE sets a per-connection maximum rate
  3. net/ipv4/tcp_bbr.c Linux kernel
    BBR sets pacing_rate from its estimated bottleneck bandwidth
  4. IP Sysctl Linux kernel
    TCP sizes TSO frames to the flow’s rate (up to 64 KB, tcp_min_tso_segs)
  5. Monitor network performance for ENA settings on your EC2 instance AWS
    bw_out_allowance_exceeded, pps_allowance_exceeded: packets queued or dropped because the instance exceeded its bandwidth or packets-per-second limit
  6. RFC 2863: The Interfaces Group MIB IETF
    ifOutDiscards: outbound packets discarded even though no error was detected, for example to free up buffer space
  7. Demonstrations of tcpretrans, the Linux eBPF/bcc version IO Visor
    Shows one line per retransmission with the remote address, port, and connection state

See also

Same layer: Root causes of TCP retransmission

Same symptom (Teleporting), other layers

View the interactive card with figures and simulations