한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › Root causes of TCP retransmission

Route change / bad ECMP path Route change / bad ECMP member

Cause ID rt-path · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)

Open the interactive card with figures and simulations →

Packets vanish for a few seconds while an internet route changes, or steadily on connections assigned to a faulty path among several ECMP paths.

Why BGP route recalculation, or faulty equipment or a bad link on one of several paths (ECMP, LAG) → Effect Temporary loss during the route switch, or steady loss only on connections using that path → On screen A sudden freeze of a few seconds then fast-forward, or “it gets better after reconnecting” (assigned to a different path)

Symptoms
Freeze, Fast-forward, Teleporting
Factors
Packet loss
Who’s affected
Specific region/ISP
When
Randomly
Owner
Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)
Game team action items
Log per-connection retransmission stats (TCP_INFO) so you can pull the IP, port, and time for affected players, don’t drop connections that stall for a few seconds right away.
Infra team action items
Monitor retransmission rate per region and ISP, check whether reconnecting changes the path, secure links from multiple ISPs, check our ECMP and LAG paths for bad links, run route measurements on the same TCP port as the game (mtr --tcp --port; the path is chosen by address and port, so an ordinary ping can take a different path and look fine).
External action items
Report the bad path to the ISP with route measurements taken on the same TCP port and a comparison from before and after reconnecting.
On the graph
Step change · RTT (ping), retransmission rate per region and ISP
Where to look
Retransmissions grouped per connection with bcc tcpretrans -c to pull the affected players’ addresses and ports; mtr on the same TCP port as the game (mtr -T -P PORT) from the server toward the player and from the player toward the server, compared. Results before and after reconnecting compared too
Confirmed if
From a certain moment, RTT for one region or ISP shifts like a step and a few seconds of loss cluster, or even within one ISP only some connections (address and port combinations) keep retransmitting and get better after reconnecting. A plain ping can look fine while only TCP mtr shows loss
Ruled out if
All connections on that ISP getting worse together at evening peak: “Bottleneck queue overflow (congestion loss).” Only one player affected, with loss already on the ping to their router: “Wireless link loss”
Check with
Infra tools (no game code needed)
Real incidents
Cloudflare 2020: Traffic loss in some cities from a Cloudflare backbone configuration error

Sources

  1. RFC 2991: Multipath Issues in Unicast and Multicast Next-Hop Selection IETF
    Diagnostic tools such as ping and traceroute are hard to trust over multiple paths; describes pinning each flow to one path by hashing it
  2. RFC 2992: Analysis of an Equal-Cost Multi-Path Algorithm IETF
    ECMP picks the next hop from a hash of the header fields that identify a flow (the same flow takes the same path)
  3. tcp(7) — Linux manual page Linux man-pages
    TCP_INFO: query per-socket state (struct tcp_info)
  4. mtr(8) manual page source mtr
    -T (--tcp) uses TCP SYN in place of ICMP, -P (--port) sets the target port
  5. Demonstrations of tcpretrans, the Linux eBPF/bcc version IO Visor
    Shows one line per retransmission with the remote address and port; -c counts retransmissions per flow

See also

Same layer: Root causes of TCP retransmission

Same symptom (Freeze), other layers

View the interactive card with figures and simulations