한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › Root causes of TCP retransmission

MTU black hole (only large packets keep getting lost) PMTU black hole

Cause ID rt-mtu · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)

Open the interactive card with figures and simulations →

If the largest packet size a link along the way can carry shrinks and the “too big” notice (ICMP) is blocked, large packets keep vanishing no matter how many times they’re resent.

Why The maximum size shrinks on a VPN or tunnel segment, and a firewall blocks the “too big” notices → Effect The sender, with no idea why, keeps retransmitting the same large packet, and the RTO doubles each time → On screen Fine normally, but when large data moves (inventory, crowded areas, loading into a zone), everything stops, including the small packets behind it, ending in a disconnect or infinite loading

Symptoms
Freeze, Disconnect, Can’t connect / infinite loading
Factors
Packet loss
Who’s affected
Specific region/ISP, Just me
When
During specific actions, Right after login or maintenance
Owner
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)
Game team action items
To lower it from the server side, set the socket’s maximum segment size (TCP_MAXSEG); just splitting messages into smaller pieces in game code won’t prevent it (TCP re-packs the outgoing data into MSS-sized segments).
Infra team action items
Network: clamp the MSS on edge devices, allow the “too big” ICMP (type 3 code 4, fragmentation needed) through firewalls and cloud network ACLs. Servers/OS: set the path MTU, make sure the server firewall and cloud security groups don’t block “too big” ICMP either, and as a last safety net set Linux tcp_mtu_probing=1.
Ballpark numbers
Usually 1,500 bytes; around 1,400 through a tunnel. If the same packet is retransmitted 5–6 times, the freeze exceeds 10 seconds.
On the graph
Outliers only · Per-connection RTO and backoff, disconnects per region and ISP
Where to look
Retransmissions on the problem connection from a server-side packet capture or bcc tcpretrans -s (shows sequence numbers), plus mss, pmtu, and backoff for that connection from ss -ti. From the server to that player’s address, a small ping compared with a 1,500-byte ping with DF set (ping -M do -s 1472)
Confirmed if
Full MSS-sized packets are retransmitted over and over with the same sequence number at doubling intervals, while smaller packets get through. No “too big” ICMP arrives (Wireshark filter icmp.type == 3 and icmp.code == 4), and the small ping gets replies while only the large DF ping vanishes without one
Ruled out if
Small packets vanishing too points to loss unrelated to size (“Bottleneck queue overflow (congestion loss),” “Route change / bad ECMP path”). If “too big” ICMP arrives and pmtu in ss -ti drops, path MTU discovery is working properly
Check with
Infra tools (no game code needed)
Learn more
tcp_mtu_probing=1 declares a black hole and lowers the MSS to 1,024 bytes only after retransmission timeouts have gone on for a few seconds (the equivalent of tcp_retries1=3). The connection is frozen until then, so keep it as a last safety net and put MSS clamping, which prevents the problem up front, first.

Sources

  1. RFC 1191: Path MTU discovery IETF
    Path MTU discovery: a packet that is too large triggers ICMP “fragmentation needed and DF set” (type 3 code 4)
  2. RFC 2923: TCP Problems with Path MTU Discovery IETF
    The PMTU black hole problem, where blocked ICMP makes only large packets keep vanishing
  3. RFC 4821: Packetization Layer Path MTU Discovery IETF
    A way for the transport layer to discover packet size without ICMP (the basis of Linux tcp_mtu_probing)
  4. IP Sysctl Linux kernel
    tcp_mtu_probing: 0 off, 1 only when a black hole is detected, 2 always (starting MSS is tcp_base_mss). tcp_retries1 defaults to 3
  5. net/ipv4/tcp_timer.c Linux kernel
    When RTO retransmissions go on for tcp_retries1 times, treats it as a detected black hole, turns on MTU probing, and lowers the MSS
  6. include/net/tcp.h Linux kernel
    TCP_BASE_MSS = 1,024 bytes
  7. RFC 6298: Computing TCP's Retransmission Timer IETF
    Doubles the RTO each time the retransmission timer expires
  8. iptables-extensions(8) — Linux manual page netfilter
    TCPMSS --clamp-mss-to-pmtu: works around large packets stalling on segments that block ICMP by adjusting the MSS in the SYN
  9. tcp(7) — Linux manual page Linux man-pages
    TCP_MAXSEG: maximum segment size for outgoing packets; set before connecting, it also changes the MSS advertised to the peer
  10. Network maximum transmission unit (MTU) for your EC2 instance AWS
    Traffic through internet gateways and VPNs uses an MTU of 1,500; PMTUD needs ICMP type 3 code 4, which doesn’t get through if security groups or network ACLs block it
  11. MTU considerations | Cloud VPN Google Cloud
    Cloud VPN gateway MTU is 1,460 bytes, and the payload MTU of an IPv4 tunnel is 1,406 bytes (around 1,400 through a tunnel)
  12. Demonstrations of tcpretrans, the Linux eBPF/bcc version IO Visor
    Shows one line per retransmission; -s also shows the sequence numbers of retransmitted packets
  13. ss(8) — Linux manual page iproute2
    mss, pmtu (path MTU), and backoff (how many times the RTO has doubled) in ss -i
  14. ping(8) — Linux manual page iputils
    -M do sets DF and won’t send packets larger than the path MTU the kernel knows; -s is the data size (the 8-byte ICMP header comes on top)
  15. Display Filter Reference: Internet Control Message Protocol Wireshark
    icmp.type and icmp.code display filters

See also

Same layer: Root causes of TCP retransmission

Same symptom (Freeze), other layers

View the interactive card with figures and simulations