한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › Root causes of TCP retransmission

NAT or load balancer mapping expires mid-connection NAT / load balancer mapping expired mid-connection

Cause ID rt-mapping · Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Network infrastructure), Infra team (Server infrastructure)

Open the interactive card with figures and simulations →

If a device along the way deletes the mapping for an idle connection (the entry that records where to forward that connection), the next packet sent can’t be delivered. The connection either retransmits over and over until it disconnects, or the device sends back a reset (RST) and it disconnects right away.

Why A connection with no packets going either way for a while (AFK, lobby) → Effect A home router’s NAT, the ISP’s CGNAT, a firewall, a load balancer, or a cloud security group deletes the idle mapping → On screen When the player moves again, retransmissions go on until a disconnect, or the disconnect is immediate

Symptoms
Disconnect, Freeze
Factors
Packet loss
Who’s affected
Just me, Specific region/ISP
When
After sitting idle
Owner
Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Network infrastructure), Infra team (Server infrastructure)
Game team action items
Client: send heartbeats at no more than half the shortest idle timeout (mappings in players’ routers and ISP CGNAT are reliably refreshed only by outbound packets, and we can’t change their timeouts, so the client sends them), reconnect automatically after a disconnect. Server: answer heartbeats, and close the connection first when none arrive for a set time (shorten the TCP keepalive interval with socket options such as TCP_KEEPIDLE, detect failures quickly with TCP_USER_TIMEOUT), resume sessions with a session token.
Infra team action items
Network: collect the idle timeouts of firewalls and load balancers on the path and share them with the game team, extend them on our own firewalls and load balancers if needed. Servers/OS: check the cloud security group’s connection tracking timeout and share it with the game team.
Ballpark numbers
How long devices keep a TCP mapping varies widely, from a few minutes to several hours. When a cloud security group tracks connections, AWS Nitro v6 instance types delete the tracking entry after 350 seconds by default (other types after 5 days; see “Cloud security group connection tracking expiry”). The Linux TCP keepalive default is “probe after 2 hours idle,” which is later than most devices.
On the graph
Mass disconnect · Disconnects, idle time before disconnect
Where to look
The last few minutes of a dropped connection in a server-side packet capture; for live connections, idle time from lastsnd and lastrcv in ss -ti (ms since the last send and receive). nstat TcpExtTCPAbortOnTimeout (connections abandoned when the timer ran out) alongside
Confirmed if
Each dropped connection had just been idle past a similar value (the idle timeout of a device on the path, e.g., 350 s for security groups on AWS Nitro v6 instances), and from the first packet after the idle period, retransmissions go on with no ACK until the connection gives up, or an RST comes back right away
Ruled out if
Disconnects during play regardless of idle time point to a different cause (“Route change / bad ECMP path,” “Firewall and connection tracking drops”). Rule this cause out for connections that exchange heartbeats at no more than half the shortest idle timeout
Check with
Infra tools (no game code needed)

Sources

  1. RFC 5382: NAT Behavioral Requirements for TCP IETF
    Recommends a TCP NAT established-connection idle timeout of at least 2 hours 4 minutes (on the premise that devices may delete idle sessions earlier)
  2. RFC 4787: Network Address Translation (NAT) Behavioral Requirements for Unicast UDP IETF
    NAT mappings must be refreshed by outbound packets (REQ-6); refresh by inbound packets is optional (for UDP)
  3. Amazon EC2 security group connection tracking AWS
    The default TCP idle tracking timeout is 350 seconds on Nitro v6 instance types and 432,000 seconds (5 days) on other types; recommends keepalives at intervals shorter than 5 minutes
  4. IP Sysctl Linux kernel
    tcp_keepalive_time defaults to 2 hours
  5. tcp(7) — Linux manual page Linux man-pages
    TCP_KEEPIDLE (idle time before keepalive starts), TCP_USER_TIMEOUT (how long to wait for unacknowledged data before closing the connection)
  6. RFC 5482: TCP User Timeout Option IETF
    TCP user timeout: how long sent data can go unacknowledged before the connection is closed
  7. ss(8) — Linux manual page iproute2
    lastsnd and lastrcv in ss -i: time since the last send and receive (ms)
  8. SNMP counter Linux kernel
    TcpExtTCPAbortOnTimeout: connections abandoned without an RST because a TCP timer ran out

See also

Same layer: Root causes of TCP retransmission

Same symptom (Disconnect), other layers

View the interactive card with figures and simulations