Cause ID rt-mapping · Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Network infrastructure), Infra team (Server infrastructure)
If a device along the way deletes the mapping for an idle connection (the entry that records where to forward that connection), the next packet sent can’t be delivered. The connection either retransmits over and over until it disconnects, or the device sends back a reset (RST) and it disconnects right away.
Why A connection with no packets going either way for a while (AFK, lobby) → Effect A home router’s NAT, the ISP’s CGNAT, a firewall, a load balancer, or a cloud security group deletes the idle mapping → On screen When the player moves again, retransmissions go on until a disconnect, or the disconnect is immediate
Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Network infrastructure), Infra team (Server infrastructure)
Game team action items
Client: send heartbeats at no more than half the shortest idle timeout (mappings in players’ routers and ISP CGNAT are reliably refreshed only by outbound packets, and we can’t change their timeouts, so the client sends them), reconnect automatically after a disconnect. Server: answer heartbeats, and close the connection first when none arrive for a set time (shorten the TCP keepalive interval with socket options such as TCP_KEEPIDLE, detect failures quickly with TCP_USER_TIMEOUT), resume sessions with a session token.
Infra team action items
Network: collect the idle timeouts of firewalls and load balancers on the path and share them with the game team, extend them on our own firewalls and load balancers if needed. Servers/OS: check the cloud security group’s connection tracking timeout and share it with the game team.
Ballpark numbers
How long devices keep a TCP mapping varies widely, from a few minutes to several hours. When a cloud security group tracks connections, AWS Nitro v6 instance types delete the tracking entry after 350 seconds by default (other types after 5 days; see “Cloud security group connection tracking expiry”). The Linux TCP keepalive default is “probe after 2 hours idle,” which is later than most devices.
On the graph
Mass disconnect · Disconnects, idle time before disconnect
Where to look
The last few minutes of a dropped connection in a server-side packet capture; for live connections, idle time from lastsnd and lastrcv in ss -ti (ms since the last send and receive). nstat TcpExtTCPAbortOnTimeout (connections abandoned when the timer ran out) alongside
Confirmed if
Each dropped connection had just been idle past a similar value (the idle timeout of a device on the path, e.g., 350 s for security groups on AWS Nitro v6 instances), and from the first packet after the idle period, retransmissions go on with no ACK until the connection gives up, or an RST comes back right away
Ruled out if
Disconnects during play regardless of idle time point to a different cause (“Route change / bad ECMP path,” “Firewall and connection tracking drops”). Rule this cause out for connections that exchange heartbeats at no more than half the shortest idle timeout
Check with
Infra tools (no game code needed)
Sources
RFC 5382: NAT Behavioral Requirements for TCPIETF Recommends a TCP NAT established-connection idle timeout of at least 2 hours 4 minutes (on the premise that devices may delete idle sessions earlier)
Amazon EC2 security group connection trackingAWS The default TCP idle tracking timeout is 350 seconds on Nitro v6 instance types and 432,000 seconds (5 days) on other types; recommends keepalives at intervals shorter than 5 minutes
IP SysctlLinux kernel tcp_keepalive_time defaults to 2 hours
tcp(7) — Linux manual pageLinux man-pages TCP_KEEPIDLE (idle time before keepalive starts), TCP_USER_TIMEOUT (how long to wait for unacknowledged data before closing the connection)