When a game sends small packets sparsely, the RTO fires before “three following packets” can pile up. The same loss causes a much longer freeze than it would on a bulk transfer.
Why Packets go out about 100 ms apart, so only a few packets are ever in flight (not yet ACKed) → Effect Collecting three duplicate ACKs takes over 300 ms, so the RTO (ping + 200 ms) fires first, doubling on consecutive losses → On screen Each loss freezes the game for about 0.3 seconds; if the retransmission is lost too, the freeze lasts close to 1 second, then fast-forward
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Game team (Client development)
Game team action items
Server: turn on TCP_NODELAY (with Nagle on, RACK has no following packets to base its decision on), send real-time packets over UDP with your own retransmission. Client: turn on TCP_NODELAY, send real-time packets the same way as the server (UDP).
Infra team action items
Use RACK-TLP (the default on current Linux), use tcp_thin_linear_timeouts so consecutive RTOs don’t double.
Ballpark numbers
With 100 ms between packets and a 60 ms ping, fast retransmit takes about 360 ms (until three later packets arrive and their acknowledgments come back), while the RTO is about 260 ms. With RACK, the packet is resent right away at about 160 ms, when the acknowledgment for the next packet comes back. Once packets are more than 200 ms apart, RACK is no faster than the RTO either.
On the graph
Gap then burst · Per-connection receive volume, RTO expirations
Where to look
Deltas of nstat TcpExtTCPTimeouts (RTO expirations), TcpExtTCPFastRetrans (fast retransmits), and TcpExtTCPLossProbes and TcpExtTCPLossProbeRecovery (TLP), compared, plus rto and backoff on game connections in ss -ti. The server’s net.ipv4.tcp_recovery, tcp_early_retrans, and tcp_sack values too
Confirmed if
Among retransmissions, RTO expirations outnumber fast retransmits, and game connections often show backoff above 0 (in the middle of an RTO). Received volume sits at 0 during the freeze and then arrives all at once on recovery
Ruled out if
Bulk transfers on the same server stalling just as long: a loss problem unrelated to the connection’s traffic pattern. Concentrated on connections missing SACK or timestamps: “Middlebox strips TCP options”
Check with
Infra tools (no game code needed)
Learn more
Linux once had an option to retransmit on a single duplicate ACK for thin streams (tcp_thin_dupack), but it was removed in 2017, and RACK now fills that role. With Nagle on (TCP_NODELAY off), no new packets go out while the sender waits for the lost packet’s acknowledgment, so RACK has no following packets to base its decision on and the connection ends up waiting for the RTO.
Sources
Thin-streams and TCPLinux kernel Thin streams that send sparsely, like games, don’t trigger fast retransmit well and rely on long timeouts; the threshold is fewer than 4 packets in flight (not yet ACKed)