When some firewalls or accelerators remove or rewrite TCP options, multiple losses get recovered only one per round trip, or the window (how much can be sent at once) shrinks, and everything slows down.
Why A firewall’s “TCP normalization” or an old accelerator strips the SACK, timestamp, and window scale options → Effect With several packets lost, recovery goes one packet per round trip, and the window is capped at 64 KB → On screen Every loss causes a much longer freeze (without SACK, RACK-TLP can’t be used either), then fast-forward when it clears. Bulk transfers such as patches are slow too
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)
Infra team action items
Network: turn off TCP normalization on the device in question, check the firewall’s sequence number randomization too, compare the options in the SYN with packet captures at both ends. Servers/OS: check in ss -ti whether connections missing sack or wscale concentrate on a specific path (Windows PCs may not use ts depending on their settings, so ts alone missing can be normal), check that the server’s net.ipv4.tcp_sack is 1.
On the graph
Always high · Recoveries started without SACK (TcpExtTCPRenoRecovery)
Where to look
Whether each connection shows sack and wscale in ss -ti, the ratio of nstat TcpExtTCPRenoRecovery (recovery started without SACK) to TcpExtTCPSackRecovery, and TcpExtTCPSACKDiscard (SACK blocks discarded as inconsistent). On a suspect path, SYNs captured at both ends with their options compared (Wireshark tcp.options.sack_perm and similar)
Confirmed if
Only connections through a specific path or device lack sack and wscale, and TcpExtTCPRenoRecovery makes up a large share. The SACK-permitted option present in the SYN as sent is missing from the SYN as received. When sequence number randomization is the cause, the options survive but TcpExtTCPSACKDiscard rises
Ruled out if
sack missing on every connection: check the server’s net.ipv4.tcp_sack value first. Options intact and TcpExtTCPSACKDiscard flat: the slow recovery has another cause (“Slow recovery on thin streams”)
Check with
Infra tools (no game code needed)
Learn more
SACK can break even when the options survive. If a firewall’s sequence number randomization rewrites only the sequence numbers in the header and leaves the numbers inside SACK blocks untouched, the sender discards the inconsistent SACKs. A server where tcp_sack=0 was set during the 2019 SACK security issue and then forgotten ends up the same way.
SNMP counterLinux kernel TcpExtTCPRenoRecovery (recovery started without SACK), TcpExtTCPSackRecovery (recovery started with SACK), TcpExtTCPSACKDiscard (invalid SACK blocks)