When internet routing information changes, packets are lost for the few seconds to tens of seconds (rarely a few minutes) it takes to converge again.
Why Routing information changes somewhere in an ISP’s network → Effect For a few seconds to tens of seconds, packets vanish or switch to a new route → On screen A sudden freeze of a few seconds, then ping settles at a different value (e.g., 40 → 70 ms)
Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)
Game team action items
Use timeouts that survive brief outages (don’t drop a connection right away just because it went silent for a few seconds).
Infra team action items
Monitor routes (watch for route and ping changes on our IP prefixes), detect failures on our own links within 1 second with BFD and fail over (the default BGP hold time is 90–180 seconds), move traffic to another link if the route switches to a long path and doesn’t come back.
External action items
Ask the ISP to investigate segments of its network where routes change often.
On the graph
Step change · RTT, traceroute path
Where to look
Compare traceroute and mtr paths from before and after the moment RTT changed, and check the BGP route change history for our prefix in RIPEstat BGPlay
Confirmed if
A freeze of a few seconds, then RTT moves to a different level, with BGP updates and AS path changes at the same time
Ruled out if
No route changes on record but high only in the evening: “Peak-hour congestion at peering links.” Only some connections bad: “One faulty ECMP path”
BGPlay (RIPEstat Data API)RIPE NCC Shows the BGP routes for an address prefix at the start time, the BGP updates observed during the period, and the ASes on the path