A text edition covering 228 causes of stutter, teleporting, and disconnects in online games, layer by layer from your screen to the server’s database. Each cause lists its symptoms, the owning team (game or infra), numbers, and authoritative sources.
The interactive original, with figures and hands-on simulations, is the Game Lag White Paper. This edition collects the same causes, glossary, and sources on a single page that works without JavaScript. Each cause also has its own page (c/ID.html). The whole thing as one Markdown file is at llms-full.txt.
Browse by symptom
Lag starts with four factors: Latency (Distance, queues, and processing time make every packet arrive late by a steady amount.) Jitter (The average looks fine, but some packets arrive early and others late. Wi-Fi, congested links, and busy CPUs cause it.) Packet loss (Overflowing queues, radio interference, and faulty equipment drop packets. A connection that cuts out for a moment is loss too, just many packets in a row.) Stall (Server ticks run late or stop (GC, locks, blocking calls, overload), or frames on your PC stop. It happens even when the network is fine.)
Stutter (Causes: 60): Movement isn’t smooth: it keeps pausing briefly and moving again.
Teleporting (Causes: 47): A character moves to a distant spot in one step, with no movement in between.
Rubber-banding (Causes: 14): Your character runs forward, then gets dragged back to where it just was.
Fast-forward (Causes: 36): A frozen screen starts moving again, and the backlog of movement, hits, and damage plays out all at once at high speed.
Slow motion (Causes: 24): Everything moves slowly. Skill casts and monster movement look stretched out. Depending on the server design, speed can stay normal while it shows up as stutter or teleporting.
Input lag (Causes: 76): It takes a while from pressing a button to seeing the result. The screen itself can still be smooth.
Freeze (Causes: 67): Everything on screen stops for a moment (0.5 s to a few seconds), then moves again.
Dropped action / rollback (Causes: 36): Something you definitely did never happened, or its result gets reversed much later.
Disconnect (Causes: 51): The connection drops mid-game and you land back at the login screen or a reconnect dialog.
Invisible / ghost entities (Causes: 20): NPCs, monsters, or players that should be there are missing only on your screen, or entities that are already gone remain only on your screen.
Owning teams and owner codes
Code
Team
Owner
Scope
cli
Game team
Client development
Game client code: frames, GC, loading, interpolation, extrapolation, prediction, and the client’s network handling (including sending heartbeats and auto-reconnect)
srv
Game team
Server development
Game server code: ticks, threads, locks, netcode design, connection handling (accept loop, listen arguments), heartbeat replies and dead-connection cleanup, socket options, query and transaction design
net
Infra team
Network infrastructure
Circuits and data center network gear (switches, routers, firewalls, load balancers, DDoS protection); cloud network ACLs, VPC routing, and load balancers; ISPs and peering
sys
Infra team
Server infrastructure
Server hardware and cloud instances (including security groups and connection tracking), OS and kernel settings, NICs, deployment and monitoring environments
dba
Infra team
DB infrastructure
DB servers and storage; DB configuration, replication, and backups; cache servers
ext
External
External
Players’ PCs and home networks, ISP segments (outside our contracts), cloud providers. We can’t fix these directly, so we respond with player guidance, requests to the provider, or workarounds
ID cg-hitch · Primary owner Game team (Client development)
One frame takes several times longer than usual to compute, so the screen freezes for a moment.
Why A burst of skill effects, a mass spawn, or a full UI refresh all land in one frame → Effect The frame can’t finish within 16.7 ms and takes 50–300 ms → On screen The screen hitches, then everyone moves at once on the next frame
When crowds gather, During specific actions, Randomly
Owner
Primary owner Game team (Client development)
Game team action items
Split heavy work across several frames, find spiking frames with a profiler, cap the number of effects.
Ballpark numbers
At 60 FPS, one frame is 16.7 ms. A single frame over 50 ms is often enough to make players feel “it stuttered.”
On the graph
Random spikes · Frame time
Where to look
Record frame time (FrameTime) and the time the CPU and GPU spent on each frame (CPUBusy, GPUBusy) with PresentMon during play. On mobile, the slow sessions and slow rendering metrics in Android vitals
Confirmed if
Frame time, normally around 16.7 ms, jumps above 50 ms at the same moments as skill effects, mass spawns, or full UI refreshes, while ping stays the same
Ruled out if
Frame time steady but other characters hitch: points to the network side (e.g., “Missing or too-short interpolation buffer”). Spikes at regular intervals: check “Client garbage collection” first
Check with
The player’s own environment
Sources (3)
Slow renderingAndroid (Google) To hit 60 FPS, a frame must render within 16 ms; late frames get skipped and show up as stutter (jank)
Slow Sessions (games only)Android (Google) Android vitals counts a game frame as slow when it takes longer than 50 ms (20 FPS) or 34 ms (30 FPS)
ID cg-gc · Primary owner Game team (Client development)
The whole game freezes while it reclaims memory that was used and thrown away (garbage). The telltale sign is stutter at regular intervals.
Why Temporary strings, arrays, and lists are created and thrown away every frame → Effect Once garbage piles up, GC pauses the main thread to reclaim it → On screen Regular stutter every few seconds to tens of seconds
Reduce allocations (avoid string concatenation, LINQ, and lambda captures), use object pools, keep incremental GC on (the default since Unity 2020), run GC ahead of time when a pause does no harm, such as on a loading screen.
Ballpark numbers
Usually a few ms to 100 ms per pause, longer on low-end phones or in games that use a lot of memory (about 150–170 ms in the simulation). Unity’s GC scans the whole heap every time it runs, so the more memory the game is using, the longer it takes.
On the graph
Periodic spikes · Frame time, GC run times
Where to look
In a development build, check the GC.Collect and GC.Alloc markers in the Unity Profiler. In Unreal, stat GC and stat Hitches (logs frames longer than the threshold set by t.HitchFrameTimeThreshold)
Confirmed if
Every spiking frame contains a GC.Collect section about as long as the spike, at regular intervals of a few seconds to tens of seconds. GC.Alloc per frame rises in crowded places
Ruled out if
No GC section in the spiking frames: “Synchronous loading and shader compilation on the main thread” or “Rendering load from large crowds”
Check with
Game server or client logs and metrics
Learn more
Common in clients that use C#, such as Unity. The usual culprit is code that builds combat log strings, damage numbers, and UI text from scratch every frame. If it only stutters in crowded places, some code is producing garbage in proportion to the player count. Incremental GC collects a little at a time in each frame (3 ms by default in Unity), but if the game makes garbage faster than it collects, it still ends up pausing all at once. Unreal Engine also has its own GC that cleans up unused game objects. It depends on the engine version and settings, but with default settings it runs about once a minute and can cause a hitch every minute. Clients that write game rules in a scripting language such as Lua also run that script’s GC separately.
Sources (5)
Garbage collection modesUnity Incremental GC is the default and collects across several frames; with it off, the main thread stops while the whole heap is scanned, which can take up to hundreds of ms
Profiler markers referenceUnity GC.Collect: time program code is paused during garbage collection (under 1 ms to hundreds of ms); GC.Alloc: managed heap allocations
Stat Commands in Unreal EngineEpic Games stat GC (garbage collection statistics), stat Hitches (logs frames that exceed t.HitchFrameTimeThreshold)
ID cg-sync-load · Primary owner Game team (Client development)
The game freezes to read files and build shaders right before it draws an area, monster, or effect for the first time.
Why Entering a new area, or a skill, piece of gear, or monster appearing for the first time → Effect The main thread waits for file reads and shader compilation → On screen A 0.1–1 s freeze the first time only; fine from the second time on
While moving or changing zones, During specific actions
Owner
Primary owner Game team (Client development)
Game team action items
Load asynchronously, preload (prewarm), precompile shaders on the loading screen or at first launch, make use of loading screens.
Ballpark numbers
Compiling one shader takes tens of ms, sometimes over 100 ms. Loading a texture takes tens to hundreds of ms depending on storage speed.
On the graph
Surge after opening · Frame spike count (right after a patch or driver update)
Where to look
Record the same route twice with PresentMon and compare the first visit with the second. In a development build, turn on r.PSOPrecache.Validation in Unreal and check stat PSOPrecache and “PSO PRECACHING MISS” in the log; in Unity, check the loading and shader sections of spiking frames in the Profiler’s Timeline view
Confirmed if
0.1–1 s spikes only at places visited or skills used for the first time, gone on the second visit. Reports surge right after a patch or graphics driver update, then taper off
Ruled out if
Spikes every time at the same place: not a shader cache problem. Repeats on every move only on PCs with slow storage: “Slow storage delays asset streaming.” Dedicated GPU memory full: “Out of graphics memory (VRAM)”
Check with
Game server or client logs and metrics
Learn more
On PC, the graphics driver stores each shader it builds in a shader cache and reuses it. So right after a graphics driver update or a game patch, that cache is invalidated, and even players who were fine stutter again for a while. The classic report is “after the patch, it hitches everywhere I go for the first time.”
Sources (5)
Shader loadingUnity The first time a shader variant is used, the graphics driver has to build it for the GPU, which can cause a noticeable freeze; once built, it’s cached and doesn’t freeze again
Direct3D 12 Return CodesMicrosoft D3D12_ERROR_DRIVER_VERSION_MISMATCH: a PSO cache built with a different driver version can’t be reused (recompiled after a driver update)
PSO Precaching for Unreal EngineEpic Games With r.PSOPrecache.Validation on, stat PSOPrecache shows missed PSO statistics and the log records “PSO PRECACHING MISS”; runtime PSO creation over 20 ms (default) counts as a hitch
ID cg-asset-stream · Primary owner Game team (Client development) · Also External (External)
On slow storage such as an HDD, reading open-world textures and models can’t keep up with movement, so objects appear late or the game stutters while it waits for reads.
Why Moving fast on a mount or by teleport, or entering a crowded area, suddenly calls for many new textures and models → Effect Slow storage such as an HDD can’t read at the needed speed, so read requests pile up, and some loads make the main thread wait until they finish → On screen Textures stay blurry for a while, buildings and characters pop in late, and the game stutters or freezes while it waits on reads
While moving or changing zones, When crowds gather
Owner
Primary owner Game team (Client development) · Also External (External)
Game team action items
Use asynchronous streaming so the main thread never waits on reads, prefetch based on movement direction and speed, show low-resolution textures (mipmaps) and simple models first and swap them in later, briefly cap movement speed or use a loading screen when streaming falls behind during fast travel, state in the minimum and recommended specs whether an SSD is required.
External action items
Tell players to install the game on an SSD and to check whether downloads or antivirus scans are running on the same disk.
Ballpark numbers
According to Microsoft, older hard drives read tens of MB per second and NVMe SSDs read several GB per second, while previous-generation games used around 50 MB per second for streaming. A game built around SSD speeds that reads far more than that can easily fall behind your movement when it runs from an HDD.
On the graph
Outliers only · Frame spike count (by storage type), disk read latency
Where to look
While moving fast, record PhysicalDisk\Avg. Disk sec/Read (average time per read) and Current Disk Queue Length in Windows Performance Monitor together with PresentMon frame time. In a development build, check Unreal’s stat Streaming and stat AsyncLoad, or AssetBundle.asset/allAssets warnings in the Unity Profiler (the result was requested before loading finished, so the main thread waits)
Confirmed if
Disk read latency and queue length shoot up when moving fast, and at the same moments frames spike or textures and objects load late. Gone when the same scene runs from an SSD
Ruled out if
Disk idle but textures blurry: “Out of graphics memory (VRAM).” Fine from the second visit to the same place: “Synchronous loading and shader compilation on the main thread”
Check with
The player’s own environment
Learn more
If it freezes only the first time and is fine afterward, it’s closer to “Synchronous loading and shader compilation on the main thread.” If it repeats on every move only on PCs with slow storage, it’s this cause. “Out of graphics memory (VRAM),” where textures get dropped and reloaded because graphics memory runs short, also causes blurry textures, so check disk read waits and graphics memory usage together. Antivirus real-time scanning can also step in every time the game opens a file and slow reads down further.
Sources (8)
DirectStorage is coming to PCMicrosoft Older hard drives read tens of MB per second and NVMe SSDs several GB per second; the asset streaming budget of previous-generation games was around 50 MB per second; open-world games read and discard distant scenery in real time as the player moves
Texture and mesh loadingUnity Synchronous upload reads and uploads in one frame on the main thread and causes a visible freeze; asynchronous upload streams over several frames
Texture Streaming Overview for Unreal EngineEpic Games The streamer raises and lowers texture resolution (mips) to match the camera view, does most of the work on async worker threads, and loads the mips visible on screen first
Profiler markers referenceUnity AssetBundle.asset/allAssets warning: the result was requested before loading finished, so the main thread stops and waits
Stat Commands in Unreal EngineEpic Games stat Streaming (memory use and count of streaming textures), stat AsyncLoad (async loading statistics)
Windows Performance Monitor Disk Counters ExplainedMicrosoft Avg. Disk sec/Read is the average time one read takes to complete (I/O latency); Current Disk Queue Length is the disk queue length at the moment of measurement
ID cg-crowd · Primary owner Game team (Client development)
When hundreds of players fill one screen, as in a siege or a world boss fight, the cost of drawing them is more than the device can handle.
Why Hundreds of players and effects overlap on one screen → Effect Animation, shadow, name tag, and effect costs grow with the player count → On screen FPS drops 60 → 15: all movement stutters, and input lag sets in
Simplify by distance (LOD), cap the number of players shown, offer a simplified effects option, lower the animation update rate.
Ballpark numbers
Even at 0.02–0.1 ms per character, 300 characters add up to 6–30 ms. The frame budget at 60 FPS is 16.7 ms, so this alone uses over 1/3 of it, and at the high end exceeds it.
On the graph
Rises with load · Frame time, characters on screen
Where to look
Compare PresentMon frame time and CPUBusy/GPUBusy before and after a siege or world boss. In a development build, Unreal’s stat Unit (game thread, rendering thread, and GPU time)
Confirmed if
Frame time climbs as more players come on screen and improves right away when a cap on displayed players or the simplified effects option is turned on
Ruled out if
Spikes regardless of player count: “Frame time spike” or “Client garbage collection.” FPS fine but only other players’ movement lags behind: “Packet processing bottleneck on the main thread”
Check with
The player’s own environment
Sources (4)
Introduction to level of detailUnity Without LOD, even objects that look tiny on screen are drawn at full complexity; LOD cuts the rendering cost
ID cg-net-mainthread · Primary owner Game team (Client development) · Also Game team (Server development)
If the client processes only a fixed amount of received packets per frame, a flood of packets keeps getting pushed to the next frame.
Why Thousands of updates per second arrive in crowded places → Effect The main thread hits its per-frame processing limit and can’t read them all → On screen Other players’ movements show up later and later, then all at once
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: receive and parse on a separate thread, merge stale position updates for the same entity and apply only the latest. Server: in crowded places, send updates for distant characters less often to cut traffic.
Ballpark numbers
Once unprocessed packets pile up, it takes only a few seconds to fall a full second behind.
On the graph
Rises with load · Unprocessed received packets, receive-to-apply delay
Where to look
Log how many packets the client leaves unprocessed each frame and the delay from packet arrival to when the game applies it, and view them alongside the number of nearby players
Confirmed if
In crowded places the leftover packet count and apply delay keep growing, while ping and the server’s send interval stay normal
Ruled out if
No apply delay but the packets themselves arrive late: the network path. Frame time rises sharply: “Rendering load from large crowds”
Check with
Game server or client logs and metrics
Sources (2)
Actor Priority in Unreal EngineEpic Games When bandwidth runs short, not every actor is replicated every time; actors are prioritized by distance from the viewer and time since their last replication
Replication Graph in Unreal EngineEpic Games Games with many players and replicated objects (such as MMORPGs) need to group them by location and send only what each player needs, or the server CPU becomes the bottleneck
ID cg-no-buffer · Primary owner Game team (Client development) · Also Game team (Server development)
If the client draws server packets the moment they arrive, jitter (variation in packet arrival times) shows up directly on screen.
Why Received positions are drawn immediately, or the buffer is shorter than the jitter → Effect Characters stop for as long as a packet is late, then jump when delayed packets arrive together → On screen Other characters move in fits and starts
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: add an interpolation buffer, adjust its length automatically to connection quality. Server: apply lag compensation (rewind-based hit registration) so hits still register correctly when players aim at positions a buffer’s length in the past.
Ballpark numbers
The usual buffer is about twice the server’s send interval (100 ms when receiving 20 updates a second).
On the graph
Always high · Packet arrival intervals, times the interpolation buffer ran empty
Where to look
On the client, record the distribution of server packet arrival intervals and the number of frames that froze or fell back to extrapolation because there was no next snapshot to interpolate to
Confirmed if
Variation in arrival intervals often exceeds the interpolation buffer length, and each time the buffer runs empty and other characters hitch. A longer buffer reduces it
Ruled out if
Hitches even with enough buffer: check whether the server’s send interval itself is irregular (tick delays)
Check with
Game server or client logs and metrics
Learn more
A longer buffer makes movement smoother, but you see opponents that much further in the past. That’s why hit registration comes paired with lag compensation, where the server rewinds to “the past that player was seeing” to check the hit.
Physics (Netcode for Entities 6.5)Unity Lag compensation: the server finds the collision world the client was seeing at that tick and decides whether the shot hit
ID cg-extrap · Primary owner Game team (Client development)
While no packets arrive, the client keeps showing characters moving at their last velocity, then snaps them back when it turns out to be wrong.
Why Packets stop arriving, so the character keeps moving in its last direction and speed → Effect In reality, the other player stopped or changed direction → On screen The other character runs on for a while, then snaps to its real position or passes through walls. With erratic packet arrival intervals, it keeps overshooting and getting pulled back, so it looks like it’s shaking
Cap extrapolation time (e.g., 200–250 ms), converge smoothly when wrong.
Ballpark numbers
At 6 m/s, being off by just 300 ms puts a character 1.8 m out of place.
On the graph
Random spikes · Extrapolation time, position correction distance
Where to look
Record how long other characters are drawn by extrapolation and how far their position is corrected when a new packet arrives
Confirmed if
Every time packets stop, extrapolation time grows with no cap, then the correction distance grows to several meters
Ruled out if
Teleporting even though extrapolation runs only briefly: packet loss or latency itself is high, so look at the connection or route
Check with
Game server or client logs and metrics
Sources (2)
Interpolation and extrapolation (Netcode for Entities 6.5)Unity Extrapolation that keeps moving in the same direction and speed when the next snapshot is late is often wrong, so it gets a cap (Unity default 20 ticks, about 1/3 s at 60 Hz)
Peeking into VALORANT's NetcodeRiot Games When the guesses that fill in late or missing data are wrong, the client drifts from the server and characters jump or slide into place when corrected
ID cg-predict · Primary owner Game team (Client development) · Also Game team (Server development)
Your client shows your character moving before the server confirms it, but if the server calculates something different, your character gets pulled back.
Why The client moves the character before the server confirms (prediction) → Effect The server calculates collisions, movement speed, or buffs differently, or never receives the command → On screen When the confirmation arrives, your character gets pulled back
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: use the same movement code as the server, send inputs redundantly, smooth out corrections. Server: use the same movement code as the client, filter duplicate inputs by input number so each one is processed only once.
Ballpark numbers
The pull-back distance is “time out of sync × movement speed”. Losing just a few commands means 1–3 m.
On the graph
Random spikes · Server corrections (mispredictions)
Where to look
Record how often and how far the server corrects position. In Unreal, count the server’s ClientAdjustPosition corrections; in Unity Netcode for Entities, count rollbacks and resimulations caused by mispredictions
Confirmed if
Corrections cluster at the times of rubber-banding reports, and correction distance is repeatedly large with specific buffs, terrain, or movement skills
Ruled out if
Corrections cluster only when loss is high: input packet loss on the connection. Only other characters look pulled back with no corrections: “Excessive extrapolation (dead reckoning)”
Check with
Game server or client logs and metrics
Sources (3)
Introduction to prediction (Netcode for Entities 6.5)Unity Client and server predict with the same simulation code; when the result differs from the server state (misprediction), the client rolls back and resimulates, and the correction becomes visible
ID cg-fixed-step · Primary owner Game team (Client development)
After one stall, the game runs its backlog of calculations all at once, and that extra work puts it behind again.
Why The game simulation runs at a fixed interval and stalls once → Effect The backlog of steps is computed in a single frame → On screen A chain of long frames causes spikes, or the cap kicks in and the world slows down
Cap catch-up per frame, handle the leftover time with interpolation.
On the graph
Random spikes · Frame time, fixed steps per frame
Where to look
In a development build profiler, check how many fixed steps ran per frame (in Unity, the number of FixedUpdate-phase markers such as FixedBehaviourUpdate) together with frame time
Confirmed if
One long frame is followed by a chain of long frames that each run several steps; once the cap (Unity’s Maximum Allowed Timestep) is reached, game time runs slower than real time
Ruled out if
A single long frame that doesn’t repeat: “Frame time spike” or “Client garbage collection”
Check with
Game server or client logs and metrics
Learn more
Unity’s physics (FixedUpdate) is the classic fixed step (0.02 s by default, 50 times a second). Maximum Allowed Timestep in the Time settings (the most time the game catches up in one frame, about 0.33 s by default) is the catch-up cap. If a frame runs longer than that, the excess time is dropped, and the game clock falls behind real time by that much.
Handling variation in timeUnity Maximum Allowed Timestep defaults to 1/3 s (0.3333333): even if the game stops for 1 s, game time advances only 0.333 s. The cap prevents the vicious cycle where catch-up steps slow things down even more
Profiler markers referenceUnity FixedBehaviourUpdate: the section where MonoBehaviour.FixedUpdate runs; physics markers are called in the FixedUpdate phase
ID cg-clock · Primary owner Game team (Client development)
If the client’s estimate of the server time is wrong, interpolation timing and cooldown checks drift out of step with the server.
Why The client syncs to server time only once when connecting and never adjusts as ping changes → Effect The point to interpolate to and the time a cooldown ends drift away from the server’s → On screen Opponents occasionally hitch; skills get rejected even after the cooldown has ended
Sync time periodically (measure round-trip time and correct), adjust gradually with no sudden jumps, switch elapsed-time measurement from the PC’s clock to a monotonic clock.
On the graph
Slow climb · Error in the estimated server time
Where to look
Periodically log the difference between the client’s estimated server time and the server time (tick number) the server puts in its packets
Confirmed if
The error grows the longer the session runs, or jumps all at once when the PC clock gets corrected, and reports of rejected skills and hitches rise around then
Ruled out if
Error stays small but skills still get rejected: points to server-side validation or latency
Check with
Game server or client logs and metrics
Learn more
If elapsed time is measured with the PC’s date and time (wall clock), game time jumps the moment Windows corrects the clock against internet time or the user changes the clock. Measure elapsed time with a monotonic clock, which never goes backward (Stopwatch and so on).
Acquiring high-resolution time stampsMicrosoft QueryPerformanceCounter (used by Stopwatch) is a clock for elapsed time that isn’t synced to external time; use system time only when UTC time is needed
ID cg-float-time · Primary owner Game team (Client development)
If the game keeps its clock in a low-precision floating-point type (float), the longer it runs, the worse its time resolution (the smallest time difference it can tell apart) becomes, and movement and effects start to shake.
Why Time elapsed since launch is accumulated in a float or passed to shaders as is → Effect The longer the game runs, the larger the smallest difference a float can represent → On screen Only clients left running for days see characters, animations, and scrolling effects shake; a restart fixes it
Store elapsed time as a double (64-bit) or an integer, wrap the time passed to shaders back around at a fixed interval, run automated tests that last several days.
Ballpark numbers
A 32-bit float has just over 7 significant digits, so after a day of running (about 86,400 s) the time resolution is about 8 ms, roughly half a frame at 60 FPS (16.7 ms), and after a week about 60 ms, more than a whole frame.
On the graph
Slow climb · Shaking reports by how long the client has been running
Where to look
Collect how long the client has been running with each shaking report and compare before and after a restart. On the development side, a test that sets the game start time several days back and runs from there
Confirmed if
Shaking only on clients left running for days, gone after a restart, and worse the longer the client has been running
Ruled out if
Shaking right after launch: “Missing or too-short interpolation buffer” or “Timer resolution”
Check with
The player’s own environment
Learn more
This shows up especially in mobile MMOs where players leave auto-hunting on for days without closing the game. Unity’s Time.time is also a float, so Unity provides a separate double version, Time.timeAsDouble, and recommends it.
Sources (2)
Time.timeAsDoubleUnity The double version of Time.time; more precise than float the longer the game runs, so it’s recommended in most cases
ID cg-vsync · Primary owner Game team (Client development) · Also External (External)
Input is delayed while several finished frames wait in a queue to be sent out in step with the monitor’s refresh.
Why The graphics driver queues 1–3 frames ahead → Effect Input takes that much longer to show up on screen → On screen Ping is low, but controls feel heavy and sluggish
Primary owner Game team (Client development) · Also External (External)
Game team action items
Support a low-latency mode, shorten the frame queue, offer a frame cap option set slightly below the refresh rate, turn on frame pacing on phones.
External action items
Tell players to pair a variable refresh rate monitor with a frame cap slightly below the refresh rate and to turn on the graphics driver’s low-latency mode.
Ballpark numbers
At 60 Hz, each frame is 16.7 ms. When the CPU runs ahead of the GPU or the display refresh and the three-frame queue (the DirectX 11 default) fills up, 50 ms is added. With double-buffered V-Sync, a frame that took 17 ms waits for the next refresh (33.3 ms), and the previous frame is shown once more in the meantime.
On the graph
Always high · Input-to-display latency
Where to look
Compare PresentMon’s MsClickToPhotonLatency and MsAllInputToPhotonLatency (from mouse or keyboard input to output on screen) and DisplayLatency while switching V-Sync, low-latency mode, and frame caps. MsPCLatency (from when the PC receives input to when it sends the frame to the display) is recorded only if the game emits PC Latency events
Confirmed if
This latency grows by one or two frames (tens of ms) with V-Sync on or with no frame cap, and shrinks with low-latency mode or a frame cap slightly below the refresh rate. Ping unchanged
Ruled out if
Controls still feel late while latency inside the PC is low: “Display, input device, and frame generation latency.” High ping: the network side
Check with
The player’s own environment
Learn more
V-Sync (vertical sync) is a setting that sends out a new frame only at the moment the monitor refreshes the screen. Screen tearing goes away, but input is delayed by the wait for that moment, and when FPS drops below 60, it bounces between 60 and 30 and stutters. A variable refresh rate monitor shortens this wait by refreshing when a frame is ready. The same thing happens on phones. If a 30 FPS game can’t pace its frames evenly on a 60 Hz screen, the average is 30 FPS, but individual frames stay on screen for uneven times such as 49, 16, and 33 ms, which shows up as stutter (an example from the Android developer docs). Android’s frame pacing library (which evens out the intervals between frames) or the equivalent engine option reduces it.
Reduce latency with DXGI 1.3 swap chainsMicrosoft Present blocks until the queue drains, so a rendered frame waits almost one extra frame before it’s displayed; a waitable swap chain reduces this
Frame Pacing libraryAndroid (Google) On a 60 Hz screen, the previous frame is shown again when there’s no new frame; example of a 30 FPS game whose frame times become uneven, such as 49, 16, and 33 ms
PresentMon Capture Application (README-CaptureApplication.md)Intel MsPCLatency (from when the PC receives input to when it sends the frame to the display), MsClickToPhotonLatency (mouse click to screen), MsAllInputToPhotonLatency (keyboard or mouse input to screen), DisplayLatency (frame submission to output to the monitor)
ID cg-leak · Primary owner Game team (Client development)
The longer the game stays open, the more memory it uses; it gets slower and slower until the game is eventually force-closed.
Why Textures, UI, and effects aren’t released when moving between areas → Effect GC runs more often, and the OS runs short of memory and starts swapping → On screen After hours of play it stutters more and more, then gets force-closed (looks like a disconnect to the player)
Measure memory usage at area transitions, find and fix textures, UI, and effects that never get released, run long automated tests (soak tests).
On the graph
Slow climb · Game process memory
Where to look
Record Process(game)\Private Bytes in Performance Monitor for a few hours. On mobile, the exit reason in Android’s ApplicationExitInfo (REASON_LOW_MEMORY) and iOS jetsam reports
Confirmed if
Memory goes up every time the player moves between areas and never comes back down, and stutter and forced shutdowns increase the longer the game runs
Ruled out if
Memory stays flat, and the only thing that grows with uptime is shaking: “Float time precision loss”
Check with
The player’s own environment
Learn more
Phones mostly hold out by compressing memory. If memory still runs short, the OS shuts the game down on the spot (players see it as getting kicked out). Devices with less RAM get shut down first.
Sources (4)
Memory allocation among processesAndroid (Google) Android holds out by compressing memory into zRAM; when that’s not enough, the low memory killer terminates processes, and a foreground app being killed looks like a crash
ApplicationExitInfoAndroid (Google) REASON_LOW_MEMORY: the system’s low memory killer terminated the app process (devices that don’t support it report REASON_SIGNALED with SIGKILL)
ID cg-crash · Primary owner Game team (Client development) · Also External (External)
An unhandled error closes the game. To the player it looks like a disconnect, but the server is fine.
Why Null reference, out of memory, graphics driver error → Effect The game process is forcibly terminated → On screen Reports of “I got kicked out” while everyone else is fine at the same moment
Primary owner Game team (Client development) · Also External (External)
Game team action items
Collect crash reports, break down stats by device and driver, fix the most frequent errors first.
External action items
If crashes cluster on a specific graphics driver version, tell players to update their driver.
On the graph
Outliers only · Crash count (by device, graphics driver, and build)
Where to look
Crash reports and the Android vitals crash rate by device, driver, and build. On players’ PCs, Event ID 1000 in the Event Viewer Application log (faulting module name) and “Display driver stopped responding and has recovered” entries
Confirmed if
A crash record exists at the reported disconnect time, and other players on the same server are fine at the same moment. Crashes cluster on specific devices, driver versions, or modules
Ruled out if
Connection dropped with no crash record: “NAT mapping expiry” or the connection side
Check with
Game server or client logs and metrics
Sources (3)
CrashesAndroid (Google) A crash is when an app exits unexpectedly because of an unhandled exception or signal (SIGSEGV and so on); tallied in Android vitals in Play Console
ID cg-anticheat · Primary owner Game team (Client development) · Also Game team (Server development)
The anti-cheat module that runs alongside the game to block cheats scans the system periodically. If a scan is heavy, or the heartbeat (a periodic keepalive signal) to the anti-cheat server is late, the game stutters or disconnects.
Why The anti-cheat module periodically scans game memory, running programs, and drivers → Effect The game thread stalls during the scan, or the heartbeat doesn’t go out on time → On screen Hitches at regular intervals; in bad cases, a disconnect with a security error message
At regular intervals, Right after login or maintenance, Randomly
Owner
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: run heavy scans off the game thread in small pieces, compare stutter and kick stats by anti-cheat module version (if they spike right after an update, report it to the anti-cheat vendor). Server: tolerate a heartbeat that’s late once or twice.
Ballpark numbers
A light scan usually takes under 1 ms, but a heavy scan running on the game thread can take tens to hundreds of ms at a time, depending on the implementation.
On the graph
Periodic spikes · Frame time, anti-cheat kicks
Where to look
Measure the spacing between spikes in PresentMon frame time, and tally the anti-cheat kick reasons the server received (for EOS, AuthenticationFailed / Authentication Timed Out and others in ClientActionReason) by anti-cheat module version and hardware
Confirmed if
Brief pauses repeat at a regular interval regardless of what’s happening in the game, and right after an anti-cheat update, stutter and authentication-timeout kicks rise on specific hardware
Ruled out if
Same interval on all hardware regardless of anti-cheat version: “Client garbage collection” or “Background processes taking up CPU”
Check with
Game server or client logs and metrics
Learn more
Anti-cheat modules sit deep in the OS as drivers, so they sometimes conflict with antivirus software, overlays, and other games’ anti-cheat modules. If stutter and kick reports surge on specific hardware right after an anti-cheat update, suspect this first.
Sources (2)
Using the Anti-Cheat InterfacesEpic Games If the server doesn’t receive the client’s anti-cheat message within the set time (RegisterTimeout), it kicks the client for an authentication timeout (a client frozen by loading is a common cause); if a recent module update is the problem, roll back to the previous module
ID co-background · Primary owner External (External) · Also Game team (Client development)
When an antivirus scan, Windows Update, streaming software, or a browser video takes over CPU cores, the game thread has to wait for CPU time.
Why Other programs hold CPU cores for a long time → Effect The game thread waits for CPU time → On screen Frames come late, and received packets are processed late too
Primary owner External (External) · Also Game team (Client development)
Game team action items
Adjust game thread priority; record whole-PC CPU usage in the logs taken when stutter happens, to tell whether another program is to blame.
External action items
Tell players to turn on Windows Game Mode and to close unneeded programs while playing (antivirus scans, Windows Update, streaming software, browser videos).
Ballpark numbers
Windows usually hands out a core for a few ms to tens of ms at a time. Getting passed over by the scheduler just once is enough to lose a frame.
On the graph
Random spikes · Whole-PC CPU usage, frame time
Where to look
Record the CPU column in Task Manager’s Processes tab and Processor Information(_Total)\% Processor Time in Performance Monitor together with PresentMon frame time. If antivirus is suspected, record with New-MpPerformanceRecording and use Get-MpPerformanceReport to find the files and processes with the longest scan times
Confirmed if
Another program’s CPU use (antivirus scan, update, streaming software) spikes at the stutter times, or files in the game folder rank near the top for scan time. Closing that program or adding an exclusion makes it go away
Ruled out if
The whole screen hitches and audio crackles even though CPU usage is low: “NIC power saving and driver issues” (DPC latency)
Check with
The player’s own environment
Learn more
Windows gives the program in the front window (foreground) a little extra priority, but when there’s more work than there are cores, the game waits too. Antivirus software gets in the way through “real-time protection” more often than through CPU use. It scans every time the game opens a file, so the freezes while assets load get longer.
Sources (6)
MultitaskingMicrosoft Windows gives each thread a time slice and moves on to the next thread when it’s used up; a time slice is about 20 ms (varies by OS and CPU)
Priority BoostsMicrosoft The process in the front window (foreground) gets its priority raised to at least that of background processes
ID co-power · Primary owner External (External) · Also Game team (Client development)
Laptop battery mode, phone power-saving mode, and device heat slow down the CPU and GPU. With heat, the telltale sign is that the game runs fine at first and slows down only after a while.
Why Battery or power-saving mode is on, or the device gets hot → Effect CPU and GPU clocks drop by 30–50%, depending on the device → On screen FPS drops and the game stutters, right from the start with power saving, or after a few minutes to about 20 minutes of play with heat
Primary owner External (External) · Also Game team (Client development)
Game team action items
Adjust graphics options automatically, manage heat with a frame cap, lower options ahead of time based on the thermal level the OS reports (iOS thermalState, Android thermal status API), mark the executable so laptops with two graphics chips use the discrete GPU (export NvOptimusEnablement and AmdPowerXpressRequestHighPerformance).
External action items
Tell players to turn off power-saving mode and to keep laptops plugged in; for reports like “it’s a good laptop but FPS is low,” check which graphics chip the game runs on and tell players to assign the game to the high-performance GPU in Windows graphics settings.
On the graph
Slow climb · FPS, CPU/GPU clocks
Where to look
Record PresentMon’s CPUFrequency, GPUFrequency, CPUTemperature, and GPUTemperature with frame time for 20–30 minutes, and check which graphics chip the game runs on with the GPU engine column in Task Manager’s Processes tab. On mobile, record Android’s thermal API (getThermalHeadroom, thermal status) and iOS thermalState together with FPS
Confirmed if
FPS drops from the point where the temperature rises and clocks fall, or clocks are low only in battery or power-saving mode. Or the game is running on integrated graphics
Ruled out if
FPS drops while clocks and temperature hold steady: “Background processes taking up CPU” or “Client memory leak”
Check with
The player’s own environment
Learn more
Laptops with two graphics chips sometimes run games on the slower integrated graphics to save power. For a report like “it’s a good laptop but FPS is low,” first check which graphics chip the game is running on.
Sources (5)
Thermal APIAndroid (Google) Devices can sustain high performance only for a limited time before heat forces throttling; recommends watching thermal status and lowering the load ahead of time
thermalStateApple The current thermal level reported by iOS; as the level rises, the app should reduce its resource use
ID co-timer · Primary owner Game team (Client development)
Windows’ default timer ticks every 15.6 ms, so “sleep for just 1 ms” actually lasts until the next timer tick, up to 15.6 ms.
Why Frame limiting and packet sending are implemented with Sleep (a short wait) → Effect The OS wakes the thread only every 15.6 ms → On screen Frame intervals and input send intervals become uneven
Use high-resolution timers, replace Sleep-based waits with event- or V-Sync-based pacing.
Ballpark numbers
15.6 ms steps can’t hit a 16.7 ms interval, so frame intervals alternate between 15.6 ms and 31.2 ms.
On the graph
Always high · Frame interval distribution
Where to look
Look at the distribution of PresentMon’s MsBetweenPresents (frame interval) and the “Platform Timer Resolution” entries in the powercfg /energy report (processes that changed the timer resolution)
Confirmed if
Frame intervals cluster at multiples of 15.6 ms, such as 15.6 ms and 31.2 ms, and the game doesn’t request a higher timer resolution
Ruled out if
Intervals spread out evenly: more likely “Background processes taking up CPU” or frame load than the timer
Check with
The player’s own environment
Learn more
In older versions of Windows, when one program set the timer to 1 ms, it applied to every program. That’s where the saying “leave a browser open and your game runs smoother” came from. Since Windows 10 version 2004, the change applies only to the program that requested it, and Windows 11 may ignore requests from windows that are minimized or fully hidden and not playing sound.
Sources (6)
_WDF_TIMER_CONFIG (wdftimer.h)Microsoft Standard timer accuracy is the system clock tick interval, 15.6 ms by default; high-resolution timers get 1 ms
timeBeginPeriod function (timeapi.h)Microsoft Before Windows 10 2004 it was a global setting; since then it applies only to the requesting process, and Windows 11 doesn’t guarantee high resolution to processes whose windows are hidden or minimized
Results for the Idle Energy Efficiency AssessmentMicrosoft Default system timer resolution is 15.6 ms; the “Platform Timer Resolution” entries in the energy report show which processes changed the timer resolution
ID co-mobile-bg · Primary owner Game team (Client development) · Also Game team (Server development)
If you leave the game for a moment to check a notification, the OS suspends the app a few seconds later, and meanwhile the server disconnects you.
Why The player leaves the game to read a message or take a call → Effect The game engine pauses gameplay, and the OS soon suspends the app and its networking → On screen Already disconnected on return, so the game reconnects
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: on return, reconnect automatically right away with a session token (resume without logging in again) without waiting on the dead connection, then fetch the latest state in one go to catch up. Server: when heartbeats stop, clean up the connection but keep the character session for a short grace period (don’t kick it right away), and resume it by session token if the player reconnects within that window.
Ballpark numbers
Game engines usually pause the moment the app goes to the background. iOS suspends the app within a few seconds, or usually within tens of seconds even with extra time granted, and Android 14 and later freezes apps that leave the screen after about 10 seconds.
On the graph
Mass disconnect · Disconnects (heartbeat timeouts), app suspend records
Where to look
Match the app suspend and resume times in the client log (OnApplicationPause in Unity) against the server’s disconnect reasons and times by session ID. On Android, also check the process exit reasons recorded in ApplicationExitInfo (REASON_LOW_MEMORY and so on)
Confirmed if
The client went into suspend just before the server’s heartbeat-timeout disconnect and reconnected right after resuming
Ruled out if
Disconnected while the app was in the foreground: “NAT mapping expiry,” “ISP-shared IP addresses (CGNAT),” or “Wi-Fi ↔ LTE/5G switching”
Check with
Game server or client logs and metrics
Learn more
When memory runs short, phones sometimes kill a backgrounded game outright. That’s why the game starts over from scratch after a trip to the camera or a payment or authentication app. It’s more common on low-end devices.
Sources (5)
Extending your app’s background execution timeApple When the app goes to the background, applicationDidEnterBackground gets 5 s before the app is suspended; if it needs more, it requests time with beginBackgroundTask (time left in backgroundTimeRemaining)
Cached apps freezerAndroid (Google) Android 14 and later freezes cached app processes after 10 s; once frozen, all threads stop
Application.runInBackgroundUnity Defaults to false, so the game pauses in the background; Android pauses in the background regardless of the setting, and iOS ignores it
ApplicationExitInfoAndroid (Google) REASON_LOW_MEMORY: the system’s low memory killer terminated the app process (devices that don’t support it report REASON_SIGNALED with SIGKILL)
ID co-netswitch · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Network infrastructure)
When you walk out of the house and your phone drops Wi-Fi for LTE or 5G, your IP address changes and the existing connection stops working.
Why The Wi-Fi signal weakens and the phone switches to the mobile network → Effect Your IP address changes, so the connection made from the old address can’t carry any more data → On screen A brief freeze, then a disconnect or a reconnect
Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Network infrastructure)
Game team action items
Server: resume the same player by session token even when the address changes, clean up the old address’s connection right away, consider a protocol that survives address changes (such as QUIC connection migration). Client: when a network switch is detected, reconnect right away with the session token without waiting for a heartbeat timeout.
Infra team action items
When using QUIC connection migration, configure the load balancer to choose the server by connection ID (choosing by address and port sends packets from a changed address to a different server).
On the graph
Mass disconnect · Disconnects and reconnects, IP changes on reconnect
Where to look
Look in the server connection log for the same session token reconnecting from a different IP, and match the times against the client’s default network change callback (registerDefaultNetworkCallback)
Confirmed if
Right after the disconnect, the reconnecting IP moves from the Wi-Fi (home connection) range to a mobile carrier range or the other way around, and a network switch callback arrives just before
Ruled out if
Disconnected while the IP stayed the same: “Cell tower handover (while moving)” or “Weak mobile signal and dead zones”
Check with
Game server or client logs and metrics
Sources (3)
Read network stateAndroid (Google) When the default network changes, new connections go over the new network and connections on the old network are eventually forced closed; registerDefaultNetworkCallback detects the switch
RFC 9000: QUIC: A UDP-Based Multiplexed and Secure TransportIETF Connection IDs keep a connection alive even when the IP address or port changes (Section 9); a load balancer that distributes by address and port alone may send packets from a changed address to a different server (Section 5.2.3)
ID co-security · Primary owner External (External) · Also Game team (Client development)
When antivirus software or a firewall inspects every packet, latency goes up, and in bad cases it mistakes the game for an attack and blocks it.
Why Security software inspects every packet sent and received, one by one → Effect Each packet picks up delay, and packets get dropped when inspection falls behind → On screen Ping spikes irregularly, or connections get blocked
Primary owner External (External) · Also Game team (Client development)
Game team action items
Maintain a security software compatibility list, register a Windows Firewall exception for the game at install time.
External action items
Tell players to add the game as an exception in their security software; if it mistakes the game for an attack, ask the security vendor to fix the false positive.
Ballpark numbers
When everything works normally, packet inspection usually takes under 1 ms. The trouble starts when the inspection module falls behind or has a bug, or when it mistakes game traffic for an attack.
On the graph
Outliers only · RTT, connection failures (per player)
Where to look
Compare after briefly turning off the security software or adding the game as an exception. On Windows, turning on Audit Filtering Platform Connection and Audit Filtering Platform Packet Drop in the audit policy logs 5157 (connection blocked) and 5152 (packet blocked) in the Security log, and Performance Monitor’s WFPv4\Packets Discarded/sec shows the number of discarded packets
Confirmed if
Block records show up for connections or packets to the game server’s address, or ping spikes and connection failures go away with the security software off
Ruled out if
Other devices in the same household behave the same way regardless of security software: the router or connection side
Check with
The player’s own environment
Sources (6)
About Windows Filtering PlatformMicrosoft Packets are allowed or blocked through hooks in the Windows network stack and a filter engine, and third-party security vendors can plug in their own filter modules (callouts)
Windows Firewall RulesMicrosoft Inbound connections are blocked by default, so apps need exception rules, which the app installer usually creates
ID co-rcvbuf · Primary owner Game team (Client development)
If the game is busy and pulls packets out of the socket (the network send/receive interface the OS provides) late, the OS buffer overflows.
Why Frames fall behind and the game reads the socket late → Effect The OS receive buffer fills up: UDP packets get dropped, and TCP shrinks the receive window so the sender stops sending → On screen Teleporting (UDP) or fast-forward (TCP)
Use a dedicated receive thread, tune the buffer size (SO_RCVBUF).
Ballpark numbers
The default receive buffer is tens to hundreds of KB, depending on the OS and settings. Updates in crowded places can reach hundreds of KB per second.
On the graph
Rises with load · UDP receive buffer drops, frame time
Where to look
Record Microsoft Winsock BSP\Dropped Datagrams (UDP datagrams dropped for lack of socket receive buffer space) and UDPv4\Datagrams Received Errors in Windows Performance Monitor along with frame time; in the game, count gaps in the sequence numbers of received packets
Confirmed if
Dropped Datagrams rises in crowded places or right after a long frame, and gaps appear in the game’s sequence numbers at the same moment. No loss on the connection at the same time
Ruled out if
Dropped Datagrams flat but sequence numbers still go missing: loss along the path
Check with
The player’s own environment
Sources (5)
socket(7) — Linux manual pageLinux man-pages SO_RCVBUF is the maximum socket receive buffer size; the default comes from rmem_default and the maximum from rmem_max (Android also runs the Linux kernel)
RFC 9293: Transmission Control Protocol (TCP)IETF The TCP window field is the number of bytes the receiver can still accept; at 0, the sender waits, sending only zero window probes
Low Latency Workloads Management and OperationsMicrosoft Dropped Datagrams and Dropped Datagrams/sec in the Microsoft Winsock BSP counter set: UDP datagrams dropped because they arrived faster than the app could process them or the receive socket buffer was too small
ID co-swap · Primary owner External (External) · Also Game team (Client development)
With dozens of browser tabs open alongside the game, the OS moves part of the game’s memory out to disk.
Why The PC runs low on RAM overall → Effect The OS moves game memory that isn’t in use right now to disk → On screen The moment that memory is used again, the game freezes for tens to hundreds of ms, depending on storage
Primary owner External (External) · Also Game team (Client development)
Game team action items
Reduce memory usage, show a warning when free memory is low.
External action items
Tell players the minimum specs and to close other programs (such as browser tabs) while playing.
On the graph
Random spikes · Hard page faults, memory usage
Where to look
Record Memory\Pages Input/sec (pages read from disk to resolve hard page faults) in Performance Monitor and memory usage and commit on Task Manager’s Performance tab, together with frame time
Confirmed if
Pages Input/sec spikes at the moment of each freeze and memory is nearly full. Closing the browser and other programs makes it go away
Ruled out if
Memory has headroom and Pages Input/sec stays quiet: “Synchronous loading and shader compilation on the main thread” or “Slow storage delays asset streaming”
Check with
The player’s own environment
Sources (3)
Introduction to the page fileMicrosoft The page file is a file on disk used to move rarely used, modified memory pages out of RAM
Working SetMicrosoft Touching a page that isn’t in RAM causes a page fault; a hard fault can only be resolved by reading from disk, such as from the page file
ID co-vram · Primary owner Game team (Client development) · Also External (External)
When the graphics settings need more memory than the graphics card has, the OS moves textures out to system memory and brings them back, and the game stutters.
Why High texture settings plus all the gear and effects in a crowded place fill up graphics card memory → Effect The OS moves textures that aren’t in use right now to system memory, then brings them back over the slow PCIe bus when needed → On screen A hitch every time a new scene or character comes into view; textures stay blurry for a while
When crowds gather, While moving or changing zones
Owner
Primary owner Game team (Client development) · Also External (External)
Game team action items
Set default options to match graphics card memory size, lower texture quality automatically when the memory budget is exceeded, simplify character textures in crowded places.
External action items
Tell players to lower texture settings, and to lower them further when running two clients.
Ballpark numbers
Graphics card memory reads hundreds of GB per second, but the PCIe bus to system memory carries roughly 16–64 GB per second depending on the generation, more than ten times slower.
On the graph
Hits a ceiling · Dedicated GPU memory, shared GPU memory
Where to look
Watch the Dedicated GPU memory and Shared GPU memory graphs under GPU on Task Manager’s Performance tab (per-process columns can also be added on the Details tab) alongside PresentMon frame time
Confirmed if
Hitches are frequent while dedicated GPU memory sits flat at its limit and shared GPU memory grows, and they go away when texture settings are lowered
Ruled out if
Dedicated memory has headroom: “Slow storage delays asset streaming” or “Synchronous loading and shader compilation on the main thread”
Check with
The player’s own environment
Learn more
If “Dedicated GPU memory” under GPU in Windows Task Manager is full and “Shared GPU memory” keeps growing, this is what’s happening. Running two clients on the same PC fills it faster (see “Streaming failure from memory or VRAM shortage”).
Sources (4)
ResidencyMicrosoft Each process has a graphics memory budget; when it’s exceeded, the kernel moves part of the discrete GPU’s heap to system memory (a last resort, so managing the budget is recommended)
GPUs in the task managerMicrosoft In Task Manager, dedicated GPU memory is the graphics card’s VRAM, and shared GPU memory is system memory used by both the GPU and the CPU
CUDA C++ Best Practices GuideNVIDIA Graphics memory bandwidth (V100: 898 GB/s) is far higher than PCIe 3.0 x16 (16 GB/s), so it recommends minimizing transfers to and from system memory
ID co-wifi-scan · Primary owner External (External) · Also Game team (Client development)
Communication pauses briefly while the OS periodically hops across channels to look for nearby Wi-Fi networks.
Why The OS or driver searches for nearby Wi-Fi networks on a fixed schedule → Effect Sending and receiving pause briefly during the scan → On screen Ping spikes at exactly regular intervals (e.g., every 60 s)
Primary owner External (External) · Also Game team (Client development)
Game team action items
Request a mode that reduces wireless scanning during play (on Android, the low-latency Wi-Fi mode WIFI_MODE_FULL_LOW_LATENCY; on Windows, the media streaming mode of WlanSetInterface; it may have no effect on some devices and drivers).
External action items
Tell players to use a wired connection, adjust location services and automatic Wi-Fi scanning settings, and update wireless drivers.
Ballpark numbers
Usually tens to hundreds of ms each time. If the spikes are suspiciously regular, suspect this first.
On the graph
Periodic spikes · RTT to the router
Where to look
During play, ping the router address (Default Gateway in ipconfig) with ping /t for a few minutes and measure the interval between spikes. Repeat the same measurement over a wired connection
Confirmed if
Ping to the router spikes by tens to hundreds of ms at exactly regular intervals (e.g., 60 s) and goes away on a wired connection
Ruled out if
Irregular spike intervals: “Wi-Fi interference and weak signal.” Fine up to the router but spiking beyond it: the connection or ISP segment
Check with
The player’s own environment
Sources (5)
WDI low latency connection qualityMicrosoft Scanning and roaming move the radio off the connected channel, so low-latency mode limits scanning and time spent off-channel
WlanSetInterface function (wlanapi.h)Microsoft Windows API that turns background scanning (wlan_intf_opcode_background_scan_enabled) and media streaming mode on and off
Wi-Fi low-latency modeAndroid (Google) Low-latency mode turns off Wi-Fi power saving; how scanning and roaming settings are optimized depends on the device maker’s implementation
pingMicrosoft /t: keeps sending echo requests until stopped
ipconfigMicrosoft Run with no parameters, it shows each adapter’s IPv4 and IPv6 addresses and default gateway
ID co-driver · Primary owner External (External) · Also Game team (Client development)
When a network card or Wi-Fi chip enters a power-saving state between packets, it takes time to wake back up.
Why Network device power saving is on, or the driver is outdated → Effect Wake-up delay, occasional device restarts → On screen Irregular delays, occasional freezes lasting several seconds
Primary owner External (External) · Also Game team (Client development)
Game team action items
On Android clients, request the low-latency Wi-Fi mode (WIFI_MODE_FULL_LOW_LATENCY) during play to turn off Wi-Fi power saving.
External action items
Tell players to update network drivers and turn off network device power saving in Device Manager; if the whole screen hitches and audio crackles, have them find the driver at fault with LatencyMon.
On the graph
Random spikes · DPC/ISR time, RTT to the router
Where to look
Record with Windows Performance Recorder (WPR), find long-running drivers (Module column) in the DPC/ISR graph in Windows Performance Analyzer (WPA), and check the network adapter’s power management (power saving) settings in Device Manager
Confirmed if
At the hitch times, network driver DPCs and ISRs run for several ms at a stretch, or turning off power saving makes the irregular delays go away
Ruled out if
DPCs are short and turning off power saving changes nothing: “Wi-Fi interference and weak signal” or “Wi-Fi background scanning”
Check with
The player’s own environment
Learn more
When a driver holds the CPU for a long time handling interrupts (Windows calls this DPC latency), the game thread can’t use that core either. The whole screen then hitches and audio crackles even though CPU usage is low. A tool such as LatencyMon can find which driver is at fault; Wi-Fi and Ethernet drivers are common culprits.
ID co-other-apps · Primary owner External (External) · Also Game team (Client development)
When cloud sync, a large download, or a game patch runs on the same PC, game packets have to wait in a queue.
Why Another app maxes out the upload or download → Effect Game packets pile up in the queues on the PC and the router → On screen Ping shoots up, input lag, fast-forward
Primary owner External (External) · Also Game team (Client development)
Game team action items
Make our own launcher and patcher pause or throttle background downloads during play.
External action items
Tell players to cap download speeds and turn off automatic updates while playing.
On the graph
Rises with load · RTT, PC traffic sent and received
Where to look
Record Network Interface\Bytes Sent/sec and Bytes Received/sec in Performance Monitor together with ping. Same approach as a bufferbloat test, where you deliberately start a large transfer with ping running
Confirmed if
Ping rises by tens to hundreds of ms while downloads or uploads run close to the connection speed, and drops back as soon as the transfer stops
Ruled out if
Ping rises while PC traffic is low: “Bufferbloat (router queue)” from another device in the same household, or the ISP segment
Check with
The player’s own environment
Sources (4)
IntroductionBufferbloat.net When network equipment such as a router buffers too much data, latency spikes sharply (bufferbloat)
Delivery Optimization referenceMicrosoft Windows Update downloads (Delivery Optimization) adjust dynamically to available bandwidth by default, and caps can be set on background and foreground download bandwidth
ID co-unfocused · Primary owner Game team (Client development)
When you switch to another window or minimize the game, the game and Windows slow it down to save power. When you come back, the backlog of packets floods in, or you’ve already been disconnected.
Why Switching to another window with Alt+Tab, or minimizing the game → Effect While it isn’t visible, the game lowers FPS sharply or pauses, and Windows also lowers the priority of programs that aren’t visible → On screen Fast-forward the moment you return; a disconnect if the game stayed minimized for a long time
Keep receiving packets and sending heartbeats on a separate thread even when the window is hidden, check the engine’s “Run in background” setting, catch up to the latest state in one go on return.
Ballpark numbers
If FPS drops to 5–10 while the window is hidden, each frame takes 100–200 ms. A game that processes packets once per frame reads them that much later.
On the graph
Gap then burst · Frame interval (before and after switching windows), packets processed
Where to look
With PresentMon running, try Alt+Tab and minimizing, and look at frame intervals while the window is hidden. Log window focus changes in the game log and match them against disconnect reasons
Confirmed if
Frame intervals stretch past 100 ms or recording stops while the window is hidden, and the moment you return, the game processes the packet backlog all at once and fast-forwards. Left minimized for long, it disconnects on a heartbeat timeout
Ruled out if
Same behavior with the window in front: “Background processes taking up CPU” or the network side
Check with
The player’s own environment
Learn more
Windows 11 doesn’t guarantee a 1 ms timer for programs whose windows are minimized or fully hidden and not playing sound. On a laptop running on battery, it slows such programs to the most power-efficient speed, and on CPUs with mixed core types it may run them on the slower efficiency cores. If only the background one of two clients on the same PC misbehaves, also see “Background window throttling.”
Sources (4)
Quality of ServiceMicrosoft Programs whose windows can’t be seen or heard get Low QoS and, on battery, are scheduled at the most efficient CPU speed and on efficiency cores
timeBeginPeriod function (timeapi.h)Microsoft Windows 11 doesn’t guarantee a higher-than-default timer resolution to processes whose windows are hidden or minimized
Application.runInBackgroundUnity Unity defaults to false, so the game loop stops when the window goes to the background
ID co-overlay · Primary owner External (External) · Also Game team (Client development)
Chat apps, launchers, recording tools, and FPS counters hook into the game’s rendering to draw their own UI on top of the game screen (hooking). That adds work to every frame and sometimes clashes with the game, causing hitches or crashes.
Why Overlays from chat apps, game launchers, graphics card tools, or recording software are turned on → Effect Every time a frame goes out to the screen, the overlay steps in and draws its own UI on top → On screen Frames get slightly later, and when a notification pops up the game hitches, shows graphics glitches, or crashes (looks like a disconnect to the player)
Primary owner External (External) · Also Game team (Client development)
Game team action items
Collect the list of running overlays along with crash reports and stutter logs.
External action items
When reports come in, tell players to turn off all overlays and test again.
On the graph
Outliers only · Frame time and crash count (players with overlays on)
Where to look
Turn off all overlays and compare PresentMon frame time in the same scene; for crashes, check the faulting module name (Faulting module name) in Event Viewer Event ID 1000
Confirmed if
Hitches and graphics glitches go away with overlays off, or the faulting module in the crash is an overlay program’s DLL
Ruled out if
Same with every overlay off: the graphics driver or “Client crash”
Check with
The player’s own environment
Learn more
When only certain players stutter or crash and their specs don’t explain it, suspect a conflict between an overlay and the anti-cheat module first.
Sources (3)
Steam Overlay (Steamworks Documentation)Valve The Steam overlay automatically hooks into games launched through Steam, and because of how it does that, it can expose memory errors in the game’s use of the rendering API and cause crashes
The application or service crashing behavior troubleshooting guidanceMicrosoft Event ID 1000 in the Application log includes the faulting module name (Faulting module name); a Windows module sometimes shows up as the faulting module because of corruption caused by another module
ID co-display-input · Primary owner External (External) · Also Game team (Client development)
If ping is normal but controls feel heavy, a TV’s video processing, a wireless controller, or frame generation may be adding delay between your input and the screen.
Why The TV’s game mode is off, a Bluetooth or wireless controller is in use, or frame generation (DLSS or FSR frame generation) is on → Effect The TV sends frames out late while it processes the picture, wireless input arrives late by its polling interval plus any interference, and frame generation waits for the next frame to create an in-between frame → On screen Ping and FPS numbers look good, but there’s a delay between pressing a button and seeing the result on screen: input lag
Primary owner External (External) · Also Game team (Client development)
Game team action items
Make frame generation optional and warn that turning it on can increase input lag, integrate the GPU vendor’s low-latency feature (NVIDIA Reflex, AMD Anti-Lag 2) when frame generation is used, show the PC-side input-to-display latency in the game, request the TV’s low-latency mode (ALLM) with Window.setPreferMinimalPostProcessing(true) in Android TV and set-top box builds.
External action items
Tell players to turn on game mode (ALLM) on their TV or monitor, use a wired controller and turn off frame generation for competitive play, keep Bluetooth devices close, and use 5 GHz Wi-Fi.
Ballpark numbers
Just sending one frame takes 16.7 ms on a 60 Hz screen and 8.3 ms at 120 Hz. Older Xbox controllers read and sent input every 8 ms. The delay added by a TV’s video processing varies by device, so no single number fits; game mode is the setting that cuts down this processing. AMD recommends using frame generation at a frame rate of 60 FPS or higher before generation.
On the graph
Always high · Input-to-display latency
Where to look
Compare PresentMon’s MsAllInputToPhotonLatency (from keyboard or mouse input to output on screen) with frame generation on and off, and use FrameType (recorded only when the driver or SDK reports it) to see whether generated in-between frames are mixed in. This value doesn’t include the controller’s wireless link or the processing inside the TV, so compare those parts by switching to TV game mode or a wired controller
Confirmed if
Ping is normal, but input-to-display latency drops with frame generation off, or the perceived delay goes away with TV game mode or a wired controller
Ruled out if
No change after switching all of these, and ping is high or spiking: the network side. PC-side latency is high because of V-Sync or the frame queue: “V-Sync and the render queue”
Check with
The player’s own environment
Learn more
Network latency shows up in ping, but this latency doesn’t. That’s why it’s the first thing to check in a “low ping but still lagging” report. Frame generation roughly doubles the FPS number on screen, but to create an in-between frame it has to wait for the next real frame, so the time until your input shows up on screen gets longer (AMD states that latency increases by design). Bluetooth devices use the same 2.4 GHz band as Wi-Fi, so interference can make input drop out or jump. For V-Sync and the render queue, which add latency inside the PC, see “V-Sync and the render queue.”
Sources (9)
Auto Low Latency Mode (ALLM)HDMI Licensing Administrator ALLM lets a device switch the display to low-latency mode (often called game mode) automatically; in low-latency mode, the TV stops some video processing to reduce delay
Xbox Series X: What’s the Deal with Latency?Microsoft Input lag is the sum of the path controller → console → HDMI → TV; older controllers read and sent input every 8 ms; sending one frame over HDMI takes 16.6 ms at 60 Hz and 8.3 ms at 120 Hz; ALLM switches the TV to game mode automatically
AMD FSR Frame GenerationAMD Frame generation is recommended at 60 FPS or higher before interpolation (below 30 FPS should be avoided); AMD Radeon Anti-Lag 2 reduces system latency by keeping CPU and GPU work in step
NVIDIA DLSSNVIDIA DLSS Frame Generation is designed to keep responsiveness when paired with NVIDIA Reflex (a low-latency feature)
PresentMon Capture Application (README-CaptureApplication.md)Intel MsAllInputToPhotonLatency (input-to-display latency), DisplayLatency (frame submission to output to the monitor), FrameType (tells frames rendered by the app apart from frames interpolated by the driver or SDK)
Window.setPreferMinimalPostProcessingAndroid (Google) Latency-sensitive windows such as games request minimal video processing from the display; over HDMI, the ALLM and Game Content Type signals switch the TV to low-latency mode
ID hn-wifi · Primary owner External (External) · Also Game team (Client development)
With a weak signal or interference, packets get resent several times over the wireless link, so they arrive unevenly.
Why Walls, distance, microwaves, Bluetooth, and neighbors’ routers degrade the radio signal → Effect Transmissions fail on the wireless link → resent several times → On screen Packets arrive unevenly (jitter), so characters move in fits and starts; with heavy loss, they teleport
Primary owner External (External) · Also Game team (Client development)
Game team action items
Adjust interpolation buffer length automatically to connection quality, show network status on screen when jitter or loss is high.
External action items
Tell players to use a wired connection, switch to 5 GHz or 6 GHz, and move the router.
Ballpark numbers
Each retransmission adds about 1–4 ms. With a weak signal, the radio resends several times at a low rate and also waits for the channel to clear, so latency can spike by 50–200 ms. The trap is that average ping looks fine.
On the graph
Random spikes · RTT to the router
Where to look
Ping the router address (Default Gateway in ipconfig) with ping /t for a few minutes, and check your router’s signal strength and channel with netsh wlan show networks mode=bssid. Compare over a wired connection from the same spot
Confirmed if
Even ping to the router spikes irregularly by tens to hundreds of ms with occasional loss, and signal strength is low. Gone on a wired connection or close to the router
Ruled out if
Steady up to the router but spiking beyond it: the connection or ISP segment. Spikes only at a fixed interval: “Wi-Fi background scanning”
Check with
The player’s own environment
Learn more
In mesh Wi-Fi, when the routers (nodes) link to each other wirelessly (wireless backhaul), a relaying node can’t send while it’s receiving and shares transmit opportunities with the hops before and after it on the same channel, so throughput can drop and latency can rise when the network is busy. Products with a dedicated wireless backhaul band may suffer less, and wiring the nodes together (Ethernet) takes this hop off the air. Powerline adapters (PLC) also check that the medium is free before sending (CSMA/CA), like Wi-Fi, and their quality changes constantly with electrical noise from appliances and with appliances switching on and off, which can cause retransmissions and jitter.
RFC 8325: Mapping Diffserv to IEEE 802.11IETF 802.11 CSMA/CA: sends only when the channel is clear; if it’s busy, defers until it clears, then waits an additional random backoff
Capacity of Ad Hoc Wireless Networks (MobiCom 2001)ACM When 802.11 relays over several wireless hops, a node can’t send while it’s receiving and adjacent hops interfere with each other, so the throughput of a chain of relays can drop to 1/3 in theory (about 1/7 in simulation)
ID hn-channel · Primary owner External (External) · Also Game team (Client development)
Where there are dozens of routers, as in an apartment building, they share the same channel and have to wait for a chance to transmit.
Why Dozens of routers use the same 2.4 GHz channel → Effect Before transmitting, a device waits until other devices finish and the channel clears → On screen In the evening, when people get home, jitter (variation in packet arrival times) rises and the game stutters
Primary owner External (External) · Also Game team (Client development)
Game team action items
Lengthen the interpolation buffer automatically when jitter rises.
External action items
Tell players to use 5 GHz or 6 GHz, a less crowded channel, or a wired connection.
On the graph
High at certain hours · RTT and jitter to the router
Where to look
Check the channels and signal strength of nearby Wi-Fi networks with netsh wlan show networks mode=bssid, and compare ping to the router in the evening and during the day
Confirmed if
Many nearby routers show up on the same 2.4 GHz channel, and jitter to the router grows only in the evening. Moving to 5 GHz or 6 GHz or to a less crowded channel reduces it
Ruled out if
Spikes regardless of time of day: “Wi-Fi interference and weak signal.” Fine up to the router but bad beyond it in the evening: “Peak-hour congestion at peering links”
Check with
The player’s own environment
Sources (4)
Recommended settings for Wi-Fi routers and access pointsApple Other routers and devices on the same channel are sources of interference; 20 MHz channel width is recommended on 2.4 GHz; interference is less of a concern on 5 GHz and 6 GHz
ID hn-bufferbloat · Primary owner External (External) · Also Game team (Client development)
When someone in the household uploads a video or downloads a large file, hundreds of ms worth of packets pile up in the router’s queue, and game packets wait behind them.
Why The connection fills up with a family member’s video upload or cloud backup, your own live stream, or a large download → Effect The router or modem holds the overflow of packets in a large queue → On screen Game packets wait at the back of the queue too, and ping shoots up to hundreds of ms
Primary owner External (External) · Also Game team (Client development)
Game team action items
Show network status on screen when ping suddenly jumps to hundreds of ms (mention a possible large transfer on the same connection).
External action items
Tell players to use a router with SQM (fq_codel, CAKE) or QoS, set the SQM speed to 90–95% of the connection speed (so the queue forms inside the router and SQM can take effect), and cap upload speeds.
Ballpark numbers
On a connection with 10 Mbps upload, a 1 MB buffer lets the queue grow to 800 ms.
On the graph
Rises with load · RTT, upload and download usage on the connection
Where to look
With ping running, saturate the connection with a speed test, or use a web test that measures latency under load (see Bufferbloat.net). Check alongside the upload and download usage shown on the router’s admin page
Confirmed if
Ping climbs to hundreds of ms while an upload or download saturates the connection and recovers when the transfer ends (suspect it when latency under load exceeds 50 ms). Gone with SQM on
Ruled out if
Ping spikes even when the connection is idle: “Wi-Fi interference and weak signal” or “Poor line quality”
Check with
The player’s own environment
Learn more
The upload side clogs especially easily, because cable and mobile connections often have far less upload bandwidth than download. In homes with plenty of fiber bandwidth, the Wi-Fi link becomes the bottleneck, and the same thing happens in the router’s wireless queue. Game packets are small and use almost no bandwidth, but they still have to wait in the queue. If only the upload direction is clogged, only your own input is late, while other players’ movement looks fine. On phones, photo backups and app updates on the same phone fill the queues in the phone’s modem and at the cell tower, with the same result.
Sources (4)
Setting up SQM for CeroWrt 3.10Bufferbloat.net Set SQM to 95% of the measured speed (85% if based on the advertised speed) to move the bottleneck from the ISP’s equipment into the router; that’s what makes it work
SQM (Smart Queue Management)OpenWrt Enter 90% of the measured download and upload speeds; cake is the recommended queue discipline (fq_codel if the CPU is weak)
Tests for BufferbloatBufferbloat.net If ping rises while a speed test saturates the connection with ping running, it’s bufferbloat; a fix is recommended when latency under load exceeds 50 ms (or the grade is below B)
ID hn-nat · Primary owner Game team (Client development) · Also Game team (Server development)
Routers remove idle connections that haven’t carried packets for a while from their NAT table. It’s a common reason the connection drops the moment you move after standing still.
Why The router records the “inside device ↔ outside server” connection in its NAT table (address translation table) → Effect If no packets pass for a while, the entry is deleted (for UDP, often after 30–120 s) → On screen Server packets can no longer get into the home, so the connection drops
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: send heartbeats at no more than half the shortest idle timeout (UDP mappings are reliably refreshed only by packets going out of the home, so the client sends them), reconnect automatically after a disconnect. Server: answer heartbeats, clean up the connection first when none arrive for a set time, and when a deleted mapping changes the outside address and port, confirm it’s the same player with the session token (an ID issued when the player connects) and resume the session.
On the graph
Mass disconnect · Disconnects (heartbeat timeouts), idle time before disconnect
Where to look
Collect the server’s disconnect reasons and the time elapsed since the last packet on that connection before it dropped (idle time), and look at the distribution. To test, increase the UDP packet interval to 30 s, 60 s, and 120 s and find the interval at which responses stop
Confirmed if
Only idle connections drop, and idle times cluster just past a specific value such as 30–120 s. A heartbeat interval shorter than that makes it go away
Ruled out if
Drops even while moving: the connection or route side. Clustered at a short value only on a specific mobile carrier: “ISP-shared IP addresses (CGNAT)”
When dozens of devices and thousands of connections pile onto a cheap router, the router itself can’t keep up.
Why Dozens of devices, plus P2P and torrent clients opening thousands of connections → Effect The router’s CPU and session table are saturated → On screen Delayed and lost packets, failed new connections
Tell players to reboot the router (temporary fix), replace the router, and clean up programs that open many connections (P2P, torrents).
On the graph
Hits a ceiling · Router CPU and connection count, RTT to the router
Where to look
Check CPU usage, connection (session) count, and number of connected devices on the router’s admin page (if the router supports it), and compare ping to the router itself before and after a reboot
Confirmed if
When connection counts are high, even ping to the router spikes or loses packets and new connections fail. After a reboot it’s fine for a while, then gets worse again
Ruled out if
Fine up to the router but bad beyond it: the connection or ISP segment
Netfilter Conntrack Sysfs variablesLinux kernel The maximum number of entries in Linux’s connection tracking table (nf_conntrack_max) and the default timeouts for each connection state
ID hn-handover · Primary owner External (External) · Also Game team (Client development), Game team (Server development)
When you travel by bus or subway, the connection drops out while your phone switches cell towers.
Why The phone switches to a different cell tower while on the move → Effect Usually a gap of tens of ms, but if the signal is bad and the switch fails, it can drop out for hundreds of ms to several seconds → On screen A freeze, then teleporting; if it lasts long, a disconnect
Primary owner External (External) · Also Game team (Client development), Game team (Server development)
Game team action items
Client: use timeouts that tolerate short dropouts, reconnect quickly. Server: use timeouts that don’t kick players right away after a few seconds of silence, resume the same session on reconnect.
External action items
Tell players that disconnects while traveling (bus, subway) are caused by cell tower switching.
On the graph
Gap then burst · Packets received, RTT
Where to look
Check whether the disconnect report came from someone traveling (bus, subway), and look at the receive gap times in the client log together with changes in network type and signal
Confirmed if
Only while traveling, receiving stops for hundreds of ms to several seconds and then packets arrive in a rush; it doesn’t reproduce when standing still
Ruled out if
Same when standing still: “Weak mobile signal and dead zones” or “Frequent 5G↔LTE switching (at 5G coverage edges)”
ID hn-rrc · Primary owner Game team (Client development)
When a phone has no traffic for a while, it drops its radio connection to a low-power state, and the next packet is delayed while it powers back up.
Why After a short period with no traffic, the phone puts its radio connection into a power-saving state → Effect To send the next packet, the connection has to be brought back up → On screen Only the first action after standing idle is noticeably slow
Keep the radio active with light periodic traffic (at a cost in battery life).
Ballpark numbers
LTE usually drops to power saving after about 10 seconds with no traffic, and coming back takes tens to hundreds of ms (measured example: about 0.3–0.6 s). On 3G it’s over 1 second.
On the graph
Outliers only · RTT of the first request after idle (mobile)
Where to look
Break down in-game RTT by the gap since the previous traffic. On mobile, compare the RTT of the first packet sent after more than 10 s of idle with that of packets sent back to back
Confirmed if
On a mobile network, only the first packet after idle is hundreds of ms late, and packets sent right after it are normal. No difference on Wi-Fi
Ruled out if
Late even when sent back to back: the signal, connection, or route side
Optimize network accessAndroid (Google) Radio state transition delay and tail time vary with the radio technology (3G, LTE, 5G) and carrier settings; 3G example: low power → full power about 1.5 s, idle → full power over 2 s
ID hn-weak-cell · Primary owner External (External) · Also Game team (Client development)
In elevators, basements, and deep inside buildings, retransmissions increase, speed drops, and eventually the connection drops.
Why Moving into a place with weak signal → Effect More radio retransmissions, lower speed, momentary dropouts → On screen Jitter and loss cause stutter and teleporting, and eventually a disconnect
Primary owner External (External) · Also Game team (Client development)
Game team action items
Polish the reconnect flow, show network quality.
External action items
Tell players the problem happens in places with weak signal (elevators, basements, deep inside buildings).
On the graph
Outliers only · RTT and loss (per mobile player)
Where to look
Check where the player was when the disconnect was reported (elevator, basement, inside a building) and the phone’s signal indicator, and compare by repeating the same actions where the signal is good
Confirmed if
RTT and loss rise and the connection drops only where the signal is weak, and it goes away after moving to a place with good signal
Ruled out if
Same even with good signal: the ISP segment or the server side
ID hn-5g-flip · Primary owner External (External) · Also Game team (Client development)
Inside buildings with weak 5G signal or at the edge of 5G coverage, the phone switches between 5G and LTE often, and each switch causes a ping spike or a brief dropout.
Why In a place where the 5G signal comes and goes (inside a building, at the edge of 5G coverage) → Effect The phone keeps switching between 5G and LTE, with a short gap each time → On screen Ping spikes at random even when standing still, with occasional freezes and teleporting
Primary owner External (External) · Also Game team (Client development)
Game team action items
Lengthen the interpolation buffer automatically when jitter rises; record network type changes (5G, LTE) in the logs taken when lag happens, to tell causes apart.
External action items
Tell players to switch to LTE-preferred mode in settings and compare, recommend using Wi-Fi.
Ballpark numbers
Each switch takes tens to hundreds of ms. Most 5G in Korea runs bundled with LTE (NSA), so the 5G part easily connects and drops over and over.
On the graph
Random spikes · RTT, network type changes (5G/LTE)
Where to look
Switch the phone to LTE-preferred mode and compare from the same spot. It’s more conclusive if the client logs changes in the network indicator from Android’s TelephonyDisplayInfo (OVERRIDE_NETWORK_TYPE_NR_NSA and so on) together with RTT
Confirmed if
RTT spikes line up with the times the 5G↔LTE indicator changes, and the spikes disappear in LTE-preferred mode
Ruled out if
Spikes even though the network indicator doesn’t change: “Weak mobile signal and dead zones” or the connection side
5G 통신서비스 품질평가 결과 발표과학기술정보통신부 As announced in 2020, 5G in Korea is offered in NSA mode, with the move to SA still planned
TelephonyDisplayInfoAndroid (Google) OVERRIDE_NETWORK_TYPE_NR_NSA: the network indicator shown when the device is on LTE and can use, or is using, dual connectivity (EN-DC) with 5G (NR)
ID hn-captive · Primary owner External (External) · Also Game team (Client development), Game team (Server development)
A café Wi-Fi login page or a corporate firewall blocks the game’s connection.
Why The login page hasn’t been completed yet, or a firewall blocks the game’s ports or UDP → Effect Connection attempts are blocked outright, or only some traffic gets through → On screen Can’t connect, or login works but entering the game fails
Primary owner External (External) · Also Game team (Client development), Game team (Server development)
Game team action items
Client: tell players why the connection is blocked (login page not completed, UDP blocked, and so on), switch to a fallback path automatically when UDP is blocked. Server: provide a fallback path such as TCP 443.
External action items
Tell players to complete the login page first on public Wi-Fi and to use a different network where traffic is restricted, such as a corporate network.
On the graph
Outliers only · Connection failures (by network)
Where to look
Have the failing player try connecting from another network such as mobile data, and check the server connection log for whether the first UDP packet arrived and whether the TCP 443 fallback path connects
Confirmed if
Fails only on specific Wi-Fi (café, office) and connects right away on other networks. The login page hasn’t been completed, or only UDP fails to reach the server
Ruled out if
Fails on every network: the account, the server, or “DNS failures and delays.” Fails across an entire country or ISP: “Country- or ISP-level UDP restrictions and packet inspection”
Check with
The player’s own environment
Sources (2)
RFC 8952: Captive Portal ArchitectureIETF Captive portal: a network that restricts access until requirements such as accepting terms or authenticating are met
ID isp-distance · Primary owner Infra team (Server infrastructure) · Also Infra team (Network infrastructure), Game team (Server development)
Even light travels only about 200,000 km per second in optical fiber. A distant server is slow no matter how good it is.
Why The server is far away (an overseas server, another continent) → Effect Round-trip time grows with distance (at least 10 ms per 1,000 km) → On screen Constant input lag on every action and a disadvantage in hit registration
Primary owner Infra team (Server infrastructure) · Also Infra team (Network infrastructure), Game team (Server development)
Game team action items
Mitigate only (code can’t change physics), offer region selection so players pick a nearby server, use lag compensation (rewind) to reduce the hit-registration disadvantage.
Infra team action items
Servers/OS: put regional servers where most players are. Network: put edge locations (PoPs) close to players, choose links and routes with fewer detours.
Ballpark numbers
Seoul–Tokyo about 30 ms, Seoul–Singapore about 75 ms, Seoul–US West Coast about 140 ms, Seoul–Europe about 230–270 ms (round trip, over real routes). Few major cables run along the direct line to Europe, so traffic goes around through Southeast Asia and Suez or through the US, and latency is much higher than the distance alone suggests.
On the graph
Always high · RTT (by country/region)
Where to look
Tag client IPs with a country and look at the RTT distribution per country. Run ping and traceroute to the server from a cloud-region VM in that area or from RIPE Atlas probes (selected by country or ASN)
Confirmed if
RTT from distant countries is always high regardless of time of day, and close to both the minimum delay computed from distance (10 ms round trip per 1,000 km) and public latency statistics
Ruled out if
Far above what distance explains: points to “Detour routing.” Rises only in the evening: “Peak-hour congestion at peering links”
ITU-T G.114: One-way transmission timeITU Planning value for optical fiber propagation delay: 5 µs/km (about 200,000 km per second, 10 ms round trip per 1,000 km)
Azure network round-trip latency statisticsMicrosoft Azure Measured median round-trip times from Seoul (Korea Central): Tokyo 30 ms, Singapore 68 ms, US West 124–136 ms, Europe 234–244 ms
Probe Selection (RIPE Atlas REST API)RIPE NCC Select RIPE Atlas probes by country, region, ASN, or address prefix and run ping and traceroute from them
ID isp-satellite · Primary owner External (External) · Also Game team (Client development), Game team (Server development)
Satellite signals have to travel to space and back. With geostationary satellites the round trip alone exceeds 0.5 seconds. Low Earth orbit satellites such as Starlink are usually fast, but latency fluctuates and the link can drop briefly at the moment routes are reassigned.
Why Connecting from home, a ship, or a plane over GEO or LEO satellite internet, or over in-flight Wi-Fi that uses satellites → Effect GEO satellites sit at about 36,000 km, so the distance itself is long. LEO systems reassign the terminal–satellite–ground station path at short intervals, with a brief burst of delay and loss at each reassignment → On screen GEO: heavy input lag on every action. LEO: fine most of the time, then stutter and teleporting at regular intervals
Primary owner External (External) · Also Game team (Client development), Game team (Server development)
Game team action items
Client: lengthen the interpolation buffer automatically to match jitter, send inputs redundantly to survive brief loss, show connection quality. Server: account for satellite latency when setting timing windows and lag compensation limits, use timeouts that don’t kick players over gaps of about 1 second.
External action items
Tell players that satellite internet can have high latency or periodic spikes, and advise a wired terrestrial connection for competitive content where possible.
Ballpark numbers
At geostationary orbit (36,000 km altitude) the signal takes 260 ms one way just to cross space, so the round trip exceeds 520 ms (ITU-T G.114). For LEO Starlink, official data (15-second averages) put the US peak-hour median at 33 ms, with even the worst 1% (p99) under 65 ms (2024). Measurement studies found that latency shifts each time routes are reassigned every 15 seconds, along with brief outages under 1 second. In 2018 in-flight internet measurements, satellite-based connections averaged 750 ms round trip.
On the graph
Outliers only · RTT/jitter (by satellite ISP ASN)
Where to look
Check whether the client IP’s ASN belongs to a satellite internet provider, and plot the RTT distribution and time series of that provider’s players separately. Ping the server for a few minutes straight from RIPE Atlas probes in that ASN, or have the player leave ping running and measure the interval between spikes
Confirmed if
GEO providers: RTT always above 500 ms. LEO providers: normally tens of ms, with RTT shifting or brief drops about every 15 seconds
Ruled out if
Not a satellite provider but RTT always high: “Propagation delay (physical distance)” or “Detour routing.” Irregular spikes: points to Wi-Fi or mobile signal
Check with
Infra tools (no game code needed)
Learn more
LEO satellites are close (one Starlink hop takes 1.8–3.6 ms), so typical latency can be similar to a terrestrial connection. If the point where the ground station hands traffic to the internet (PoP) is far from the game server, though, the path gets longer by that much, and routing through laser links between satellites adds more latency. Measurement studies attribute the 15-second fluctuation to route reassignment that happens at the same moment worldwide; it is unrelated to switching between satellites. In-flight Wi-Fi latency varies widely by technology (satellite or ground-based cell towers), and systems that use geostationary satellites have the same long round trip described above.
Sources (5)
ITU-T G.114: One-way transmission timeITU One-way propagation delay planning values for satellite links: 12 ms at 400 km altitude, 110 ms at 14,000 km, 260 ms at 36,000 km (geostationary)
Improving Starlink’s LatencyStarlink US peak-hour median 48.5 ms→33 ms, slowest 1% (p99) over 150 ms→under 65 ms (2024); one satellite hop 1.8–3.6 ms; routing over laser links adds latency, and so does the distance from the ground station to the internet access point (PoP)
A Multifaceted Look at Starlink Performance (WWW 2024)ACM Starlink reassigns routes every 15 seconds at the same moment worldwide; at these boundaries latency and throughput fluctuate and brief outages under 1 second occur (unrelated to switching between satellites); terminal↔satellite↔ground station latency about 40 ms
Probe Selection (RIPE Atlas REST API)RIPE NCC Select RIPE Atlas probes by country, region, ASN, or address prefix and run ping and traceroute from them
ID isp-routing · Primary owner Infra team (Network infrastructure) · Also External (External)
Because of interconnection agreements between ISPs, traffic to even a nearby server can take a long way around.
Why Your ISP and the server’s ISP aren’t directly connected → Effect Traffic passes through another country or city, adding distance and hops → On screen Only players on certain ISPs have unusually high ping
Primary owner Infra team (Network infrastructure) · Also External (External)
Infra team action items
Connect to multiple ISPs (multihoming), monitor ping per ISP to find the ones taking detours, negotiate route changes with ISPs.
External action items
Ask the ISP to adjust its routing.
Ballpark numbers
Even within one country, ping can differ two- or threefold depending on the route.
On the graph
Always high · RTT (by ISP/ASN)
Where to look
Compare RTT by ISP (ASN), and use traceroute or mtr from RIPE Atlas probes on the slow ISP, or from players, to see which countries and cities the route passes through. Measure IPv4 and IPv6 separately (mtr -4, -6)
Confirmed if
Within the same region, one ISP is always higher and its route passes through another country or a distant city. Or only one address family (IPv4 or IPv6) is high
Ruled out if
All ISPs similarly high: “Propagation delay (physical distance).” High only in the evening: “Peak-hour congestion at peering links”
Check with
Infra tools (no game code needed)
Learn more
IPv4 and IPv6 routes are chosen separately, so for the same server one of them can take a long detour and be slow (a 2016 APNIC measurement found separate groups of users within a single ISP for whom IPv6 was 15, 25, or 75 ms slower than IPv4). Apps that use Happy Eyeballs (RFC 8305) go with whichever of IPv6 and IPv4 connects first. They try IPv6 first, and if it connects within the recommended 250 ms, they never try IPv4. So they tend to end up on IPv6 even when that path is a little slower. If ping is high only on a particular ISP, measure IPv4 and IPv6 separately.
The Internet at the Speed of Light (HotNets 2014)ACM Actual router paths are about 1.5 times the straight fiber distance (median), and packets between two nearby points sometimes travel around the far side of the globe (hairpinning)
Probe Selection (RIPE Atlas REST API)RIPE NCC Select RIPE Atlas probes by country, region, ASN, or address prefix and run ping and traceroute from them
IPv6 Performance – RevisitedAPNIC Comparison of IPv6 and IPv4 round-trip times for the same dual-stack users: some access networks handle IPv6 packets completely differently, producing groups within one ISP where IPv6 is 15, 25, or 75 ms slower
ID isp-peak · Primary owner Infra team (Network infrastructure) · Also External (External)
Around 9–11 PM, video traffic surges and the links between ISPs (peering) tend to get congested.
Why Evening streaming and downloads pile up → Effect Queues build and packets drop on peering links → On screen Players on certain ISPs get stutter and teleporting only in the evening
Primary owner Infra team (Network infrastructure) · Also External (External)
Infra team action items
Add more direct connections with the affected ISP, route around congested paths, monitor evening loss and ping per ISP.
External action items
Ask the ISP to add capacity on the peering link.
On the graph
High at certain hours · RTT/loss (by ISP)
Where to look
Plot RTT and loss per ISP (ASN) by time of day, and get mtr runs in the evening and during the day from RIPE Atlas probes on that ISP or from players to find the hop where loss starts
Confirmed if
Only one ISP sees RTT and loss rise around 9–11 PM every evening, and in mtr the loss and delay persist from the inter-ISP link all the way to the destination
Ruled out if
All ISPs rise together: points to our own links or servers. Only one household is bad in the evening: “Congested Wi-Fi channel”
Probe Selection (RIPE Atlas REST API)RIPE NCC Select RIPE Atlas probes by country, region, ASN, or address prefix and run ping and traceroute from them
ID isp-cable · Primary owner External (External) · Also Infra team (Network infrastructure)
When a submarine cable is cut, traffic takes long detours for weeks (sometimes months) until it is repaired, and the remaining links get congested.
Why Cable cut or equipment failure → Effect Traffic crowds onto long detour routes and the remaining links → On screen Ping surges and packet loss for overseas players that last days to weeks
Primary owner External (External) · Also Infra team (Network infrastructure)
Infra team action items
Secure links on other routes, move traffic onto them during an outage.
External action items
Tell overseas players the cause and the expected recovery time, ask the link provider for its repair schedule.
On the graph
Step change · RTT (by overseas country)
Where to look
Find when RTT and loss rose on the per-country graph, match that time against Cloudflare Radar’s internet outage summaries and submarine cable operators’ notices, and check with traceroute whether the route now goes around another continent
Confirmed if
From a certain moment, RTT for a specific overseas region steps up and stays there for days to weeks, with cable outage reports from the same period. The route has switched to an unusually long detour
Ruled out if
Back to normal within a few days with no outage reports: “BGP route changes and convergence” or a problem in the ISP’s segment
Q2 2024 Internet disruption summaryCloudflare Red Sea cables damaged in February 2024 were still under repair in July (conflict zone); the EASSy and Seacom cuts in May were repaired in 19 days
Q1 2024 Internet disruption summaryCloudflare West African cable cuts (March 14) were repaired 3–6 weeks later, with traffic moved to other cables in the meantime
ID isp-bgp · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)
When internet routing information changes, packets are lost for the few seconds to tens of seconds (rarely a few minutes) it takes to converge again.
Why Routing information changes somewhere in an ISP’s network → Effect For a few seconds to tens of seconds, packets vanish or switch to a new route → On screen A sudden freeze of a few seconds, then ping settles at a different value (e.g., 40 → 70 ms)
Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)
Game team action items
Use timeouts that survive brief outages (don’t drop a connection right away just because it went silent for a few seconds).
Infra team action items
Monitor routes (watch for route and ping changes on our IP prefixes), detect failures on our own links within 1 second with BFD and fail over (the default BGP hold time is 90–180 seconds), move traffic to another link if the route switches to a long path and doesn’t come back.
External action items
Ask the ISP to investigate segments of its network where routes change often.
On the graph
Step change · RTT, traceroute path
Where to look
Compare traceroute and mtr paths from before and after the moment RTT changed, and check the BGP route change history for our prefix in RIPEstat BGPlay
Confirmed if
A freeze of a few seconds, then RTT moves to a different level, with BGP updates and AS path changes at the same time
Ruled out if
No route changes on record but high only in the evening: “Peak-hour congestion at peering links.” Only some connections bad: “One faulty ECMP path”
BGPlay (RIPEstat Data API)RIPE NCC Shows the BGP routes for an address prefix at the start time, the BGP updates observed during the period, and the ASes on the path
ID isp-ecmp · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)
ISPs and data centers keep several paths to the same destination and send each connection down one of them. If a single path fails, only the players assigned to it keep lagging.
Why One link or device in a bundle of links is faulty or congested → Effect The path is chosen from the address and port combination (hash), so only connections assigned to that path see loss and delay → On screen Same region and ISP, but only some players keep teleporting. Reconnecting sometimes fixes it
Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)
Game team action items
Record per-connection loss and retransmission stats so you can pull the IP, port, and time for affected players (for TCP, the retransmission count from TCP_INFO; for UDP, compute it from missing packet sequence numbers).
Infra team action items
Collect affected players’ IPs, ports, and times and pass them to the ISP or data center, monitor loss per path, measure paths with the same protocol and port as the game (mtr --tcp or --udp with --port), remove the faulty link or device from the bundle if the path runs over our equipment.
External action items
Ask the ISP to check and replace the faulty path, tell players they can work around it for now by reconnecting (when reconnecting changes the port).
Ballpark numbers
With 4 paths, only about a quarter of players are affected. A ping test may take a different path from the game and come back perfectly fine.
On the graph
Outliers only · Per-connection loss/retransmissions (by IP/port)
Where to look
Split per-connection loss and retransmissions by source IP and port. Run mtr in UDP mode (-u) against the game port (-P) with a fixed source port (-L), and repeat several times with different source ports. With -P and no -L, the source port changes on every probe and several paths get mixed together
Confirmed if
Within the same region and ISP, only certain source port (or address) combinations keep losing packets, and the problem goes away when a reconnect changes the port
Ruled out if
Bad no matter which port: congestion or failure across a whole segment
Check with
Infra tools (no game code needed)
Learn more
To keep a connection’s packets in order, network devices (ECMP, LAG) pin each connection to one path using a value computed from its addresses and ports (or only its addresses, depending on device settings). Where only addresses are used, reconnecting lands on the same path and doesn’t help. So when reports like “ping is fine but the game lags” and “reconnecting fixed it” come in together, suspect this cause.
mtr(8) manual page sourcemtr The -u (UDP), -P (destination port), and -L (UDP source port) options; with only -P, the probe sequence number goes into the source port, so it changes on every probe
ID isp-shaping · Primary owner External (External) · Also Game team (Server development), Infra team (Network infrastructure)
When you go over your data allowance, or on plans that manage certain kinds of traffic, packets get delayed or dropped.
Why Speed throttled after the plan’s data runs out, or certain traffic restricted → Effect Packets wait in a queue or get dropped → On screen Lag after a certain amount of usage, especially on mobile
Primary owner External (External) · Also Game team (Server development), Infra team (Network infrastructure)
Game team action items
Reduce game traffic (compression, send only what’s needed).
Infra team action items
If game traffic is delayed or dropped only on a particular ISP, gather evidence and escalate to that ISP.
External action items
Tell players to check whether their plan’s data has run out or is being throttled and whether other apps on the same phone are using data, ask the ISP whether it restricts game traffic.
Ballpark numbers
Korean mobile plans usually throttle to 1–5 Mbps once the data allowance runs out, and cheap plans to hundreds of kbps. The game itself uses little data, but when other apps on the same phone use data, a queue forms in front of the throttling equipment.
On the graph
Hits a ceiling · Throughput, RTT
Where to look
Have the player check remaining data and throttling status in the carrier’s app and run a speed test to see the top speed. On the server side, compare loss and RTT per ISP
Confirmed if
Throughput stops rising at one value such as 1–5 Mbps or a few hundred kbps, and from then on RTT and loss grow whenever other apps on the same phone use data. Goes away after topping up data or switching to Wi-Fi
Ruled out if
No throttling but one ISP is bad: “Peak-hour congestion at peering links” or “Detour routing”
Check with
The player’s own environment
Sources (2)
SKT, 고객 선택권과 혜택 강화한 신규 5G 요금제 출시SK텔레콤 Examples of speed control after a 5G plan’s base data runs out: up to 400 kbps, 1 Mbps, 3 Mbps
SKT, 요금제 개편SK텔레콤 Service continues at up to 400 kbps after the included data runs out (Korea’s nationwide “safety net data” program)
ID isp-udp-block · Primary owner External (External) · Also Game team (Client development), Game team (Server development), Infra team (Network infrastructure)
Some networks block specific UDP addresses and ports or throttle UDP, and packet inspection equipment filters out protocols it doesn’t recognize. Games that communicate over UDP can’t connect on those networks or disconnect often.
Why Connecting from an ISP network that throttles UDP, or from a network with country- or ISP-level traffic inspection (censorship) equipment → Effect Specific UDP addresses and ports are blocked, UDP is throttled at busy hours, ports or protocols not on an allowlist are filtered out, or the first few packets get through before the flow is blocked → On screen Only players in certain countries or on certain ISPs can’t connect or get infinite loading, disconnect soon after connecting, or teleport from packet loss at busy hours
Right after login or maintenance, Always, Evening peak hours
Owner
Primary owner External (External) · Also Game team (Client development), Game team (Server development), Infra team (Network infrastructure)
Game team action items
Client: fall back automatically to a TCP/TLS 443 path if UDP doesn’t connect within a few seconds, also detect connections that drop soon after succeeding and retry them on the fallback path, log which path was used. Server: accept the same game protocol over TCP 443 (TLS) as well, adjust timeouts because the fallback path can add latency.
Infra team action items
Before launching in a new country, measure UDP reachability and peak-hour loss on local ISP networks, place relays or gateways that accept the TCP 443 fallback close to the region, monitor UDP and TCP connection success rates per country and ASN, gather evidence and escalate for ISPs confirmed to throttle UDP.
External action items
Ask the ISP or agency about its UDP restriction criteria and any relief, tell players to try connecting from another network to compare.
Ballpark numbers
Measurements cited by an IETF document show that 3–5% of networks block UDP entirely. When Google reviewed its 2016 QUIC (UDP-based) usage, 4.4% of clients couldn’t use UDP/QUIC because it was blocked or the path MTU was too small, mostly behind corporate firewalls, and it saw no case of an entire ISP blocking it. Another 0.3% were on networks where loss rose sharply at peak hours, apparently from UDP throttling; Google brought that down from 1% in 2015 by working with the ISPs.
On the graph
Outliers only · UDP connection success rate (by country/ASN)
Where to look
Split UDP connection success rate and TCP 443 fallback success rate by country and ASN. From a cloud VM or a player’s PC on that ISP’s network, test connections to the game’s UDP port and to TCP 443 separately, and compare mtr -u -P (game port) with mtr -T -P 443 to see at which hop responses disappear
Confirmed if
Only in a specific country or ASN, UDP gets no first response or drops within a few seconds, while TCP 443 works from the same place. With throttling, UDP loss rises clearly only at peak hours and TCP is less affected
Ruled out if
TCP fails too: points to a path outage, IP blocking, or “DNS failures and delays.” Same in every country: our own server or firewall configuration. Loss only during traffic bursts, for UDP and TCP alike: “Policer drops excess traffic”
Check with
Infra tools (no game code needed)
Learn more
According to an IRTF survey document, packet inspection equipment may pick out UDP flows by address, port, and protocol and block them, or block everything except the protocols it allows (an allowlist). If the equipment decides based on only a few fields of the packet, even a small protocol change can get traffic blocked. In the early days of QUIC, one firewall let the first few packets through after a single header bit changed and then blocked the rest, so the client’s logic for falling back to TCP never kicked in. When launching in a new country, this can surface as reports that “it works fine in Korea, but players on some ISPs in that country can’t connect.” If it’s blocked only on one location’s network, such as a café or an office, see the “Public Wi-Fi and corporate network restrictions” entry.
Sources (4)
RFC 9308: Applicability of the QUIC Transport ProtocolIETF Measurement studies show 3–5% of networks block UDP entirely, so UDP-based apps must accept connection failures or provide a TCP (TLS) fallback; firewalls may block ports that aren’t tied to a registered service
The QUIC Transport Protocol: Design and Internet-Scale Deployment (SIGCOMM 2017)ACM 2016: 4.4% of clients couldn’t use QUIC over UDP because UDP/QUIC was blocked or the path MTU was small (mostly behind corporate firewalls; no ISP-wide blocking observed); 0.3% were on networks that appeared to throttle UDP (higher loss at peak hours; down from 1% in 2015 after requests to ISPs); a case where a firewall let only the first few packets through after a 1-bit header change and blocked the rest, defeating the TCP fallback logic
RFC 9505: A Survey of Worldwide Censorship TechniquesIRTF Inspection equipment in the network can pick out and block TCP and UDP flows by address, port, and protocol (blocking of UDP endpoints has been observed with QUIC); allowing only approved protocols leads to overblocking, and throttling of specific traffic is also used
mtr(8) manual page sourcemtr Sends UDP with -u or TCP SYN with -T and sets the destination port with -P, so the route is measured with the same protocol and port as the game
Loose connectors, old wiring, or a faulty modem cause steady packet loss and periodic line drops.
Why Damaged cable, poor contact, faulty modem or optical network terminal → Effect Packets dropped from bit errors; now and then the line drops for a few seconds to about a minute while it reconnects → On screen Steady low-level packet loss, occasional freezes of a few seconds or disconnects
Tell players to check whether other games and video calls also drop, and if so, to ask their ISP for a line check.
On the graph
Random spikes · Loss rate, line reconnect log
Where to look
Measure loss up to the ISP’s first hop for a few minutes with pathping (or mtr), and check reconnect times in the internet (WAN) connection log on the router’s admin page
Confirmed if
Steady loss from the ISP’s first hop even when the line is idle, and reconnect times in the router log line up with freezes and disconnects. Other games and video calls drop at the same time
Ruled out if
Loss starts on the wireless hop to the router: “Wi-Fi interference and weak signal.” Loss starts far into the ISP network: the ISP’s route
ID isp-dns · Primary owner External (External) · Also Game team (Client development)
If DNS, which turns server names into addresses, is slow or fails, the game can’t find its login or patch servers.
Why ISP DNS outage or misconfiguration → Effect The login or patch server address can’t be resolved → On screen A long wait after pressing Connect, or no connection at all. Players already connected are fine
Primary owner External (External) · Also Game team (Client development)
Game team action items
Cache addresses (remember the last server address that connected successfully), set up multiple DNS resolvers (query another one if one fails).
External action items
Tell players to try switching to another DNS resolver, such as a public DNS.
On the graph
Outliers only · Login failures (by ISP), DNS lookup time
Where to look
Query the login server name against the ISP’s DNS and a public DNS separately with Resolve-DnsName -Server (or nslookup), and compare response times and results
Confirmed if
Only the ISP’s DNS fails to respond or takes a long time, and switching to a public DNS connects right away. Players already connected are fine
Ruled out if
Every DNS returns the address right away but connecting still fails: points to the route, a firewall, or the server
Cloudflare 1.1.1.1 Incident on July 14, 2025Cloudflare When a public DNS resolver stopped for 62 minutes, users who could no longer resolve names effectively lost access to all internet services
ID isp-ddos-path · Primary owner Infra team (Network infrastructure) · Also External (External)
Massive attacks aimed at the game company, or at someone else on the same network, fill up shared links.
Why A flood of attack traffic → Effect Legitimate traffic on the same links gets delayed and dropped too → On screen Many players teleport, disconnect, or can’t connect at the same time
Primary owner Infra team (Network infrastructure) · Also External (External)
Infra team action items
Use a DDoS protection service, reroute traffic during attacks, hide server addresses (keep servers behind protection equipment and don’t expose their real addresses).
External action items
If the attack targets someone else on the same network, ask the ISP to block it upstream.
On the graph
Hits a ceiling · Link inbound traffic (bps/pps), interface drops
Where to look
Inbound traffic and dropped packet counts on our links and equipment, plus the DDoS protection service’s attack detection log, lined up with the times when disconnects cluster
Confirmed if
Inbound traffic flattens out at link capacity and drops rise, while players across many regions and ISPs teleport or disconnect at the same moment
Ruled out if
Links have headroom but only some ISPs are bad: congestion or routing problems in the ISP segment
Check with
Infra tools (no game code needed)
Sources (3)
Infrastructure layer attacksAWS Volumetric attacks such as UDP reflection and SYN floods overwhelm network capacity or tie up firewall and load balancer resources
ID isp-cgnat · Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Network infrastructure)
Mobile networks and some ISPs have many subscribers share one IP address, and they delete the mappings of idle connections after a short time.
Why ISP equipment manages the session table for huge numbers of subscribers → Effect Session table limits, short idle timeouts → On screen Disconnects after sitting idle, and false positives that block everyone sharing the same IP at once
Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Network infrastructure)
Game team action items
Client: send heartbeats at no more than half the shortest idle timeout (which can be about 30 seconds on mobile networks), send them from the client because ISP CGNAT mappings are reliably refreshed only by outgoing packets, reconnect automatically on disconnect. Server: respond to heartbeats and close the connection proactively if none arrive for a set time, use a session token to keep the same player when a changed mapping gives them a new address and port, be careful with IP-based blocking because many people can share one IP (decide together with account and device signals).
Infra team action items
Adjust per-IP connection limits and new-connections-per-second limits on firewalls and DDoS protection equipment to allow for ISP-shared IPs (raise the thresholds for mobile carrier ranges or exempt them).
Ballpark numbers
On mobile networks the UDP idle timeout can be as short as about 30 seconds.
On the graph
Mass disconnect · Disconnects (heartbeat timeouts), idle time before disconnect (by ISP)
Where to look
In the connection logs, check how many accounts connect from one IP at the same time and their ISP (ASN), and collect per ISP the idle time of connections that dropped after idling. On the player side, check the internet (WAN) address on the router’s admin page
Confirmed if
Multiple accounts connect from one IP in mobile carrier ranges, and idle times before disconnect cluster at short values around 30–60 seconds. The router’s WAN address is in 100.64.0.0/10 (shared address space for carrier NAT) or differs from the address the server sees
Ruled out if
Concentrated among home router users regardless of ISP: “NAT mapping expiry”
ID isp-vpn · Primary owner External (External) · Also Infra team (Network infrastructure), Game team (Server development)
With a VPN or game booster on, packets go through that company’s relay servers. If the relay is far away or busy, the connection can actually get slower.
Why The VPN or booster sends every game packet through its relay servers → Effect Distance to the relay and its congestion add up, and tunnel headers shrink the MTU (the largest packet size that can be sent at once) → On screen Higher ping and packet loss; can’t connect when the relay address gets blocked along with everyone else using it
Primary owner External (External) · Also Infra team (Network infrastructure), Game team (Server development)
Game team action items
Keep UDP packets at 1,200 bytes or less (so they don’t fragment even when tunnel headers shrink the MTU), account for shared VPN and booster relay addresses in IP-based blocking and decide together with account and device signals.
Infra team action items
If you have many overseas players, set up your own locations (PoPs) near them, check the routes of ISPs with a cluster of reports saying “turning on the booster made it better.”
External action items
Tell players to turn off the VPN or booster and compare.
Ballpark numbers
A nearby relay adds a few ms; a detour through another country adds tens of ms to over 100 ms.
On the graph
Outliers only · RTT (per player), network operator of the client IP
Where to look
Check whether the client IP’s ASN belongs to a VPN, booster, or hosting provider, and have the player turn off the VPN or booster and compare ping and traceroute
Confirmed if
RTT and loss rise, or connections get blocked, only with the VPN or booster on, and traceroute shows hops through the relay server
Ruled out if
Same with it on or off: the line or the ISP segment. Better with it on: a problem in the original ISP route (“Detour routing,” “Peak-hour congestion at peering links”)
Check with
The player’s own environment
Learn more
Conversely, when the ISP’s route is bad, a booster can take a better route and lower ping. Reports that “turning on the booster made it better” are a clue to an ISP route problem such as detour routing or evening congestion.
Azure network round-trip latency statisticsMicrosoft Azure Round-trip latency added when going through a location (PoP) in another country: Seoul–Tokyo 30 ms, Seoul–Hong Kong 39 ms, Seoul–Singapore 68 ms
ID dc-firewall · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), Game team (Client development)
A firewall tracks every connection it lets through by recording it in a session table. Once the table is full, it can’t accept new connections.
Why A connection surge or an attack pushes the session count to its limit → Effect No free entry to record a new connection, so it’s refused → On screen Players trying to get in can’t connect or get infinite loading, and some existing connections disconnect too
Right after login or maintenance, When crowds gather
Owner
Primary owner Infra team (Network infrastructure) · Also Game team (Server development), Game team (Client development)
Game team action items
Server: smooth out connection surges with a login queue, reuse connections to avoid opening short ones over and over, proactively close connections whose heartbeats have stopped (so dead connections don’t hold session table entries for long). Client: send heartbeats at no more than half the shortest idle timeout, reconnect automatically on disconnect with growing, randomized retry intervals (so everyone doesn’t pile back in at once).
Infra team action items
Enlarge the session table, clean up short-lived connections quickly (shorten the timeout for closed sessions), tell the game team whenever you shorten the idle session timeout so they can match the heartbeat interval, block attacks, alert on session table utilization.
On the graph
Hits a ceiling · Firewall session count, new connection failures
Where to look
Graph the firewall’s concurrent session count together with its session limit, and search the device log for packets dropped because a session couldn’t be created. On a Linux firewall, compare nf_conntrack_count with nf_conntrack_max and check dmesg for “nf_conntrack: table full, dropping packet”; on an AWS instance, check conntrack_allowance_exceeded in ethtool -S
Confirmed if
New connection failures rise from the moment the session count flattens at the limit, along with session creation failure logs or drop counters
Ruled out if
Session count well below the limit but connections fail: “Connection queue (listen backlog) overflow” or the login server. Only idle connections drop: “Cloud security group connection tracking expiry”
Check with
Infra tools (no game code needed)
Sources (5)
Netfilter Conntrack Sysfs variablesLinux kernel Maximum entries in the connection tracking table (nf_conntrack_max), how long closing connections are kept (TIME_WAIT and FIN_WAIT default 120 seconds), established TCP default 5 days, current entry count (nf_conntrack_count)
Amazon EC2 security group connection trackingAWS Once an instance exceeds the number of connections it can track, packets for new connections are dropped; idle connections can exhaust the tracking table
net/netfilter/nf_conntrack_core.c (Linux v6.12)Linux kernel When the connection tracking table is full, the kernel logs “nf_conntrack: table full, dropping packet” and drops packets for new connections
ID dc-ddos · Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
Diverting traffic to a scrubbing center to stop attacks makes the route longer, and legitimate players are sometimes mistaken for attackers and blocked.
Why After an attack is detected (or all the time), inbound traffic is diverted to a scrubbing center → Effect The route gets longer, and some legitimate packets are flagged as attack traffic → On screen Ping rises for everyone; players in certain regions or on certain ISPs can’t connect
Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
Game team action items
Document the game’s traffic pattern (ports, packet sizes, packets per second) and share it with the infra team, keep UDP packets at 1,200 bytes or less.
Infra team action items
Write protection rules that fit the game’s traffic pattern, use regional scrubbing locations, reduce TCP packet size on tunnel segments (MSS clamping), check for false positives with connection failure rates per region and ISP.
Ballpark numbers
A scrubbing location in the same country adds a few ms; going through a location in another country adds 30–100 ms or more. Usually only inbound traffic takes the detour, and the server’s responses go straight out. If the filtered traffic comes back through a tunnel, the largest packet size that can be sent at once (MTU) also shrinks, which can lead to a problem where only large packets vanish.
On the graph
Step change · RTT (ping), connection failure rate by region/ISP
Where to look
Put the protection device’s or service’s diversion (scrubbing) start and end records and its block logs on the same timeline as the RTT graph and the connection failure rates per region and ISP. From the affected region, check with mtr or traceroute whether a scrubbing location shows up in the path
Confirmed if
RTT steps up when diversion turns on, stays there, and comes back down when it turns off. Or legitimate player addresses show up in the block log, and only that region or ISP sees its connection failure rate rise
Ruled out if
RTT rises at times with no diversion or block records: “Detour routing” or “BGP route changes and convergence.” Only large packets vanish: “MTU mismatch (only large packets vanish)”
Check with
Infra tools (no game code needed)
Sources (2)
Maximum transmission unit and maximum segment sizeCloudflare Inbound traffic is delivered after filtering over a GRE tunnel (MTU 1,476) while outbound responses go straight to the internet (DSR); limiting TCP MSS to 1,436 or less is recommended, and without it large packets are dropped or fragmented
ID dc-lb-idle · Primary owner Infra team (Network infrastructure) · Also Game team (Client development), Game team (Server development)
A load balancer deletes idle connections after a set time. The game assumes the connection is still alive, and then the player gets disconnected.
Why The player sends no packets for a while (chat window open, away from keyboard) → Effect The load balancer cleans up the idle connection (common defaults are 60–350 seconds) → On screen Disconnect the moment the player moves again
Primary owner Infra team (Network infrastructure) · Also Game team (Client development), Game team (Server development)
Game team action items
Client: send heartbeats at no more than half the shortest idle timeout (30 seconds or less behind a 60-second ALB), reconnect automatically on disconnect. Server: respond to heartbeats and close the connection proactively if none arrive for a set time, resume the session with a session token.
Infra team action items
Check the idle timeout of every load balancer on the path, share the values with the game team, and raise them if needed.
Ballpark numbers
Defaults are 60 seconds for AWS ALB, 350 seconds for TCP and 120 seconds for UDP on NLB, and 4 minutes for TCP on Azure Load Balancer. The ALB and NLB TCP values can be changed, but the NLB UDP value of 120 seconds can’t. When the time runs out, ALB also closes the server-side connection, while NLB deletes it silently, so the server often never finds out.
On the graph
Mass disconnect · Disconnects, idle time before disconnect
Where to look
Check the idle timeout setting of each load balancer on the path, and collect for each dropped connection the time from its last packet to the disconnect. On AWS NLB, also check TCP_ELB_Reset_Count in CloudWatch (number of RSTs sent by the load balancer)
Confirmed if
Idle times of dropped connections cluster just past the setting (60 s for ALB, 350 s for NLB TCP, and so on), and it reproduces when you sit still longer than that and then move. On NLB, TCP_ELB_Reset_Count rises at those times
Ruled out if
Disconnects regardless of idle time: not this cause. Clusters near 350 seconds on a server that clients reach directly without a load balancer: “Cloud security group connection tracking expiry.” On the player’s home router: “NAT mapping expiry”
Check with
Infra tools (no game code needed)
Sources (4)
Edit attributes for your Application Load BalancerAWS ALB idle timeout default 60 seconds (1–4,000 seconds); the load balancer closes the connection if the client or target connection is silent for that long
Network Load BalancersAWS NLB TCP idle default 350 seconds (60–6,000 seconds); after that it only stops tracking and answers later data with RST; UDP flows are fixed at 120 seconds
Configure load balancer TCP reset and idle timeoutMicrosoft Azure Azure Load Balancer idle timeout default 4 minutes (4–100 minutes), no guarantee the session is kept beyond that, TCP reset is optional
ID dc-cloud-conntrack · Primary owner Infra team (Server infrastructure) · Also Game team (Client development), Game team (Server development)
The firewall attached to a cloud server (security group) also tracks connections, and tracking entries for idle connections expire after a set time. Even on servers that clients reach directly without a load balancer, players who sat idle can get disconnected.
Why The security group is set up so that it tracks game connections (only certain addresses allowed, restricted outbound rules, traffic through an NLB, and so on) → Effect The tracking entry for a connection that sat idle for a while expires, and the security group silently drops packets that arrive after that → On screen After being away, the player moves again, gets no response, then disconnects. The server program doesn’t notice for a long time
Primary owner Infra team (Server infrastructure) · Also Game team (Client development), Game team (Server development)
Game team action items
Client: send heartbeats at no more than half the shortest idle timeout (175 seconds or less for 350 seconds on TCP, 90 seconds or less for 180 seconds on UDP streams), reconnect automatically on disconnect. Server: respond to heartbeats and close the connection proactively if none arrive for a set time, resume the session with a session token.
Infra team action items
Check the instance’s connection tracking timeout (TcpEstablishedTimeout) and raise it if needed (UDP can’t be raised because 180 seconds is already the maximum), review a security group setup that creates no tracking (game ports open to all addresses, all outbound allowed; connections through an NLB are still tracked), run idle tests when moving to a new instance generation.
Ballpark numbers
On AWS, Nitro v6 instance types delete tracking entries for idle TCP connections after 350 seconds by default (5 days on other types). For UDP, the defaults are 180 seconds for flows with several request/response exchanges (streams) and 30 seconds for flows that went only one way or had a single request and response.
On the graph
Mass disconnect · Disconnects, idle time before disconnect
Where to look
Check the instance’s connection tracking timeout setting and the security group rules (whether the setup creates tracking), and collect the idle times of dropped connections. Right after a disconnect, use ss -tnoi on the server to see whether the connection stays ESTABLISHED with the retransmission timer (timer:(on,…)) running and backoff growing
Confirmed if
Idle times of dropped connections cluster just past 350 s for TCP, 180 s for UDP streams, or 30 s for one-way UDP, and the server-side socket stays ESTABLISHED without noticing the disconnect (if the server has data to send, it just keeps retransmitting)
Ruled out if
The security group setup doesn’t track (game ports open to all addresses, all outbound allowed, no NLB in the path): not this cause. Traffic goes through an NLB: compare the values with “Load balancer idle timeout”
Check with
Infra tools (no game code needed)
Sources (3)
Amazon EC2 security group connection trackingAWS TCP idle tracking default 350 seconds (Nitro v6; 432,000 seconds = 5 days on others), UDP one-way 30 seconds and stream 180 seconds (180 max); rules that allow all addresses aren’t tracked; connections through an NLB are always tracked
ss(8) — Linux manual pageiproute2 In -o, timer:(on,…) is the retransmission timer; in -i, backoff is the number of times the retransmission wait has doubled
ID dc-nat-gateway · Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
When servers in a private subnet connect out (platform authentication, payments, external APIs), a NAT gateway rewrites their address and port. If concurrent connections to the same destination exceed the gateway’s port limit, new connections fail.
Why Servers open many short connections to the same external address, such as platform authentication or payments, or keep connections open for a long time → Effect The NAT gateway can’t allocate any more source ports for that destination, so new connections fail → On screen The game itself is fine, but only features that call external services, such as login, payments, and reward delivery, fail or slow down (can’t connect / infinite loading, dropped action / rollback)
Right after login or maintenance, Evening peak hours, When crowds gather
Owner
Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
Game team action items
Reuse connections to external APIs (HTTP keep-alive, connection pools) and don’t open a new connection per request, send keepalives on idle pooled connections more often than the NAT idle timeout (350 seconds on AWS) or close them first, retry failures with growing, randomized intervals, record failure rates and latency per external call.
Infra team action items
Add IP addresses to the NAT gateway (an AWS public NAT gateway takes only 2 Elastic IPs by default, so request a quota increase for more), split gateways per availability zone and subnet, alert on port allocation failure metrics (AWS ErrorPortAllocation, Failed in Azure SNAT Connection Count, OUT_OF_RESOURCES in Google Cloud dropped_sent_packets_count), raise the minimum ports per VM or use dynamic port allocation on Google Cloud NAT.
Ballpark numbers
An AWS NAT gateway can open up to 55,000 concurrent connections to the same destination (IP, port, protocol) per IP address, and you can attach up to 8 IPs to raise that. It deletes connections that stay silent for 350 seconds and answers later packets on them with RST. Azure NAT Gateway has 64,512 SNAT ports per public IP (up to 16 IPs). Google Cloud NAT divides 64,512 ports per NAT IP among VMs, and the default minimum is 64 ports per VM (static allocation), so with default settings a single VM is usually limited to 64 concurrent connections to the same destination.
On the graph
Hits a ceiling · NAT gateway concurrent connections, port allocation failures
Where to look
Line up the NAT gateway metrics ErrorPortAllocation, ActiveConnectionCount, and PacketsDropCount in AWS CloudWatch (on Azure, SNAT Connection Count filtered by the Failed state and Dropped Packets; on Google Cloud, dropped_sent_packets_count with reason OUT_OF_RESOURCES) against the times the game server’s external calls failed
Confirmed if
ErrorPortAllocation (Failed SNAT Connection Count on Azure, OUT_OF_RESOURCES drops on Google Cloud) goes above 0 when external calls fail, and the failures concentrate on calls to one or two heavily used destinations such as authentication or payment servers
Ruled out if
Port allocation failures at 0, but the game server’s connect fails with EADDRNOTAVAIL and TIME_WAIT is close to the size of the ephemeral port range: “Ephemeral port exhaustion on server-to-server connections.” Connections succeed but responses are slow: “External service dependency”
Check with
Infra tools (no game code needed)
Learn more
“Ephemeral port exhaustion on server-to-server connections” is about one server running out of ephemeral ports. This limit sits on the NAT gateway and is shared by all the servers behind it (Google Cloud NAT divides it per VM). If only external calls fail while the servers still have plenty of room in TIME_WAIT and the ephemeral port range, this is the cause. Ports from closed connections also aren’t reused for the same destination right away (Azure applies a cooldown; Google Cloud blocks them during TIME_WAIT), so the more you repeat short connections, the sooner you hit the limit.
Sources (7)
NAT gateway basicsAWS 55,000 concurrent connections per IPv4 address to the same destination (destination IP, port, protocol), expandable by attaching up to 8 IPs (public NAT gateways get 2 Elastic IPs by default, more through a quota increase request); bandwidth scales automatically from 5 to 100 Gbps and throughput from 1 million to 10 million packets per second, and packets beyond that limit are dropped
NAT gateway metrics and dimensionsAWS ErrorPortAllocation: number of times a source port couldn’t be allocated (above 0 means too many concurrent connections), ActiveConnectionCount, IdleTimeoutCount (connections cleaned up after 350 seconds idle), PacketsDropCount
Troubleshoot NAT gatewaysAWS Connections expire after 350 seconds idle and later sends get an RST; keepalives shorter than 350 seconds recommended; when hitting the connection limit, add gateways per availability zone, add IPs, or reduce connections
Source Network Address Translation (SNAT) with Azure NAT GatewayMicrosoft Azure 64,512 SNAT ports per public IP (up to 16 IPs); each connection to the same destination needs a different port; closed ports go through a cooldown before reuse for the same destination
Metrics and alerts for Azure NAT GatewayMicrosoft Azure SNAT Connection Count filtered by the Failed state above 0 suggests SNAT port exhaustion; Dropped Packets
IP addresses and portsGoogle Cloud 64,512 ports each for TCP and UDP per NAT IP; default minimum ports per VM is 64 (static allocation) or 32 (dynamic allocation); the number of ports reserved for a VM caps its concurrent connections to the same destination; ports of closed connections can’t be used during TIME_WAIT
Logs and metricsGoogle Cloud dropped_sent_packets_count with reason OUT_OF_RESOURCES: packets dropped for lack of NAT IPs or ports
ID dc-lb-imbalance · Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
Connections pile onto one server, or players keep getting sent to a server that’s already dead.
Why The distribution rule is a poor fit, or the health check can’t see the real state → Effect One server alone is overloaded, or players try to connect to a dead server → On screen Only some channels or some players get slow motion, can’t connect, or get infinite loading
Right after login or maintenance, When crowds gather
Owner
Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
Game team action items
Implement a health check that answers the load balancer’s probes based on the real game state (tick progress, DB connections), report server load along with it.
Infra team action items
Switch to health checks that verify real game responses, distribute by server load, monitor differences in connection counts between servers.
On the graph
Outliers only · Connections/CPU utilization per server
Where to look
Overlay connection counts (ss -s) and CPU utilization of each server behind the load balancer on one graph, and compare the load balancer’s target health status (HealthyHostCount and UnHealthyHostCount in CloudWatch on AWS) with the game servers’ actual state
Confirmed if
Only one or two servers have far higher connections and CPU than the rest, or a server whose tick has stopped stays “healthy” and keeps taking new connections
Ruled out if
Connection counts even across servers but one channel is slow: load inside that channel (“Single-threaded zone overload (hotspot)”)
Load Balancing in the DatacenterGoogle Plain round robin lets CPU usage differ by up to 2× between tasks; weighted distribution where backends report their load in responses and health checks; a lame duck state in which a backend asks not to be sent new requests
Health checks for Network Load Balancer target groupsAWS Default health check every 30 seconds, target removed after 2 failures; UDP services are checked with TCP or HTTP health checks, so configuring them to reflect the real service state is recommended
ID dc-microburst · Primary owner Game team (Server development) · Also Infra team (Network infrastructure), Infra team (Server infrastructure)
When several servers send packets to thousands of players at the same instant, the small buffer on the switch port where that traffic converges overflows in less than 1 ms.
Why A world boss spawn or massive skills, or ticks on several servers lining up so they all send at once → Effect Buffers where several ports feed into one, or where a fast port feeds a slower one (hundreds of KB to a few MB per port), fill up in an instant → On screen Some packets dropped; many players teleport or have skills fail to go off at the same moment
Primary owner Game team (Server development) · Also Infra team (Network infrastructure), Infra team (Server infrastructure)
Game team action items
Spread sends evenly across the tick (pacing), offset each server’s tick start time slightly.
Infra team action items
Network: use switches with bigger buffers, spread traffic (place servers across several switches and ports), watch drop counters per switch port. Servers/OS: cap each server’s total send rate (Linux tc shaper).
Ballpark numbers
A 10 Gbps port can send about 1.25 MB in 1 ms. If traffic from two ports converges on one port at once, 1.25 MB piles up every 1 ms. Even at 10% average utilization over 1 second, the port can overflow at the 1 ms scale.
On the graph
Rises with load · Switch port output drops
Where to look
Collect output drop counters (ifOutDiscards, or output drops depending on the device) at the shortest interval you can on the switch ports the servers connect to and on the ports where their traffic converges, and match them against boss spawns and big battles. 1-second or 1-minute average utilization graphs won’t show it
Confirmed if
Average utilization is low, but output drops rise every time players crowd into one place, and at those moments many players report teleporting and skills not going off
Ruled out if
Drops rise steadily during hours of high average utilization: “Data center link saturation.” Input errors (CRC) rise: “Bad cables and port errors”
Data Center TCP (DCTCP) (SIGCOMM 2010)ACM Commodity switches have shallow buffers (48 ports share 4 MB, and one port can use up to about 700 KB); loss occurs when many flows converge on one port for a brief moment
RFC 2863: The Interfaces Group MIBIETF ifOutDiscards: number of packets discarded even though no error occurred, for reasons such as freeing buffer space
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)
Infra team action items
Network: prioritize game traffic (QoS), separate links for bulk transfers, alert on link utilization. Servers/OS: rate-limit backups, log shipping, and deployments, and run them during quiet hours.
On the graph
Hits a ceiling · Link utilization, RTT (ping)
Where to look
Put the utilization of the data center link (uplink) interface (computed from SNMP ifHCInOctets and ifHCOutOctets) and its output drops (ifOutDiscards) on the same timeline as the backup, deployment, and log shipping schedules
Confirmed if
RTT and drops rise across the whole server when link utilization flattens at the bandwidth limit, and those times overlap with bulk transfer jobs
Ruled out if
Per-minute utilization well below the limit but drops still present: “Switch microbursts”
RFC 2863: The Interfaces Group MIBIETF ifHCInOctets and ifHCOutOctets: bytes received and sent on the interface (64-bit); ifOutDiscards: packets discarded without being sent
ID dc-failover · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), Game team (Client development)
When a router or firewall fails and traffic switches to the standby unit (failover), everyone freezes for a few seconds.
Why Switchover to standby equipment because of a failure or maintenance → Effect The switchover takes a few seconds, and connections reset if session state isn’t synced → On screen Every player on the server freezes at once; mass disconnects
Primary owner Infra team (Network infrastructure) · Also Game team (Server development), Game team (Client development)
Game team action items
Server: use timeouts that survive brief outages (a few seconds), let players resume their session with a session token when they reconnect after a disconnect. Client: reconnect automatically on disconnect (randomize retry intervals so everyone doesn’t pile in at once).
Infra team action items
Use redundancy that shares connection state, detect failures within 1 second with BFD, test failover regularly.
Ballpark numbers
About 1–3 seconds if the equipment detects the failure immediately. Without fast failure detection (BFD), relying only on default BGP timers, the route can be down for 90–180 seconds before neighboring equipment notices.
On the graph
Mass disconnect · Connections, total server traffic in/out
Where to look
Router and firewall event logs (VRRP role changes, BFD and BGP sessions going down, failover records) next to total server connections and traffic at the same time
Confirmed if
At the failover time in the device log, traffic for every server behind that device drops to 0 for a few seconds, or connection counts fall together
Ruled out if
Only one server’s connections drop: “Server crash” or “NIC driver and firmware problems.” Device logs clean and the frozen server is a single cloud VM: “Cloud host maintenance and live migration”
ID dc-bad-cable · Primary owner Infra team (Network infrastructure)
Bad optics or a bad cable corrupt a steady share of the packets that pass through that path.
Why Bit errors from bad optics or cables → Effect The equipment silently drops corrupted packets → On screen Only some servers or players using that path teleport or rubber-band from steady packet loss
Monitor and alert on port error (CRC) counters, replace parts such as optics and cables, take the problem link out and route around it until it’s replaced.
On the graph
Outliers only · CRC errors per port, loss rate per server/path
Where to look
Check the CRC counters on both ends of the link. On switches, the port’s FCS errors (dot3StatsFCSErrors) and input errors (ifInErrors); on servers, crc under RX errors in ip -s -s link (kernel stat rx_crc_errors)
Confirmed if
CRC errors on one port keep rising regardless of traffic volume or time of day, and only servers and players passing through that port see loss
Ruled out if
No CRC errors and only output drops rising: congestion (“Switch microbursts,” “Data center link saturation”)
ID dc-mtu · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)
If the MTU (the largest size that can be sent at once) shrinks somewhere along the path and the “packet too big” messages are blocked, only large packets keep vanishing.
Why The MTU shrinks on a tunnel or VPN segment → Effect A firewall blocks the “packet too big” messages (ICMP), so the sender never finds out → On screen Freezes and then disconnects only when opening large screens such as the inventory or character list
During specific actions, Right after login or maintenance
Owner
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)
Game team action items
To lower it directly on the server side, set the socket’s maximum segment size with TCP_MAXSEG (splitting messages into smaller pieces in game code alone won’t prevent it), keep UDP packets at 1,200 bytes or less.
Infra team action items
Network: reduce TCP packet size on tunnel segments (MSS clamping), allow “packet too big” messages (ICMP) through firewalls and cloud network ACLs. Servers/OS: allow “packet too big” messages (ICMP) in server firewalls and cloud security groups too, turn on MTU probing in the server kernel (tcp_mtu_probing=1) as a last safety net that only kicks in after a freeze of a few seconds.
Ballpark numbers
Usually 1,500 bytes, shrinking to around 1,400 through a tunnel.
On the graph
Outliers only · Disconnects by region/ISP, failed large responses
Where to look
From the affected player’s PC, ping the server with the don’t-fragment flag (DF) set, varying the size. On Windows, ping /f /l 1472 SERVER_IP; on Linux, ping -M do -s 1472 SERVER_IP (1,472 is the 1,500 MTU minus the 20-byte IP header and the 8-byte ICMP header). Lower the size step by step to find the largest size that gets through, and check whether the server-side security groups and firewalls allow ICMP “packet too big” messages (Fragmentation Needed)
Confirmed if
Small pings get through but the 1,472-byte DF ping fails (no reply, or an error saying fragmentation is needed), and the largest size that passes is small, around 1,400. Players in the same region freeze only when opening large screens
Ruled out if
The 1,472-byte DF ping also gets through: not a path MTU problem. Even small pings fail: ICMP itself is blocked, so this method can’t tell
Check with
The player’s own environment
Sources (5)
RFC 2923: TCP Problems with Path MTU DiscoveryIETF If a firewall blocks ICMP (Fragmentation Needed), path MTU discovery fails and only large packets keep vanishing (black hole); pings and small messages still work, which makes it hard to diagnose
IP SysctlLinux kernel tcp_mtu_probing=1 is normally off and turns on TCP path MTU probing only when it detects an ICMP black hole
ping(8) — Linux manual pageiputils -M do sets the DF flag and refuses packets larger than the path MTU; -s sets the data size (default 56 bytes, plus the 8-byte ICMP header)
pingMicrosoft /f sets the DF flag and is used to find path MTU problems; /l sets the data size
ID nic-irq · Primary owner Infra team (Server infrastructure)
If the NIC sends every packet-arrival interrupt to a single CPU core, that core becomes the bottleneck.
Why A single receive queue, or RSS (which spreads packets across cores) turned off → Effect One core hits 100% and can’t pull packets off in time → On screen Packet loss and latency across the whole server when players crowd in (teleporting, input lag)
Configure RSS (spreading by the NIC) and RPS (spreading by the kernel), spread interrupts across several cores, make UDP queue selection include ports (rx-flow-hash udp4 sdfn in ethtool -N), keep interrupt-handling cores separate from the game tick thread’s cores, watch %soft per core.
Ballpark numbers
One core can push roughly hundreds of thousands of packets per second through the kernel, depending on packet size and settings. If per-core utilization shows receive processing (%soft in mpstat) piled onto a single core, this is what’s happening.
On the graph
Hits a ceiling · %soft per core, packets received per second
Where to look
Check %soft (share of time spent on software interrupts) per core with mpstat -P ALL 1, which core each NIC queue’s interrupts go to in /proc/interrupts, the number of queues with ethtool -l, and packets per queue with ethtool -S (names vary by driver)
Confirmed if
One core’s %soft sits near 100% while the rest are idle, and interrupts and packets pile into one queue. From then on, packets received per second can’t climb any higher
Ruled out if
%soft spread evenly across cores: not this cause. CPU idle but loss present: “Cloud PPS limit exceeded” or “Ring buffer too small”
Check with
Infra tools (no game code needed)
Learn more
Even with multiple queues, if most traffic comes from a handful of addresses, such as gateways or proxies, it all lands in one queue. For UDP, some NICs pick the queue from addresses only by default, and traffic spreads evenly only after you change that to include ports.
Sources (4)
Scaling in the Linux Networking StackLinux kernel RSS (the NIC spreads packets across multiple receive queues) and RPS (the kernel spreads them), giving each queue its own interrupt and spreading those across cores; RSS is recommended when receive interrupt handling is the bottleneck
How to receive a million packets per secondCloudflare Measurements where a receive queue served by a single core topped out at about 350,000–430,000 packets per second; a case where the NIC hashed UDP by IP address only and everything piled into one queue
ID nic-ring · Primary owner Infra team (Server infrastructure)
If the NIC’s ring buffer, which briefly holds incoming packets, is small, a sudden burst overflows it and packets get dropped.
Why The ring buffer is left at its small default (256–2,048 slots depending on the driver) → Effect During a burst, the buffer overflows before the CPU can pull packets off → On screen Loss only at burst moments (teleporting, skills not going off). No trace in the game server logs
Enlarge the ring buffer (ethtool -G), watch drop counters (such as rx_missed_errors in ethtool -S; names vary by driver).
Ballpark numbers
At 1 million packets per second, 1,024 slots fill in about 1 ms. If the CPU is late even once in that window, the buffer overflows. Most NICs can be raised to several thousand slots.
On the graph
Random spikes · NIC receive drop counters
Where to look
Collect the receive drop counters in ethtool -S (rx_missed_errors, rx_fifo_errors, and so on; names vary by driver) and missed in ip -s -s link at short intervals, and check the current and maximum ring size with ethtool -g
Confirmed if
Drop counters rise at burst moments, and the current ring size is far below the maximum. Enlarging the ring reduces the drops
Ruled out if
Drop counters flat but loss present: the next stage in the kernel (“Kernel socket buffers too small”) or the network path. One core’s %soft at 100%: “NIC interrupts concentrated on one core”
Interface statisticsLinux kernel Packets the device drops for lack of buffers are counted in rx_missed_errors; driver-specific stats are visible with ethtool -S
ethtool(8) — Linux manual pageethtool -g shows ring size (current and maximum), -G changes it, -S shows driver-specific stats
ID nic-coalesce · Primary owner Infra team (Server infrastructure)
When the NIC collects packets and notifies the CPU once per batch to reduce CPU load, packets arrive later by the time spent collecting.
Why The NIC collects packets for a set time or count before raising an interrupt → Effect Packets wait while the batch fills → On screen A small rise in latency. Usually tiny, but ms-scale if overdone
Use adaptive coalescing, tune the values for game servers (ethtool -C).
Ballpark numbers
Typically tens to hundreds of µs. That’s negligible for most games, but aggressive settings can push it into the ms range.
On the graph
Always high · Round-trip time within the same data center
Where to look
Check the current coalescing settings (adaptive-rx, rx-usecs, rx-frames) with ethtool -c, and compare ping round-trip time to another server in the same data center before and after changing them
Confirmed if
rx-usecs is set high (hundreds of µs or more), and lowering it cuts round-trip time within the data center by about the same amount
Ruled out if
Round-trip time unchanged after lowering it: not this cause
ID nic-cloud-pps · Primary owner Infra team (Server infrastructure) · Also Game team (Server development), External (External)
Each cloud instance type has limits on packets per second and bandwidth, and traffic over them is silently dropped.
Why Rising CCU pushes packets per second over the instance limit → Effect The cloud network drops the excess → On screen Teleporting and skills not going off from unexplained packet loss. Server CPU has headroom
Primary owner Infra team (Server infrastructure) · Also Game team (Server development), External (External)
Game team action items
Combine packets (one tick’s messages in one packet), avoid sending tiny packets frequently.
Infra team action items
Check and alert on limit-exceeded counters (pps_allowance_exceeded, conntrack_allowance_exceeded, and so on for AWS), move to a larger instance, avoid connection tracking limits with a security group setup that creates no tracking.
External action items
Ask the cloud provider for packets-per-second and connection tracking limits per instance type.
Ballpark numbers
Limits differ by instance size, and the packets-per-second limit is often not published. The “up to 10 Gbps” on small instances is a burst speed available only while credits last (usually 5–60 minutes); the normal baseline speed is much lower.
On the graph
Hits a ceiling · Packets per second, allowance exceeded counters
Where to look
Collect the ENA counters pps_allowance_exceeded, bw_in_allowance_exceeded, bw_out_allowance_exceeded, and conntrack_allowance_exceeded from ethtool -S at short intervals and view them alongside packets per second. You can also publish these counters with the CloudWatch agent and set alarms on them
Confirmed if
Allowance exceeded counters rise at the times of loss, and packets per second stops climbing at a fixed value. Server CPU has headroom
Ruled out if
Exceeded counters flat: not this cause. One core’s %soft at 100%: “NIC interrupts concentrated on one core”
Check with
Infra tools (no game code needed)
Learn more
conntrack_allowance_exceeded means the connection tracking table was full and new connections were dropped. If the table has room and only idle connections drop because their tracking expired, see the “Cloud security group connection tracking expiry” entry.
Amazon EC2 instance network bandwidthAWS The “up to N Gbps” of instances with 16 vCPUs or fewer is a burst that spends network I/O credits (usually 5–60 minutes), falling back to baseline bandwidth when the credits run out
ID nic-saturate · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Running a 1 Gbps or 10 Gbps card at its limit makes the transmit queue grow until packets get dropped.
Why More broadcasts push traffic to the card’s limit → Effect The transmit queue grows, and packets are dropped when it overflows → On screen Latency and loss across the whole server (input lag, teleporting)
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Reduce traffic (area of interest, compression, send only what changed).
Infra team action items
Upgrade the card (a faster NIC, or a larger instance in the cloud), alert on NIC utilization.
On the graph
Hits a ceiling · NIC transmit volume, transmit drops
Where to look
Compare txkB/s and %ifutil (utilization relative to interface speed) from sar -n DEV 1 with the NIC speed or instance bandwidth, alongside TX dropped from ip -s link
Confirmed if
Transmit volume flattens near the NIC or instance bandwidth, and from then on transmit drops and server-wide latency rise
Ruled out if
Bandwidth has headroom: not this cause. Many small packets and loss: “Cloud PPS limit exceeded”
Check with
Infra tools (no game code needed)
Sources (3)
Interface statisticsLinux kernel tx_dropped: number of packets dropped during transmission for lack of resources
ID nic-noisy · Primary owner Infra team (Server infrastructure) · Also External (External)
When other virtual machines on the same physical server use a lot of network or CPU, your server’s processing gets delayed at irregular times.
Why Other VMs on the same physical server use a lot of resources → Effect Packet processing on your VM is delayed at irregular times → On screen Occasional jitter (variation in packet arrival times) with no obvious cause, showing up as stutter
Primary owner Infra team (Server infrastructure) · Also External (External)
Infra team action items
Use dedicated hosts or instances with guaranteed performance, stop and restart instances with persistent jitter to move them to another host.
External action items
Report the problem host to the cloud provider.
On the graph
Random spikes · Round-trip jitter within the data center, %steal
Where to look
Ping another server in the same data center continuously to record round-trip jitter, and compare it, along with %steal from mpstat, against other instances with the same configuration
Confirmed if
Only this instance shows irregular spikes in round-trip jitter or %steal, while others with the same configuration stay quiet. Stopping and restarting it to move it to another host makes the problem go away
Ruled out if
All instances with the same configuration spike the same way: not a host problem. Look at game server load or the network path
ID nic-host-maintenance · Primary owner Infra team (Server infrastructure) · Also Game team (Server development), External (External)
When a cloud provider performs maintenance on a physical server (host), it moves VMs to another host (live migration) or pauses them briefly. The whole server freezes during that time, and if the pause is long, connections drop.
Why The provider moves the VM to another host, or pauses it briefly, for host maintenance or a predicted failure → Effect During the move, CPU, memory, and network slow down, and at the end the VM stops completely for a moment (from under 1 second to around 30 seconds, depending on the provider and method) → On screen Everyone on the server freezes at once and then sees fast-forward and teleporting; if the freeze outlasts the timeout, mass disconnects
Primary owner Infra team (Server infrastructure) · Also Game team (Server development), External (External)
Game team action items
Use timeouts that survive pauses of a few seconds, cap how many ticks the server catches up after a pause, compute elapsed time with a monotonic clock, have a procedure that saves progress and moves players to another server when a maintenance notice arrives.
Infra team action items
Subscribe to and alert on maintenance notices (Google Cloud maintenance-event, AWS scheduled events and AWS Health, Azure Scheduled Events), replace servers ahead of time during low-traffic hours when a notice arrives, reschedule maintenance where the provider allows it (Azure Maintenance Configuration, AWS scheduled events depending on type), compare maintenance records with incident records.
External action items
Ask the cloud provider about maintenance schedules and impact, report instances that keep pausing.
Ballpark numbers
Google Compute Engine says live migration pauses are usually much shorter than 1 second, and the system clock can jump forward by up to 5 seconds during the pause. The maintenance-event metadata value changes 60 seconds before the move (if you have queried it at least once beforehand). On Azure, maintenance that doesn’t need a reboot almost always pauses the VM for under 10 seconds, and rarely (no more than once every 18 months for general-purpose sizes) for about 30 seconds; live migration usually takes no more than 5 seconds. Azure Scheduled Events gives notice of these pauses (Freeze) at least 15 minutes ahead. If host hardware fails suddenly, though, recovery starts right away with no notice.
On the graph
Gap then burst · Server packets sent/received, tick interval
Where to look
Match the freeze time against the provider’s records. Google Cloud: compute.instances.migrateOnHostMaintenance in the audit logs; AWS: scheduled events in describe-instance-status and AWS Health; Azure: Microsoft.Compute/virtualMachines/liveMigration/action in the Activity Log and the time the VM availability metric (VmAvailabilityMetric) dropped to 0. Inside the server, check whether metrics and logs have a gap during the freeze and whether the clock jumped right after (time sync logs)
Confirmed if
The time the whole server froze overlaps with a maintenance or migration time in the provider’s records, and every metric and log inside the server is blank for those few seconds
Ruled out if
Not in the provider’s records and short freezes recur often: “CPU steal (virtual machines).” NIC reset entries in the kernel log: “NIC driver and firmware problems”
Check with
Infra tools (no game code needed)
Learn more
AWS gives notice through scheduled events. system-reboot means the instance will be rebooted and moved to a new host; system-maintenance means network or power maintenance may affect it briefly. Even if the pause lasts only a few seconds, clients that got no ACK for packets sent to the server during that time keep doubling their retransmission wait, so TCP connections can stay stalled for longer after the pause ends (“TCP RTO and exponential backoff”). When the VM wakes up, its clock can jump and lead to a “System clock jump (NTP step),” and failed load balancer health checks may take the server out of rotation for a while. Instances that can’t be moved (such as Google Cloud bare metal instances) are stopped or restarted during maintenance.
Sources (6)
Live migration process during maintenance eventsGoogle Cloud Live migration pauses are usually much shorter than 1 second; the system clock jumps forward by up to 5 seconds during the pause; disk, CPU, memory, and network performance drop briefly during the move; VMs that don’t live-migrate are terminated for maintenance (bare metal instances don’t support live migration)
Query metadata server for maintenance event noticesGoogle Cloud The maintenance-event metadata value changes 60 seconds before live migration (when the VM is set to live-migrate and the value was queried at least once since the last maintenance)
Scheduled events for Amazon EC2 instancesAWS Scheduled event types (system-reboot reboots and moves to a new host; system-maintenance means brief impact from network or power maintenance), notified by email and AWS Health, checked with describe-instance-status, reschedulable for some types
Maintenance and updatesMicrosoft Azure Maintenance without a reboot almost always pauses for under 10 seconds, rarely (no more than once every 18 months for general-purpose sizes) for about 30 seconds, and live migration usually 5 seconds or less; the clock syncs automatically after the pause; long-lived TCP connections may drop, or recovery may take longer as peers retransmit data sent to the paused VM with exponential backoff; load balancer health checks mark the VM unhealthy within about 10 seconds; confirm with Microsoft.Compute/virtualMachines/liveMigration/action in the Activity Log and VmAvailabilityMetric dropping to 0 during the pause; pick when maintenance applies with Maintenance Configuration
Scheduled Events for Linux VMs in AzureMicrosoft Azure Freeze (a pause of a few seconds; CPU and network may stop) is announced at least 15 minutes ahead; for host hardware failures, recovery starts right away with no notice period
ID nic-reset · Primary owner Infra team (Server infrastructure)
When a driver bug or a malfunctioning feature hangs the card, all traffic in and out stops while it restarts.
Why Driver bug, malfunctioning offload feature → Effect The NIC hangs and restarts (a few seconds) → On screen Everyone on that server freezes together and then teleports or disconnects
Check and alert on “transmit queue … timed out” and “Link is Down” entries in the kernel log, update drivers and firmware, turn off the problem feature (such as an offload).
On the graph
Gap then burst · Server packets sent/received
Where to look
Search the kernel log with dmesg for “NETDEV WATCHDOG … transmit queue N timed out”, driver resets, and “Link is Down” or “Link is Up” entries, and check the server’s packets sent and received at those times
Confirmed if
At the time of the freeze, the kernel log has a transmit queue timeout or link down/up entries, and packets sent and received drop to 0 for those few seconds
Ruled out if
Kernel log clean and the switch-side port fine: a stall in the game server process (“Server GC stop-the-world pause,” “Deadlock”) or “Network equipment failover”
Check with
Infra tools (no game code needed)
Sources (3)
net/sched/sch_generic.c (Linux v6.12)Linux kernel When a transmit queue stalls, the kernel watchdog logs “NETDEV WATCHDOG … transmit queue N timed out” and calls the driver’s reset function
ID nic-offload · Primary owner Infra team (Server infrastructure)
GRO and LRO bundle several packets into one to reduce CPU load. Depending on settings, a small game packet may wait briefly for the next packet to bundle with.
Why The NIC and kernel bundle arriving packets together for processing → Effect With hardware aggregation (LRO) or a batching wait time setting turned on, packets wait briefly for the next one → On screen A small rise in latency (usually tens of µs or less)
Tune for game traffic (turn off LRO, check the batching wait time setting), check it after other causes since the effect is usually small.
On the graph
Always high · Round-trip time within the same data center
Where to look
Check lro and gro status with ethtool -k and the device’s gro_flush_timeout sysfs setting, and compare round-trip time for small packets within the data center before and after changing them
Confirmed if
LRO is on or gro_flush_timeout is above 0, and turning it off or setting it to 0 reduces round-trip time for small packets
Ruled out if
Difference stays within a few µs after the change: not this cause
Check with
Infra tools (no game code needed)
Sources (3)
NAPILinux kernel A large gro_flush_timeout batches more work together but adds latency under low load
ID so-backlog · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Game team (Client development)
When tens of thousands of players connect at once right after maintenance, the kernel’s connection queue (listen backlog) overflows and connection attempts are dropped.
Why As maintenance ends, connections pour in faster than the game server can accept them → Effect The kernel’s connection queue (listen backlog: the smaller of the value the server code passes to listen and the kernel cap) fills up → On screen Connection attempts are dropped and retried again and again: can’t connect / infinite loading
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Game team (Client development)
Game team action items
Server: raise the listen value passed in code, keep the thread that accepts connections from stalling on other work, add a login queue system. Client: lengthen the retry interval (randomized to spread retries out).
Infra team action items
Raise the kernel’s somaxconn (it only helps when the listen value in the server code goes up too), keep SYN cookies on, monitor the overflow count (TcpExtListenOverflows in nstat).
Ballpark numbers
The Linux kernel cap (somaxconn) defaults to 4,096 since 5.4 (128 before that), but if the server code passes a smaller value to listen, that value is the limit. When the queue is full, Linux silently drops connection requests without returning an error. The client OS resends a few times, starting after 1 second, so players just see a long loading screen with no “Connection failed” message. A Windows server sends back a refusal, so the client sees “Connection failed” right away.
On the graph
Surge after opening · Connection queue overflows (ListenOverflows), connection attempts
Where to look
Increase in TcpExtListenOverflows and TcpExtListenDrops from nstat -az; with ss -ltn, the listening socket’s Recv-Q (connections waiting for accept) compared with its Send-Q (the backlog limit)
Confirmed if
ListenOverflows rises when connections pile in, and the listening socket’s Recv-Q sits at its Send-Q value
Ruled out if
ListenOverflows unchanged: not this cause. Connection established but loading never finishes: “Login storm and N+1 queries.” Blocked at an exact player count: “File descriptor limit”
Check with
Infra tools (no game code needed)
Sources (5)
listen(2) — Linux manual pageLinux man-pages A listen backlog larger than somaxconn is silently truncated; somaxconn defaults to 4,096 (since 5.4; 128 before); when the queue is full, requests may be ignored and left to client retries
IP SysctlLinux kernel tcp_syn_retries: the connection request (SYN) is resent several times, with a 1 s wait before the first retransmission; tcp_abort_on_overflow off by default (no refusal sent on overflow); tcp_syncookies on by default
listen function (winsock2.h)Microsoft On Windows, when the queue is full the client gets a WSAECONNREFUSED error
SNMP counterLinux kernel TcpExtListenOverflows: number of connection requests (SYN) dropped because the accept queue was full; TcpExtListenDrops goes up along with it
net/ipv4/tcp_diag.c (Linux v6.12)Linux kernel Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed
ID so-fd · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Every connection needs a file descriptor (fd: the number the OS gives an open file or socket), and the number of fds one process can open is capped.
Why Concurrent users reach the process’s file descriptor limit → Effect The server can’t accept new connections (Too many open files). Opening log files and DB connections fails too → On screen Past an exact player count, nobody gets in: can’t connect / infinite loading
Right after login or maintenance, When crowds gather
Owner
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Always close the socket when a connection ends (to prevent fd leaks), when accept fails with EMFILE (out of fds), pause accepting briefly or accept with a spare fd kept in reserve and close it right away (so the server doesn’t burn CPU handling the same new-connection notification over and over).
Infra team action items
Check ulimit and the service settings (LimitNOFILE in systemd), alert when usage nears the limit.
Ballpark numbers
On Linux, the limit is still often 1,024 unless the service is configured otherwise. Game servers usually raise it to tens of thousands or hundreds of thousands. Windows has no default limit this low.
On the graph
Hits a ceiling · Open fds of the process, concurrent users
Where to look
fd-nr (open file descriptors) of the game server process from pidstat -v, the open-files limit in /proc/PID/limits, and accept failures (EMFILE, Too many open files) in the server log
Confirmed if
The fd count flattens at the limit, and from that moment accept fails with EMFILE
Ruled out if
fd count well below the limit: not this cause. Connection requests dropped in the kernel: “Connection queue (listen backlog) overflow.” Connection tracking involved: “Server conntrack table full”
Check with
Infra tools (no game code needed)
Learn more
Connections that weren’t accepted stay in the kernel’s connection queue (listen backlog), so depending on the code, the server may keep getting “new connection” notifications and waste CPU on them.
ID so-sockbuf · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
With small send and receive buffers, a burst of traffic makes the kernel drop packets arriving over UDP, and TCP sends block because the buffer has no room left.
Why SO_SNDBUF and SO_RCVBUF left at their defaults or set too small → Effect During a burst, or while the receiving thread pauses briefly, the UDP receive buffer overflows and drops packets; TCP waits because the send buffer has no room → On screen Teleporting (UDP loss) or fast-forward (TCP waiting)
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Set buffer sizes (SO_SNDBUF, SO_RCVBUF) in code to match the traffic, remember that setting a TCP buffer size explicitly turns off Linux autotuning, keep sizes moderate because oversized buffers let stale data pile up and add delay, keep the receiving thread from stalling.
Infra team action items
Tune the kernel caps (rmem_max, wmem_max; buffer sizes set in code can’t exceed them either) and the default (rmem_default), monitor the buffer overflow counter (RcvbufErrors).
Ballpark numbers
The default UDP receive buffer on Linux is about 208 KB. Even a small packet takes up far more kernel memory than its actual size, so tens to hundreds of packets fill it. On a server receiving 100,000 packets per second, it overflows if the receiving thread stalls for just a few ms.
On the graph
Random spikes · UDP receive buffer overflows (UdpRcvbufErrors)
Where to look
Increase in UdpRcvbufErrors from nstat -az, and skmem in ss -uamn (rb is the receive buffer size, d is packets dropped because they couldn’t be queued to the socket); for TCP, whether send-queue memory (w) in the skmem of ss -tm has reached the send buffer size (tb)
Confirmed if
UdpRcvbufErrors (or the socket’s d) rises during bursts or when the receiving thread stalls, with rb near the default (about 208 KB). For TCP, w stays pinned at tb and send blocks
Ruled out if
Counters unchanged but loss present: the NIC stage (“Ring buffer too small”) or the network path
Check with
Infra tools (no game code needed)
Sources (6)
socket(7) — Linux manual pageLinux man-pages SO_RCVBUF and SO_SNDBUF default to rmem_default and wmem_default and are capped at rmem_max and wmem_max; the kernel doubles the value you set
include/net/sock.h (Linux v6.18)Linux kernel The default socket buffer is defined as 256 packets of 256 bytes including sk_buff overhead (SKB_TRUESIZE(256)×256); even a small frame counts as sk_buff + MTU (about 208 KB is the computed value on x86-64)
IP SysctlLinux kernel tcp_rmem, tcp_wmem: setting SO_RCVBUF or SO_SNDBUF directly turns off autotuning for that socket
net/ipv4/udp.c (Linux v6.12)Linux kernel When the UDP receive queue exceeds the socket buffer size, packets are dropped immediately and RcvbufErrors goes up
net/ipv4/proc.c (Linux v6.12)Linux kernel Counter names shown by nstat: RcvbufErrors and SndbufErrors in the Udp group
ss(8) — Linux manual pageiproute2 skmem in -m: rb receive buffer size, tb send buffer size, w send-queue memory, d packets dropped before reaching the socket
ID so-context · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Running far more threads than there are cores makes the OS spend CPU just switching between them.
Why Hundreds to thousands of threads, for example one thread per connection → Effect Higher context-switching cost (swapping out the running thread) and more cache misses → On screen CPU is busy but throughput is low and ticks are uneven: stutter, slow motion
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Match thread count to core count, use asynchronous I/O (epoll, IOCP).
Infra team action items
Monitor context switches and runnable threads (cs and r in vmstat).
Ballpark numbers
One context switch costs a few µs, and more once you add the cache misses that follow.
On the graph
Rises with load · Context switches per second, runnable threads
Where to look
cs (context switches per second) and r (running or waiting for CPU) from vmstat 1 compared with the core count; voluntary (cswch/s) and involuntary (nvcswch/s) context switches per game server thread from pidstat -w -t
Confirmed if
As concurrent users grow, r climbs far above the core count, cs spikes with it, and hundreds of threads show many involuntary context switches
Ruled out if
r stays at or below the core count: not this cause. Mostly voluntary switches: threads are waiting on locks or I/O (“Lock contention,” “Blocking I/O design”)
vmstat(8) — Linux manual pageprocps-ng The cs (context switches per second) and r (processes running or waiting to run) fields
I/O Completion PortsMicrosoft Handle many asynchronous I/Os with a pre-created thread pool and IOCP, and match the number of concurrently running threads to CPU concurrency
pidstat(1) — Linux manual pagesysstat In -w, cswch/s counts voluntary context switches (the task stopped on its own to wait for a resource) and nvcswch/s counts involuntary ones (forced out after using up its time slice); -t shows them per thread
ID so-steal · Primary owner Infra team (Server infrastructure) · Also External (External)
While the physical server (hypervisor) briefly gives a virtual machine’s CPU time to another VM (CPU steal), the game server stalls.
Why Other VMs on the same host use a lot of CPU → Effect The game server’s VM loses its turn on the CPU for a few ms to tens of ms at a time → On screen Unexplained tick-time spikes: stutter, freeze
Primary owner Infra team (Server infrastructure) · Also External (External)
Infra team action items
Monitor steal (st in top and vmstat), use dedicated cores or hosts, avoid burstable instances that slow down when CPU credits run out, stop and start instances with persistently high steal to move them to another host.
External action items
Report hosts with persistently high steal to the cloud provider.
On the graph
Random spikes · %steal, server tick time
Where to look
%steal from mpstat -P ALL 1 on the same time axis as server tick time
Confirmed if
%steal spikes whenever the tick spikes, and drops after a stop and start moves the instance to another host
Ruled out if
%steal near 0 while ticks still spike: a cause inside the game server (“Server GC stop-the-world pause,” “Lock contention”). In a container: “Container CPU throttling (CFS quota)”
ID so-cpu-quota · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
With a CPU limit on a container, the moment it uses up its quota within a set period (usually 100 ms), it is forced to stop for the rest of that period (throttling).
Why A CPU limit is set on the game server container (in Kubernetes, for example) → Effect A burst of tick work uses up the quota, and the server stops for tens of ms until the next period → On screen Average CPU is low, yet ticks spike periodically: stutter, slow motion
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Match the worker thread count to the CPU limit (so the runtime doesn’t create as many threads as the host has cores).
Infra team action items
Set a generous CPU limit or remove it and assign dedicated cores, monitor the throttle count (nr_throttled).
Ballpark numbers
On a server limited to 2 cores, 8 threads working at once use up the quota for a 100 ms period in 25 ms and then stop for 75 ms.
On the graph
Rises with load · Throttle count (nr_throttled), server tick time
Where to look
Increase in nr_throttled and throttled_usec (nr_throttled and throttled_time on cgroup v1) in the container cgroup’s cpu.stat, alongside server tick time
Confirmed if
Average CPU utilization stays below the limit, yet nr_throttled and throttled_usec keep rising, lining up with tick spikes
Ruled out if
nr_throttled not rising: not this cause. The VM itself being held back: “CPU steal (virtual machines)”
Check with
Infra tools (no game code needed)
Sources (3)
CFS Bandwidth ControlLinux kernel Threads that use up the quota for a period stop until the next period (throttling), default period 100 ms, nr_throttled statistic
Control Group v2Linux kernel cpu.max takes the form “$MAX $PERIOD” (quota, period), and the default is “max 100000” (100 ms period)
ID so-cstate · Primary owner Infra team (Server infrastructure)
Idle CPU cores drop into deep power-saving states (C-states) and lower their frequency to save power. Waking up and raising the frequency when a packet or timer arrives takes time, which adds delay to handling small packets.
Why The OS frequency scaling policy (governor) or the BIOS power settings allow deep C-states and low frequencies → Effect An idle core is up to hundreds of µs late every time it wakes from a deep power-saving state, and a frequency pinned low slows the tick computation itself → On screen Usually hard to notice, but with many server-to-server calls it adds up to input lag that gets worse when the server is quiet. With the frequency pinned low, ticks fall behind when crowds gather: slow motion
Set the BIOS power settings to performance, set the OS governor to performance (scaling_governor in cpufreq), limit deep C-states on latency-sensitive servers (tuned latency-performance profile, /dev/cpu_dma_latency in PM QoS, the intel_idle.max_cstate kernel parameter), and after the change compare round-trip time within the data center, tick-time jitter, and power usage.
Ballpark numbers
According to the intel_idle driver tables in Linux 6.12, Intel server CPUs take 1–2 µs to wake from the shallow C1 state and 133 µs (Skylake-SP) to 290 µs (Sapphire Rapids) from the deep C6 state. One wakeup is small, but when a request passes through several servers, the delays add up. The kernel picks deeper states the longer it expects to stay idle, so this shows up more on quiet servers where packets arrive only now and then. The generic cpufreq powersave governor pins the frequency at the lowest allowed value (the intel_pstate algorithm of the same name adjusts it to the load).
On the graph
Always high · Round-trip time within the same data center, core frequency
Where to look
Per-core C-state residency and actual frequency from cpupower monitor; for each state under /sys/devices/system/cpu/cpu0/cpuidle/, its name, latency (µs to wake up), and usage; scaling_governor in cpufreq; and the current profile from tuned-adm active
Confirmed if
Idle cores sit in the deepest C-state for long stretches or the frequency is pinned near the minimum, and switching to the performance governor and shallow C-states reduces the round-trip time and jitter of small requests
Ruled out if
Difference after the change stays within tens of µs: safe to ignore this cause. Spikes in the ms range: “CPU steal (virtual machines)” or another layer
Check with
Infra tools (no game code needed)
Learn more
For bare-metal servers in a data center, check the BIOS (firmware) power settings together with the OS settings. In the cloud, only some instance types let the OS change C-states and frequency, and AWS defaults to maximum performance, so most instances can be left as they are. The tuned latency-performance profile on Red Hat-based systems sets the governor to performance and uses PM QoS to allow only shallow C-states. Turning off power saving raises power consumption, so apply it only to latency-sensitive servers.
Sources (7)
CPU Idle Time ManagementLinux kernel Each power-saving state has a wakeup time (exit latency) and a minimum stay (target residency), and deeper states are chosen based on expected idle time; per-state latency, usage, and time in sysfs; PM QoS (/dev/cpu_dma_latency) and intel_idle.max_cstate limit deep states
drivers/idle/intel_idle.c (Linux v6.12)Linux kernel Wakeup times of C-states on Intel server CPUs: Skylake-SP C1 2 µs, C1E 10 µs, C6 133 µs; Ice Lake C6 170 µs; Sapphire Rapids C1 1 µs, C6 290 µs
CPU Performance ScalingLinux kernel Check and change the governor with scaling_governor; performance requests the highest allowed frequency, powersave the lowest
intel_pstate CPU Performance Scaling DriverLinux kernel The intel_pstate powersave algorithm differs from the generic powersave governor and scales with load (similar to schedutil and ondemand)
Chapter 2. Getting started with TuneDRed Hat The latency-performance profile turns off power-saving features, sets the governor to performance, and uses PM QoS to allow only shallow C-states; check the current profile with tuned-adm active
Processor state control for Amazon EC2 Linux instancesAWS Only some instance types let the OS control C-states and P-states, which can be changed to reduce latency; the default settings give maximum performance and suit most workloads; Graviton runs at a fixed frequency, so the OS doesn’t control it
ID so-oom · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
When memory runs out, Linux picks the process using the most memory and kills it. Usually that’s the game server.
Why Memory runs out from a leak or a surge in usage, or the container hits its memory limit → Effect The kernel kills the game server process → On screen Everyone on that server disconnects at once, and recent progress may be rolled back
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Fix leaks, set a memory usage ceiling with a procedure that saves and shuts down cleanly as usage approaches it.
Infra team action items
Set memory alerts, size the container memory limit to actual usage, adjust which process gets killed first (oom_score_adj).
Ballpark numbers
The kernel log (dmesg) records “Out of memory: Killed process”, and Kubernetes shows OOMKilled. Windows has no OOM killer; there the server usually dies with an error when a memory allocation fails.
On the graph
Mass disconnect · Connection count, memory usage
Where to look
“Out of memory: Killed process” entries in dmesg, OOMKilled in the pod status on Kubernetes, or an increase in oom_kill in memory.events on cgroup v2, lined up with the time connections dropped
Confirmed if
At the moment connections dropped all at once, a record shows the game server process being killed, and memory usage had been climbing to the limit right before
Ruled out if
No OOM record but the process died: check the crash log and core dump (see “Server crash”)
Check with
Infra tools (no game code needed)
Sources (4)
mm/oom_kill.c (Linux v6.12)Linux kernel The process using the most memory gets the highest score (oom_score_adj factored in); the kernel logs “Out of memory: Killed process …” when it kills
Pushing the Limits of Windows: Virtual MemoryMicrosoft On Windows, once the commit limit is reached, allocations that commit memory fail, which can lead to app errors or system failures
Control Group v2Linux kernel oom_kill in memory.events: number of processes in this cgroup killed by the OOM killer
ID so-reclaim · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
The process stalls while the OS compacts memory to build huge pages or reclaims memory to free it up.
Why Free memory runs low, or transparent huge pages (THP) trigger memory compaction → Effect The thread that asked for memory waits until reclaim or compaction finishes → On screen Irregular server stalls (a few ms to hundreds of ms)
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Cut down large memory allocations at runtime (allocate up front at startup and reuse).
Infra team action items
Configure huge pages (THP) so only regions that ask for them use them (madvise), raise the free-memory threshold (vm.min_free_kbytes and so on).
On the graph
Random spikes · Server tick time, memory PSI
Where to look
some and full in /proc/pressure/memory (share of time stalled waiting for memory) and the increase in compact_stall from /proc/vmstat, alongside server tick time; the /sys/kernel/mm/transparent_hugepage/defrag setting
Confirmed if
Memory PSI rises and compact_stall increases when ticks spike. defrag is set to always
Ruled out if
PSI and compact_stall unchanged: not this cause. Swap usage rising: “Swap”
Check with
Infra tools (no game code needed)
Sources (3)
Transparent Hugepage SupportLinux kernel With defrag=always, a failed THP allocation reclaims and compacts memory on the spot and stalls; with madvise, only regions that asked for it do
Documentation for /proc/sys/vm/Linux kernel min_free_kbytes: the minimum free memory (watermark) the kernel keeps in reserve
PSI - Pressure Stall InformationLinux kernel some (share of time some tasks stalled waiting for memory) and full (share of time all tasks stalled) in /proc/pressure/memory
ID so-timejump · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
When the server clock is moved forward or back by several seconds in one step, timers that depend on the system clock fire all at once or stop.
Why Time sync moves the clock by a large amount in one step → Effect Timers fire in a batch or stop, and timeouts are misjudged → On screen Buff and cooldown glitches, mass disconnects, fast-forward
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Compute elapsed time, timeouts, and cooldowns with a monotonic clock that never jumps or goes backward, use the wall clock only for display and logging.
Infra team action items
Adjust the clock gradually (chrony’s makestep only right after startup), monitor time sync status (clock offset).
Ballpark numbers
ntpd steps the clock in one go when the offset exceeds 0.128 s; below that, it slews gradually at a rate that takes a little over 30 minutes to remove a 1-second offset. With the recommended setting (makestep), chrony, now widely used, steps the clock only a few times right after startup, then slews gradually. The clock also jumps when a VM pauses briefly and resumes.
On the graph
Random spikes · Timer firings and disconnects, clock adjustment log
Where to look
Entries in the time sync service’s log where the clock was stepped by a large amount, lined up with when the problems occurred. With chrony, any adjustment larger than logchange (default 1 s) is written to syslog
Confirmed if
Clock adjustments appear in the log at the times of buff and cooldown glitches, mass disconnects, and fast-forward, and the size of the adjustment matches the size of the glitch
Ruled out if
No clock adjustments logged: not this cause. On a VM, also check for a pause and resume (“Cloud host maintenance and live migration”)
Check with
Infra tools (no game code needed)
Sources (4)
ntpd - Network Time Protocol (NTP) daemonNetwork Time Foundation Offsets above the 128 ms step threshold are corrected in one step and smaller ones gradually, at 0.5 ms per second, so correcting 1 second takes 2,000 s (about 33 minutes)
chrony – Frequently Asked Questionschrony Recommended: allow steps only a few times right after startup, as in makestep 1 3; a VM that was paused and resumed can wake up with the wrong time
clock_gettime(2) — Linux manual pageLinux man-pages CLOCK_MONOTONIC is not affected by discontinuous jumps in the system clock and never goes backward
chrony.conf(5)chrony logchange: clock adjustments larger than this value (default 1 s) are written to syslog
ID so-cron · Primary owner Infra team (Server infrastructure)
Log compression, backups, and security scans that run at the same time every day take up CPU and disk.
Why An OS job runs at a scheduled time → Effect It shares CPU and disk with the game server → On screen Stutter and slow motion at a fixed time, such as 4 a.m. every day
Stagger job times, lower their priority (nice, ionice), move them off the game server (run them on a separate server).
On the graph
Periodic spikes · CPU utilization, disk queue, server tick time
Where to look
Run times of scheduled jobs collected from crontab and systemctl list-timers, and which processes use CPU and disk when ticks spike, from pidstat -u -d
Confirmed if
Ticks spike at the same time every day (or every hour), and at that time a scheduled job process takes up CPU and disk
Ruled out if
Spikes don’t line up with a fixed time of day: not this cause. Spikes every few seconds or minutes: “Server GC stop-the-world pause,” “Timers firing all at once”
Check with
Infra tools (no game code needed)
Sources (4)
ionice(1) — Linux manual pageutil-linux Jobs in the idle class get disk I/O only when no other program is using the disk
ID so-os-update · Primary owner Infra team (Server infrastructure)
The game code hasn’t changed, but the server has been slower since an OS, kernel, driver, or firmware update. Updates can change defaults, the scheduler, CPU vulnerability mitigations, and driver behavior.
Why A routine security patch or a new server image changes the kernel, drivers, or firmware → Effect Changed defaults or scheduler, or newly enabled vulnerability mitigations, make the same work take more CPU time and change the order in which threads get the CPU → On screen A server that ran fine is a little slower all the time from the day of the update: input lag, plus stutter and slow motion when crowds gather
Apply updates to a few servers first and compare tick time, latency, and CPU utilization against the previous version before rolling out wider, deploy on a different day from game patches, record kernel, driver, and firmware versions and key sysctl values before and after the update, boot into the previous kernel to confirm when problems appear, weigh the security risk before turning mitigations off (mitigations=off).
Ballpark numbers
A new kernel version brings new default behavior. For example, Linux began moving its scheduler from CFS to EEVDF in 6.6, and the default connection queue cap (somaxconn) changed from 128 to 4,096 in 5.4. CPU vulnerability mitigations add work, such as flushing internal CPU buffers when returning from the kernel to a program (at the end of every system call) and on context switches and VM transitions, so network servers that make a system call per packet are hit harder. Fully blocking some vulnerabilities requires turning off SMT (the feature that runs one core as two threads), and turning off SMT can cut performance sharply depending on the workload. The kernel parameter mitigations=off turns all these mitigations off and recovers the performance, but leaves the system exposed to the vulnerabilities.
On the graph
Step change · Server tick time, CPU utilization, latency under the same load
Where to look
Update history from the package manager and reboot times, the kernel version from uname -r, and NIC driver info from ethtool -i, lined up with when latency rose. Updated and non-updated servers compared under the same load with mpstat and pidstat, along with the mitigation status in /sys/devices/system/cpu/vulnerabilities/
Confirmed if
Latency and CPU utilization step up from the reboot after the update and stay there, and under the same load only the updated servers run high. Booting into the previous kernel or driver brings them back
Ruled out if
Updated and non-updated servers are equally slow under the same load: not this cause. A game patch went out the same day and packets per player or packet size changed: “Patch changes the traffic pattern”
Check with
Infra tools (no game code needed)
Learn more
Check mitigation status in the files under /sys/devices/system/cpu/vulnerabilities/. The default (mitigations=auto) mitigates with SMT left on, but auto,nosmt turns SMT off on vulnerable CPUs, so the logical core count can drop by half after a kernel upgrade. Updating the OS on the same day as a game patch makes it hard to tell which one caused a problem, so deploy them separately.
Sources (6)
The kernel’s command-line parametersLinux kernel mitigations=: off disables all CPU vulnerability mitigations for more performance but leaves the system exposed, the default auto mitigates with SMT on, auto,nosmt turns SMT off when needed
MDS - Microarchitectural Data SamplingLinux kernel Mitigations flush CPU buffers when returning from the kernel to user space and when entering a VM; the files under /sys/devices/system/cpu/vulnerabilities/ show vulnerability and mitigation status; many CPUs need SMT off for full protection, and turning SMT off can have a large performance impact depending on the workload
Spectre Side ChannelsLinux kernel As a mitigation, branch prediction buffers are flushed on context switches and VM transitions, and stronger mitigations add overhead to every program
EEVDF SchedulerLinux kernel Linux began moving from CFS to the EEVDF scheduler in 6.6
ID so-conntrack · Primary owner Infra team (Server infrastructure) · Also Game team (Server development), Game team (Client development)
When the connection tracking (conntrack) table, where the Linux firewall records every connection, reaches its limit, new packets are dropped.
Why Connection surges and repeated short-lived connections pile up connection entries → Effect The table fills up, and new connections and some packets are dropped → On screen Can’t connect, and teleporting from unexplained packet loss
Right after login or maintenance, When crowds gather
Owner
Primary owner Infra team (Server infrastructure) · Also Game team (Server development), Game team (Client development)
Game team action items
Server: cut down short-lived connections (reuse connections for server-to-server calls). Client: when a connection fails or drops, retry with growing, randomized intervals.
Infra team action items
Raise the table size (nf_conntrack_max), exclude game ports from tracking (NOTRACK in the raw table), alert on usage.
Ballpark numbers
The default limit is about 60,000 to 260,000 entries depending on server memory. When it overflows, the kernel log shows “nf_conntrack: table full, dropping packet”.
On the graph
Hits a ceiling · conntrack entry count (nf_conntrack_count)
Where to look
net.netfilter.nf_conntrack_count (current entries) from sysctl on the same graph as nf_conntrack_max, and “nf_conntrack: table full, dropping packet” in dmesg
Confirmed if
nf_conntrack_count flattens at max, and from that moment the kernel log shows table full
Ruled out if
Entry count well below max: not this cause. For the AWS instance’s own connection tracking limit, check conntrack_allowance_exceeded (see “Cloud PPS limit exceeded”)
Check with
Infra tools (no game code needed)
Sources (3)
Netfilter Conntrack Sysfs variablesLinux kernel nf_conntrack_max defaults to the number of hash buckets (nf_conntrack_buckets), which is set by memory size
net/netfilter/nf_conntrack_core.c (Linux v6.12)Linux kernel Default size is 65,536 with more than 1 GB of memory and 262,144 with more than 4 GB (64-bit); when full, it logs “nf_conntrack: table full, dropping packet” and drops
ID so-ports · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
When a game server opens and closes short connections to the DB or other servers very often, closed connections hold their ports for a while, and new connections can’t be opened.
Why A new connection is opened and closed for every request → Effect The side that closes first holds the port for about 60 seconds on Linux (TIME_WAIT), and the pool of usable ports runs dry → On screen Internal requests fail: failed saves, broken features
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Reuse connections (connection pool), stop opening and closing a new connection for every request.
Infra team action items
Widen the port range (ip_local_port_range), consider TIME_WAIT reuse for outgoing connections (tcp_tw_reuse on Linux), monitor the TIME_WAIT count.
Ballpark numbers
The default Linux port range (32768–60999) holds about 28,000 ports. More than 470 new connections per second to the same destination address use them up. Windows has about 16,000 ports by default (49152–65535) and a longer TIME_WAIT, so it runs out even faster.
On the graph
Hits a ceiling · TIME_WAIT sockets, internal connection failures
Where to look
TIME_WAIT sockets counted per destination address with ss -tan state time-wait, and connect failures (EADDRNOTAVAIL) in the game server log
Confirmed if
TIME_WAIT sockets to the same destination (the DB, for example) flatten near the size of the ephemeral port range (about 28,000 by default), and connect fails with EADDRNOTAVAIL
Ruled out if
Few TIME_WAIT sockets, but only outbound connections to the outside fail: “Cloud NAT gateway connection and port limits”
Check with
Infra tools (no game code needed)
Learn more
On Linux, the 60-second TIME_WAIT is hard-coded in the kernel. Lowering tcp_fin_timeout, which has a similar name, doesn’t shorten TIME_WAIT.
Sources (5)
IP SysctlLinux kernel ip_local_port_range default 32768–60999; tcp_tw_reuse; tcp_fin_timeout is how long the FIN_WAIT_2 state is kept
TCP/IP port exhaustion troubleshootingMicrosoft Windows dynamic ports default to 49152–65535, and a closed connection holds its port in TIME_WAIT for 4 minutes by default
ID sk-hol · Primary owner Game team (Server development) · Also Game team (Client development)
To keep data in order, TCP holds back every packet that arrived after a lost one until the lost packet is received again.
Why One packet goes missing → Effect The packets behind it have arrived but wait in the receive buffer → On screen Everything stops, then releases all at once: fast-forward
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: send real-time positions over UDP, use reliable delivery only for what truly needs it, split traffic into multiple streams. Client: change network handling to match the server (UDP, separate channels).
Ballpark numbers
Losing one packet stalls things for at least one round-trip time plus a little more; if the retransmission is lost too, the stall lasts hundreds of ms to several seconds.
On the graph
Gap then burst · Bytes received per connection, retransmissions
Where to look
Retransmitted packets and the gaps around them on that player’s connection in a server-side packet capture (tcpdump, Wireshark); for the whole server, the increase in TcpRetransSegs from nstat -az
Confirmed if
Each stall starts with the retransmission of a single packet, and right after the retransmitted packet arrives, the backlog is processed all at once (bytes received sit at 0, then surge)
Ruled out if
Games that talk over UDP: not applicable. Stalls with no retransmissions: look at the server tick (“Tick overrun”)
ID sk-rto · Primary owner Game team (Server development) · Also Game team (Client development)
Each time a retransmission fails again, the wait doubles, so a brief connection drop turns into a long stall.
Why The connection drops briefly, and retransmissions fail one after another → Effect The wait before the next attempt doubles each time: 0.3 → 0.6 → 1.2 → 2.4 s (at 100 ms ping) → On screen The connection was down for 1 second, but the game stalls for over 2 seconds. A longer drop eventually ends in a disconnect
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: answer heartbeats and close the connection yourself if nothing arrives for a set time (bring the give-up point forward with TCP_USER_TIMEOUT), resume sessions with a session token, use reliable UDP. Client: send heartbeats at short intervals and reconnect quickly when responses stop, without waiting on TCP retransmission.
Ballpark numbers
The Linux RTO (retransmission timeout) has a floor of “ping + 200 ms”, and starts at 1 second while a connection is being set up. With the default setting (tcp_retries2=15), the connection is abandoned only after about 15 minutes of failed retransmissions.
On the graph
Gap then burst · Per-connection RTO and backoff, RTO expirations
Where to look
rto (retransmission wait in ms) and backoff (consecutive expirations) of the stalled connection from ss -ti; for the whole server, the increase in TcpExtTCPTimeouts (retransmission timer expirations) from nstat -az
Confirmed if
The stalled connection shows backoff of 1 or more with rto grown to seconds, and TCPTimeouts rises at that time
Ruled out if
If retransmissions finish as fast retransmits with no RTO expiration, stalls are short. That points to “TCP head-of-line blocking”
net/ipv4/tcp_input.c (Linux v6.12)Linux kernel Linux RTO = smoothed RTT + RTT variation, and the variation term has a floor of tcp_rto_min (200 ms), so RTO is at least RTT + 200 ms
IP SysctlLinux kernel tcp_rto_min_us defaults to 200 ms, initial RTO for connection requests is 1 s, with tcp_retries2=15 it takes at least 924.6 s (about 15 minutes) to give up
ss(8) — Linux manual pageiproute2 rto (retransmission timer, ms) and backoff (exponential backoff count) in -i
net/ipv4/tcp_timer.c (Linux v6.12)Linux kernel Each time the retransmission timer expires, TCPTimeouts goes up, backoff increases by one, and RTO doubles (up to the maximum)
ID sk-nagle · Primary owner Game team (Server development) · Also Game team (Client development)
Nagle’s algorithm, which batches small packets, and delayed ACK, which sends ACKs late, interact so that each message written in pieces is delayed by 40–200 ms.
Why Small messages are written in pieces without TCP_NODELAY turned on → Effect The sender waits for an ACK, and the receiver sends its ACK late → On screen Ping is low, yet every action is consistently sluggish: input lag
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: turn on TCP_NODELAY, collect one tick’s worth of messages and write them at once, don’t rely on turning off delayed ACK on the receiving side (Linux TCP_QUICKACK only lasts briefly, and Windows needs a registry change on every PC) because the game can’t reliably control it. Client: turn on TCP_NODELAY, collect one frame’s worth of messages and write them at once.
Ballpark numbers
Linux usually delays ACKs by 40 ms (up to 200 ms depending on the situation). Older Windows versions used 200 ms, and current versions use 40 ms (40 ms in the Windows Server 2019 default template). The receiving OS decides the delayed ACK, so if the server sends messages in pieces with Nagle on, each can be delayed by 40–200 ms depending on the receiving PC.
On the graph
Always high · Action response time (in-game RTT)
Where to look
Gaps between requests and responses in a server-side packet capture (tcpdump, Wireshark), and whether the server and client code turn on TCP_NODELAY
Confirmed if
Ping is low, but gaps of around 40 ms (200 ms on older Windows) keep appearing between small packets, and each gap ends right after the other side’s ACK arrives. Turning on TCP_NODELAY makes them disappear
Ruled out if
Response gaps close to ping: not this cause. Game server slow to produce responses: server processing (“Message queue backlog”)
Check with
Infra tools (no game code needed)
Sources (5)
RFC 9293: Transmission Control Protocol (TCP)IETF Nagle holds back small data while unacknowledged data is outstanding; it must be possible to turn off per connection; delayed ACK under 0.5 s; the problem of the two interacting
include/net/tcp.h (Linux v6.12)Linux kernel Linux delayed ACK minimum TCP_DELACK_MIN (HZ/25 = 40 ms), maximum TCP_DELACK_MAX (HZ/5 = 200 ms)
ID sk-block-send · Primary owner Game team (Server development)
When one player on a slow connection has a full send buffer and the server sends in blocking mode (a send call that doesn’t return until the buffer has room), the server thread waits on that one player.
Why A slow client’s send buffer is full → Effect With blocking sends, the server thread waits until the buffer has room → On screen Everyone that thread handles gets a freeze or slow motion
Use non-blocking sends, cap the send queue per client, drop stale updates.
On the graph
Random spikes · Server tick time, Send-Q per connection
Where to look
Connections whose Send-Q (bytes not yet ACKed or not yet sent) has filled up to the send buffer size, found with ss -tn; in a thread dump (stacks) of the game server taken when ticks spike, any thread stuck in a send call
Confirmed if
While a slow connection has a full Send-Q, the thread handling it is stuck in send, and only the players on that same thread stall with it
Ruled out if
Stalled threads waiting outside send (locks, DB calls): “Lock contention,” “Blocking calls on the game thread”
Check with
Infra tools (no game code needed)
Sources (3)
send(2) — Linux manual pageLinux man-pages send() blocks when the send buffer has no room, and returns immediately with EAGAIN in non-blocking mode
send function (winsock2.h)Microsoft Winsock send also blocks when buffer space runs out, unless the socket is in non-blocking mode
net/ipv4/tcp_diag.c (Linux v6.12)Linux kernel Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed
ID sk-slow-client · Primary owner Game team (Server development)
When a client’s outgoing data keeps piling up, the server drops stale updates or disconnects it.
Why The client’s connection can’t keep up with what the server sends → Effect The server drops stale updates, or disconnects the client once a limit is exceeded → On screen Just that player sees teleporting or gets a disconnect
Send less (update rate by distance), keep sending at lower quality, keep less data queued in the kernel (TCP_NOTSENT_LOWAT on Linux).
Ballpark numbers
With a 256 KB send buffer, a 30 KB/s connection builds up more than 8 seconds of backlogged data. Linux may also grow this buffer automatically to several MB.
On the graph
Outliers only · Send-Q per connection, dropped updates per client
Where to look
Per-client send queue length, dropped updates, and disconnect reasons logged by the game server; on the server, that connection’s Send-Q and cwnd together from ss -tni
Confirmed if
Only the connections of players who teleport or disconnect keep a full Send-Q, and the game log shows dropped updates or a disconnect for exceeding the send queue for those players
Ruled out if
Send-Q empty while the player still teleports: not a server send-side problem. Look at that player’s connection loss (“Wireless link loss”) or on-screen interpolation
Check with
Game server or client logs and metrics
Sources (2)
IP SysctlLinux kernel tcp_wmem: the maximum auto-tuned send buffer defaults to 64 KB–4 MB (depending on memory); tcp_notsent_lowat and TCP_NOTSENT_LOWAT limit the amount of unsent data
net/ipv4/tcp_diag.c (Linux v6.12)Linux kernel Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed
ID sk-keepalive · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)
When the other side vanishes without a close signal, TCP notices only much later. Keepalive (a TCP feature that checks whether an idle connection is still alive) is off by default, and even when it’s on, checks start only after 2 hours of idle time.
Why The client vanishes without a close signal because its power went off or its connection dropped → Effect The server assumes the connection is still alive (keepalive default 7,200 s; if data was being sent, about 15 minutes until retransmission gives up) → On screen A ghost character stays behind, and reconnecting fails with an “Already logged in” error
After sitting idle, Right after login or maintenance
Owner
Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)
Game team action items
Server: answer heartbeats and close the connection yourself if nothing arrives for a set time (tune TCP_KEEPIDLE and TCP_USER_TIMEOUT), and on reconnect use the session token to replace the old session and resume it. Client: send game-level heartbeats every few seconds to tens of seconds (half the shortest idle timeout or less), reconnect automatically when disconnected.
Infra team action items
Lower the kernel defaults (tcp_keepalive_time and so on) for sockets whose code doesn’t set its own values (applies only to sockets with SO_KEEPALIVE on).
Ballpark numbers
By default, Linux starts checking after 7,200 seconds of idle time, sends 9 probes 75 seconds apart, and drops the connection if none are answered. That adds up to about 2 hours 11 minutes. Windows also waits for 2 hours of idle time by default before it starts checking.
On the graph
Outliers only · Time since last receive, per connection
Where to look
lastrcv (ms since the last receive) and the keepalive timer (timer:(keepalive,…)) per connection from ss -tnoi, lined up with the game server’s “Already logged in” rejections
Confirmed if
ESTABLISHED connections with lastrcv of minutes to hours are still there, and reconnects for those accounts are rejected with “Already logged in”
Ruled out if
No long-silent connections, yet “Already logged in” still appears: the game server’s session cleanup code
Check with
Infra tools (no game code needed)
Sources (4)
tcp(7) — Linux manual pageLinux man-pages 9 probes 75 s apart after 7,200 s of idle time (about 11 more minutes), applies only to sockets with SO_KEEPALIVE on, TCP_KEEPIDLE and TCP_USER_TIMEOUT
ID sk-fragment · Primary owner Game team (Server development)
A UDP packet larger than the MTU (the largest size that can be sent in one piece) is fragmented at the IP layer, and losing just one fragment throws away the whole packet.
Why Snapshots in crowded areas exceed 1,500 bytes → Effect They go out split into several fragments, and losing any one of them discards the whole packet → On screen Large packets are lost several times as often. Teleporting only in crowded areas
Split packets yourself to 1,200 bytes or less, send only what changed.
Ballpark numbers
On a connection with 2% loss, about 8% of packets split into 4 fragments are lost. Some firewalls and ISPs drop fragmented packets outright, so those players never receive a single large packet.
On the graph
Rises with load · IP fragments created (IpFragCreates), snapshot size
Where to look
On the server, the increase in IpFragCreates (fragments created while sending) from nstat -az; on the receiving side, IpReasmFails (reassembly failures). UDP packet size distribution from game server logs or a packet capture
Confirmed if
IpFragCreates rises where crowds gather, UDP packets larger than 1,500 bytes show up, and teleporting reports go up at the same time
Ruled out if
IpFragCreates not rising: no fragmentation on the server’s sending side
Check with
Infra tools (no game code needed)
Sources (3)
RFC 8085: UDP Usage GuidelinesIETF Losing one fragment means the packet can’t be reassembled and is lost entirely; UDP apps should avoid IP fragmentation
net/ipv4/proc.c (Linux v6.12)Linux kernel Counter names shown by nstat: FragCreates (fragments created) and ReasmFails (reassembly failures) in the Ip group
ID sk-reliable-udp · Primary owner Game team (Server development) · Also Game team (Client development)
When the retransmission rules you built on top of UDP are too conservative, recovery is slow; when they’re too aggressive, they clog the connection even more.
Why Retransmission interval, retry count, and window size don’t suit the connection → Effect Slow recovery, or duplicate sends that make congestion worse → On screen Skills not going off, fast-forward, worse lag during congestion
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: base retransmission on measured round-trip time, separate channels by importance. Client: apply the same retransmission and channel settings as the server.
On the graph
Random spikes · Reliable UDP retransmission rate, in-game RTT
Where to look
Per-connection statistics from the library in use (retransmissions, estimated round-trip time, retransmission timeout), logged on server and client and compared with the same player’s actual connection loss rate (measured with mtr)
Confirmed if
Retransmission rate several times the actual loss rate: settings too aggressive. Retransmission timeout several times the measured round-trip time: settings too conservative
Ruled out if
Retransmission rate close to the loss rate and timeout in line with round-trip time: not a settings problem. Look at the connection loss itself
Check with
Game server or client logs and metrics
Sources (1)
RFC 8085: UDP Usage GuidelinesIETF Retransmissions can add to congestion, so they fall under congestion control; round-trip time is estimated as an average of several measurements (EWMA), initial value 1 s, lower the sending rate when the timer expires
ID sk-slowstart · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
When a connection has been idle for a while, TCP shrinks the congestion window (how much it can send at once) again, so a sudden large send goes out in several rounds.
Why A large burst of data (on entering a town, for example) goes out over a connection that was idle → Effect The congestion window has shrunk, so the data is spread over several round trips → On screen Right after entering, nearby characters and NPCs appear a few round trips late (more noticeable on distant servers)
While moving or changing zones, After sitting idle
Owner
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Shrink the entry data (send what’s essential first).
Infra team action items
Turn off tcp_slow_start_after_idle (Linux, server-wide setting).
Ballpark numbers
After an idle period longer than the RTO, the congestion window starts shrinking, and after a long idle period it drops to about 14 KB (10 packets). Then 100 KB can’t go out at once and takes 3 round trips.
On the graph
Outliers only · Transfer time right after entering (players with long RTT)
Where to look
The value of sysctl net.ipv4.tcp_slow_start_after_idle, and whether cwnd (congestion window) in ss -ti for that connection has shrunk at the moment a player enters an area after being idle
Confirmed if
The setting is 1 (the default), and on entering after idle, cwnd drops to around 10 and the transfer is split over several round trips. The longer a player’s RTT, the later things appear, and setting it to 0 makes the problem go away
Ruled out if
cwnd stays large, yet things still appear late: server-side entry handling (“Spawn burst when entering a crowded area”)
Check with
Infra tools (no game code needed)
Sources (4)
IP SysctlLinux kernel tcp_slow_start_after_idle on by default, the congestion window shrinks after an RTO of idle time (RFC 2861 method)
RFC 5681: TCP Congestion ControlIETF If no data has been sent for longer than the RTO, the congestion window is cut to at most the restart window min(IW, cwnd) and slow start begins again
ID sk-congestion · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
TCP treats loss as a sign of congestion and cuts its sending rate by 30–50%. It reacts the same way to Wi-Fi loss.
Why A little loss on Wi-Fi or the connection while there’s a lot to send → Effect TCP cuts its sending rate sharply and recovers slowly (CUBIC, the Linux and Windows default, cuts by 30%) → On screen Updates fall behind in crowded areas: fast-forward, input lag
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Send less (area of interest, changes only), spread sends out so they don’t go out in one burst.
Infra team action items
Switch to a congestion control algorithm such as BBR (tcp_congestion_control).
On the graph
Sawtooth · Per-connection congestion window (cwnd) and sending rate
Where to look
Several ss -ti snapshots of a lagging player’s connection for changes in cwnd and ssthresh and the congestion control name (cubic, bbr), plus whether Send-Q builds up
Confirmed if
After each loss, cwnd drops sharply and climbs back slowly, and Send-Q builds up while it’s low, lining up with the times of fast-forward and input lag reports
Ruled out if
cwnd is ample, yet updates still fall behind: the receiver’s window (“Zero window (a stall that looks like retransmission)”) or the server’s sending side
net/ipv4/tcp_diag.c (Linux v6.12)Linux kernel Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed
ID sk-linger · Primary owner Game team (Server development)
When the server cuts a connection abruptly, the final notice or save-complete signal it sent is lost.
Why The server closes the connection with an abortive close (RST). This happens when SO_LINGER is set to 0 seconds, or when the socket is closed before all received data has been read → Effect The kick reason and final data still in transit are thrown away → On screen An unexplained “Connection closed due to an unknown error”
Send the reason, then close only the sending direction (shutdown), read incoming data until the other side closes and only then close the socket, avoid SO_LINGER of 0 seconds.
On the graph
Random spikes · Connections ended by RST
Where to look
Increase in TcpExtTCPAbortOnData (closed with RST while data was left to send, SO_LINGER 0 s) and TcpExtTCPAbortOnClose (closed with unread data left) from nstat -az, and whether a server-side packet capture at the moment of disconnect shows an RST going out where a FIN should be
Confirmed if
At the times of “Connection closed due to an unknown error” reports, the server sends RST, and AbortOnData and AbortOnClose rise
Ruled out if
Server closed normally with FIN, yet the reason still doesn’t show: the client’s close handling
Check with
Infra tools (no game code needed)
Sources (4)
closesocket function (winsock.h)Microsoft Turning SO_LINGER on with a time of 0 makes close an abortive close that resets the connection immediately, and unsent data is lost
SNMP counterLinux kernel TcpExtTCPAbortOnData: closed with RST while data was left to send (SO_LINGER 0 s, for example); TcpExtTCPAbortOnClose: closed with unread data left, sending RST
ID sk-blocking-io · Primary owner Game team (Server development)
In a design where a thread can’t do anything else while it waits on one socket, everything slows down as the player count grows.
Why Each connection waits on its own reads and writes → Effect A delay on one connection spreads to the other connections on the same thread → On screen As concurrent users grow, everyone gets slow motion and input lag
Move to asynchronous I/O based on epoll, IOCP, or io_uring.
On the graph
Rises with load · Response time, thread count
Where to look
Game server thread count and per-thread voluntary context switches (cswch/s, times a thread stopped to wait for a resource) from pidstat -w -t, compared with response time as concurrent users grow
Confirmed if
Response time climbs steeply as concurrent users grow, and most of the threads (one per connection) show many voluntary switches while barely using CPU (waiting on sockets)
Ruled out if
Threads use CPU continuously without waiting: compute overload (“Tick overrun”)
I/O Completion PortsMicrosoft The Windows approach of handling many asynchronous I/Os with a pre-created thread pool
pidstat(1) — Linux manual pagesysstat cswch/s in -w: voluntary context switches where the task stopped on its own to wait for a resource; -t shows them per thread
ID sk-reuseport · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
When several processes share one port, the kernel assigns each connection to a process by address hash and never reassigns it. If one of those processes stalls, only the players assigned to it wait.
Why A gateway or login server runs several processes on one port with SO_REUSEPORT → Effect Even when one process stalls from GC or overload, the new connections and UDP packets assigned to it don’t move to another process → On screen Only some players can’t connect or freeze. During a restart that changes the process count, some UDP sessions drop
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Never let the receiving thread stall, implement a procedure for handing over sessions on restart.
Infra team action items
Monitor the connection queue per process (Recv-Q in ss), make sure deploys that change the process count follow the session handover procedure.
On the graph
Outliers only · Connection queue (Recv-Q) per listening socket
Where to look
Recv-Q (connections waiting for accept) and the owning process of each listening socket on the same port from ss -ltnp, plus throughput compared per process
Confirmed if
Of the sockets on the same port, only one keeps building up Recv-Q, and its process is stalled or its throughput is near 0
Ruled out if
Recv-Q builds up evenly on all sockets: overall overload (“Connection queue (listen backlog) overflow”)
Check with
Infra tools (no game code needed)
Sources (4)
socket(7) — Linux manual pageLinux man-pages SO_REUSEPORT lets several sockets bind to the same address and share incoming TCP connections and UDP packets
net/core/sock_reuseport.c (Linux v6.12)Linux kernel Without a BPF program, the packet hash is mapped onto the number of sockets in the group to pick the socket
Why does one NGINX worker take all the load?Cloudflare SO_REUSEPORT splits queues per worker with a simple hash, so if one worker is blocked, every connection queued to it stalls
net/ipv4/tcp_diag.c (Linux v6.12)Linux kernel Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed
ID sk-udp-connreset · Primary owner Game team (Server development)
When a Windows server sends UDP to a client that has already left, a “port unreachable” (ICMP) message comes back. That message makes the next receive call fail with an error, and if the server code treats the error as a failure of the socket itself, everyone using that socket is affected.
Why UDP keeps going to the address of a client that just left, and a “port unreachable” (ICMP) message comes back → Effect Windows fails the next receive call with WSAECONNRESET (10054), and the server code stops receiving or closes the socket → On screen Everyone who was using that socket freezes or disconnects at once
Turn off SIO_UDP_CONNRESET with WSAIoctl so these messages aren’t reported, just log receive errors and keep receiving.
On the graph
Mass disconnect · Connection count, receive error log
Where to look
UDP receive error codes (WSAECONNRESET, 10054) in the game server log and records of the receive loop stopping or the socket being closed; in a server-side packet capture, whether an ICMP Port Unreachable arrived just before
Confirmed if
A WSAECONNRESET receive error is logged right before connections drop all at once, and before that an ICMP Port Unreachable arrives from the address of a client that just left
Ruled out if
Linux server, or code that turns off SIO_UDP_CONNRESET: not applicable
Check with
Game server or client logs and metrics
Sources (3)
Winsock IOCTLsMicrosoft SIO_UDP_CONNRESET turns reporting of UDP “port unreachable” (PORT_UNREACHABLE) messages on and off
recvfrom function (winsock.h)Microsoft WSAECONNRESET on a UDP socket means a previous send got an ICMP Port Unreachable back
ID sp-tick-overrun · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
When one tick has more work than its budget, the server’s tick interval stretches, and the whole area slows down or stutters.
Why One tick (e.g., 50 ms) has more work than its budget → Effect Game state meant to update 20 times a second updates only 8 times → On screen Slow motion across the zone (stutter on some server designs), sluggish skill response
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Cut expensive computation, split the tick across threads, spread players out (channels), record tick processing time as a metric.
Infra team action items
Add tick time and per-core CPU utilization to monitoring and alerts, consider CPUs and instances with high single-core performance (clock speed).
Ballpark numbers
The budget is 50 ms on a 20-tick server, 33 ms at 30 ticks, and 16.7 ms at 60 ticks. To be ready for sudden crowds, it’s safer to leave headroom and normally use only about half the budget.
On the graph
Rises with load · Server tick time, player count per zone/channel, game thread CPU
Where to look
Tick time (p99) and tick-overrun count logged by the server, on one graph with player count per zone/channel. Without tick metrics, CPU utilization of the game thread from pidstat -t 1
Confirmed if
Tick time goes over budget (50 ms at 20 ticks) when players crowd in, while the game thread sits near 100% CPU
Ruled out if
Ticks overrun while game-thread CPU is low: points to waiting (GC pause, locks, blocking calls). Long run queue latency in bcc runqlat means threads aren’t getting CPU time: points to a CPU shortage or too many threads
Check with
Game server or client logs and metrics
Learn more
What a late tick looks like depends on the server design. A server that advances game state by a fixed amount of time each tick (e.g., 50 ms) slows down game time itself, so players see slow motion. A server that moves everything by the actual elapsed time in one step keeps the game running at normal speed, but packets become sparse and each move is large, so it shows up as stutter or teleporting. Either way, input response gets slower. If one game thread runs the whole server, the whole server slows down; if each area has its own thread, only that area does. Some games, such as EVE Online, deliberately slow game time by up to 10 times during large battles (Time Dilation) so the computation can keep up.
VALORANT's 128-Tick ServersRiot Games A 128-tick server must finish each frame within 7.8125 ms; server frame time is measured per subsystem and the budget is split among them
HED-GP Technical Retrospective: What a HED-acheCCP Games Under overload, EVE Online slows game time with Time Dilation down to a floor of 10% (10 times slower); node CPU normally stays below 80%
Handling variation in timeUnity When a fixed-step simulation falls behind, it runs the catch-up steps in a batch and discards time beyond the limit, so game time runs slower than real time
pidstat(1) — Linux manual pagesysstat -t shows per-thread statistics (CPU utilization and so on) for the threads in a process
ID sp-aoi · Primary owner Game team (Server development)
If you compare everyone against everyone to work out who can see whom, 10 times the players means 100 times the computation.
Why Every character’s distance is checked against every other character, or, even with a grid, hundreds of players crowd around one cell → Effect About 10,000 comparisons for 100 players, about 1 million for 1,000 → On screen Ticks spike where crowds gather, such as world bosses and sieges: slow motion, stutter
Split the world into a grid or regions and compare only nearby entities, update distant entities less often, cap how many players one player can see.
Ballpark numbers
If a distance check plus the update to the visible and not-visible lists takes 0.1 µs (one ten-millionth of a second) per pair of players, 1,000 players (about 1 million pairs) cost 100 ms per tick. That’s twice the 20-tick budget (50 ms).
On the graph
Rises with load · Server tick time, players gathered in one place
Where to look
Player count and tick time per zone/channel on the same graph, plus the time spent on AOI calculation within the tick, measured separately. Without a separate measurement, per-function CPU share of the game process from perf top -p
Confirmed if
When the number of players in one place doubles, tick time grows nearly 4 times, and view range and distance functions take most of the CPU time
Ruled out if
Tick time grows in proportion to player count, or send and serialization functions take a large share: points to “Broadcast fan-out overload” or serialization and compression cost
Replication Graph in Unreal EngineEpic Games The default approach of checking every connection for each actor becomes a server CPU bottleneck with many players and actors; MMORPGs and similar games split the world into a grid and reuse per-cell lists
ID sp-broadcast · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Sending one player’s movement to everyone who can see them creates updates on the order of the square of the crowd size.
Why Each player’s changes are sent to everyone who can see them → Effect 1,000 players who all see each other means 1 million updates per tick → On screen The send queue and bandwidth saturate, causing delay and loss (input lag, fast-forward, teleporting)
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Lower the update rate by distance and importance (nearby enemies every tick, distant players a few times a second), cap how much is sent to each player and fill it with the most important updates first, pack several updates into one packet, cap the number of players displayed.
Infra team action items
Alert on each server’s outbound bandwidth and packets per second against the NIC and instance network limits, check headroom before large events.
Ballpark numbers
1,000 players × 1,000 players × 20 ticks = 20 million updates per second. At 40 bytes each, that’s about 6.4 Gbps for the whole server and about 6.4 Mbps per receiving player. Capping the visible player count at 150 brings it down to about 1 Gbps total and about 1 Mbps per player.
On the graph
Rises with load · Server outbound packets and bytes, players gathered in one place
Where to look
txpck/s and txkB/s (packets and KB sent per second by the server NIC) from sar -n DEV 1 alongside the player count graph. On cloud instances, the allowance-exceeded counters in ethtool -S (bw_out_allowance_exceeded and pps_allowance_exceeded on AWS ENA)
Confirmed if
As the crowd grows, outbound packets and bytes rise faster than the player count (close to its square), and from the moment they hit the limit, the allowance-exceeded counters or transmit drops increase
Ruled out if
Outbound volume unchanged while only tick time grows: points to AOI calculation or game logic
Actor Priority in Unreal EngineEpic Games When a connection’s bandwidth is saturated, each actor gets a priority (distance, line of sight, time since last sent) and bandwidth goes to the most important first
Detailed Actor Replication Flow in Unreal EngineEpic Games NetUpdateFrequency sets the update rate per actor; actors are sent in priority order, and once the connection is saturated, the rest wait for the next tick
sar(1) — Linux manual pagesysstat rxpck/s and txpck/s (packets received and sent per second), rxkB/s and txkB/s (KB received and sent per second) in sar -n DEV
ID sp-hotzone · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
When each area runs on a single thread and players crowd into one place, only that one core hits 100%.
Why One thread runs each area (channel) → Effect When players crowd into one place, only that core saturates while the other cores have room to spare → On screen Only that area lags; other areas are fine
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Spread players across channels, parallelize work within an area, cap player count.
Infra team action items
Add per-core CPU utilization to monitoring and alerts (a single saturated core gets buried in the server-wide average).
Ballpark numbers
On a 16-core server, one core at 100% shows up as only about 6% total CPU utilization. You have to look at per-core utilization to find it.
On the graph
Rises with load · Per-core CPU utilization, per-thread CPU
Where to look
Per-core utilization from mpstat -P ALL 1 and per-thread CPU of the game process from pidstat -t 1, compared with the player count in the zone the busiest thread runs
Confirmed if
Total server CPU is low, but one thread (one core) sits near 100%, and at that moment players are crowded into the zone that thread runs
Ruled out if
Several cores evenly high: overall server overload. Only one core’s %soft (receive processing) high: points to NIC interrupts concentrated on one core
Check with
Infra tools (no game code needed)
Learn more
Other areas stay fine only when each area runs its own tick independently. If the area threads wait for each other every tick and move on to the next tick together, the busiest area slows the tick for the whole server.
Time Dilation – How’s That Going?CCP Games EVE Online’s Time Dilation works per node, so even distant star systems on the same node slow down; big battles run on reinforced nodes that host only 4 star systems
mpstat(1) — Linux manual pagesysstat Shows per-processor utilization separately from the overall average (-P ALL); %soft is the share of time spent handling software interrupts
pidstat(1) — Linux manual pagesysstat -t shows per-thread statistics (CPU utilization and so on) for the threads in a process
ID sp-lock · Primary owner Game team (Server development)
When several threads wait on one lock to write the same data, they run one at a time no matter how many threads you add.
Why Several threads use shared data at once, such as the auction house or guild storage → Effect The others wait until the thread holding the lock finishes → On screen Only certain features are slow; in bad cases, the whole tick is delayed
Split locks into finer-grained ones, do less work inside locks, move to a message-based design (give each piece of data an owning thread, and have other threads only send it requests as messages).
Ballpark numbers
If 20% of the work happens inside the lock, throughput tops out at 5 times that of one thread no matter how many threads you add; at 40%, it stops at 2.5 times.
On the graph
Rises with load · Request processing time, per-thread CPU and context switches
Where to look
Per-thread voluntary context switches (cswch/s, times a thread stopped to wait for a resource) from pidstat -w -t 1; where threads wait after leaving the CPU (wait time per call stack) from bcc offcputime -p. For .NET, the lock contention count in dotnet-counters (dotnet.monitor.lock_contentions on .NET 9 and later, Monitor Lock Contention Count on 8 and earlier)
Confirmed if
Processing time grows as load rises while CPU utilization stays low, most of the wait time is concentrated in call stacks trying to acquire a lock, and the lock contention count rises along with it
Ruled out if
CPU maxed out: a compute problem (tick overrun, single-threaded zone overload). Waiting on DB or file calls: points to blocking calls on the game thread
Check with
Infra tools (no game code needed)
Learn more
This happens in designs where several threads modify game data together. A design where one thread owns each area or feature and threads communicate only through messages has almost no locks, but you have to watch for work piling up on one thread (single-threaded zone overload). If the game thread waits for a lock held by a slow save operation, that whole tick stalls.
Amdahl's Law in the Multicore EraIEEE IEEE Computer 2008 paper (authors’ copy). If the fraction that can’t be parallelized is 1−f, the speedup can never exceed 1/(1−f) no matter how many cores you add (Amdahl’s law)
Request schedulingMicrosoft Orleans grains (actors) use a single-threaded execution model that runs each request to completion one at a time, so state is never modified concurrently; grains waiting on each other’s responses can deadlock
pidstat(1) — Linux manual pagesysstat cswch/s in -w is the number of voluntary context switches from stopping to wait for a resource; -t shows it per thread
Well-known EventCounters in .NETMicrosoft Monitor Lock Contention Count (monitor-lock-contention-count): number of times contention occurred when trying to acquire a monitor lock
.NET runtime metrics.NET dotnet.monitor.lock_contentions since .NET 9: number of times contention occurred when trying to acquire a monitor lock since the process started
ID sp-deadlock · Primary owner Game team (Server development)
When two threads each wait for a lock the other holds, both stop forever.
Why Thread A holds lock 1 and waits for lock 2, while B holds lock 2 and waits for lock 1 → Effect Both stop forever, and related threads stop one after another → On screen The whole server stops, and everyone disconnects when the watchdog restarts it
Set lock ordering rules, use locks with timeouts, add a watchdog that captures a thread dump at the moment of the hang.
On the graph
Mass disconnect · Connection count, server outbound traffic
Where to look
Call stacks of every thread captured during the hang: jstack for the JVM (finds and flags deadlocks automatically), dotnet-stack for .NET, gdb’s thread apply all bt for native servers, or dump a core file with gcore, restart, and analyze it afterward
Confirmed if
Two or more threads are stuck with stacks waiting for locks held by each other, and process CPU utilization is near 0 the whole time
Ruled out if
One thread spinning at 100% CPU during the hang: infinite loop. Threads waiting on DB or external responses: points to blocking calls or thread pool exhaustion
Check with
Infra tools (no game code needed)
Sources (6)
Runtime locking correctness validatorLinux kernel Taking two locks in opposite orders causes circular waiting and a deadlock (lock inversion deadlock); the Linux kernel checks lock ordering and warns in advance
Liveness, Readiness, and Startup ProbesKubernetes A liveness probe catches a deadlock, where the app is running but can’t make progress, and restarts the container
ID sp-sync-call · Primary owner Game team (Server development)
If the server waits for a DB response or a file write in the middle of a tick, all game progress on the server stops for that long.
Why The tick waits on DB reads and writes, log writes, or external API calls → Effect If the DB takes 100 ms, the tick stalls for 100 ms too → On screen Every time the DB or disk slows down, the whole field hitches
Hand all slow work (DB reads and writes, log writes, external API calls) off to run asynchronously and apply the results on the next tick (a timeout alone still leaves the tick stalled while it waits).
Ballpark numbers
Even a 0.5 ms DB round trip within the same data center adds up to 50 ms if it’s called 100 times in one tick. That alone uses up the entire 20-tick budget.
On the graph
Random spikes · Server tick time, DB query latency
Where to look
The tick time graph on the same time axis as DB query latency (slow query log and so on) and disk latency. Without tick metrics, bcc offcputime -p to see where the game thread waits
Confirmed if
Tick spikes coincide with DB or file latency spikes, and the game thread’s wait time is concentrated in call stacks that receive DB responses or write files
Ruled out if
DB and disk latency quiet while ticks spike: points to a GC pause or lock contention
ASP.NET Core Best PracticesMicrosoft Call data access, I/O, and long-running work asynchronously; synchronous blocking calls lead to thread pool exhaustion and slow responses
ID sp-queue · Primary owner Game team (Server development)
When requests arrive faster than they’re processed and pile up in the queue, the ones at the back get processed only seconds later or are dropped.
Why Requests arrive faster than they can be processed → Effect The queue grows, and messages are dropped once it passes its limit → On screen Skills and trades respond late or are dropped
Monitor queue length, adopt a policy that drops the oldest requests first, parallelize processing.
On the graph
Hits a ceiling · Queue length and age of the oldest message, messages processed per second
Where to look
Per-queue length, age of the oldest message, and messages received, processed, and dropped per second, as logged by the server. Without code metrics, the Recv-Q of the game socket from ss (or netstat): data the kernel has received but the process hasn’t read yet
Confirmed if
While arrivals exceed processing, the processing rate stays flat at one value, and queue length, message age, and drops keep growing
Ruled out if
Queue short and messages young, yet responses are slow: points to connection latency or delay in the tick itself
Avoiding insurmountable queue backlogsAWS Amazon Builders’ Library. Monitor backlog by the age of waiting messages; real-time systems process the newest data first (closer to LIFO) and may drop old messages
ID sp-timer-burst · Primary owner Game team (Server development)
When every monster respawn, every buff expiry, and the on-the-hour reward all land on the same tick, that one tick becomes tens of times heavier.
Why Respawn, expiry, reward, and autosave timers are all set to the same moment → Effect That one tick has tens of times its usual work → On screen A hitch at each of those scheduled times
Spread timer times out with a little randomness, split the work across several ticks.
On the graph
Periodic spikes · Server tick time
Where to look
The times ticks spiked, collected to check the intervals (on the hour, every 5 minutes, and so on), and cross-checked against the list of respawn, buff expiry, reward, and autosave timers that run at the same moment
Confirmed if
Ticks spike at the same times or intervals every time, and some game timer work fires all at once at those times
Ruled out if
Periodic, but lining up with pause times in the GC log or the server’s cron and backup times: points to “Server GC stop-the-world pause” or “Scheduled jobs”
Check with
Game server or client logs and metrics
Sources (1)
Timeouts, retries, and backoff with jitterAWS Amazon Builders’ Library. Add jitter to all timers, periodic jobs, and delayed work to spread out load that would otherwise land at the same moment; a case where one-minute requests from many servers piled into the first few seconds of every minute
ID sp-pathfinding · Primary owner Game team (Server development)
When hundreds of monsters chase players and compute paths at the same time, it takes a lot of CPU.
Why Large mob pulls or mass spawns send many monsters chasing players at once → Effect Each monster runs its own pathfinding → On screen Slow motion in that hunting ground only
Cache paths, cap the number of path computations, split the work across several ticks.
On the graph
Rises with load · Server tick time, monsters per zone
Where to look
Number of monsters chasing players and tick time, per zone. Without a separate count, per-function CPU share of the game process from perf top -p
Confirmed if
Tick time grows during large mob pulls and mass spawns, and pathfinding (path search) functions take a large share of CPU time
Ruled out if
Ticks grow when there are few monsters but many players: points to AOI calculation or broadcast fan-out
Check with
Game server or client logs and metrics
Sources (2)
AI.NavMesh.pathfindingIterationsPerFrameUnity Pathfinding processes only a set number of nodes per frame, spreading the work over several frames, so the game stays smooth even with long paths or many requests at once
ID sp-serialize · Primary owner Game team (Server development)
Turning outgoing data into bytes and compressing it takes CPU too, and with many players this cost explodes.
Why Structs are converted to bytes and compressed for every update → Effect Cost grows with the square of the player count → On screen Sends go out late: input lag
Reuse a packet built once for many players, use a lightweight format.
On the graph
Rises with load · Server CPU utilization, CPU of the thread that builds packets
Where to look
Share of the game process’s CPU time spent in serialization, compression, and encryption functions (including library functions such as zlib, LZ4, and OpenSSL) from perf top -p, compared between quiet and crowded times
Confirmed if
The more players crowd in, the bigger the share of serialization, compression, and encryption functions, and the thread that builds packets saturates first
Ruled out if
These functions take a small share: points to AOI calculation or game logic
Check with
Infra tools (no game code needed)
Learn more
If the connection encrypts packets (TLS, DTLS, and so on), encryption and decryption use CPU too. Encryption is done separately for each connection, so even when one built packet is reused for many players, the encryption cost scales with the number of recipients. Symmetric ciphers such as AES-GCM are fast enough for one core to handle several GB per second, so their share is usually small, but their speed varies widely with the size of the unit encrypted at once (the record), and with many small packets, as in games, the cost per byte goes up. In the handshake done once per connection, the server signs with its certificate key and computes the key exchange (ECDHE). One core can do about 1,100 (RSA 2048) to 18,000 (ECDSA P-256) signatures and about 9,000 key exchanges per second, which becomes a burden when logins surge.
Sources (4)
Introduction to Iris in Unreal EngineEpic Games Holds the state to replicate as a single quantized copy to cut expensive work, and shares that work across connections
VALORANT's 128-Tick ServersRiot Games Comparing replicated variables for each client every frame and bundling the changed values reads memory all over the place, which is slow and costs a lot of server CPU
How "expensive" is crypto anyway?Cloudflare BoringSSL measurements: AES-128-GCM about 3.7 GB per second (varies widely with record size); per core per second, 1,120 RSA 2048 signatures, 18,477 ECDSA P-256 signatures, and 9,394 P-256 ECDHE operations; on Cloudflare edge servers, the TLS library used about 1.8% of CPU
ID sp-crash · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
When the server process dies from an unhandled error, everyone on that server disconnects at the same time.
Why A fatal error such as a reference to something that doesn’t exist (null reference), bad data, or running out of memory → Effect The server (or zone) process exits → On screen Everyone disconnects at once, and progress since the last save may be rolled back
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Analyze crash dumps and fix the root cause, save often.
Infra team action items
Restart the process automatically, set up crash dump collection and retention, alert the moment a server goes down.
On the graph
Mass disconnect · Connection count, process restarts
Where to look
Core dump records in coredumpctl list (time, PID, terminating signal) and the service manager’s (systemd) records of abnormal exits and restarts. On Windows servers, dump files saved by WER
Confirmed if
At the moment the connection count dropped to near 0, the game server process exited abnormally and left a core dump
Ruled out if
Process stayed alive but connections dropped: points to network equipment or an idle timeout. A watchdog restart after a long hang: points to an infinite loop or deadlock
Check with
Infra tools (no game code needed)
Sources (3)
Collecting User-Mode DumpsMicrosoft Configure Windows Error Reporting (WER) to collect full or mini dumps locally when a user-mode program crashes
systemd.service(5) — Linux manual pagesystemd Restart=on-failure automatically restarts the service after an abnormal exit, a kill by signal (including core dumps), or a watchdog timeout; recommended for long-running services
coredumpctl(1) — Linux manual pagesystemd list queries core dumps saved by systemd-coredump, showing crash time, PID, and the signal that caused the crash
ID sp-threadpool · Primary owner Game team (Server development)
When every worker thread is tied up in slow work, new requests just wait with no end in sight.
Why Worker threads are tied up waiting on external API or DB responses → Effect No thread is free to take a new request → On screen Infinite loading in specific features such as login or the shop
When crowds gather, Right after login or maintenance
Owner
Primary owner Game team (Server development)
Game team action items
Put timeouts on slow calls, separate thread pools by feature, make calls asynchronous.
On the graph
Hits a ceiling · Thread pool thread count and queue length, request processing time
Where to look
For .NET, thread pool thread count and queue length in dotnet-counters monitor (dotnet.thread_pool.thread.count and dotnet.thread_pool.queue.length on .NET 9 and later, ThreadPool Thread Count and ThreadPool Queue Length on 8 and earlier), and where the worker threads are waiting, from dotnet-stack. For JVM and native servers, the same from a thread dump
Confirmed if
CPU utilization is well below 100%, yet the thread count keeps creeping up or sits at its cap, the queue builds up, and most workers are waiting on responses from the same external call (DB, HTTP)
Ruled out if
Queue empty yet still slow: the called service itself is slow, which points to a cascading failure or an external service dependency
Check with
Infra tools (no game code needed)
Learn more
If packet receiving and game logic share the same worker thread pool, the moment a few slow jobs occupy every worker, packet processing stops for the whole server.
Debug ThreadPool StarvationMicrosoft When the pool has no threads left and new work has to wait, responses slow down; blocking code that holds threads is the cause. In dotnet-counters, dotnet.thread_pool.thread.count creeping up while CPU is well below 100% signals exhaustion (dotnet.thread_pool.queue.length is often large too); dotnet-stack shows where threads are waiting
Avoiding insurmountable queue backlogsAWS Concurrent requests = arrival rate × latency (Little’s law). At 100 requests per second, latency growing from 100 ms to 10 s turns 10 threads into 1,000 and the pool runs dry
Bulkhead PatternMicrosoft Azure With separate connection and thread pools for each called service, a failure in one service blocks only its own pool
.NET runtime metrics.NET dotnet.thread_pool.thread.count (thread pool thread count) and dotnet.thread_pool.queue.length (queued work items) since .NET 9
Well-known EventCounters in .NETMicrosoft ThreadPool Thread Count (threadpool-thread-count) and ThreadPool Queue Length (threadpool-queue-length) on .NET 8 and earlier
ID sp-infinite-loop · Primary owner Game team (Server development)
When a bug keeps a tick from ever finishing, the server stops, and the watchdog forces a restart.
Why A loop that never ends because of a wrong condition, or runaway recursion → Effect The tick never finishes, and the server stops → On screen A freeze, then everyone disconnects
Cap loop iterations, add a watchdog, write tests that reproduce the problem input.
On the graph
Mass disconnect · Connection count, per-thread CPU
Where to look
Per-thread CPU from pidstat -t 1 during the hang, and which function the thread spinning at 100% is in, from perf top -t (thread ID) or gdb. If it has already restarted, the watchdog timeout record (WatchdogSec in systemd, a failed Kubernetes liveness probe)
Confirmed if
While the server is stopped, one game thread sits at 100% CPU, and its stack keeps looping inside the same function or loop
Ruled out if
CPU near 0 during the hang: points to a deadlock or waiting on an external response
Check with
Infra tools (no game code needed)
Sources (4)
systemd.service(5) — Linux manual pagesystemd WatchdogSec=: if the service doesn’t send a keep-alive signal (WATCHDOG=1) within the set time, it’s treated as failed and stopped, then restarted automatically depending on the Restart= setting
Liveness, Readiness, and Startup ProbesKubernetes A liveness probe catches a state where the app is running but can’t make progress and restarts it; by default it checks every 10 seconds and restarts after 3 consecutive failures
pidstat(1) — Linux manual pagesysstat -t shows per-thread statistics (CPU utilization and so on) for the threads in a process
ID sp-hot-entity · Primary owner Game team (Server development)
When hundreds of players hit one boss at the same time, the computation for that single boss piles up in one place, and hit information goes out to everyone watching.
Why Hundreds of players use skills, buffs, and debuffs on one boss nonstop → Effect The boss’s HP, aggro list, and debuff calculations pile up in one place, and every hit sends damage number and effect packets to everyone watching → On screen Skills land late and damage numbers pop up in bursts; slow motion only around the boss
Bundle or skip other players’ damage numbers and effects, cap the number of debuffs on one target, split hit processing across several ticks.
Ballpark numbers
800 players hitting twice a second makes 1,600 hits per second. Telling all 800 watchers about every hit means 1.28 million messages per second.
On the graph
Rises with load · Server tick time, messages sent
Where to look
Tick time and packets sent during the boss fight, alongside the number of players around the boss, plus events per second (hits, buffs, debuffs) per target if available
Confirmed if
As players gather around the boss, tick time and outbound traffic climb steeply, and the boss alone has tens of times more events per second than any other target
Ruled out if
Things slow down just as much whenever players gather, boss or no boss: points to “AOI calculation blowup (N²)” or “Broadcast fan-out overload”
Check with
Game server or client logs and metrics
Sources (1)
HED-GP Technical Retrospective: What a HED-acheCCP Games Even a single attack must be reported to every client watching, creating an O(n²) load of n players notifying n players, and message-heavy drone attacks make this load grow even faster
ID sp-spawn-burst · Primary owner Game team (Server development) · Also Game team (Client development)
When you teleport into a town packed with players, the server has to send the appearance, gear, and status of hundreds of newly visible players all at once.
Why You suddenly appear somewhere crowded by teleporting, logging in, or switching channels → Effect Full data for hundreds of players is built and sent at once, and your PC also loads it all at once → On screen A brief pause right after arrival, characters pop in late one by one, and input lags
While moving or changing zones, Right after login or maintenance
Owner
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: send in order of distance, spread over several ticks, cache appearance data. Client: preload during the loading screen, create received characters over several frames.
Ballpark numbers
If one player’s appearance, gear, and buff data is 300 bytes, 500 players come to about 150 KB. Tens of times the usual per-tick traffic (a few KB) arrives in a single moment.
On the graph
Surge after opening · Bytes sent per connection, client frame time
Where to look
Bytes and packets sent over that connection in the first few seconds after arriving somewhere crowded (server log), and client frame time (net graph, client log)
Confirmed if
Right after arrival, outbound traffic on that connection shoots up to tens of times a normal tick and then settles, and client frame time spikes at the same moment
Ruled out if
The same pause when moving somewhere quiet: points to a zone transfer (handoff between servers) or client loading
Check with
Game server or client logs and metrics
Sources (2)
Detailed Actor Replication Flow in Unreal EngineEpic Games When an actor channel is first opened, initial data such as position and rotation is sent along with it, and once the connection is saturated, the remaining actors wait for the next tick
Actor Priority in Unreal EngineEpic Games Actors are prioritized by distance and view direction, and the nearest visible ones are sent first
ID sp-entity-buildup · Primary owner Game team (Server development)
When ground items that should have disappeared, summons, and finished timers pile up without being cleaned up, every tick has more work to do the longer the server stays up.
Why Ground items, summons, expired timers, and empty party data aren’t removed on time → Effect The lists walked every tick grow longer day by day → On screen Fine right after maintenance, then after a few days only that server or area gets more and more sluggish
Record entity count per area as a metric and watch the trend, give each entity a lifetime and a count cap, clean up periodically.
Ballpark numbers
On a server that walks every entity once per tick, doubling the entity count doubles that part of the tick time.
On the graph
Sawtooth · Entity count per zone, server tick time
Where to look
Entity count (ground items, summons, timers) and tick time per zone and server, over a period longer than the maintenance cycle (several weeks)
Confirmed if
Entity count and tick time start low after maintenance, climb every day, and drop sharply at each maintenance or restart, over and over, while memory stays ample
Ruled out if
Ticks unchanged while memory keeps climbing: points to a memory leak
Check with
Game server or client logs and metrics
Learn more
Like a memory leak, it gets worse the longer the server runs, but here memory is fine and only tick time grows. If the entity count graph forms a sawtooth with each maintenance cycle, this is the cause.
Sources (2)
Actor Ticking in Unreal EngineEpic Games Actors and components tick once every frame unless given their own interval, and ticking can be turned off when not needed
AActor::SetLifeSpanEpic Games Setting a lifespan on an actor destroys it automatically when it expires
ID sp-patch-traffic · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure)
When new content, effects, or synced fields raise packet size and frequency, a server that ran fine starts hitting MTU, bandwidth, and packet-rate limits after the patch.
Why The patch adds new skill effects, synced fields, or item data, making packets bigger or more frequent → Effect Large packets exceed the MTU and get fragmented, and the extra volume runs into bandwidth limits, cloud PPS limits, and send buffers → On screen Teleporting, skills not going off, and input lag in crowded places, starting right after the patch. Nothing changed in the infra, yet loss goes up
Whole server, Specific zone/channel, Specific region/ISP
When
When crowds gather, Evening peak hours, Always
Owner
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure)
Game team action items
Split packets yourself to 1,200 bytes or less, send only changes for new synced fields and lower their rate by distance and importance, compare packets and bytes per second per player and the largest packet size against the previous build on a test server before deploying, tag traffic metrics with the build version.
Infra team action items
Servers/OS: mark deploy times on graphs and compare packets and bytes per second per player and average packet size before and after the deploy, alert on instance allowance-exceeded counters, move to a bigger instance if needed. Network: check the capacity limits of firewalls, load balancers, and DDoS protection gear, and whether they block fragments.
Ballpark numbers
UDP packets are safe at 1,200 bytes or less; the path MTU on the internet is usually 1,500 bytes, and smaller through tunnels (1,476 bytes through a GRE tunnel). Packets larger than the path MTU are fragmented or dropped, and a fragmented packet is lost entirely if even one fragment is lost. If packets per second per player rise 20%, the server total rises 20% too, and an instance running close to its limit overflows right away.
On the graph
Step change · Packets and bytes per second per player, average packet size
Where to look
Packets and bytes per second on the server NIC (rxpck/s, txpck/s, rxkB/s, and txkB/s from sar -n DEV; NetworkPacketsOut and NetworkOut on EC2) divided by concurrent users, compared before and after the deploy. Average packet size is bytes ÷ packets; for the size distribution, run Wireshark’s Packet Lengths statistics on a packet capture
Confirmed if
From right after the deploy, packets and bytes per player or average packet size step up and stay there, and from the same moment the number of fragments the server creates (fragcrt/s in sar -n IP) or the instance allowance-exceeded counters (pps_allowance_exceeded and bw_out_allowance_exceeded on AWS ENA) rise
Ruled out if
Traffic pattern the same before and after the deploy, but only latency and loss went up: look at infra changes made at the same time (configuration, routing, equipment, OS or kernel updates)
Check with
Infra tools (no game code needed)
Learn more
When a report says “it worked fine before the patch,” this is the game-side cause to check first, along with infra changes. Even if the patch notes list no network changes, one new effect or synced field gets multiplied by hundreds of players in crowded places. Where the extra traffic actually hits a limit is covered in “IP fragmentation of UDP packets,” “Cloud PPS limit exceeded,” “NIC bandwidth saturation,” “Kernel socket buffers too small,” and “Middlebox over capacity (firewall, IPS, DDoS protection).” This entry covers the case where a game patch is what pushed traffic into those limits, so cut the traffic the patch added before raising the limits. If the OS or kernel was also updated at the same time, tell this cause apart from “Performance changes after OS, kernel, driver, or firmware updates” by whether per-player traffic changed.
Sources (7)
RFC 8085: UDP Usage GuidelinesIETF UDP apps SHOULD NOT send datagrams larger than the path MTU; losing one fragment loses the whole fragmented packet, and some NATs and firewalls drop all fragments
sar(1) — Linux manual pagesysstat rxpck/s and txpck/s (packets per second) and rxkB/s and txkB/s (KB per second) in sar -n DEV; fragcrt/s (IP fragments created per second, ipFragCreates) in sar -n IP
ID mem-gc · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
While a Java or C# server halts every thread to collect garbage (stop-the-world), the whole server stalls.
Why The heap fills up and GC starts → Effect Every game thread stops while GC collects (longer the more live data there is) → On screen Everyone on the server freezes at the same moment, then the game fast-forwards
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Explicitly select a low-pause GC (ZGC, Shenandoah, or G1 with a lower pause target) in the launch options, reduce allocations, tune heap size.
Infra team action items
Use instances with enough memory for a generous heap, give containers at least 2 CPUs and about 1.8 GB of memory (with less, JDK 26 and earlier pick Serial GC by default), monitor GC pause times.
Ballpark numbers
A minor GC that collects only new objects (the young generation) takes a few to tens of milliseconds. A full GC over a whole heap holding several GB of live data can take over 1 second. ZGC stays under 1 ms almost regardless of heap size, and Shenandoah pauses are also short because they don’t grow with heap size.
On the graph
Periodic spikes · Server tick time, GC pause time
Where to look
Turn on GC logging and overlay pause times and lengths on the server tick-time graph. Java: Pause lines from the startup option -Xlog:gc* (-XX:+PrintGCDetails on JDK 8 and earlier). .NET: the GC pause metric in dotnet-counters (dotnet.gc.pause.time on .NET 9 and later, % Time in GC since last GC on 8 and earlier). Go: the line GODEBUG=gctrace=1 writes for each GC
Confirmed if
Tick spikes line up with GC pauses, and pause length matches spike length. Every zone and channel on the server spikes at the same moment
Ruled out if
Ticks spike with no long pauses in the GC log: another cause such as locks, blocking calls, or disk writes. Only one zone spikes: the script engine’s GC (mem-script-gc) or that zone’s load
Check with
Infra tools (no game code needed)
Learn more
Java’s G1 has a default target of 200 ms per pause, which is 4 ticks on a 20-tick server. If a container gets fewer than 2 CPUs or less than about 1.8 GB of memory, Java on JDK 26 and earlier picks Serial GC, which collects on a single thread, as the default GC, and pauses get much longer. C# (.NET) servers usually turn on server GC and background GC, but collections of generations 0 and 1 (Gen0/1), where new objects live, and full GCs with compaction still stop every thread. Go pauses are usually under 1 ms, but under heavy allocation the code requesting memory has to take on part of the GC work, so ticks slow down. With any GC, if allocation outpaces collection, the game thread eventually stops. G1 falls back to a full GC, and ZGC stalls the thread requesting memory until the collection finishes.
Garbage-First (G1) Garbage CollectorOracle G1 default pause target 200 ms (MaxGCPauseMillis); if memory runs out during collection, it falls back to a full GC that stops and compacts the whole heap
JEP 439: Generational ZGCOpenJDK ZGC pauses are 1 ms or less regardless of heap size; G1 pauses range from a few ms to a few seconds. Risk of allocation stalls when allocation outpaces reclamation
Background garbage collection.NET Background GC applies only to gen2 collections; gen0 and gen1 collections (foreground GC) stop all managed threads
A Guide to the Go Garbage CollectorGo Go’s GC runs mostly concurrently with only short stop-the-world pauses; under heavy allocation, goroutines take on GC work (assist), which adds latency
JEP 271: Unified GC LoggingOpenJDK JDK 9 reimplemented GC logging on unified logging (-Xlog); -Xlog:gc writes one line per GC, like the old -XX:+PrintGC
The java CommandOracle Table mapping old GC log options to -Xlog: -XX:+PrintGCDetails becomes -Xlog:gc*
dotnet-counters diagnostic tool.NET .NET 9 and later show System.Runtime meters (dotnet.gc.pause.time and others); .NET 8 and earlier show the older EventCounters (% Time in GC since last GC and others)
runtime packageGo GODEBUG=gctrace=1: one line per GC with wall-clock time per phase, heap size at GC start and end, and the heap goal
ID mem-script-gc · Primary owner Game team (Server development)
Even on a C++ server, if quests, AI, and skills run in a scripting language such as Lua, the zone stops while the script engine’s GC runs.
Why Each zone’s script engine creates large numbers of temporary objects while running quests, AI, and events → Effect When the script engine’s GC collects a lot at once, that zone’s tick stops → On screen Periodic hitches only in certain zones or during certain events
Configure incremental or generational GC, advance GC a little every tick, reduce temporary objects in scripts.
Ballpark numbers
Once a script heap grows to hundreds of MB, an all-at-once collection (with incremental collection turned off, or a full collection in generational mode) can take tens to hundreds of milliseconds.
On the graph
Periodic spikes · Tick time per zone, script engine memory
Where to look
Record each zone’s tick time and the memory use of that zone’s script engine (collectgarbage("count") in Lua) every tick and overlay them on one graph
Confirmed if
The moments script memory drops sharply (an all-at-once collection) line up with that zone’s tick spikes, while other zones are fine
Ruled out if
Spikes with no change in script memory: that zone’s load or locks. Every zone on the server spikes at once: server GC (mem-gc) or swap (mem-swap)
Check with
Game server or client logs and metrics
Sources (1)
Lua 5.4 Reference ManualLua.org Incremental mode splits collection into small steps interleaved with execution (large steps make it stop-the-world); a major collection in generational mode is a stop-the-world pass over every object; collectgarbage("count") returns the total memory Lua is using (KB)
ID mem-alloc · Primary owner Game team (Server development)
Creating large numbers of temporary objects during an event makes GC run far more often than usual.
Why Item drops, combat logs, and event rewards create a flood of temporary objects → Effect GC runs several times as often, and objects not yet discarded get promoted to the old generation, so full GCs come sooner too → On screen Periodic hitches only during events
Use object pools and reusable buffers, profile allocations.
On the graph
Rises with load · GC count, allocation rate
Where to look
Count GCs per minute from GC logs (Java -Xlog:gc, Go GODEBUG=gctrace=1); on .NET, check allocation volume and GC count in dotnet-counters (dotnet.gc.heap.total_allocated and dotnet.gc.collections on .NET 9 and later, Allocation Rate and Gen 0 GC Count on 8 and earlier). Overlay with concurrent users and event times
Confirmed if
When the event starts, allocation rate and GC count climb faster than player count, and short pauses become frequent. Back to normal once the event ends
Ruled out if
GC count unchanged but each pause gets longer: live data has grown (mem-gc-thrash, mem-leak)
Check with
Infra tools (no game code needed)
Sources (3)
Garbage Collector ImplementationOracle Minor GC when the young generation fills; some surviving objects move to the old generation, and when the old generation fills, the whole heap is collected (much slower than a minor GC); -Xlog:gc writes one line per GC
A Guide to the Go Garbage CollectorGo The higher the allocation rate, the more often GC cycles run; GODEBUG=gctrace=1 prints GC trace output
dotnet-counters diagnostic tool.NET .NET 9 and later show dotnet.gc.heap.total_allocated and dotnet.gc.collections; .NET 8 and earlier show Allocation Rate and Gen 0 GC Count
ID mem-leak · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Memory that is never freed piles up little by little and, days later, leads to GC storms, swapping, or the process getting killed.
Why Data for logged-out characters and event handlers is never freed → Effect Free memory shrinks over several days → On screen Fine right after maintenance, laggier every day, and eventually the server goes down
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Analyze heap dumps, run long-duration load tests.
Infra team action items
Monitor and alert on per-process memory usage trends.
On the graph
Slow climb · Process memory (RSS), heap after GC
Where to look
Watch the game server process’s memory (RSS from pidstat -r) over several days; on servers with GC, watch the heap left right after GC. Java: the after-GC value in the before/after usage of -Xlog:gc lines. .NET: the post-GC heap size in dotnet-counters (dotnet.gc.last_collection.heap.size on .NET 9 and later, GC Heap Size on 8 and earlier)
Confirmed if
The heap left right after GC (the baseline) climbs every day after a restart and doesn’t come down even in the quiet early-morning hours
Ruled out if
Heap baseline flat while only RSS climbs: fragmentation (mem-fragment) or native memory. Rises and falls with player count: normal usage
Check with
Infra tools (no game code needed)
Learn more
Weekly restarts during scheduled maintenance hide a leak, so it can go unnoticed for a long time. It often shows up suddenly when maintenance is postponed once or an event brings in more players.
Sources (5)
Troubleshoot Memory LeaksOracle Suspect a leak when execution gets gradually slower; memory eventually runs out and the program terminates abnormally. The key data for leak analysis is a heap dump
Debug a memory leak in .NET.NET Even with GC, holding references to objects no longer needed is a leak, leading to performance degradation and OutOfMemoryException. Check memory trends and analyze dumps
ID mem-gc-thrash · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
When live data gets close to the heap limit, each GC reclaims almost nothing, so GC runs over and over without a break.
Why An event crowd or a leak pushes live data close to the heap limit → Effect GC reclaims only a little, so another full GC follows right away and GC uses most of the CPU → On screen The whole server alternates between slow motion and freezes for several minutes, then dies from running out of memory
Evening peak hours, When crowds gather, The longer it runs
Owner
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Size the heap well above peak live data (usually 2× or more), reduce long-held references and leaks.
Infra team action items
Alert on GC time ratio, restart promptly without waiting for it to recover, use instances with enough memory to grow the heap.
Ballpark numbers
GC taking more than 10% of run time is usually treated as a warning sign. Some Java GCs throw an out-of-memory error when they spend 98% of the time in GC and still reclaim almost nothing.
On the graph
Hits a ceiling · Heap after GC, GC time ratio
Where to look
How close the heap left after GC is to the max heap, and the share of time spent in GC. Java: “after GC (heap size)” in -Xlog:gc lines and how often Pause Full lines appear. .NET: dotnet-counters (% Time in GC since last GC on .NET 8 and earlier, growth of dotnet.gc.pause.time on .NET 9 and later). Go: the interval between GODEBUG=gctrace=1 lines
Confirmed if
The heap stays near its max even right after GC, full GCs run back to back, and the GC time ratio climbs well above normal (usually past 10%). Ticks slow down across the whole server meanwhile
Ruled out if
Plenty of heap left after GC but pauses are long: GC type or settings (mem-gc)
Check with
Infra tools (no game code needed)
Sources (5)
The Parallel CollectorOracle Parallel GC throws OutOfMemoryError if it spends more than 98% of total time in GC and recovers less than 2% of the heap
Garbage-First Garbage Collector TuningOracle By default (GCTimeRatio=12), G1 sizes the heap to keep GC time at about 8% of total time or less; full GCs caused by excessive heap occupancy show up in the log as Pause Full (G1 Compaction Pause)
A Guide to the Go Garbage CollectorGo With the default GOGC=100, the heap goal is about 2× the live heap; near the memory limit, GC runs nonstop (thrashing); GODEBUG=gctrace=1 prints GC trace output
Garbage Collector ImplementationOracle -Xlog:gc lines show the GC type (Pause Young, Pause Full), “before-GC usage->after-GC usage(heap size)”, and the pause time
dotnet-counters diagnostic tool.NET .NET 9 and later show dotnet.gc.pause.time; .NET 8 and earlier show % Time in GC since last GC
ID mem-swap · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
When memory runs short and the OS moves part of it out to disk, every access to that memory waits on a disk more than 1,000 times slower.
Why Memory in use exceeds physical RAM → Effect The OS moves part of it to disk and reads it back when needed → On screen Ticks balloon to hundreds of ms, and every player on the server sees slow motion and freezes
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Put a cap on process memory use (heap size and so on), check for leaks.
Infra team action items
Configure game servers not to use swap, act on memory alerts, provision RAM well above peak usage, since with swap off the process is killed (OOM) the moment memory runs out.
Ballpark numbers
A RAM read takes about 100 ns; reading back from an SSD takes about 100 µs (1,000×), a cloud disk over the network about 1 ms (10,000×), and an HDD 10 ms (100,000×).
On the graph
Slow climb · Swap usage, swap-in/out
Where to look
Overlay on tick time: the si and so columns of vmstat 1 (amount swapped in and out per second), some and full in /proc/pressure/memory (share of time stalled waiting for memory), and pidstat -r majflt/s for the game server process (page faults that had to read from disk)
Confirmed if
si above 0 at the time of the lag, with the game server’s majflt/s and the memory full value rising together
Ruled out if
si and so at 0 and memory pressure (PSI) near 0: swap isn’t the cause. No swap but majflt/s and PSI rising: memory is running out and code pages are being reread, so free up memory first
Check with
Infra tools (no game code needed)
Learn more
A server with GC reads all over the heap when it collects, so if even part of the heap is swapped out, a single GC can stretch to seconds or tens of seconds. With swap off, there is no slow swapping stage and the process goes straight to being killed (OOM), so secure spare memory first. Even without swap, when memory is nearly exhausted the OS may drop the executable’s code pages from memory and read them back again, so the whole server can slow down badly for a while before the OOM kill.
Documentation for /proc/sys/vm/Linux kernel swappiness: relative cost of swapping vs. reclaiming file pages; swap is random I/O and therefore expensive
Concepts overviewLinux kernel Reclaims page cache backed by files on disk and swappable pages; if that’s still not enough, the OOM killer kills a process
Solidigm™ D7-P5520 and D7-P5620 Product BriefSolidigm 99.99th percentile latency (four-nines latency) of 130 µs for server NVMe SSDs: basis for a single SSD read taking around 100 µs
PSI - Pressure Stall InformationLinux kernel some (share of time some tasks were stalled) and full (share of time all tasks were stalled at once) in /proc/pressure/memory
ID mem-cache-miss · Primary owner Game team (Server development)
When data is scattered all over memory, the CPU has to go all the way out to slow RAM and wait every time.
Why Objects scattered behind pointers and accessed in no particular order → Effect Data isn’t in the CPU cache, so every read goes to RAM (roughly 100 times slower) → On screen The same work costs several times more tick time; in bad cases, slow motion
Lay out data that is used together contiguously (data-oriented design).
On the graph
Always high · Tick time, CPU utilization
Where to look
Attach perf stat -d -p PID to the game server process to measure instructions per cycle (insn per cycle) and L1/LLC cache misses, and view them alongside tick time and CPU utilization
Confirmed if
CPU stays busy, but insn per cycle is low and LLC misses are high. Confirmed if a build with a changed data layout cuts tick time sharply at the same player count
Ruled out if
Low CPU utilization but slow ticks: a cause that waits outside the CPU, such as locks or I/O waits
ID mem-fragment · Primary owner Game team (Server development)
When repeated allocation and freeing chops free space into small pieces, the process holds far more memory than it actually uses.
Why Many threads allocate and free blocks of varying sizes over a long time → Effect Free space ends up scattered in small pieces that can’t be returned to the OS, so usage keeps growing like a leak → On screen The longer it runs, the slower it gets from swapping and memory shortage, until it gets killed
Use per-size memory pools and fragmentation-resistant allocators (jemalloc, mimalloc, and so on).
On the graph
Slow climb · Process memory (RSS)
Where to look
Run two servers on the same build, and on one of them either reduce the number of glibc arenas with the MALLOC_ARENA_MAX environment variable or switch to another allocator such as jemalloc. Compare RSS from pidstat -r over several days
Confirmed if
With similar player and entity counts, only the changed server’s RSS stops growing or grows much more slowly
Ruled out if
Still climbs the same way after switching allocators: memory that is never freed (mem-leak)
Check with
Infra tools (no game code needed)
Learn more
It looks just like a leak on the graph, yet heap analysis finds no leak site. The default Linux allocator (glibc) is especially bad on servers with many threads, and just switching allocators can cut memory use significantly.
Sources (3)
mallopt(3) — Linux manual pageLinux man-pages To reduce thread contention, glibc malloc creates arenas up to a multiple of the CPU count, and more arenas mean more memory use (limit with M_ARENA_MAX, or set the MALLOC_ARENA_MAX environment variable)
jemalloc memory allocatorjemalloc A general-purpose malloc implementation that emphasizes fragmentation avoidance and scalable concurrency
ID mem-numa · Primary owner Infra team (Server infrastructure)
On a server with two CPUs, using memory attached to the other CPU slows access down.
Why Threads and their memory sit on different CPU sockets → Effect Memory access slows down (1.5–2× depending on hardware) → On screen Same specs, but performance differs from process to process
Pin processes and their memory to one socket (numactl), on a two-socket machine run separate game server processes per socket.
On the graph
Outliers only · Tick time per process, memory per node
Where to look
Use numastat -p PID to see which NUMA node holds the game server process’s memory, check whether numa_miss and other_node in numastat are rising, and compare with the node of the CPU the process runs on
Confirmed if
Only the slow processes have most of their memory on a different node from the CPU they run on, and the gap disappears after restarting them with CPU and memory pinned to one node via numactl
Ruled out if
Slow even with the same node placement as the fast processes: another cause such as a noisy neighbor, CPU throttling, or that process’s load
Check with
Infra tools (no game code needed)
Sources (3)
What is NUMA?Linux kernel Memory in the same cell is faster and has higher bandwidth; memory in another (remote) cell is slower to access
numastat(8) — Linux manual pagenumactl numa_miss (allocated on a node other than the intended one) and other_node (allocated on this node by a process running on another node) counters; -p shows a process’s memory per node
ID dk-sync-log · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
If the game thread waits for the disk to finish every log line, the game stalls too whenever the disk is busy.
Why Combat and trade logs are written straight to a file from the game thread → Effect When a durable write (fsync) is required or the OS write buffer (page cache) hits its limit, a single write takes tens of ms while the disk is busy → On screen Hitches in log-heavy fights
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Log asynchronously (memory buffer + separate thread), reduce log volume, don’t fsync on the game thread.
Infra team action items
Run log rotation and compression at low I/O priority, keep logs on a different disk from data, monitor disk write latency.
On the graph
Random spikes · Server tick time, disk write latency
Where to look
Overlay w_await and aqu-sz from iostat -x 1 on tick time, and use perf trace -p PID --duration 10 to find write and fsync calls in the game server that took over 10 ms, along with their threads
Confirmed if
At the tick spikes, the game thread’s write and fsync calls take tens of ms, and disk write latency spikes at the same moment. Often lines up with log rotation or compression
Ruled out if
Ticks spike with no slow system calls on the game thread: another cause such as GC, locks, or tick overrun. Only the dedicated logging thread is slow: no effect on gameplay
Check with
Infra tools (no game code needed)
Learn more
Normally the OS accepts writes into memory (the page cache) first and flushes them to disk later, so a log line usually completes right away. Stalls happen when fsync demands a durable write, when backed-up writes exceed the limit and the OS blocks the write call, and when log files are rotated or compressed. That’s why it’s fine most of the time and spikes only at moments when the disk is busy.
Sources (5)
fsync(2) — Linux manual pageLinux man-pages fsync flushes modified data all the way to the disk (including the disk cache) and blocks until the device reports completion
Documentation for /proc/sys/vm/Linux kernel When backed-up (dirty) writes reach dirty_ratio, the writing process has to do the disk writeback itself
ionice(1) — Linux manual pageutil-linux A job at idle I/O priority gets disk time only when no other program is using the disk
iostat(1) — Linux manual pagesysstat -x: w_await (average time per write request, including time waiting in the queue), aqu-sz (average queue length, formerly avgqu-sz)
perf-trace(1) — Linux manual pageperf -p traces the system calls of a running process; --duration shows only calls that took longer than the given ms
ID dk-fsync · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (DB infrastructure)
Asking for data to be written to disk “for sure” takes 0.1 ms to tens of ms per request depending on the disk, and when requests pile up, the queue grows.
Why Scheduled saves and logout rushes send a flood of durable write requests → Effect The disk queue grows → On screen Lag at every save time, slow logouts and channel changes
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (DB infrastructure)
Game team action items
Batch saves (many saves in one fsync), spread out save times.
Infra team action items
Servers/OS: use server SSDs with power-loss protection, monitor disk queue length and fsync latency. DB hosts: if saves go to the DB, put the DB log disk on the same kind of SSD, monitor commit latency.
Ballpark numbers
The time per call varies by hardware, but roughly: server SSD (with power-loss protection) 0.1 ms, regular SSD 1 to a few ms, cloud disk 1–2 ms, HDD 10 ms or more. With one thread waiting on each call in turn, an HDD can’t manage even 100 per second.
On the graph
Periodic spikes · Disk queue length, flush and write latency
Where to look
Overlay f/s and f_await (flushes the disk handled and how long they took), plus w/s, aqu-sz, and w_await, from iostat -x 1 on scheduled save and logout times. Older sysstat shows aqu-sz as avgqu-sz. On cloud disks, check EBS VolumeQueueLength and VolumeAvgWriteLatency
Confirmed if
At every save time and logout rush, flush count and queue length spike together, and w_await and f_await reach several times normal. Saves and channel changes slow down at those moments
Ruled out if
Queue spikes at times unrelated to saves or logouts: backup or compression (dk-backup) or the IOPS limit (dk-iops). Slower at the same flush count: points to burst credit depletion (dk-burst)
Check with
Infra tools (no game code needed)
Sources (6)
fsync(2) — Linux manual pageLinux man-pages fsync empties even the disk cache and blocks until the device reports completion
Reliability (PostgreSQL Documentation)PostgreSQL Ordinary SATA disks and many SSDs have write caches that are lost on power failure; durable writes need a cache with battery or power-loss protection
iostat(1) — Linux manual pagesysstat -x: f/s and f_await (flush requests the disk handled and their average time), w/s, w_await, aqu-sz (formerly avgqu-sz)
Amazon CloudWatch metrics for Amazon EBSAWS VolumeQueueLength (requests waiting to complete), VolumeAvgWriteLatency (1-minute average write latency, Nitro instances)
ID dk-burst · Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)
Some cloud disks and small server sizes have burst credits that let them run faster than baseline for a while, so when a busy period drags on and the credits run out, speed drops suddenly.
Why Sustained use above baseline performance → Effect Burst credits run out and performance drops sharply to baseline → On screen Lag starts a few hours into every evening
Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)
Infra team action items
Servers/OS: use disks with provisioned performance (gp3, provisioned IOPS), alert on credit balance, also check the instance’s disk bandwidth burst limit and CPU credits. DB hosts: move DB disks, including managed databases, to provisioned performance too, alert on credit balance.
Ballpark numbers
An AWS gp2 100 GB disk normally gets 300 IOPS, bursts to 3,000, and lasts about 30 minutes on a full credit balance. gp3 has no credits and always gets 3,000. Small Azure Premium SSDs also burst on credits for up to 30 minutes.
On the graph
Hits a ceiling · IOPS, burst credit balance
Where to look
In CloudWatch, EBS BurstBalance (gp2, st1, sc1), the instance’s EBSIOBalance% and EBSByteBalance% (some instances that burst), and CPUCreditBalance on burstable instances. On Azure, burst credit usage metrics such as Data Disk Used Burst IO Credits Percentage
Confirmed if
From the moment the balance drops near 0, IOPS (VolumeReadOps, VolumeWriteOps) flattens at baseline, and VolumeQueueLength and lag rise together. Starts after the peak has lasted a few hours
Ruled out if
All balances are healthy but IOPS is flat: a fixed volume or instance limit (dk-iops)
Check with
Infra tools (no game code needed)
Learn more
Even with a healthy disk, small virtual servers have a burst limit on the instance’s own disk bandwidth (for example, at least 30 minutes a day), which produces the same pattern. Low-cost servers that run on CPU credits also slow down to baseline performance once the credits run out.
Sources (7)
Amazon EBS General Purpose SSD volumesAWS gp2 baseline is 3 IOPS per GiB (minimum 100), bursting to 3,000 IOPS on I/O credits; 5.4 million credits last at least 30 minutes. gp3 always gets 3,000 IOPS with no burst
Managed disk burstingMicrosoft Azure Premium SSD P20 and smaller use credit-based bursting; a full credit balance gives 30 minutes at max burst speed
Amazon EBS-optimized instance typesAWS Some instances sustain maximum EBS performance for only 30 minutes once every 24 hours, then return to baseline
Standard mode for burstable performance instancesAWS In standard mode, a burstable instance that runs out of CPU credits lowers CPU utilization to the baseline level (gradually, without a sudden drop)
Amazon CloudWatch metrics for Amazon EBSAWS BurstBalance: remaining I/O credits for gp2 and throughput credits for st1 and sc1 (%); VolumeReadOps, VolumeWriteOps, VolumeQueueLength
CloudWatch metrics that are available for your instancesAWS EBSIOBalance% and EBSByteBalance%: remaining EBS credits on some instances that burst for 30 minutes once every 24 hours; CPUCreditBalance: remaining CPU credits on burstable instances
Disk metricsMicrosoft Azure Disk and VM burst credit usage (5-minute intervals), such as Data Disk Used Burst IO Credits Percentage
ID dk-iops · Primary owner Infra team (Server infrastructure) · Also Game team (Server development), Infra team (DB infrastructure)
When requests exceed what the disk can handle per second, the queue grows and latency explodes.
Why Read and write requests approach the disk’s capacity → Effect The queue grows (usually exploding above 90% utilization) → On screen Slow saves and loading; freezes if the calls are blocking
Primary owner Infra team (Server infrastructure) · Also Game team (Server development), Infra team (DB infrastructure)
Game team action items
Merge requests, cache frequently read data, make disk access asynchronous so the game thread never waits on the disk.
Infra team action items
Servers/OS: use faster disks, check disk bandwidth and IOPS limits per instance type, alert on disk utilization and queue, run large file copies during quiet hours. DB hosts: alert on IOPS and throughput utilization for DB disks too, check the disk limits of the DB instance size.
Ballpark numbers
HDD about 150 IOPS, SATA SSD tens of thousands, NVMe hundreds of thousands. The default cloud disk (AWS gp3) gets 3,000 IOPS and 125 MiB per second. The per-second throughput limit is separate from IOPS, and when a large file copy fills it, even small writes get stuck behind it.
On the graph
Hits a ceiling · IOPS, disk queue length
Where to look
r/s and w/s, rkB/s and wkB/s, aqu-sz, and r_await and w_await from iostat -x 1. In the cloud, EBS VolumeReadOps, VolumeWriteOps, and VolumeQueueLength, the limit-exceeded checks VolumeIOPSExceededCheck and VolumeThroughputExceededCheck, and on the instance side InstanceEBSIOPSExceededCheck and InstanceEBSThroughputExceededCheck
Confirmed if
Requests per second or throughput flatten at the limit while aqu-sz and await shoot up together. In the cloud, the exceeded-check metric reads 1
Ruled out if
Even at 100% %util, low await means there may still be headroom. On SSDs and RAID that process requests in parallel, %util doesn’t mean the limit has been reached. High await without hitting the limit: latency of the disk itself (dk-hdd) or fsync (dk-fsync)
Check with
Infra tools (no game code needed)
Learn more
In the cloud, each server size (instance type) has its own disk bandwidth and IOPS limits, separate from the disk’s limits. Even with an expensive disk attached, a small server gets capped at the instance limit.
Sources (8)
Exos X18 Data SheetSeagate 4K random reads on a 7,200 rpm server HDD: 170 IOPS (QD16)
D3-S4520 SSDSolidigm Server SATA SSD 4 KB random read/write up to 92K/48K IOPS
iostat(1) — Linux manual pagesysstat -x: r/s and w/s, rkB/s and wkB/s, aqu-sz (formerly avgqu-sz), r_await and w_await, %util. On RAID and modern SSDs that process requests in parallel, %util doesn’t indicate the performance limit
Amazon CloudWatch metrics for Amazon EBSAWS VolumeIOPSExceededCheck and VolumeThroughputExceededCheck: 1 if the volume tried to exceed its IOPS or throughput limit (Nitro instances); VolumeQueueLength
ID dk-full · Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure), Game team (Server development)
When logs and dumps pile up and fill the disk, writes fail, and without safeguards the server crashes.
Why Logs, dumps, and temp files pile up to 100% → Effect Writes fail. Crash if there’s no error handling, failed saves if there is → On screen Disconnects, rolled-back progress
Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure), Game team (Server development)
Game team action items
Handle write failures (retry the save and raise an alert so the server doesn’t crash), cut unnecessary logs and dumps.
Infra team action items
Servers/OS: log rotation, capacity alerts, separate log and data disks. DB hosts: watch that DB transaction logs (WAL, binlog) don’t pile up because replication stopped or a log backup was missed.
On the graph
Slow climb · Disk usage
Where to look
Usage from df -h and inode usage from df -i, plus ENOSPC errors in server and DB logs. For the DB: slots whose active is false in PostgreSQL pg_replication_slots, file count and size from MySQL SHOW BINARY LOGS, log_reuse_wait_desc in SQL Server sys.databases, and FreeStorageSpace on RDS
Confirmed if
Usage climbs steadily over several days, the moment it hits 100% lines up with crashes or failed saves, and ENOSPC shows up in the logs
Ruled out if
Writes fail with plenty of space left: another cause such as permissions or a file size limit
Check with
Infra tools (no game code needed)
Learn more
A DB’s transaction logs (WAL, binlog, and so on) are never deleted and keep piling up if a replica stops or a log backup is missed. When that disk fills, every write on the DB stops, and saves and trades fail all at once.
Troubleshoot a full transaction log (SQL Server Error 9002)Microsoft SQL Server When the log is full, the DB is read-only and can’t be modified; missed log backups, replication lag, and long transactions are common causes that block log truncation; see what is blocking it in log_reuse_wait_desc of sys.databases
Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)
Infra team action items
Servers/OS: lower the I/O priority of backup, compression, and scan jobs, stagger their times. DB hosts: take backups from a replica.
On the graph
Periodic spikes · Disk utilization, disk wait time
Where to look
Overlay %util, await, and aqu-sz by day from the past few days of sar -d history (daily files in /var/log/sa; sadc must collect disk data with -S DISK), find the processes with the highest kB_rd/s and kB_wr/s at that time with pidstat -d 1, and match them against cron and systemd timer schedules
Confirmed if
await and %util spike at the same time every day, and backup, compression, or scan processes account for most disk reads and writes at that time
Ruled out if
Spikes at a different time each day: a scheduled job is unlikely. The game server itself does most of the I/O at that time: points to saves or logging (dk-fsync, dk-sync-log)
Check with
Infra tools (no game code needed)
Sources (4)
ionice(1) — Linux manual pageutil-linux A job run in the idle class gets disk time only when no other program is using the disk
sar(1) — Linux manual pagesysstat -d: per-device await, aqu-sz, and %util from daily history files (default /var/log/sa); disk data must be collected with sadc’s -S DISK option
ID dk-lazy-load · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
If the server reads dungeon or map data from disk the first time it’s requested, everyone freezes for that tick.
Why Someone enters a dungeon or area for the first time → Effect The server reads the data from disk on the game thread → On screen Everyone on that server freezes briefly
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Preload at server startup, load asynchronously.
Infra team action items
For servers just created from a snapshot, warm up the disk before putting them into service (read every block once) or use fast snapshot restore.
On the graph
Random spikes · Server tick time, disk reads
Where to look
Match freeze times to first-entry records for dungeons and areas in the game server log, and at those moments check the game server’s disk reads (kB_rd/s from pidstat -d) and slow read and open calls with perf trace --duration. On a newly launched cloud server, compare EBS VolumeAvgReadLatency with older servers
Confirmed if
Freezes only on the first entry, with no freeze on the second entry to the same place. During the freeze, the game thread is waiting on a file read
Ruled out if
Freezes just the same in areas already loaded: another cause such as tick overrun or GC
Check with
Game server or client logs and metrics
Learn more
In the cloud, a server just created from a snapshot (a copy of a disk) fetches every block from remote storage the first time it reads it, so reads are far slower than usual. Suspect this if first entry takes unusually long only on servers newly launched by autoscaling.
Sources (5)
Initialize Amazon EBS volumesAWS Volumes created from snapshots have higher latency and lower performance while blocks are fetched from S3; initialize them in advance by reading every block with dd or fio
Amazon EBS fast snapshot restoreAWS Fast snapshot restore provides volumes that are fully initialized at creation, removing first-access latency
ID dk-coredump · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
When the server crashes, writing several GB of memory to disk can delay the restart by several minutes.
Why A server crash writes all of memory to a file → Effect No restart until several GB have been written → On screen After the server dies and players disconnect, they can’t connect again for a long time
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Consider small dumps holding only the needed memory (minidumps), fix the cause of the crash.
Infra team action items
Limit dump size (OS core dump settings), use fast disks, decouple the restart from the dump (compress and upload dumps separately after the restart).
On the graph
Mass disconnect · Connection count, server restart times
Where to look
Line up the crash time, core file size (coredumpctl list and info, or the file where core_pattern points), the time the file finished writing, and the time the service came back up, and check wkB/s from iostat -x during that window
Confirmed if
After the crash, disk writes stay near the limit while a multi-GB core file is written, and the restart begins only after the write finishes
Ruled out if
Restart is still slow with core dumps off or finished small: points to the server startup process, such as map loading or a DB cold cache (db-cold-cache)
Check with
Infra tools (no game code needed)
Sources (4)
core(5) — Linux manual pageLinux man-pages RLIMIT_CORE caps core file size, coredump_filter selects which memory regions to include, and core dumps can be piped to a program for separate handling
Minidump FilesMicrosoft A minidump holds only a useful subset of crash dump information, so it is fast and small
coredumpctl(1) — Linux manual pagesystemd list: core dumps recorded in the journal (TIME is the crash time reported by the kernel); info: details for each dump and the size written to disk
Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure), Game team (Server development)
Game team action items
Design around sequential writes.
Infra team action items
Servers/OS: replace with SSDs (starting with storage disks that see the most scattered reads and writes). DB hosts: replace the DB disks with the most scattered reads and writes with SSDs first.
On the graph
Always high · Disk read/write latency (r_await, w_await)
Where to look
Check whether it’s a rotating disk (HDD) with lsblk -d -o NAME,ROTA, and look at r/s, w/s, r_await, and w_await from iostat -x 1. For virtual servers, check the disk type in the cloud or storage specs
Confirmed if
A rotating disk, with r_await and w_await always at a few ms to tens of ms even at only tens to about a hundred requests per second
Ruled out if
High latency on an SSD: points to queue saturation (dk-iops) or burst credit depletion (dk-burst)
ID db-no-index · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Without an index, finding the rows that match a condition means reading the entire table (a full table scan).
Why A new feature ships with a search on a condition that has no index → Effect Scanning millions of rows makes a single query take hundreds of ms to several seconds → On screen Mailbox and trade history load slowly, and tied-up connections make other requests wait too
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Review the query plan of every new query before deploying, add indexes, check that modifying queries (UPDATE, DELETE) use an index too.
Infra team action items
Watch the slow query log, find full-table-scan queries and share them with the game team, add indexes in production with online methods that hold locks only briefly.
Ballpark numbers
With an index, a few ms. Without one, the query slows down in proportion to data size, which on a large table means hundreds to tens of thousands of times slower.
On the graph
Step change · DB query latency, rows read
Where to look
MySQL: Rows_examined and Rows_sent in the slow query log (with log_queries_not_using_indexes on, queries that don’t use an index are logged too), SUM_NO_INDEX_USED and SUM_ROWS_EXAMINED in performance_schema events_statements_summary_by_digest, then run EXPLAIN. PostgreSQL: seq_scan and seq_tup_read in pg_stat_user_tables, then run EXPLAIN
Confirmed if
A query that appeared after the deploy reads thousands of times more rows (Rows_examined) than it returns (Rows_sent), and EXPLAIN shows a full table scan (MySQL type ALL, PostgreSQL Seq Scan). seq_tup_read on a large table climbs steeply from the deploy time
Ruled out if
Slow even while using an index: lock waits (db-hot-row, db-ddl-lock) or a query plan change (db-plan-flip). A full scan of a small table can be normal
Check with
Infra tools (no game code needed)
Learn more
Reads aren’t the only thing that slows down. A modifying query (UPDATE, DELETE) without an index can, depending on the DB, lock every row it scans and block saves for unrelated players.
Sources (8)
How MySQL Uses IndexesMySQL Without an index, the whole table is read starting from the first row, and the bigger the table, the higher the cost
The Slow Query LogMySQL Logs queries that exceed long_query_time (default 10 seconds); queries that don’t use indexes can also be logged separately
CREATE INDEX (PostgreSQL Documentation)PostgreSQL Building with CONCURRENTLY creates the index without blocking writes; a regular build blocks writes until it finishes
Statement Summary TablesMySQL events_statements_summary_by_digest: SUM_NO_INDEX_USED (times executed without an index) and SUM_ROWS_EXAMINED per normalized query
EXPLAIN Output FormatMySQL type ALL means a full table scan, usually avoided by adding an index
ID db-hot-row · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
When everyone tries to modify the same row (a guild vault, a popular auction house item, a server-wide counter), only one request at a time gets the lock.
Why An event or a popular item concentrates updates on the same row → Effect Requests wait until they get the lock → On screen Failed trades, “Please try again later” messages, timeouts
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Split the row (sharded counters), keep transactions short, aggregate in memory and apply in one write.
Infra team action items
Monitor row lock wait time and count, find the rows where contention concentrates, and share them.
Ballpark numbers
If one request holds the lock for 10 ms, that row can be modified at most 100 times a second. A round trip to another server inside the transaction cuts that number further.
On the graph
Rises with load · Row lock wait count and time
Where to look
MySQL: growth of Innodb_row_lock_waits and Innodb_row_lock_time plus Innodb_row_lock_current_waits, and sys.innodb_lock_waits to find who is waiting on whom. PostgreSQL: sessions whose wait_event_type is Lock in pg_stat_activity and requests whose granted is false in pg_locks; turning on log_lock_waits (off by default) logs long lock waits
Confirmed if
Lock waits climb steeply with events and player count, and most waiting requests point to the same row (same key) in the same table
Ruled out if
Waits spread evenly across many tables and rows: points to disk or CPU saturation. One session holds a lock for a long time without releasing it: a long-open transaction (db-long-tx)
Check with
Infra tools (no game code needed)
Sources (7)
InnoDB LockingMySQL When one transaction locks a row (index record), other transactions can’t modify that row and wait
How to Minimize and Handle DeadlocksMySQL Recommends keeping transactions small and short and committing right after related changes to reduce conflicts
Server Status VariablesMySQL Innodb_row_lock_waits and Innodb_row_lock_time give the count and total time of row lock waits; Innodb_row_lock_current_waits gives how many are waiting right now
ID db-deadlock · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
When two transactions (groups of DB operations processed as one unit) each wait for a row the other has locked, the DB forcibly cancels one of them.
Why Trade A locks in item→currency order, trade B in currency→item order → Effect The DB detects the deadlock and rolls back one side → On screen Trades and crafting fail now and then, items revert
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Use a consistent lock order, keep transactions short, retry automatically on failure.
Infra team action items
Keep deadlock detection on, collect and share deadlock records, lower the lock wait timeout (default 50 seconds) on MySQL servers with detection turned off.
Ballpark numbers
Detection is almost instant in MySQL (InnoDB), takes 1 second by default in PostgreSQL, and up to about 5 seconds in SQL Server. Both requests stay stuck in the meantime. On a MySQL server with detection turned off because of very high concurrency, requests wait until the lock wait timeout (default 50 seconds).
On the graph
Random spikes · Deadlock count, failed trade count
Where to look
MySQL: LATEST DETECTED DEADLOCK in SHOW ENGINE INNODB STATUS (the most recent one only), every deadlock in the error log with innodb_print_all_deadlocks on, and lock_deadlocks in INFORMATION_SCHEMA.INNODB_METRICS. PostgreSQL: deadlocks in pg_stat_database. SQL Server: xml_deadlock_report from the system_health session, which is on by default. Error codes on the game server side: MySQL 1213, PostgreSQL 40P01, SQL Server 1205
Confirmed if
Deadlock count rises at the times trades and crafting fail, and the two recorded transactions lock the same tables in opposite order
Ruled out if
Failures with an unchanged deadlock count: lock wait timeout exceeded (MySQL error 1205) or a hot row (db-hot-row)
Check with
Infra tools (no game code needed)
Sources (10)
InnoDB Startup Options and System VariablesMySQL With detection on (the default), InnoDB detects deadlocks immediately and rolls one back; innodb_lock_wait_timeout defaults to 50 seconds
Deadlock DetectionMySQL At very high concurrency, detection itself can slow things down, so it is sometimes turned off in favor of the lock wait timeout
Deadlocks guideMicrosoft SQL Server Deadlock checks run every 5 seconds by default, dropping to as little as 100 ms when deadlocks are frequent; the system_health session, on by default, collects xml_deadlock_report; the victim gets error 1205
How to Minimize and Handle DeadlocksMySQL Always modify multiple rows and tables in the same order, retry on failure, log every deadlock with innodb_print_all_deadlocks
InnoDB Standard Monitor and Lock Monitor OutputMySQL LATEST DETECTED DEADLOCK: the two transactions in the most recent deadlock, the locks they held and waited for, and which one was rolled back
ID db-pool · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
The number of connections open to the DB is fixed, so when slow queries hold connections, every other request waits.
Why Slow queries or a flood of requests put every connection in use → Effect New requests wait until a connection frees up → On screen Infinite loading at login, slow saves, timeouts
Right after login or maintenance, When crowds gather
Owner
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Remove slow queries, tune pool size and wait timeout (don’t just keep growing the pool), separate pools per feature.
Infra team action items
Check the DB’s max connections and CPU/IOPS headroom, confirm that server count × pool size stays within max connections before adding servers or autoscaling, add connection wait and lock wait metrics to monitoring.
Ballpark numbers
Estimate the connections you need as “requests per second × time each request holds a connection”. At 2,000 requests a second and 5 ms each, 10 connections are busy on average at all times. To handle bursts, pools are usually sized at two to three times that. If queries slow to 150 ms, the same load needs 300.
On the graph
Hits a ceiling · DB connections in use, connection wait time
Where to look
Count connection states per game server on the DB side. MySQL: Host, Command (Sleep for idle connections), and Time in SHOW PROCESSLIST, plus Threads_connected, Threads_running, and refused connections in Connection_errors_max_connections. PostgreSQL: pg_stat_activity grouped by client_addr and state. If the game server’s connection pool library exports wait count and wait time, view those too
Confirmed if
All of one game server’s connections, up to the pool size, are running queries with 0 idle, and logins and saves wait meanwhile. Or the DB’s total connection count hits max_connections and new connections are refused
Ruled out if
Slow even with plenty of idle connections: latency of the queries themselves (db-no-index, db-hot-row) or saturated DB resources
Check with
Infra tools (no game code needed)
Learn more
Blindly growing the pool only adds DB CPU load and lock contention, and everyone slows down together. Also, if server count × pool size exceeds the DB’s max connections, newly added or restarted servers can’t even open a connection. This is common right after autoscaling or maintenance.
Sources (6)
Number Of Database ConnectionsPostgreSQL Once DB resources are fully used, adding connections actually lowers throughput; matching active connections to resources and queueing the rest gives better latency and throughput
Too many connectionsMySQL When all max_connections are in use, new connections are refused with a Too many connections error
Server Status VariablesMySQL Threads_connected, Threads_running, Connection_errors_max_connections (connections refused because max_connections was reached)
ID db-replica-lag · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Writes go to the primary and reads come from replicas, so when a replica falls behind, data that was just written isn’t visible yet.
Why A burst of writes on the primary puts replicas several seconds behind → Effect Reading just-saved data from a replica finds it missing → On screen An item you just bought doesn’t show up, marketplace prices are stale, duplicate-reward bugs
Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Game team action items
Read just-written data from the primary, check and grant rewards in one transaction on the primary (block duplicates with a unique key or a conditional UPDATE).
Infra team action items
Alert on replication lag, give replicas specs equal to or better than the primary and enable parallel replication, run bulk deletes in small chunks, manage long-running aggregate queries on replicas.
On the graph
Rises with load · Replication lag (seconds)
Where to look
MySQL: Seconds_Behind_Source from SHOW REPLICA STATUS on the replica (SHOW SLAVE STATUS on versions before 8.0.22). PostgreSQL: write_lag, flush_lag, and replay_lag in pg_stat_replication on the primary. RDS: ReplicaLag
Confirmed if
Lag is several seconds or more at the times of “it’s not showing up” reports, and things look normal once the lag clears. Lag grows during write bursts, bulk deletes, or long aggregate queries on the replica
Ruled out if
Lag near 0 but data still not showing up: points to the game server’s cache or sync
Check with
Infra tools (no game code needed)
Learn more
Replicas can fall behind even without heavy writes. A single bulk delete that took 10 minutes on the primary puts a replica that far behind while it replays, and long-running aggregate queries on a replica also slow its catch-up.
Sources (6)
SHOW REPLICA STATUS StatementMySQL Seconds_Behind_Source: time elapsed since the event the replica is currently applying was written on the primary (replication lag)
Replica Server Options and VariablesMySQL replica_parallel_workers lets multiple threads apply transactions in parallel (default 4; 0 means a single thread applies them in order)
Log-Shipping Standby Servers (PostgreSQL Documentation)PostgreSQL Streaming replication is asynchronous by default, so there is a delay between commit and the replica applying it (usually under 1 second if the replica can keep up)
The Cumulative Statistics System (PostgreSQL Documentation)PostgreSQL write_lag, flush_lag, and replay_lag in pg_stat_replication: time from the primary writing WAL until the replica reports it has written, flushed to disk, and applied it
ID db-checkpoint · Primary owner Infra team (DB infrastructure)
Queries slow down at the moments the DB periodically writes its accumulated in-memory changes to disk in bulk.
Why Changes pile up and are periodically written to disk → Effect The disk gets busy at that moment and queries slow down → On screen Saves and loading slow down periodically
Spread checkpoints out in small, even steps, size the transaction log (redo log, WAL) generously, use fast disks.
On the graph
Periodic spikes · DB query latency, disk write volume
Where to look
PostgreSQL: checkpoint times and buffers written from the log_checkpoints log (on by default in recent versions), checkpoint counts (num_timed and num_requested in pg_stat_checkpointer on 17 and later, checkpoints_timed and checkpoints_req in pg_stat_bgwriter on 16 and earlier), and checkpoint_warning messages. MySQL: the gap between Log sequence number and Last checkpoint at in the LOG section of SHOW ENGINE INNODB STATUS. Overlay the server’s disk write volume and write latency
Confirmed if
Query latency spikes line up with checkpoint times, with disk write volume and write latency spiking at the same moments. In PostgreSQL, far more requested checkpoints (num_requested) than timed ones (num_timed) means WAL keeps hitting max_wal_size and checkpoints come early
Ruled out if
Spikes on a cycle unrelated to checkpoint times: backups or batch jobs (dk-backup, db-batch)
Check with
Infra tools (no game code needed)
Learn more
If the transaction log that records changes (the redo log in MySQL, WAL in PostgreSQL) is too small, the DB has to rush a checkpoint every time the log fills, and write throughput drops sharply for short periods.
Sources (7)
WAL Configuration (PostgreSQL Documentation)PostgreSQL By default, a checkpoint runs every 5 minutes or every 1 GB of WAL (max_wal_size) and is expensive because it writes all dirty pages. checkpoint_completion_target spreads the writes out to avoid I/O bursts. If checkpoints come closer together than checkpoint_warning, the log suggests raising max_wal_size
Configuring Buffer Pool FlushingMySQL When the redo log fills, a sharp checkpoint briefly drops throughput; adaptive flushing spreads the writes out evenly
ID db-cold-cache · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
After a DB restart, the memory cache is empty, so for a while every lookup reads from disk.
Why The DB restarts for maintenance → Effect Frequently used data isn’t in memory, so it’s read from disk → On screen Logins and loading are slow for a while right after maintenance
Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Game team action items
Open gradually (use a login queue to ramp up the number of players logging in step by step).
Infra team action items
Warm the cache after a restart (check the buffer pool save/restore settings), pre-read the disk too on a DB restored from a snapshot.
On the graph
Surge after opening · Disk reads, buffer cache hit ratio
Where to look
MySQL: the ratio of Innodb_buffer_pool_reads (reads that missed the buffer pool and went to disk) to Innodb_buffer_pool_read_requests, and warm-up progress in Innodb_buffer_pool_load_status. PostgreSQL: blks_read and blks_hit in pg_stat_database. Also check disk reads on the DB server
Confirmed if
Right after the restart, disk reads spike and the hit ratio is low, recovering over time, and logins and loading are slow during that window
Ruled out if
Hit ratio normal but slow right after maintenance: a login storm and N+1 queries (db-login-storm) or the connection pool (db-pool)
Check with
Infra tools (no game code needed)
Learn more
MySQL saves the list of buffer pool pages at shutdown and reloads them in the background at startup, but filling the pool takes time. If a cloud DB was restored from a snapshot (a copy of a disk), the disk itself is also slow for every block read for the first time, so the slowdown lasts even longer.
Saving and Restoring the Buffer Pool StateMySQL Saves the list of recently used pages (25% by default) at shutdown and reads them back at startup to shorten warm-up after a restart; both are on by default
Initialize Amazon EBS volumesAWS Volumes created from snapshots have higher latency and lower performance until all blocks have been fetched
Server Status VariablesMySQL Innodb_buffer_pool_reads (logical reads that missed the buffer pool and read straight from disk), Innodb_buffer_pool_read_requests, Innodb_buffer_pool_load_status (warm-up progress)
ID db-login-storm · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
If loading one character takes dozens of separate queries, tens of thousands of simultaneous logins turn into millions of queries.
Why Character loading queries items, skills, and quests one by one → Effect Simultaneous logins right after maintenance make the query count explode → On screen Infinite loading at login, and even saves for players already in the game get held up
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Fetch in batched queries, use a login queue and caching, check how many queries ORM lazy loading generates.
Infra team action items
Pull a ranking of the most frequently called queries and share it, monitor query and connection counts during the login window right after maintenance.
On the graph
Surge after opening · DB queries per second, login count
Where to look
Overlay the login count right after maintenance with the DB’s queries per second (growth of Questions in MySQL) and compute queries per login. Pull the most frequently called queries from COUNT_STAR in MySQL events_statements_summary_by_digest or calls in PostgreSQL pg_stat_statements
Confirmed if
Dozens of queries per login, and the top queries are short queries of the same shape that look up by a single character ID. If queries per login went up after a patch, that patch is the starting point
Ruled out if
Few queries per login but each one is slow: cold cache (db-cold-cache) or indexes (db-no-index)
Check with
Infra tools (no game code needed)
Learn more
Lazy loading in an ORM (a library that builds DB queries for you) generates queries like these without developers even noticing. On a dev server with only a few characters it goes unnoticed, and it first shows up with simultaneous logins on live servers.
Sources (5)
Efficient Querying.NET ORM lazy loading creates the N+1 problem, sending one more query per item and badly hurting performance; recommends loading in one batch (eager loading)
ID db-batch · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Running ranking aggregation, mass mail sends, or old-data cleanup during live service ties up locks and the disk.
Why Bulk jobs run during service hours → Effect Wide-range locks, disk and CPU tied up → On screen Failed trades and saves, slow loading at certain times of day
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Split jobs into small chunks and run them a little at a time, run aggregation on a replica.
Infra team action items
Provide a replica for aggregation, schedule batches for quiet hours, watch for lock escalation and gap lock waits.
On the graph
Periodic spikes · DB query latency, lock waits
Where to look
Find long queries running at the time of the lag. MySQL: slow query log. PostgreSQL: query_start and query in pg_stat_activity. Match them against lock wait metrics at the same time and the batch schedule (cron, DB event scheduler). SQL Server: record lock escalation with the lock_escalation extended event
Confirmed if
Large UPDATE, DELETE, or aggregation queries run at the same time every time, and lock waits and disk utilization rise together meanwhile
Ruled out if
No long queries at that time: checkpoints (db-checkpoint) or a server backup (dk-backup)
Check with
Infra tools (no game code needed)
Learn more
SQL Server converts to a table lock when one statement holds more than about 5,000 row locks (lock escalation). At that moment, every request using the same table stops. MySQL, with default settings, also locks the gaps between rows when modifying by a range condition (gap locks), blocking inserts of new rows.
Sources (4)
Transaction Locking and Row Versioning GuideMicrosoft SQL Server Lock escalation when one statement holds 5,000 or more locks on one table (or index); recorded with the lock_escalation extended event
InnoDB LockingMySQL At InnoDB’s default isolation level, REPEATABLE READ, searches and scans use next-key locks, so gap locks block inserts of new rows into those gaps
The Slow Query LogMySQL Logs queries exceeding long_query_time with execution time (Query_time), lock time (Lock_time), and rows read
ID db-failover · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
When the primary DB dies, writes stop while it fails over to a standby, and the last data that hadn’t been replicated yet can be lost.
Why The primary DB fails and a standby is promoted → Effect No writes for seconds to minutes during the switch; with asynchronous replication, unreplicated data may be lost → On screen Every save fails for a moment, items and XP roll back
Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Game team action items
Make saves retryable, configure the connection pool and DNS cache to drop broken connections quickly and reconnect to the new address, verify reconnection during failover drills.
Infra team action items
Use synchronous or semi-synchronous replication (at the cost of write latency), run failover drills, monitor failover time and replication lag.
Ballpark numbers
Automatic failover on a managed DB usually takes tens of seconds to 2 minutes. With asynchronous replication, you can lose recent saves equal to the replication lag (under 1 second to a few seconds).
On the graph
Mass disconnect · DB connection count, write error count
Where to look
Put the DB’s failover records (on RDS, events RDS-EVENT-0013 failover started and RDS-EVENT-0049 failover completed; on self-managed DBs, promotion logs) and the game server’s DB connection count and connection error count on one graph. With asynchronous replication, also check replication lag just before the failure (RDS ReplicaLag, replay_lag in PostgreSQL pg_stat_replication)
Confirmed if
Save failures cluster in one window that lines up with the span between failover start and completion. The amount rolled back is close to the replication lag just before the failure. A game server whose errors continue after the failover finishes is still using connections opened to the old address
Ruled out if
Disconnects at times with no failover record: points to the network or DB overload
Failing over a Multi-AZ DB instance for Amazon RDSAWS Multi-AZ failover usually takes 60–120 seconds; connections must be reestablished afterward, and a JVM DNS cache TTL of 60 seconds or less is recommended
High availability for Amazon AuroraAWS Reads and writes fail during the outage, and recovery usually takes under 60 seconds (often under 30)
Semisynchronous ReplicationMySQL With asynchronous replication, committed transactions may be missing from the replica when the primary dies; semi-synchronous replication narrows this by waiting for one replica to acknowledge receipt, at the cost of higher latency
ID db-save-interval · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
If the server saves only once every few minutes to reduce load, progress is lost when the server dies in between.
Why Character state is saved once every few minutes → Effect A server crash or outage hits in between → On screen After reconnecting, the character is back to where it was minutes ago (rollback)
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Save important events (trades, rare drops) immediately, keep a change log.
Infra team action items
Confirm the DB has enough IOPS and CPU headroom for the extra writes a shorter save interval brings.
On the graph
Mass disconnect · Connection count, rollback reports
Where to look
Line up crash and outage times with the last save time of the characters that reported rollbacks (the game server’s save log or the DB’s last-modified column)
Confirmed if
The point the character reverted to matches the last save before the crash, and the lost time is shorter than the save interval
Ruled out if
Reverted even though the game server log says the save completed: data loss from a DB failover (db-failover) or a stale value read from a replica (db-replica-lag)
Check with
Game server or client logs and metrics
Sources (2)
Asynchronous Commit (PostgreSQL Documentation)PostgreSQL Batching writes and flushing them late raises throughput, but the most recent transactions can be lost in a failure (the same tradeoff)
Redis persistenceRedis Taking RDB snapshots every few minutes means accepting the loss of the last few minutes of data on an abnormal shutdown
ID db-cache-stampede · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
When cache entries for popular data expire at the same time, thousands of requests hit the DB all at once.
Why Popular data stored in Redis or a similar cache expires at the same time → Effect Requests trying to rebuild the same data rush to the DB all at once → On screen DB overload makes one feature after another slow down or freeze
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Randomize expiry times, let a single request refresh the value while the rest keep using the old value.
Infra team action items
Set up Redis replicas and automatic failover so the cache doesn’t empty entirely on a restart or failure, confirm the DB has enough headroom to survive an empty cache.
On the graph
Periodic spikes · Cache hit ratio, DB queries per second
Where to look
Overlay keyspace_hits and keyspace_misses (hit ratio), expired_keys, and restarts (uptime_in_seconds) from Redis INFO on the DB’s queries per second, and count how many copies of the same query run at once on the DB at that moment (MySQL SHOW PROCESSLIST, PostgreSQL pg_stat_activity)
Confirmed if
At the moment cache misses shoot up, the DB query count spikes with them, and most concurrent queries are the same query reading the same data. Lines up with the expiry cycle of popular keys or a Redis restart
Ruled out if
Cache misses normal but only DB queries rise: a login storm (db-login-storm) or a batch job (db-batch)
Check with
Infra tools (no game code needed)
Learn more
The same thing happens when a Redis restart or failure empties the whole cache. The more an architecture relies on the cache and keeps the DB small, the greater the risk.
Sources (6)
Scaling Memcache at Facebook (NSDI '13)USENIX When a hot key is invalidated, many reads rush to the DB (thundering herd); prevented with leases (only one client refreshes) and by returning stale values; clusters with an empty cache are warmed up separately
Optimal Probabilistic Cache Stampede PreventionVLDB Endowment When a popular item expires, many requests regenerate it at once (cache stampede); prevented by probabilistically refreshing early, before expiry
INFORedis keyspace_hits and keyspace_misses (successful and failed key lookups), expired_keys (number of expired keys), uptime_in_seconds (time since startup)
SHOW PROCESSLIST StatementMySQL The statement each session is running (Info) and the time spent in its current state (Time, seconds)
ID db-long-tx · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
When a transaction stays open for a long time, it keeps holding its locks and the DB can’t clean up (purge) old versions of data, so everything gradually slows down.
Why A transaction stays open while waiting for another server’s response, or a long aggregate query runs on the primary during service → Effect Its locks are never released, and old row versions awaiting cleanup keep piling up → On screen Features that use those rows time out, and saves and lookups slow down across the board over several hours
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Don’t wait on network calls or user input inside a transaction, run aggregate queries on a replica.
Infra team action items
Alert on and kill long-open transactions, provide a replica for aggregation, watch undo log and dead row growth.
On the graph
Slow climb · Undo log length (History list length), dead row count
Where to look
MySQL: find the oldest transaction by trx_started in INFORMATION_SCHEMA.INNODB_TRX, and check History list length (undo log not yet purged) in the TRANSACTIONS section of SHOW ENGINE INNODB STATUS. PostgreSQL: xact_start in pg_stat_activity, sessions whose state is idle in transaction, and n_dead_tup in pg_stat_user_tables
Confirmed if
A transaction minutes to hours old exists, History list length or n_dead_tup keeps climbing while it’s open, then falls as cleanup (purge, VACUUM) runs after that transaction ends
Ruled out if
Slow across the board with no old transactions: checkpoints (db-checkpoint) or the disk
Check with
Infra tools (no game code needed)
Learn more
The DB keeps old versions so readers can see data as it was before a change (MVCC). These records can only be removed once the oldest transaction ends, so if one transaction stays open for hours, MySQL piles up undo logs and PostgreSQL piles up dead rows (dead tuples) that VACUUM can’t clean up. In SQL Server, the transaction log can’t shrink and may fill the disk.
Sources (7)
InnoDB Multi-VersioningMySQL While a transaction that can see old versions remains, update undo logs can’t be discarded and the rollback segment grows; recommends committing often, even for read-only transactions
Routine Vacuuming (PostgreSQL Documentation)PostgreSQL Old row versions can’t be removed while other transactions can still see them; long-open transactions must be ended or their sessions terminated
Purge ConfigurationMySQL Purge cleans up the list of undo logs from committed transactions (history list); the backlog shows as History list length in the TRANSACTIONS section of SHOW ENGINE INNODB STATUS
ID db-redis-block · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Redis processes commands one at a time, so a single slow command blocks every request behind it.
Why A full key search with KEYS in production, or reading or deleting a ranking or list with millions of elements in one go → Effect Every other request waits until that command finishes (tens of ms to several seconds) → On screen Features that use sessions, rankings, or the cache all hitch at once, slow logins
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Replace KEYS with SCAN, split large keys, delete with UNLINK (background deletion), spread out expiry times that cluster in the same second.
Infra team action items
Watch the slow command log (SLOWLOG), block dangerous commands such as KEYS on production servers, check for large keys regularly, turn off THP and keep enough spare memory for fork, run RDB and AOF persistence on replicas.
Ballpark numbers
A typical command takes under 1 ms. Handling millions of elements at once can take anywhere from hundreds of ms to several seconds.
On the graph
Random spikes · Redis response latency, slow command count
Where to look
SLOWLOG GET for commands over slowlog-log-slower-than; turn on the latency monitor (off by default) with CONFIG SET latency-monitor-threshold, then check per-event latency such as fork and expire-cycle with LATENCY LATEST and LATENCY DOCTOR. Also check fork time and large keys with latest_fork_usec in INFO and redis-cli --bigkeys
Confirmed if
At the time of the stall, SLOWLOG shows KEYS or commands handling a whole large key, or LATENCY records fork or expire-cycle events of tens of ms or more at the same time
Ruled out if
SLOWLOG and LATENCY are empty but it’s slow only as seen from the game server: the network or waiting inside the game server (SLOWLOG measures only command execution time, excluding time spent talking to the client)
Check with
Infra tools (no game code needed)
Learn more
Redis also stalls the moment it forks the process to create a save file (RDB snapshot) or rewrite the AOF. On modern servers this takes about 10 ms per GB of memory, so about 300 ms for 30 GB. With transparent huge pages (THP) turned on, every write after a fork copies an entire huge page (copy-on-write), sharply increasing stalls and memory use, so THP is usually turned off and plenty of spare memory is kept. Redis also stalls briefly to delete keys when a very large number of them expire in the same second.
Sources (7)
Diagnosing latency issuesRedis One thread processes requests in turn, so a slow command blocks everything behind it; use SCAN in place of KEYS; fork measured at about 9–13 ms per GB on physical servers and modern VMs; THP causes latency and memory spikes from copying after fork; mass expiry in the same second causes stalls
KEYSRedis Use with extreme care in production; can ruin performance on large databases (40 ms for 1 million keys on an entry-level laptop)
UNLINKRedis Asynchronous deletion: unlinks the key immediately and reclaims its memory in another thread
SLOWLOGRedis Slow command log that records commands exceeding slowlog-log-slower-than; execution time excludes I/O with the client
Redis latency monitoringRedis latency-monitor-threshold defaults to 0 (off); LATENCY LATEST and LATENCY DOCTOR; records latency per event such as fork and expire-cycle
INFORedis latest_fork_usec: time taken by the last fork (microseconds)
Redis CLIRedis --bigkeys: scans the keyspace to find large keys
ID db-plan-flip · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Even with the code unchanged, if the DB changes how it executes a query (its query plan), a query that took 2 ms yesterday takes hundreds of ms today.
Why Automatic statistics updates, a DB restart, or shifts in data distribution make the DB build a new query plan → Effect A plan that skips the index gets picked, the same query becomes tens to hundreds of times slower, and connections get tied up → On screen With no deploy at all, loading for a specific feature suddenly slows down and other requests wait too
Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Game team action items
For queries whose result count varies widely by parameter value, split them or consider plan hints, design queries that reliably use an index.
Infra team action items
Watch slow queries and query plan history, pin good plans (Query Store in SQL Server, and so on), manage when statistics get updated.
On the graph
Step change · Average execution time per query
Where to look
Collect the average time per normalized query periodically and watch the trend. MySQL: AVG_TIMER_WAIT in events_statements_summary_by_digest. PostgreSQL: mean_exec_time in pg_stat_statements (mean_time on 12 and earlier). Compare query plans from before and after the slowdown with EXPLAIN or PostgreSQL auto_explain, and in SQL Server with the Regressed Queries view in Query Store
Confirmed if
With no deploy at the time, one query’s average time steps up tens of times, the moment lines up with a statistics update or a DB restart, and the query plan has changed
Ruled out if
Slower with the query plan unchanged: data growth, lock waits (db-hot-row), or the disk
Check with
Infra tools (no game code needed)
Learn more
SQL Server reuses a plan built for the first value it sees (parameter sniffing). A plan built for a new character with only a few items gets very slow when used for an old character with tens of thousands of items, and the reverse is common too. It may recover when a restart clears the plan, then go bad again.
Sources (7)
Query Processing Architecture GuideMicrosoft SQL Server Parameter sniffing: the query plan is built for the parameter values passed at compile or recompile time
Monitor performance by using the Query StoreMicrosoft SQL Server Plans change with statistics, schema, and index changes, and the plan cache keeps only the latest plan; Query Store’s plan forcing pins a good plan; the Regressed Queries view compares slowed-down queries and their plans
Statement Summary TablesMySQL events_statements_summary_by_digest: COUNT_STAR and AVG_TIMER_WAIT (average time) per normalized query
ID db-ddl-lock · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Adding a column or index to a table during live service can make every request that uses that table wait, all because of one lock that’s needed only briefly.
Why A hotfix adds a column or index to a table in live use → Effect The schema change waits for a long transaction opened earlier, and every request that comes after waits for the schema change → On screen Features that use that table (inventory, mail, and so on) stop entirely and time out
Primary owner Infra team (DB infrastructure) · Also Game team (Server development)
Game team action items
Coordinate the timing of hotfixes that include schema changes with DB infrastructure, deploy code that works without the new column first.
Infra team action items
Set a short lock wait timeout and retry on failure, run it when there are no long transactions, use online schema change tools, change large tables during maintenance.
On the graph
Step change · Sessions waiting on locks, query latency on that table
Where to look
MySQL: count sessions in SHOW PROCESSLIST whose State is Waiting for table metadata lock, and find the blocking session (blocking_pid) with sys.schema_table_lock_waits. PostgreSQL: requests whose granted is false and AccessExclusiveLock in pg_locks, and find the blocking session with pg_blocking_pids()
Confirmed if
From the moment the schema change starts, every query on that table piles up waiting on the lock, with an unfinished transaction or the schema change statement at the front
Ruled out if
Waits concentrate on specific rows while other rows in the same table go through fine: a hot row (db-hot-row)
Check with
Infra tools (no game code needed)
Learn more
MySQL briefly takes a metadata lock when changing a schema, and PostgreSQL briefly takes its strongest table lock. Even if the change itself is instant, a single unfinished transaction ahead of it makes every request behind it wait.
Sources (8)
Online DDL Performance and ConcurrencyMySQL Even online DDL briefly needs an exclusive metadata lock to finish; it waits if there is a long transaction, and the waiting lock request blocks every transaction behind it
Server System VariablesMySQL lock_wait_timeout: metadata lock wait limit, default 31,536,000 seconds (1 year)
ID in-gateway · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Game team (Client development)
Putting an intermediate server between the client and the game server adds processing time at every hop, and that server becomes a single point of failure.
Why Client ↔ gateway ↔ game server architecture → Effect The intermediate server adds processing and queueing time, and when it’s overloaded everyone is affected → On screen Higher ping for everyone; if a gateway fails, every player routed through it disconnects
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Game team (Client development)
Game team action items
Server: make the gateway tier scale out to more machines, let a character carry on unchanged when it reconnects through another gateway after its gateway dies (session reconnection). Client: reconnect automatically when the gateway connection drops.
Infra team action items
Scale gateways horizontally (add machines), monitor CPU, connection count, and processing latency per gateway.
Ballpark numbers
Inside the same data center, each hop normally adds less than 1 ms. When the gateway is overloaded, that grows to tens to hundreds of ms.
On the graph
Rises with load · Gateway processing latency, gateway CPU and connection count
Where to look
Gateway CPU and connection count, Recv-Q on the gateway’s sockets (ss, netstat), and the latency difference before and after the gateway. For HTTP/gRPC calls through a service mesh, compare the Istio standard metric istio_request_duration_milliseconds split by sender (reporter=source) and receiver (reporter=destination)
Confirmed if
Game server processing time is unchanged but latency grows only across the gateway hop, and at the same time gateway CPU is saturated or Recv-Q builds up
Ruled out if
Paths that skip the gateway (direct connection, another gateway) are just as slow: points to the connection or the game server
Check with
Infra tools (no game code needed)
Learn more
With a service mesh such as Istio, the sidecar proxy (Envoy) running next to each server adds one more hop. A request between services passes through the sender’s sidecar and then the receiver’s sidecar, and every feature you add to the proxy, such as log and metric collection, adds processing and queueing time.
Performance and ScalabilityIstio In sidecar mode, a request passes through the sender’s sidecar proxy and then the receiver’s; every added feature lengthens the processing path inside the proxy, and telemetry collection adds queueing time to the next request
What is EnvoyEnvoy Envoy is a separate process running alongside every application server, and the app sends and receives through Envoy on localhost
Istio Standard MetricsIstio istio_request_duration_milliseconds (distribution of HTTP/gRPC request duration); the reporter label separates the sending (source) and receiving (destination) proxy
ID in-zone-transfer · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Entering another area or dungeon means handing the character’s data to another server, and that handoff can be slow or fail.
Why Entering a dungeon or traveling to another continent changes which server is responsible → Effect Save → transfer → load, with a wait if the target server is busy or has no free dungeon instance → On screen Long loading screens, failed entry, disconnects mid-transfer
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Shrink the transferred data, reserve the target server in advance, send the player back to where they were if the transfer fails.
Infra team action items
Monitor free instance headroom on dungeon and zone servers, add servers before peak hours.
On the graph
Rises with load · Zone transfer duration and failure count
Where to look
Server-side timing for each transfer stage (save, transfer, load) and failure reasons, plus player count and free instance count on the target server
Confirmed if
At the times long loading or failed entry is reported, transfer time rises or failures cluster, and the target server is crowded or out of free instances
Ruled out if
The transfer finishes quickly but the game freezes after arrival: points to a spawn burst when entering a crowded area, or client-side loading
Check with
Game server or client logs and metrics
Learn more
Seamless worlds without loading screens also switch the responsible server when you cross a server boundary. Near the boundary you may see a brief hitch or rubber-banding.
ID in-cascade · Primary owner Game team (Server development) · Also Infra team (Network infrastructure)
When one service slows down, the servers that call it get tied up waiting for responses, and even unrelated features stop.
Why One service, such as the DB or authentication, slows down → Effect Threads and connections on the calling servers are tied up waiting for responses, and retries of failed requests add more load → On screen Everything slows down or stops, even features that look unrelated
Primary owner Game team (Server development) · Also Infra team (Network infrastructure)
Game team action items
Put a timeout on every call, add circuit breakers and per-feature isolation (bulkheads), retry with growing intervals and a capped count, keep health check responses separate from busy work.
Infra team action items
Give load balancer health checks slack in failure count and interval so a briefly slow server isn’t pulled right away, limit how many servers can be pulled at once.
On the graph
Hits a ceiling · Per-service response time and error rate, thread and connection usage
Where to look
Per-service response time, error rate, and retry count on one screen with aligned time axes, to find what slowed down first. Behind a load balancer: target response time (TargetResponseTime on AWS ALB), target 5xx count (HTTPCode_Target_5XX_Count), and number of targets pulled as unhealthy (UnHealthyHostCount)
Confirmed if
One service’s latency rises first, then thread and connection usage on its callers hits the limit, errors spread to other services, and retry count and pulled-target count rise together
Ruled out if
Several services slowed down at the same instant: check shared resources (DB, network, hosts) first
Check with
Infra tools (no game code needed)
Learn more
Health checks (probes that confirm a server is alive) also make cascades worse. When a busy server answers a check late, the load balancer pulls a server that is actually working, its traffic piles onto the remaining servers, and the next server falls behind too.
Circuit Breaker PatternMicrosoft Azure Requests blocked until their timeout hold threads and DB connections and make unrelated features fail; once failures pile up within a set time, calls are rejected immediately
Timeouts, retries, and backoff with jitterAWS Amazon Builders’ Library. With 3 retries at each layer of a 5-deep call chain, DB load grows 243 times; retry at only one layer and cap retries with a token bucket
CloudWatch metrics for your Application Load BalancerAWS TargetResponseTime (time from the request leaving the load balancer until the target starts responding), HTTPCode_Target_5XX_Count (5xx responses generated by targets), UnHealthyHostCount (number of unhealthy targets)
ID in-subservice · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
When a server that runs separately from the game server, such as chat, party, or auction house, fails, only that feature stops working.
Why A server dedicated to one feature slows down or dies → Effect Only requests for that feature get no response → On screen Chat doesn’t work, party invites do nothing, the marketplace loads forever (combat is fine)
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Design the game to keep running when a feature fails, show status per feature, don’t pile many features onto one central server.
Infra team action items
Set up health checks and alerts for each auxiliary server, add redundancy and automatic restarts.
On the graph
Mass disconnect · Request success rate per feature, auxiliary server connection count and health checks
Where to look
Health checks, process state, and connection count for each auxiliary server (chat, party, auction house), plus request success rate and response time per feature. Behind a load balancer, UnHealthyHostCount for the target group
Confirmed if
Only the server behind the reported feature fails health checks or shows a sharp drop in connections, while game server ticks and combat are normal
Ruled out if
Several features stopped at once: points to a central server that relays them all, or a cascading failure
Check with
Infra tools (no game code needed)
Learn more
If one central server (a world or manager server) relays parties, guilds, whispers, and cross-server moves, several features stop at once when that one server slows down.
Sources (3)
Bulkhead PatternMicrosoft Azure Isolating components into pools lets the rest keep working when one fails and keeps the failure from spreading
ID in-deploy · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
If you restart a server for an update without moving its connections, everyone on it disconnects, and the final saves before shutdown and the reconnects all hit at once.
Why Servers restart one after another to roll out a hotfix → Effect Each server shuts down without moving its connections, and saves for every player on it hit the DB at once → On screen Disconnects without notice, a surge of reconnects
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Add draining (block new connections only and wait for current players to leave), move characters to another server, spread out saves before shutdown, report ready after a restart only once cache loading and JIT warm-up are done, for hot reload load the new data ahead of time on a separate thread and swap it in all at once between ticks.
Infra team action items
Have the deploy tool wait for each server to drain before restarting it, route traffic to a restarted server only after confirming it’s ready (warm-up done), announce deploy times.
Ballpark numbers
With 5,000 players on one server, 5,000 saves hit the DB within a few seconds before shutdown.
On the graph
Mass disconnect · Connections per server, DB write count
Where to look
Deploy tool job history (restart time per server) overlaid as vertical lines (annotations) on graphs of connections, disconnects, DB writes, and login requests
Confirmed if
Connections per server drop sharply one server at a time at each restart, with DB writes spiking just before and login requests right after
Ruled out if
Disconnect times don’t line up with the deploy or restart history: points to a server crash or network equipment
Check with
Infra tools (no game code needed)
Learn more
Servers are also slow for a few minutes right after they come back. The cache is empty so DB queries pile up, and Java and C# servers haven’t yet finished optimizing code as it runs (JIT warm-up), so the same work takes longer. Reloading scripts or data tables without shutting down (hot reload) also stops the tick while it reads, causing a brief pause.
Sources (3)
Site Reliability Engineering, Chapter 20: Load Balancing in the DatacenterGoogle A server that receives SIGTERM goes into lame duck state, sending new requests to other servers and finishing only the ones in progress; for the first few minutes after a restart, before JIT optimization, it uses more resources, so it takes traffic only after warming up
Liveness, Readiness, and Startup ProbesKubernetes A readiness check holds back traffic until connections are established, files are loaded, and caches are warmed
ID in-autoscale · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
When players flood in, servers are added automatically, but getting them ready takes several minutes, and the existing servers are overloaded in the meantime.
Why Connections spike when an event starts → Effect Several minutes pass before new servers boot and are ready → On screen Slow motion, and players who can’t connect, for the first few minutes after the event starts
When crowds gather, Right after login or maintenance
Owner
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Spread players across channels (players already in a crowded channel can’t be moved to a new server), cut startup and data loading time for new servers.
Infra team action items
Scale out ahead of events, keep warmed-up spare servers, when scaling in shut a server down only after its remaining players leave.
Ballpark numbers
1 to a few minutes to detect the load (because metrics are averaged over several minutes), then several more minutes to boot a new server, read game data, and fill caches.
On the graph
Surge after opening · Instance count, CPU utilization, queued logins
Where to look
Autoscaling activity history (when scale-out was decided, when new instances went into service) overlaid on CPU utilization and connection graphs. On AWS, the Auto Scaling group metrics (must be enabled to appear) GroupDesiredCapacity (target count), GroupPendingInstances (starting up), and GroupInServiceInstances (in service)
Confirmed if
For several minutes after a connection spike, only the desired count and pending instances rise while existing servers sit at their CPU limit, and things ease from the moment in-service instances increase
Ruled out if
Still slow after new instances come in: points to a cause other than server count (a shared resource such as the DB, a cascading failure)
Check with
Infra tools (no game code needed)
Learn more
Autoscaling is mostly used where a new server can simply take new players, such as login, gateway, and dungeon servers. Scaling in causes trouble too. If you scale in during the early-morning lull and shut servers down without waiting for the remaining players to leave, those players disconnect.
Amazon CloudWatch metrics for Amazon EC2 Auto ScalingAWS Group metrics are published every minute only when enabled; GroupDesiredCapacity (the count the group tries to maintain), GroupPendingInstances (instances not yet in service), GroupInServiceInstances (instances in service)
ID in-monitoring · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
During an outage, log volume explodes, and servers that ship logs synchronously get even slower because of the logging.
Why Errors make log and metric volume explode → Effect The log collector falls behind, and servers that send synchronously wait on it → On screen Stutter and freezes during an outage get worse because of logging
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Send asynchronously, sample, drop when the buffer overflows, batch repeated error logs into one.
Infra team action items
Size log collector capacity for the burst volume seen during outages, alert on collector backlog.
On the graph
Random spikes · Log volume, log collector queue
Where to look
Log lines and bytes per second on the server, plus the log collection agent’s queue and dropped count, alongside tick time. If a thread is stalled, use bcc offcputime -p to see whether it’s waiting on log writes or shipping
Confirmed if
When ticks spike, log volume jumps to tens of times normal, and the game thread’s wait time concentrates in log write/send call stacks
Ruled out if
Log volume is normal or the game thread isn’t waiting on logging: the log burst is only a result of the outage, so look separately for whatever caused the first errors
Check with
Infra tools (no game code needed)
Sources (3)
Logging in C#Microsoft .NET logging methods are synchronous, so with slow storage it recommends writing to fast storage first and moving the logs later
Asynchronous loggersApache Software Foundation Asynchronous logging absorbs short bursts in a queue, but if output stays slow, the queue fills and logging drops to the speed of the slowest output, or logs are dropped (Discard) depending on policy
ID in-clock-skew · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
When each server’s clock is slightly off from the others, cooldown, buff, and event start checks disagree from server to server.
Why A server whose time sync stopped drifts hundreds of ms to several seconds away from the other servers → Effect Passing absolute times, such as when a buff ends, between servers makes their checks disagree → On screen A buff disappears or a cooldown starts over after moving to another server
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
Pass remaining time between servers in place of absolute timestamps.
Infra team action items
Monitor time sync (NTP, chrony), alert on clock differences between servers.
Ballpark numbers
With healthy time sync (NTP, chrony), servers in the same data center usually stay within a few ms of each other. If sync stops, or a virtual machine is paused for a long time and then resumes, the gap grows to hundreds of ms to several seconds.
On the graph
Slow climb · Clock offset per server
Where to look
chronyc tracking on every server, comparing System time (difference between the system clock and NTP time), Last offset, and Ref time (when a measurement from the time source was last applied)
Confirmed if
The problem server’s offset is hundreds of ms or more away from the other servers, or its Ref time stopped long ago, and the mismatched checks happen only on moves to or from that server
Ruled out if
All servers’ offsets are within a few ms: points to the game’s own time calculation or client clock sync error
Check with
Infra tools (no game code needed)
Learn more
A single server’s clock jumping forward or backward all at once is covered in the server OS layer under “System clock jump (NTP step).”
chrony – Frequently Asked Questionschrony Ordinary computer clocks drift less than 100 ppm, but virtual machines can drift more; a VM that was paused and resumed can be far enough off to need a step correction
clock_gettime(2) — Linux manual pageLinux man-pages CLOCK_REALTIME can jump discontinuously from manual changes or NTP adjustments; CLOCK_MONOTONIC is not affected by such jumps
chronyc(1)chrony chronyc tracking fields: System time (difference between NTP time and the system clock), Last offset (offset estimated at the last update), Ref time (when the last measurement from the time source was applied)
ID in-bots · Primary owner Game team (Server development) · Also Infra team (Network infrastructure)
Bots send requests far more often than people do and eat into server capacity.
Why Large numbers of bots connected, farming, moving, and trading nonstop → Effect Server processing load and DB load go up → On screen A specific farming spot or the whole server slows down (slow motion, input lag)
Primary owner Game team (Server development) · Also Infra team (Network infrastructure)
Game team action items
Detect bots, rate-limit requests per account and character.
Infra team action items
Rate-limit connections and requests per IP (loosely, since many people share one IP at PC bangs and on mobile networks), block bot IP ranges at the firewall/WAF.
On the graph
Outliers only · Requests per second per account and IP
Where to look
Distribution and top list of requests per second per account and character from game server logs. Without code metrics, requests per IP at the firewall/WAF
Confirmed if
A handful of accounts or IPs send requests nonstop at rates no human could produce, and limiting them visibly cuts server load
Ruled out if
Requests are spread evenly across accounts: points to normal player growth (tick overrun, autoscaling delay)
ID in-external · Primary owner External (External) · Also Game team (Server development)
When an external service such as platform login, payments, or identity verification is slow or down, players get stuck at that step.
Why An external authentication or payment service is down or slow → Effect That step waits for a response → On screen Can’t log in, payments fail. Players already in the game are fine
Right after login or maintenance, During specific actions
Owner
Primary owner External (External) · Also Game team (Server development)
Game team action items
Put timeouts and a friendly message on external calls, cache authentication results, set up a retry and compensation process for payments.
External action items
Ask the authentication, payment, or platform provider to confirm the outage and restore service, tell players the problem is an external service outage.
On the graph
Step change · External call response time and error rate, successful logins
Where to look
Response time, error rate, and timeout count per external call (platform login, payments, identity verification), plus the provider’s status page
Confirmed if
From the time login and payment failures pile up, errors and timeouts for one specific external call step up and stay there, and the provider’s status page shows an outage at the same time
Ruled out if
External calls are healthy but logins are blocked: points to the login server itself (thread pool exhaustion, DB) or the OS connection queue (backlog)
Timeouts, retries, and backoff with jitterAWS Amazon Builders’ Library. Waiting for a response holds resources such as threads and connections, so set timeouts, and retry APIs with side effects only when they’re idempotent
Circuit Breaker PatternMicrosoft Azure Calls likely to fail are rejected immediately without waiting for the timeout, which protects response time
ID in-region-match · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Network infrastructure), External (External)
When a player lands on a server in a distant region while a closer region exists, that player’s ping stays high even though their connection is fine.
Why Bad GeoIP data, a VPN, assigning a whole party by the members’ average ping, rules that widen the search to distant regions when there aren’t enough players, assignment by DNS resolver location → Effect The player connects to a server across the ocean even though a nearby region exists → On screen In a game with servers in several regions, only you (or only your party) always have high ping, with input lag, rubber-banding, and skills that don’t go off
Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Network infrastructure), External (External)
Game team action items
Server: assign by the per-region ping the client measures in place of GeoIP, cap ping in the rules that widen to distant regions, for parties check the highest member ping as well as the average, log the assigned region and the ping at that moment. Client: measure per-region ping over UDP and send it with the matchmaking request, show the connected region and ping on screen, offer a manual region choice.
Infra team action items
If DNS picks the region, confirm the authoritative DNS supports EDNS Client Subnet (if the player’s resolver doesn’t send it, assignment goes by resolver location), update the GeoIP database regularly, tag connection logs on each region’s servers with GeoIP country and ASN to find the countries and ISPs that end up in distant regions.
External action items
Tell players to turn off VPNs and game boosters and reconnect, tell players on a corporate or overseas DNS to switch to their ISP’s DNS, ask the GeoIP provider to correct wrong locations.
Ballpark numbers
A Seoul player sent to the US West region when Tokyo is available sees ping rise from about 30 ms to about 130 ms. GeoIP is about 99.8% accurate at the country level, but at the city level only about 66% of lookups fall within 50 km, even in the US, and with a VPN it returns the VPN server’s location in place of the player’s.
On the graph
Outliers only · RTT (ping) per player, distribution of assigned regions
Where to look
Tag client IPs from each region’s connection records (load balancer access logs, VPC flow logs) with GeoIP country and ASN, and count which region each country and ISP connects to. For a single player, compare the region they actually connected to with their ping to the nearest region (measured by the player, or mtr from that region’s server to the player’s IP)
Confirmed if
High-RTT players or countries are connected to a distant region even though a nearby one exists, and ping measured to the nearby region is low
Ruled out if
Correctly assigned to the nearby region but ping is still high: points to detour routing or the player’s own connection or Wi-Fi
Check with
Infra tools (no game code needed)
Learn more
Picking the region by DNS (geolocation or latency-based DNS) guesses location from the address of the player’s DNS resolver in place of the player’s own address. If the resolver doesn’t support EDNS Client Subnet, which passes along part of the player’s address, players on a corporate DNS or a distant DNS are assigned by where the resolver is. Matchmaking systems may also judge a party by the average of its members’ ping, or widen the ping limit after a long wait and assign a distant region. AWS GameLift Servers also uses the average as the default for party ping, and its example setting widens the ping cap from 50 ms to 100 ms and then 200 ms. For a player on a VPN, the extra latency of going through the relay server (“Routing through a VPN or game booster”) can overlap with a distant-region assignment; tell them apart by whether the assigned region changes when the player turns off the VPN and reconnects. Connecting to a distant region because there is no nearby region at all is covered in “Propagation delay (physical distance).”
Sources (8)
RFC 7871: Client Subnet in DNS QueriesIETF DNS that answers differently by location guesses location from the address of the resolver sending the query, and gives inappropriate answers when the player uses a central resolver far away. EDNS Client Subnet (an optional feature) passes along part of the player’s address
How Amazon Route 53 uses EDNS0 to estimate the location of a userAWS If the resolver doesn’t support edns-client-subnet, the player’s location is guessed from the resolver’s address and the answer is based on the resolver’s location (same for geolocation and latency-based routing)
Geolocation accuracyMaxMind About 99.8% at the country level, about 66% at the US city level (within 50 km); with a VPN it gives the VPN server’s location in place of the end user’s; mobile network IPs are used across wide areas, so fine-grained location is unknown; the database needs continuous updates; corrections can be requested
FlexMatch rule typesAWS The latency rule (maxLatency) looks at player latency per location; for parties it uses the members’ average by default (partyAggregation avg); queues can place games in regions that don’t meet the latency rule
Create a player latency policyAWS Places the game at the location with the lowest average latency across all players, but players with extreme latency get placed too; example policy widens the ping cap from 50 ms to 100 ms and then 200 ms
Amazon GameLift Servers UDP ping beaconsAWS The game client measures latency to a UDP endpoint at each hosting location and uses it for placement and matchmaking; closer to real game traffic than ICMP ping
ID in-cert · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Client development)
When the certificate on a login, API, or patch server expires or is missing its intermediate certificate, every client that connects from that moment on fails the TLS connection.
Why The certificate is past its validity period, the server sends it without the intermediate certificate, or the date and time on the player’s device are wrong → Effect The client fails certificate validation and drops the TLS connection → On screen Can’t connect / infinite loading at the login or patch step, or only HTTPS features such as the store fail. Players already connected are usually fine
Right after login or maintenance, During specific actions
Owner
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Client development)
Game team action items
Log certificate errors under an error code distinct from other connection failures and show a message, if the date is wrong tell players to set their device’s date and time automatically, if you use certificate pinning include a backup key and coordinate the certificate rotation schedule with the infra team.
Infra team action items
Network: when TLS terminates at the load balancer or CDN, alert on managed certificate auto-renewal status and days remaining (DaysToExpiry on ACM), keep the validation DNS records in place. Servers/OS: when TLS terminates on the servers, automate renewal and reload the config after renewal, configure the full chain including intermediate certificates, check the remaining validity of every login, API, and patch address from outside on a schedule and alert on it.
Ballpark numbers
Let’s Encrypt certificates last 90 days and renewal every 60 days is recommended; AWS Certificate Manager checks DNS-validated certificates 45 days before expiry and renews them automatically. If automatic renewal fails silently, new connections are all blocked at exactly the expiry time.
On the graph
Mass disconnect · Successful logins, TLS handshake errors
Where to look
openssl s_client -connect HOST:443 -showcerts to see the certificate list the server actually sends, then each certificate’s expiry date (notAfter) with openssl x509 -noout -enddate. If TLS terminates at the load balancer, TLS negotiation error count (ClientTLSNegotiationErrorCount on AWS ALB and NLB) and successful logins
Confirmed if
The expiry date has passed or the intermediate certificate is missing from the list the server sends, and errors started rising at the expiry time or when the certificate was changed
Ruled out if
Certificate list and expiry date are fine but only some players fail: check the date and time on those players’ devices or the root certificate list of an old OS
Check with
Infra tools (no game code needed)
Learn more
A config missing the intermediate certificate can look fine when you open it in a desktop browser. Browsers remember intermediate certificates picked up from other sites and fill the gap, but clients without that memory, such as Android apps, fail. Validity periods are also getting shorter. Let’s Encrypt plans to cut the default validity to 64 days in 2027 and 45 days in 2028, so a setup hard-coded to renew every 60 days leaves only four days of margin with a 64-day certificate and runs past expiry with a 45-day one. AWS Certificate Manager also doesn’t auto-renew imported certificates, and renewal fails if you delete the validation DNS record. Blocked logins look similar to “DNS failures and delays,” but a certificate problem fails at the TLS handshake after the server address has been resolved, and it starts at the expiry time or when the certificate was changed.
FAQLet's Encrypt Default certificate lifetime of 90 days, renewal every 60 days recommended
Decreasing Certificate Lifetimes to 45 DaysLet's Encrypt Default lifetime cut to 64 days in February 2027 and 45 days in February 2028; a fixed 60-day renewal interval will no longer be enough, so renewal at about two-thirds of the lifetime is recommended
Renewal for domains validated by DNSAWS 45 days before expiry, checks whether the certificate is in use by an AWS service and whether the validation CNAME record exists, then renews automatically; if it can’t validate, sends notices 30, 15, 7, 3, and 1 days before expiry
Supported CloudWatch metricsAWS DaysToExpiry: days left until the certificate expires, published twice a day until expiry
Security with network protocolsAndroid (Google) If the server omits the intermediate certificate, Android apps fail with SSLHandshakeException, while desktop browsers may fill it in from cached intermediates and show no error; check the chain the server sends with openssl s_client
Network security configurationAndroid (Google) With certificate pinning, you must include backup keys to prepare for key rotation or CA changes; otherwise connections break until the app is updated
openssl-s_clientOpenSSL -showcerts: shows the certificates the server sent, in the order it sent them (not a validated chain)
openssl-x509OpenSSL -enddate: prints the certificate’s expiry date (notAfter); -checkend: checks whether it expires within the given number of seconds
CloudWatch metrics for your Application Load BalancerAWS ClientTLSNegotiationErrorCount: number of connections that failed to establish a TLS session, for example because the client dropped the connection after failing to validate the server certificate
ID in-login-queue · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)
When players flood in right after launch or maintenance, the login queue hits its cap and turns away new arrivals, and players who were already waiting lose their place during a brief disconnect and go back to the end of the line.
Why More people try to connect than the login server can take at once, so it keeps a queue, and when the queue gets too long it refuses new entries to protect the server → Effect The longer the queue, the longer the wait, and a brief Wi-Fi or mobile network drop during that time costs the player their place → On screen Can’t connect / infinite loading, the game quits with an error while waiting, the player starts over at the back of the line
Right after login or maintenance, Evening peak hours
Owner
Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)
Game team action items
Server: set the queue cap to what the login server can actually handle, hold a disconnected player’s place for a set time (reconnect grace period), show queue position and estimated wait, record queue length, refusals, and disconnects while waiting as metrics. Client: when disconnected while waiting, reconnect automatically to the same place without quitting the game, spread out retries with exponential backoff and jitter.
Infra team action items
Servers/OS: measure login and lobby server capacity with load tests before launch, prepare spare machines that can be added quickly at launch, graph queue metrics alongside connection attempts.
Ballpark numbers
At the 2021 FINAL FANTASY XIV expansion launch, new entries were refused once the queue passed 17,000 players per logical data center (Error 2002). If a player disconnected while waiting, the lobby server waited tens of seconds to 1 minute, and a player who reconnected within that time kept their place in the queue.
On the graph
Hits a ceiling · Login queue length, refusals at the cap, disconnects while waiting
Where to look
Queue length, average wait time, refusals at the cap, and disconnects while waiting, as recorded by the login and lobby servers, on the same graph as connection attempts
Confirmed if
Right after launch or maintenance, refusals rise while queue length flattens at the cap, and disconnects while waiting concentrate on Wi-Fi and mobile network players
Ruled out if
The queue is short but login is slow: points to the DB (db-login-storm) or the OS connection queue (so-backlog)
Check with
Game server or client logs and metrics
Learn more
A login rush that slows down the DB is covered in “Login storm and N+1 queries,” and an OS connection queue overflow in “Connection queue (listen backlog) overflow.” This entry is about the design of the login queue the game keeps on purpose. The queue cap is a safeguard that protects the login server, so you can’t remove it: refusing excess requests early is what keeps the server processing the requests it can handle. What matters is reducing what refusals and disconnects cost players, and the longer the queue gets, the more the errors fall on players with unstable connections such as Wi-Fi and mobile networks.
Response to Congestion (as of Dec. 11)Square Enix Once the queue passed 17,000 players per logical data center, new entries were refused so the login servers wouldn’t go down (Error 2002); a player disconnected while waiting got tens of seconds to 1 minute from the lobby server to reconnect and resume mid-queue, and went to the back of the line after that
Using load shedding to avoid overloadAmazon Builders' Library Load shedding: refusing excess requests early so the server keeps processing the requests it can handle
ID sy-request-response · Primary owner Game team (Client development) · Also Game team (Server development)
Press a button and there’s no animation or sound until the server answers. Your ping becomes your response time.
Why Skills, movement, and item pickups play only after the server confirms them → Effect From the moment you press, nothing happens for a full round trip plus the tick wait → On screen At 150 ms ping, every action feels 0.2 s sluggish
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: start animations, sounds, and effects the moment the player presses (client-side feedback), show only results (damage, rewards) after the server confirms, predict movement and basic attacks and apply them immediately, when a position correction arrives from the server re-apply the unconfirmed inputs from that position. Server: compute movement from the received inputs and send a correction only when the difference from the client’s predicted position exceeds a threshold.
Ballpark numbers
Response time ≈ ping + half the tick interval + one frame. At 20 ticks and 150 ms ping, about 190 ms.
On the graph
Always high · Time from input to start of feedback, RTT (ping)
Where to look
Client log in a development build with the button press time, the time the first animation or sound starts, and the time the server response arrives, next to in-game RTT. Add latency with the engine’s network emulation (Unreal NetEmulation.PktLag) or Linux tc netem on a test server and measure at different pings
Confirmed if
Feedback always starts at the same instant the server response arrives, the time from input to feedback equals RTT plus the tick wait, and it grows by exactly the latency you add
Ruled out if
Feedback starts on the press and only results such as damage numbers come late: normal design. Delay longer than a tick interval even at low ping: points to a double tick wait or a client frame problem
Check with
Game server or client logs and metrics
Learn more
For games that don’t need fast reactions, such as turn-based, card, and idle games, this approach is the simplest and safest. The problem is real-time action games that build even movement and basic attacks this way.
Using Gameplay Abilities in Unreal EngineEpic Games Local Predicted runs immediately on press and the server makes the final decision; Server Initiated has no prediction, so the player sees the delay
Using Network Emulation in Unreal EngineEpic Games Test with minimum and maximum latency and a packet loss percentage on server and client; set from the console, e.g., NetEmulation.PktLag
tc-netem(8) — Linux manual pageiproute2 A test tool that adds delay and jitter (delay TIME JITTER) and loss (loss random PERCENT) to outgoing packets to mimic a real network
ID sy-chatty · Primary owner Game team (Server development) · Also Game team (Client development)
If one action needs several server round trips one after another, your ping is multiplied by that many.
Why Open shop → request list → check price → buy → refresh inventory, each as a separate request → Effect Each request is sent only after the answer to the previous one arrives → On screen At 150 ms ping, a single purchase takes close to 1 s. Loading takes unusually long
During specific actions, Right after login or maintenance
Owner
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: change the protocol so several steps go in one request and response (e.g., include the updated inventory in the purchase response). Client: fetch the data you’ll need ahead of time, use UI that doesn’t wait for results.
Ballpark numbers
Time taken ≈ number of round trips × (ping + server processing + tick wait). With 5 round trips at 150 ms ping, about 0.85–1 s.
On the graph
Always high · Completion time per feature, round trips per action
Where to look
Server-side packet capture (Wireshark) while a test account performs one action such as a shop purchase or login, counting how many times requests and responses alternate and the gaps between them. With server request logs, group by session ID and look at the request count and each request’s arrival and response times
Confirmed if
One action sends several requests in turn, each waiting for the previous response, completion time is roughly round trips × RTT, and the same feature is proportionally slower for players in high-ping regions
Ruled out if
Only one or two round trips but one response takes a long time: points to server processing or the DB. All players equally slow regardless of ping: check server load
Check with
Infra tools (no game code needed)
Sources (1)
Chatty I/O antipatternMicrosoft Azure Many small I/O requests add up to latency that badly hurts responsiveness; recommends fewer, larger requests
ID sy-no-queue · Primary owner Game team (Client development) · Also Game team (Server development)
If you can’t press the next skill until the server confirms the previous one has finished, a round trip gets inserted between every skill in a rotation.
Why The next skill input is accepted only “after the previous skill is confirmed” → Effect A gap as long as your ping opens between every skill → On screen Gaps between skills in a rotation, and the higher the ping, the lower the DPS
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: an input buffer window that accepts input within a set time before the cooldown ends (e.g., 0.3–0.4 s) and sends it to the server right away. Server: don’t reject input that arrives slightly early, run it the moment the cooldown ends.
Ballpark numbers
In a rotation with a 1-second cooldown at 150 ms ping, each gap between skills is 0.15 s or more, so you cast over 13% fewer skills in the same time.
On the graph
Always high · Gap between skills, RTT (ping)
Where to look
Server log per character of when a skill cooldown ended, when the next skill request arrived, and when it ran, with the gap compared against the player’s RTT
Confirmed if
There is always a gap of about one RTT between the cooldown ending and the next skill running, and higher-ping players have longer gaps and cast fewer skills in the same time
Ruled out if
The gap is constant regardless of ping: global cooldown or animation length by design. The gap spikes only occasionally: check jitter/loss or tick overrun
Check with
Game server or client logs and metrics
Learn more
World of Warcraft, for example, has an input buffer window that players can adjust in the settings. When the window is longer than the round trip, ping barely shows up between skills in a rotation.
ID sy-short-window · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)
When the time you have to react is short, as with dodges, parries, and guards, ping eats up that time and some attacks become impossible to avoid.
Why Short timing windows, such as a 0.5 s boss attack telegraph or a 0.2 s parry window → Effect You see the telegraph late (downstream latency + interpolation), and your input also arrives late (upstream latency + tick wait) → On screen You get hit even though you clearly dodged, parries don’t go off
Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)
Game team action items
Server: schedule attack telegraphs at a server time and send them in advance, widen the timing window by the player’s ping (lag compensation). Client: play received telegraphs at their scheduled server time.
Infra team action items
Put servers close to regions with many players (regional servers) to cut ping itself.
Ballpark numbers
At 150 ms ping and 100 ms interpolation, the telegraph takes about 0.18 s to appear on your screen and your input takes about 0.1 s to reach the server. Add 0.25 s of human reaction time, and a 0.5 s telegraph is nearly impossible.
On the graph
Outliers only · Dodge/parry failure rate (by ping bracket)
Where to look
Server log of the timing window’s start and end times, when the player’s input reached the server, and that player’s RTT, with failure rate broken down by ping bracket (e.g., 50 ms steps)
Confirmed if
Failure rate is clearly higher in higher ping brackets, and failed inputs arrive shortly after the window closes (within RTT plus interpolation time)
Ruled out if
Failure rate is similar across ping brackets: the pattern is just hard. Inputs arrive inside the window but still count as failures: check the window-checking code or server validation
ID sy-no-lagcomp · Primary owner Game team (Server development) · Also Game team (Client development)
If the server checks hits only against where targets are on the server right now, what you saw on your screen and the server’s call disagree.
Why The opponent on your screen is at a position about 0.2 s in the past (at 150 ms ping and 100 ms interpolation) → Effect The server checks against the current position, so the target has already left the spot you aimed at → On screen A clear hit counts as a miss. You have to lead moving targets
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: rewind to what the attacker was seeing and check the hit there (lag compensation), or switch to a target-lock system. Client: send the moment you were seeing (the server time being interpolated) with each attack.
On the graph
Outliers only · Hit rate on moving targets (by ping bracket)
Where to look
Server hit log with the attack time, the target position on the attacker’s screen (sent by the client), the target position the server used, and the attacker’s RTT. In a development build, drawing the position the server used over the client’s screen shows it right away
Confirmed if
On misses, the gap between the two positions is roughly target speed × (attacker RTT + interpolation time), and higher ping lowers the hit rate on moving targets only
Ruled out if
Stationary targets get missed too: hitbox or collision check problem. Still mismatched with rewind in place: check whether the client reports its interpolation time to the server incorrectly
Peeking into VALORANT's NetcodeRiot Games The server rewinds to the game state the player saw at the moment of the shot to check the hit; the client sends the simulation time it was seeing
ID sy-lagcomp-overreach · Primary owner Game team (Server development)
If the server rewinds too far in the attacker’s favor, the target gets hit even after they’ve already taken cover.
Why The server rewinds a long way to check hits for a high-ping attacker → Effect On the target’s screen, they were already behind cover → On screen “I got shot behind a wall,” high-ping players have the advantage
Cap the rewind (e.g., 200–250 ms), for attackers with higher ping rewind only up to the cap and let them lead their shots for the rest.
On the graph
Outliers only · Rewind time per hit (by attacker ping)
Where to look
Server hit log with the rewind time for each hit, attacker RTT, and the server time the target entered cover. In a development build, draw the rewound hitboxes on screen (sv_showlagcompensation in the Source engine)
Confirmed if
Hits in “I got shot behind a wall” reports come mostly from attackers with long rewind times, and rewind time grows with attacker ping with no cap
Ruled out if
Behind-the-wall hits also show up on hits with short rewind times: hitbox or collision check problem. If the target has high ping, their own movement reached the server late
Check with
Game server or client logs and metrics
Learn more
Rewound hit checks follow “favor the shooter.” A “favor the target” exception has also been proposed: skip the rewind if the target had already reached safety on their own screen.
Sources (3)
Peeking into VALORANT's NetcodeRiot Games Without a rewind cap, a player with 500 ms latency could land hits 0.5 s after the target took cover, so a cap is set
Source SDK 2013: player_lagcompensation.cppValve Source engine rewind cap sv_maxunlag defaults to 1 second (maximum 1 second); sv_showlagcompensation draws the rewound hitboxes on screen
ID sy-client-auth · Primary owner Game team (Server development) · Also Game team (Client development)
When each client decides its own results, your own screen feels responsive, but results disagree with other players’ screens and the game is easy to hack.
Why The client decides position and hits, and the server only relays them → Effect Two players each claim they hit first, and the server can’t verify either claim → On screen Opponents teleport or pass through walls, “I hit them but it didn’t count”
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: validate important results (such as hits) on the server, check movement speed and distance. Client: when the server rejects or corrects a result, roll back to the server’s value.
On the graph
Always high · Impossible movement speeds and conflicting hit reports
Where to look
On the server, log the positions and hits clients report as-is, compute movement speed from consecutive position reports, and count reports over maximum speed and cases where two players both claim they hit first
Confirmed if
The server passes reports on to other clients without validation, and impossible speeds or conflicting hit reports show up steadily regardless of patch or region
Ruled out if
The server computes or validates results itself: not this cause. Teleporting in that case points to packet loss or the interpolation buffer
ID sy-lockstep · Primary owner Game team (Server development) · Also Game team (Client development)
When everyone computes the same turn together, one player’s late input makes everyone wait.
Why Each turn can be computed only after every player’s input has arrived → Effect One player’s input arrives late because of jitter or packet loss → On screen Everyone hitches at the same time, and in bad cases a “Waiting for players” window appears
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: adjust input delay automatically to ping, drop only the lagging player for a moment so everyone else keeps going without waiting. Client: apply the agreed input delay, in P2P without a relay server have the host client also handle input delay adjustment and lagging players.
Ballpark numbers
If you set input delay shorter than “time for an input to reach the other player + jitter,” freezes become frequent. That time is half the ping when players exchange inputs directly, and about half the sum of both players’ pings through a relay server.
On the graph
Random spikes · Turn wait time, input arrival delay per player
Where to look
Per turn, each player’s input arrival time and how long the turn stalled waiting, and whose input each stalled turn was waiting for. With a relay server, also visible in a server-side packet capture as the arrival interval of each player’s input packets
Confirmed if
In every stalled turn, the same one player’s input arrived later than the input delay, and that player’s jitter/loss spiked at the same time
Ruled out if
All inputs arrived on time but it still stalls: points to the slowest PC’s computation time or server processing. No stalls but results differ between two screens: a desync, so check “Pathfinding mismatch in command sync”
Check with
Game server or client logs and metrics
Sources (3)
Deterministic LockstepGaffer On Games Frame n can be computed only after all its inputs arrive, so a late input means waiting; a small playout delay buffer for absorbing jitter causes hitches
ID sy-rollback · Primary owner Game team (Client development)
The game predicts the opponent’s input and shows it early, then rewinds and recomputes if the guess was wrong. The higher the ping, the further it rewinds.
Why The opponent changes their input (different from the prediction) → Effect The real input arrives half a ping late, so the game rewinds that far and recomputes → On screen The opponent’s animation skips a few frames or changes suddenly
Mix in 1–3 frames of input delay to shorten rollbacks, cap the rollback length.
Ballpark numbers
At 100 ms ping (50 ms one way), that’s about 3 frames of rollback at 60 FPS. With 2 frames of input delay, it drops to 1 frame.
On the graph
Random spikes · Frames rolled back, RTT (ping)
Where to look
Client log of the frame count for each rollback, the RTT at that time, the input delay setting, and the time spent rolling back and recomputing
Confirmed if
Rollback frame count is large at the moments the opponent’s animation jumped, and the average rollback is roughly (one-way latency − input delay) ÷ frame time, growing with ping
Ruled out if
Stutter even with short rollbacks: a performance problem where recomputation takes longer than one frame. Results still differ between the two screens after rollback: a desync
Check with
Game server or client logs and metrics
Sources (2)
GGPO Rollback Networking SDKGGPO The game predicts the opponent’s input and moves ahead; if the real input differs, it recomputes from the point where they diverged up to the present
ID sy-no-timestamp · Primary owner Game team (Client development) · Also Game team (Server development)
If server events carry no timestamp and play as soon as they arrive, network jitter carries straight through into uneven animation timing.
Why “Attack start” and “play effect” events run as soon as they arrive → Effect Each packet arrives at a different time, so the intervals are uneven → On screen Chained attack animations speed up and slow down, and boss pattern timing differs every time
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: play events at the time attached to them (scheduled events, interpolation buffer). Server: attach the time the event happened (server time) when sending.
On the graph
Random spikes · Event playback interval, packet arrival interval
Where to look
Event times from server logs and arrival and playback times from client logs, matched by event number, comparing intervals. Reproduce in a development build by adding jitter (the jitter value in tc netem, min/max latency in Unreal network emulation)
Confirmed if
Events happen at steady intervals on the server, but playback intervals follow the uneven arrival intervals exactly
Ruled out if
Arrival intervals are even but playback is uneven: client frame problem (frame time spikes). Intervals are already uneven on the server: tick overrun
Snapshot InterpolationGaffer On Games Drawing received snapshots immediately stutters because of jitter; holding them briefly in an interpolation buffer before drawing makes motion smooth
tc-netem(8) — Linux manual pageiproute2 A test tool that adds delay and jitter (delay TIME JITTER) and loss (loss random PERCENT) to outgoing packets to mimic a real network
Using Network Emulation in Unreal EngineEpic Games Test with minimum and maximum latency and a packet loss percentage on server and client; set from the console, e.g., NetEmulation.PktLag
ID sy-double-tick · Primary owner Game team (Server development)
If requests wait for the next tick to be processed and the results wait for the tick after that to be sent, the tick interval is added twice.
Why Received requests are processed on the next tick → Effect Results are also batched and sent on the next send tick → On screen Ping on the connection is low, but responses are consistently late by about 1.5 times the tick interval. On a 10-tick server, 0.15 s on average and 0.2 s at worst
Send the response on the same tick that processes the request, raise the tick rate, send important responses immediately.
Ballpark numbers
On a 10-tick server one tick is 100 ms, so the tick waits alone add 150 ms on average and 200 ms at worst. With a single wait, it’s 50 ms on average.
On the graph
Always high · Time from request arrival to response send
Where to look
Server-side packet capture while a test account repeats the same action (e.g., using an item), measuring the gap between the request packet’s arrival and the response packet’s departure. With server logs, request arrival time, the tick number that processed it, and the time the response was sent
Confirmed if
Time spent inside the server averages about 1.5 times the tick interval, up to about 2 times, and stays constant regardless of RTT
Ruled out if
Time inside the server is around half the tick interval on average: only one tick wait. Longer than the tick interval and uneven: check tick overrun
Check with
Infra tools (no game code needed)
Sources (3)
Peeking into VALORANT's NetcodeRiot Games An arriving input waits up to one tick for the tick boundary, and applying and sending take another frame; this shrinks as the tick rate goes up
VALORANT's 128-Tick ServersRiot Games Part of the latency comes from the network, part from the server tick rate
ID sy-strict-check · Primary owner Game team (Server development)
If the server checks movement speed, cooldowns, and range too strictly, it rejects even valid inputs that arrive bunched together because of jitter.
Why Strict rules such as “max distance per tick” or “0 ms cooldown tolerance” → Effect When jitter makes two commands arrive in the same tick, they’re judged as rule violations → On screen Rubber-banding, skills rejected even though the cooldown is up
Check against an accumulated allowance (token bucket), leave slack for ping and jitter.
On the graph
Random spikes · Server validation rejections and position corrections
Where to look
Server log for each validation rejection or position correction with the reason, how many commands from that player arrived in that tick, and the arrival gap from the previous command
Confirmed if
Rejections and corrections cluster at moments when 2 or more commands arrived in one tick, while movement and use counts summed over a few seconds stay within the rules
Ruled out if
Still over the limit even when summed over a few seconds: real speed hacking or cheating is possible. Rejections concentrated on one ISP in the evening: points to validation false positives concentrated on one ISP’s players
Check with
Game server or client logs and metrics
Sources (2)
Source SDK 2013: player.cppValve A per-tick command processing budget that accumulates (up to sv_maxusrcmdprocessticks, 24 ticks) lets bunched-up commands through; a developer comment says stricter restrictions caused stutter even for legitimate players
ID sy-host · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)
When one player’s PC acts as the server, that player’s connection and PC performance decide how the game feels for everyone.
Why The host’s PC acts as the server (P2P, listen server) → Effect If the host’s connection or PC is slow, everyone feels it, while the host has zero ping → On screen Only the host has the advantage, and when the host leaves everyone freezes or disconnects
Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)
Game team action items
Server: move to dedicated servers that make the calls, until then pick a player with a good connection and PC as host during matchmaking. Client: support host migration, measure and report ping to other participants, upload speed, and PC performance during matchmaking.
Infra team action items
Secure server machines or instances for dedicated servers, place them close to regions with many players.
On the graph
Outliers only · Lag reports and disconnects per host
Where to look
Match logs with the host’s upload speed, each participant’s RTT to the host, the host PC’s frame time, and when the host left, with lag and disconnect reports grouped by host. Players can also confirm it by playing again with the same group and only the host changed
Confirmed if
Lag and disconnects concentrate in one host’s matches, all participants get worse together when that host’s upload speed is low or frame time is long, and without host migration everyone disconnects the moment the host leaves
Ruled out if
Only participants in the same region are affected, regardless of host: connection or route problem. With dedicated servers, this isn’t the cause
Check with
Game server or client logs and metrics
Sources (3)
Networking Overview for Unreal EngineEpic Games A listen server’s host has an advantage over other clients and carries a heavy load, running both the server and rendering
ID sy-optimistic-reject · Primary owner Game team (Client development) · Also Game team (Server development)
When the server later refuses a hit or skill your screen already showed, the result you clearly saw never happened.
Why Hit effects and skill animations play before the server confirms (client-side feedback) → Effect The server rechecks range, target position, cooldown, and resources and rejects the action → On screen Blood sprays but there’s no damage, the skill animation plays with no effect, only the cooldown runs
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: show only what needs confirming (damage numbers, deaths, rewards) from server results, check common rejection reasons locally first, when rejected refund the cooldown and resources and show the reason. Server: leave slack for ping in range and target position checks, include the reason in rejection responses, collect rejection rate per skill as a metric.
Ballpark numbers
The rejection arrives ping + tick wait after the press. At 150 ms ping, you think you hit for about 0.2 s.
On the graph
Outliers only · Server rejection rate per skill (by ping bracket)
Where to look
Server-side rejection rate and rejection reasons (range, target position, cooldown, resources) per skill, broken down by player RTT bracket. On the client, a count of actions shown with client-side feedback that were then rejected
Confirmed if
Rejections concentrate on specific skills with range or target position as the reason, and the rejection rate rises with ping
Ruled out if
Rejection reasons are cooldown or resources and unrelated to ping: check whether client and server data values (cooldown, cost) differ. No rejections but feedback starts only after the server responds: points to “Feedback only after the server responds (request-response)”
Check with
Game server or client logs and metrics
Learn more
Client-side feedback is the best way to hide ping. However, the more the information the client and server use to decide (enemy position, remaining resources) differs, the more often rejections happen. Collecting rejection rate per skill as a metric makes it easy to find where the calls disagree.
Sources (2)
Using Gameplay Abilities in Unreal EngineEpic Games Local Predicted abilities run immediately on the client, but the server makes the final decision and can reverse the result
ID sy-path-mismatch · Primary owner Game team (Server development) · Also Game team (Client development)
When only “go here” is exchanged and each side computes the path itself, even a small difference in the calculation sends a character or monster down a different path until it gets pulled back into place.
Why With click-to-move and monster chasing, only the destination is sent and the client computes the path on its own → Effect Differences in terrain data, collisions with other characters, or calculation order make it move along a different path from the server’s → On screen Monsters walk through walls and then snap to another spot, clicked characters change direction as if sliding
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: send the path’s intermediate points (waypoints) as well, sync positions periodically. Client: converge smoothly when positions diverge, use the same terrain data as the server.
On the graph
Random spikes · Position correction count and distance per entity
Where to look
Per entity, the difference between the position the server sent and the position the client computed, with the coordinates of corrections plotted on the map. Periodically comparing checksums of both sides’ path results or positions pinpoints when they started to diverge
Confirmed if
Corrections cluster at specific terrain (ledges, narrow passages, slopes) or crowded spots and repeat at the same place even for players with normal network metrics
Ruled out if
Corrections happen only when loss or jitter spikes, regardless of place: connection problem. A single monster jumping on several players’ screens at once: check whether a slow client has control of the monster
Check with
Game server or client logs and metrics
Learn more
This approach is one reason click-to-move and tab-target games are less sensitive to ping. The trade-off is that nothing guarantees both sides get the same result, so a mechanism that syncs positions now and then is essential. Floating-point results can differ slightly depending on CPU type, compiler, and optimization settings (including debug versus release builds). In designs that exchange only inputs and assume both sides compute identical results, such as lockstep and rollback, these tiny differences can accumulate until the game state on the two screens splits apart (desync).
Sources (6)
Deterministic LockstepGaffer On Games Even if deterministic on the same machine, floating-point results can differ across compilers, OSes, and CPUs
State SynchronizationGaffer On Games Sending state along with inputs keeps both sides in sync without perfect determinism
Peeking into VALORANT's NetcodeRiot Games Packet loss, or two characters trying to move to the same spot, makes server and client simulations diverge and need correction
Floating Point DeterminismGaffer On Games The same floating-point code can give different results depending on compiler, CPU architecture, and debug or release build; includes a case where AMD and Intel CPUs returned slightly different values from transcendental functions
/fp (Specify floating-point behavior)Microsoft /fp:fast may reorder or combine floating-point operations and give results different from other /fp settings, and operations fused with FMA can also differ from a separate multiply and add
ID sy-low-send-rate · Primary owner Game team (Server development) · Also Game team (Client development)
If the server sends position updates (snapshots) only a few times a second, the interpolation buffer has to be that much longer, and you see other characters further in the past.
Why Position updates sent only 5–10 times a second to save bandwidth → Effect Smooth rendering needs a buffer of twice the packet interval (200–400 ms); with a shorter buffer, a single missed packet causes a freeze → On screen Opponents’ direction changes show up late and disagree with hit registration. With a short buffer: stutter, plus teleporting on packet loss
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: send nearby or in-combat targets often and distant ones rarely, send only what changed (delta compression) to shrink each update and raise the rate. Client: adjust the interpolation buffer length automatically to the packet interval.
Ballpark numbers
At 10 per second, the packet interval is 100 ms and the buffer 200 ms. Add the 75 ms one-way latency of a 150 ms ping, and you see opponents about 0.3 s in the past.
On the graph
Always high · Packet arrival interval per client, interpolation buffer length
Where to look
Server-side packet capture filtered to the flow going to one player, with packets per second and intervals in Wireshark I/O Graphs. With game logs, update interval per entity together with the client’s interpolation buffer headroom (time left until the next snapshot arrives)
Confirmed if
Position updates are always sparse at 5–10 per second (100–200 ms apart), and the interpolation buffer is set above 200 ms or its headroom often hits 0
Ruled out if
Updates go out densely but only the arrival intervals wobble: points to jitter or loss. Only distant entities arrive rarely when it’s crowded: per-connection send budget and priority
ID pt-slow-burst · Primary owner Game team (Server development) · Also External (External)
Inputs from a player with a bad connection reach the server unevenly, in bunches. If the server applies whatever arrived on each tick, other players see that character hitch and then cover several steps at once.
Why The lagging player’s move commands arrive 0 at a time on some ticks and 2–3 at a time on others → Effect The server applies them all on the tick they arrive, so that character’s position changes in steps → On screen On other players’ screens, only that character hitches and then moves several steps at once. Everyone else looks fine
Primary owner Game team (Server development) · Also External (External)
Game team action items
Spread commands evenly with a per-player input buffer, apply them at their original spacing using input sequence numbers, lengthening only the interpolation buffer on other players’ screens isn’t enough (the server’s own position history is already stair-stepped).
External action items
Tell lagging players to use a wired connection and check their Wi-Fi and router.
Ballpark numbers
With 80 ms of jitter on a 20-tick (50 ms) server, commands per tick swing between 0 and 3.
On the graph
Outliers only · Commands applied per tick per player, jitter per player
Where to look
Server-side packet capture filtered to packets from the reported player, counting how many arrived in each tick interval (e.g., 50 ms) and comparing with other players. With server logs, move commands applied per tick and input sequence numbers per player
Confirmed if
Only the reported player’s packets arrive in bunches, alternating between 0 and 2–3 per tick, that player’s jitter/loss is high, and other players’ packets arrive evenly. It improves when that player switches to wired
Ruled out if
Several characters move in bursts at once: server tick delay or the viewer’s own connection. Arrival and application are even but that character still looks jumpy: interpolation or display problem on the viewer’s side
Check with
Infra tools (no game code needed)
Learn more
In a server-authoritative design, this is normal behavior. One lagging player’s lag shows up to others only as “that player moving strangely” and doesn’t affect anyone else’s controls or monster movement. Anything that involves that player directly (trades, party mechanics, PvP hit registration), however, is delayed along with them.
Sources (3)
Peeking into VALORANT's NetcodeRiot Games The server puts arriving inputs into a per-player move queue in tick order and fills gaps with prediction; corrections are visible only to that player, and everyone else sees smooth motion
State SynchronizationGaffer On Games Even packets sent 60 times a second arrive bunched, e.g., 2 in one frame and 0 in the next
Source SDK 2013: player.cppValve Commands that arrive bunched are spread out (metered out) across server ticks
ID pt-event-server · Primary owner Game team (Server development) · Also Game team (Client development)
On a server that processes and broadcasts packets as soon as they arrive, a lagging player’s bunched-up actions run back to back immediately.
Why A lagging player’s skill and move requests arrive in a bunch → Effect The server runs them in order the moment they arrive and tells everyone right away → On screen Others see that player use several skills in an instant or move as if fast-forwarding
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: run actions at the spacing of their attached input times (accepting those times only within an allowed range), or run bunched actions one after another spaced by a minimum interval (global cooldown) without rejecting them, don’t check cooldowns by arrival time alone (valid inputs get dropped). Client: attach the input time to each action.
On the graph
Outliers only · Action execution interval per player
Where to look
Server log of each player’s action arrival time, execution time, and client-attached input time (if any), comparing execution intervals with input intervals. Also the arrival intervals of that player’s packets in a server-side packet capture
Confirmed if
Input intervals are normal, but server arrival and execution intervals are bunched within a few ms, and the bunches line up with the times other players reported fast-forward
Ruled out if
Intervals are already bunched in the input times: points to the client or a macro. Server execution intervals are even but look bunched only on other players’ screens: the viewer’s connection
Deterministic LockstepGaffer On Games Applying inputs as they arrive gives uneven results even when they’re sent at 60 Hz, because the spacing isn’t even
ID pt-input-buffer · Primary owner Game team (Server development) · Also Game team (Client development)
If the server holds a few inputs per player and takes out one per tick, others see smooth motion, but your own actions are confirmed on the server that much later.
Why The server collects a lagging player’s inputs in a buffer and applies one per tick → Effect A small buffer often runs empty, so the character stands still or the server guesses from the last input; a large buffer confirms the player’s own inputs late → On screen Too small: others see hitches. Too large: your own skill results come late (input lag)
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: size each player’s buffer automatically to their connection, take two inputs at a time to catch up when behind, tell clients whose buffers often run empty to send inputs earlier. Client: send inputs slightly earlier as the server instructs (client time adjustment).
Ballpark numbers
It varies by game, but 1–3 ticks’ worth is typical. VALORANT keeps the server buffer shorter, averaging half a frame (about 4 ms) on 128-tick servers. Adaptive buffers that grow only for players with high jitter are common.
On the graph
Outliers only · Input buffer length and empty count per player
Where to look
Server-side, per player per tick: inputs left in the input buffer, times the buffer ran empty and was filled with a guess from the last input, and time from input arrival to application
Confirmed if
Players with small buffers run empty often and briefly freeze on others’ screens at those moments, and players with large buffers have input-to-application times longer by the buffer length
Ruled out if
The buffer almost never runs empty but others see stutter: interpolation problem on the viewer’s side. Short buffer but still high input lag: RTT itself or a double tick wait
Check with
Game server or client logs and metrics
Sources (3)
Peeking into VALORANT's NetcodeRiot Games The server adjusts the client’s time base so the input queue stays just long enough to absorb uneven arrivals with minimal latency; the server buffering target averages half a frame
ID pt-isp-validation · Primary owner Game team (Server development) · Also Infra team (Network infrastructure)
Players on high-jitter connections have their inputs arrive in bunches, so they often trip the server’s speed and cooldown checks.
Why Jitter on a specific ISP’s or region’s connections rises in the evening → Effect The server judges valid inputs that arrived in a bunch as speeding or cooldown violations → On screen Only that ISP’s players get rubber-banding and rejected skills, and in bad cases the server kicks them (disconnect)
Evening peak hours, While moving or changing zones
Owner
Primary owner Game team (Server development) · Also Infra team (Network infrastructure)
Game team action items
Check against an allowance accumulated over several seconds, loosen thresholds based on connection quality (ping, jitter), add a warning stage before a forced disconnect, spread bunched inputs evenly across ticks with a per-player input buffer to cut false positives at the source.
Infra team action items
Check loss rate and jitter distribution per ISP by time of day and share them with the game team, check the path through that ISP (bidirectional mtr), reroute or escalate to the ISP if needed.
On the graph
High at certain hours · Validation rejections and forced disconnects per ISP (ASN), jitter per ISP
Where to look
Server logs of validation rejections, corrections, and forced disconnects tagged with the ISP (ASN) of the connecting IP and the time, counted by ISP and time of day. The infra team runs bidirectional mtr toward that ISP at the same time to check jitter and loss
Confirmed if
Rejections and forced disconnects concentrate on one ISP and rise in the evening, that ISP’s jitter is high at the same time, and movement summed over a few seconds stays within the rules
Ruled out if
Only specific accounts repeat regardless of ISP: real cheating possible. Rising across all ISPs together: a server-side cause where lagging server ticks apply commands in bunches (tick overrun)
Check with
Game server or client logs and metrics
Sources (3)
Source SDK 2013: player.cppValve A per-tick command budget stops speed hacks; a developer comment says stricter limits cause stutter for legitimate players too
ID pt-raid-member · Primary owner Game team (Server development) · Also Game team (Client development)
In raid mechanics where everyone has to react together at a set moment, one lagging player’s late reaction fails the whole party.
Why Group mechanics such as “everyone spread out at once” or “one player presses the button” → Effect The lagging player sees the telegraph late, and their input also arrives late → On screen One player causes a wipe, and the rest of the party feels it happened “because of the laggy player”
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: give mechanic timing windows slack for ping, send telegraphs ahead with server time, design mechanics so one player’s failure doesn’t wipe the party. Client: play received telegraphs in step with server time.
On the graph
Outliers only · RTT of each player who caused a mechanic failure
Where to look
Server mechanic log with the player who caused the failure, their input arrival time, the timing window, and their RTT and loss
Confirmed if
Most of the inputs that caused wipes come from the same one player, whose RTT is clearly higher than the party average and whose inputs arrive just after the timing window
Ruled out if
Failures are spread evenly across party members: the window itself is too short (“Short timing windows eaten up by ping”). The lagging player’s input arrived inside the window and still failed: the server’s window-checking code
ID pt-mob-control · Primary owner Game team (Server development)
Some games hand monster movement to one nearby player’s client to reduce server load. If that player’s connection is bad, the monster moves strangely on everyone’s screen.
Why The server hands monster movement to the client of the nearest (or first-arriving) player → Effect That player’s reports reach the server late or in bunches → On screen Only that monster hitches and then teleports on every nearby screen. It looks fine on the controlling player’s own screen
Hand control to a player with a good connection (by ping and loss), have the server take control back immediately when reports stop, have the server compute important monsters such as bosses itself.
On the graph
Outliers only · Position report interval per monster (by controlling client)
Where to look
Server-side, per monster: the client with control and that client’s report interval, RTT, and loss. The arrival interval of that client’s packets is also visible in a server-side packet capture
Confirmed if
Every monster that moves strangely is controlled by the same one player, that player’s reports are uneven or stop, and handing control to someone else fixes it right away
Ruled out if
Monsters the server computes itself jump the same way: server tick delay or the viewer’s connection. Still jumping after control moves: pathfinding mismatch in command sync
Check with
Game server or client logs and metrics
Learn more
The controlling player sees nothing wrong, so reports only say “the monster is acting weird.” If everyone except one player sees the same monster moving strangely, first check who has control of that monster.
Sources (2)
Authority (Netcode for GameObjects 2.5)Unity In a distributed authority model, each game instance (client) takes authority over some network objects and simulates them
ID pt-heavy-char · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
A character with thousands of items or mails piled up, or an unusually large friend list, block list, or set of buffs, has several times more to load, save, and announce to nearby players than others. It’s slow only on that character, regardless of connection.
Why Thousands of items or event rewards pile up in the inventory and mailbox of a long-played character → Effect Every login, zone move, and save reads and writes that much from the DB, and the equipment and buff data sent to nearby players is large too → On screen Only that character has long loading screens and hitches when opening the inventory or mail. On a server where the game thread waits on saves, nearby players freeze briefly too
Right after login or maintenance, During specific actions, While moving or changing zones
Owner
Primary owner Game team (Server development) · Also Infra team (DB infrastructure)
Game team action items
Cap inventory and mail storage and auto-clean old mail, load only the parts you need in pieces, save only what changed and do it off the game thread.
Infra team action items
Find slow queries that repeat for the same character in the slow query log and pass them to the game team, provide a top list of characters with the most item and mail rows.
Ballpark numbers
If one item is one DB row, a character with 5,000 items reads 5,000 rows on every login. That’s tens of times more than an ordinary character.
On the graph
Outliers only · Login and save time per character, DB rows read per character
Where to look
Slow reads and saves that repeat for the same character ID in the DB slow query log (MySQL slow query log, PostgreSQL log_min_duration_statement), plus a top list of per-character row counts in the item and mail tables
Confirmed if
Slow queries concentrate on a few character IDs, those characters have tens of times the average number of item and mail rows, and it’s just as slow from another PC or connection
Ruled out if
Other characters on the same account or other players are slow too: DB hardware or locking. The character is fine on another PC: the player’s environment
Check with
Infra tools (no game code needed)
Learn more
If the same character is just as slow from a different PC and connection while other characters on the same account are fine, suspect the character data. That’s why reports need the character name.
ID pt-phase · Primary owner Game team (Server development) · Also Game team (Client development)
If two characters are in different channels or instances, or in different “phases” where the visible NPCs depend on quest progress, they see different worlds.
Why The second character is assigned to a different channel, or its quest stage differs → Effect The server doesn’t send that NPC to that character (working as intended) → On screen The NPC is missing on one side only. It looks like a bug but is by design
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: send channel and phase info to the client, add “check both characters’ channel and quest stage” to the QA checklist. Client: show the channel and phase on screen.
On the graph
Outliers only · Nearby entity count per client, channel/phase
Where to look
The two characters’ channel numbers and progress on the relevant quest compared in the game, then checked again with both on the same channel and stage. With server entity send logs, why that NPC wasn’t sent to that character (channel, phase)
Confirmed if
The two characters differ in channel or quest stage, and the NPC shows up once they match
Ruled out if
Same channel and stage but the NPC is missing on one side only: points to spawn messages dropped during loading, spawn data lost in the burst after entering, or an AOI registration race
Check with
The player’s own environment
Learn more
Also check whether quest progress is saved per account or per character. With two characters on the same account, one character’s progress can change the other’s phase.
Sources (2)
Actor Relevancy in Unreal EngineEpic Games The server replicates only relevant actors to each connection and doesn’t send irrelevant ones
ID pt-loading-drop · Primary owner Game team (Client development) · Also Game team (Server development)
Right after you enter a zone, the server sends spawn messages for nearby NPCs, but the client is still loading the map and throws them away.
Why The server sends spawn messages for nearby entities right after processing the entry → Effect The client is still loading and has no message handler yet, so it drops the messages → On screen The server treats them as sent and never resends them. The NPC stays invisible until it leaves view and comes back
Right after login or maintenance, While moving or changing zones
Owner
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: send “ready” when loading finishes, or hold packets received during loading and process them afterward. Server: send nearby information only after receiving “ready.”
Ballpark numbers
If two clients load at the same time on the same PC, or the loading one is a background window, they share CPU and disk and processing is throttled, so that client’s loading can take several times longer. The same bug also surfaces when the server gets faster at processing entries.
On the graph
Outliers only · Loading time per client, messages dropped during loading
Where to look
Count and type of messages the client received and dropped during loading and the time loading finished, compared with the time the server sent the spawn messages. Easy to reproduce by loading two clients at once on the same PC or leaving the loading one as a background window
Confirmed if
The server sent the spawn message for the missing NPC, it arrived before loading finished, and the dropped-message count rose at that time. Happens only on the client with the longer loading
Ruled out if
Spawn message arrived after loading finished but the NPC is still invisible: lost baseline snapshot or entity ID reuse mix-up. The server never sent that NPC’s message at all: AOI registration race or a channel/phase difference
Actor Relevancy in Unreal EngineEpic Games Actors that are no longer relevant are destroyed on the client and replicated anew when they become relevant again
ID pt-aoi-race · Primary owner Game team (Server development)
If a character registers in the AOI grid at the same moment an NPC moves between grid cells, that NPC’s spawn message can be missed.
Why Processing an entry, channel change, or teleport happens at the same instant an NPC moves → Effect That NPC is left out of the “newly visible entities” calculation → On screen A few specific NPCs are invisible, or NPCs that already left are still there
Handle AOI updates on one thread in one order, periodically resync the entire “visible list.”
On the graph
Random spikes · Mismatches between the server’s visible list and the client’s entity list
Where to look
Server log of AOI grid registration, entity cell moves, and spawn/despawn message sends with tick numbers, plus periodic comparison of the server’s “visible list” with the list the client holds
Confirmed if
The missing NPC changed cells in the same tick as that character’s entry or teleport, and there’s no record of a spawn message sent for that NPC
Ruled out if
The spawn message was sent but the client didn’t receive it or dropped it: delivery side (spawn data lost in the burst after entering, spawn messages dropped during loading). Always the same NPC missing: phase or display settings difference
Check with
Game server or client logs and metrics
Sources (2)
Replication Graph in Unreal EngineEpic Games MMORPGs and similar games split the world into a grid with an actor list per cell and send based on the cell the client is in
Actor Relevancy in Unreal EngineEpic Games Relevancy is decided per connection, and actors that are no longer relevant are destroyed on the client
ID pt-baseline · Primary owner Game team (Server development) · Also Game team (Client development)
When the server sends “only what changed since last time,” losing the full state sent once at the start (the baseline) means later changes can’t be applied.
Why The packet with an entity’s full state (baseline) is lost or dropped before processing → Effect The client has nothing to apply later changes to, so it ignores them → On screen That entity is invisible, or suddenly appears much later
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: always resend the baseline until an acknowledgment (ACK) arrives, build changes only against a baseline the client has acknowledged. Client: send the ACK for a baseline only after actually applying it, ask the server again when changes arrive for an unknown entity.
On the graph
Random spikes · Changes received for unknown entities
Where to look
Times the client dropped changes that arrived without a baseline, with entity IDs, matched against when the server sent that entity’s baseline and when the ACK came back. Reproduce in development by adding loss (loss in tc netem, packet loss percentage in Unreal network emulation)
Confirmed if
For the invisible entity, the server sent a baseline and never got an ACK, yet kept sending only changes, which the client dropped
Ruled out if
The baseline was ACKed and applied on the client but the entity is still invisible: missed despawn message or entity ID reuse mix-up
Check with
Game server or client logs and metrics
Sources (4)
Snapshot CompressionGaffer On Games Changes must be built only against a baseline the other side has acknowledged (acked), and the initial state is sent separately
Quake III Arena source: code/server/sv_snapshot.cid Software Delta compression uses the snapshot the client acknowledged as its baseline, and a full snapshot is sent when the baseline gets too old
tc-netem(8) — Linux manual pageiproute2 A test tool that adds delay and jitter (delay TIME JITTER) and loss (loss random PERCENT) to outgoing packets to mimic a real network
Using Network Emulation in Unreal EngineEpic Games Test with minimum and maximum latency and a packet loss percentage on server and client; set from the console, e.g., NetEmulation.PktLag
ID pt-ghost · Primary owner Game team (Server development) · Also Game team (Client development)
The reverse case: if the “it’s gone” message is missed, NPCs or players that already died or left stay on your screen only.
Why Death, leave, or out-of-view messages are lost or arrive out of order → Effect The client thinks the entity is still there → On screen A monster that doesn’t react when hit, a player who already left still standing there
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: periodically send the “currently visible list.” Client: remove entities not on the list, hide entities that should be moving but haven’t updated in a long time.
On the graph
Random spikes · Entities that exist only on the client
Where to look
The server’s “currently visible list” compared with the client’s entity list, counting entities that exist only on the client, and despawn message send and receive logs matched by entity ID
Confirmed if
The server sent the ghost entity’s despawn message but the client has no record of receiving it, or the despawn arrived before the spawn and the order is reversed
Ruled out if
The entity is still on the server’s visible list too: the server failed to clean up the entity. Right after a new entity appeared with the same ID: entity ID reuse mix-up
ID pt-spawn-burst · Primary owner Game team (Server development) · Also Game team (Client development)
The moment you enter a zone, the server sends spawn data for tens to hundreds of nearby entities all at once. If it goes over an unreliable channel, or the receive buffer overflows while the client is loading and can’t read the socket, part of it disappears and never comes back.
Why Spawn data arrives in a short burst right after entering → Effect A loading client reads the socket late and the OS receive buffer overflows, or a large UDP packet is fragmented and losing a single fragment loses the whole packet. On an unreliable channel, nothing is resent either → On screen A few NPCs are missing only on the slower-loading client. They show up after leaving view and coming back
Right after login or maintenance, While moving or changing zones
Owner
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: always send spawn and despawn messages over a reliable channel with guaranteed retransmission, send initial data in chunks. Client: receive on a thread separate from loading, increase the receive buffer size.
Ballpark numbers
The default UDP receive buffer on a PC varies by OS but is usually tens to hundreds of KB. If the entry data for a crowded town is bigger than that, even a brief pause in reading the socket during loading overflows it.
On the graph
Surge after opening · Data received right after entering, missed spawn messages
Where to look
Spawn messages the server sent right after entry compared with the number the client received, and which channel (reliable or unreliable) they went over. In a server-side packet capture, the volume sent to that player right after entry and fragmented packets (Wireshark filter ip.flags.mf == 1 || ip.frag_offset > 0)
Confirmed if
The client received fewer than were sent, the missing ones cluster in the burst right after entry, and they went over an unreliable channel or large packets were fragmented. Happens more often on the slower-loading client
Ruled out if
Sent and received counts match but entities are still invisible: dropped after receipt (spawn messages dropped during loading) or an AOI calculation problem. Missing at random times unrelated to entry: packet loss on the connection
ID pt-id-reuse · Primary owner Game team (Server development) · Also Game team (Client development)
If the server reuses the same entity ID when a dead NPC respawns, a client that missed the despawn message in between mistakes the new NPC for the old one.
Why An NPC dies and respawns with the same entity ID → Effect A client that missed the despawn message ignores the spawn message as an “already known entity,” or leaves the entity in its dead state → On screen The NPC is missing on one screen only or appears lying dead, and sometimes shows up looking like a different NPC
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: add a generation number to entity IDs to tell reuses apart. Client: when a spawn message arrives for a known ID, delete the existing entity and create a new one.
On the graph
Random spikes · Spawn messages for already known IDs
Where to look
Server log of creation and deletion times per entity ID (with generation number if any), plus the count of spawn messages the client received for known IDs and of delete-and-recreate cases treated as “no change” during AOI updates
Confirmed if
The invisible or dead-looking NPC has the same ID as an NPC that just died, and in between the client didn’t receive the despawn message or the server sent neither the despawn nor the spawn message
Ruled out if
IDs carry a generation number that is also used in comparisons: not this cause. Invisible even though the ID wasn’t reused: points to a lost spawn message
Check with
Game server or client logs and metrics
Learn more
It happens on the server side too. If the list of entities in view is compared by ID only, an NPC that died and respawned with the same ID between two AOI updates looks like “no change,” and neither the despawn nor the spawn message is sent. If AOI update timing differs from player to player, only the clients whose update lands on that moment are affected.
Sources (2)
Entity struct (Entities 1.3)Unity An Entity consists of an Index and a generation number (Version), which tells whether a reused Index is still valid
ID pt-port-collision · Primary owner Game team (Client development) · Also Game team (Server development)
If the client is built to use a fixed local port, a second client on the same PC either can’t get the port or ends up splitting packets with the first.
Why Two clients try to open the same local UDP port (forcing a share with a reuse option) → Effect The OS delivers incoming packets to only one socket, or doesn’t guarantee which one gets them. The router and server also see both clients as the same address → On screen One client misses world packets, so NPCs and other players are invisible, or it disconnects
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: let the OS pick the local port automatically (bind to port 0). Server: tell connections apart by a session token issued per connection.
On the graph
Outliers only · Packets received per client
Where to look
On the player’s PC with both clients running, netstat -ano -p udp in a command prompt to see the local UDP ports each game process (PID) has open. On the server side, whether the two sessions come in from the same public IP and the same port
Confirmed if
Both game processes are bound to the same local port, or the server sees both sessions as the same IP and port. Fine with only one running
Ruled out if
The two clients use different local ports but one still misbehaves: sessions keyed by IP or device, or a multi-client restriction
ID pt-session-key · Primary owner Game team (Server development) · Also Game team (Client development)
If the server or an intermediate server identifies connections by IP or device ID, it treats two clients on the same PC (same public IP) as one person.
Why The session table is keyed by IP, or by IP + device ID → Effect The second client’s data overwrites or gets mixed into the first session → On screen One side can’t see NPCs, and the other disconnects or receives someone else’s data
One client on the same PC, Same household, Specific region/ISP
When
Right after login or maintenance
Owner
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: identify every connection by a unique session token on both the server and intermediate servers, and be sure to fix this, because several people in the same home (behind router NAT) and players on mobile networks where the carrier shares one IP among many subscribers (CGNAT) hit the same problem. Client: use the session token each running client received separately.
On the graph
Outliers only · Concurrent sessions from the same public IP, session overwrites
Where to look
Server and intermediate server logs of the key used to look up the session, the session token, and the client IP and port, checking whether the existing session changed the moment a second connection came in from the same IP. Reproduces by starting two clients one after another on the same PC
Confirmed if
The moment the second client connects, the first session’s address or character data changes, and the same disconnects appear for other players behind the same router or mobile network (CGNAT)
Ruled out if
Two sessions from the same IP are kept apart with different tokens: not this cause. Two processes use the same local port: fixed UDP port collision
ID pt-multiclient · Primary owner Game team (Client development) · Also Game team (Server development)
If an anti-cheat module or server policy limits multiple clients on one PC, the second client is blocked from launching or connecting, or the first one gets disconnected. Some games only block features on the extra client.
Why The anti-cheat module detects a duplicate launch, or the server limits extra connections from the same device → Effect The second launch or connection is refused, or one side is disconnected. Rarely, only some features on the extra client are blocked → On screen Can’t connect, or one side disconnects. In games that only block features, NPCs or shops are invisible on one side only
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: if you restrict it, show a clear message, add a QA exception to the anti-cheat module. Server: add a QA exception to the same-device connection limit too.
On the graph
Outliers only · Connection refusals and disconnects by reason (duplicate login)
Where to look
The message shown when the second client launches and the disconnect message on the first. Whether the server’s connection refusal and forced disconnect logs record reason codes such as duplicate login or same device
Confirmed if
A refusal message appears the moment the second client launches or connects, or the first one is disconnected with a duplicate login reason, and there’s no problem with only one client running
Ruled out if
Both connect without any refusal or disconnect reason but NPCs are invisible on one side: fixed UDP port collision, sessions keyed by IP or device, or a loading or display cause
Check with
The player’s own environment
Sources (1)
CreateMutexW function (synchapi.h)Microsoft If a named mutex already exists, ERROR_ALREADY_EXISTS is returned, which is used to detect duplicate launches and allow only a single instance
ID pt-background · Primary owner Game team (Client development) · Also External (External)
For a client in a background window, the game, engine, and OS cut its frame rate and processing. Received packets aren’t processed in time, so they back up or overflow.
Why Background frame limits in game options or the graphics driver (e.g., the NVIDIA driver lets you pick 20–200 per second), power saving, the engine’s background pause setting. The OS also gives CPU and GPU priority to the window in front (foreground) → Effect Fewer packets are processed per frame, so the queue builds up, and packets are dropped when the receive buffer overflows → On screen When the window comes to the front, things appear all at once, or some NPCs never show up
Primary owner Game team (Client development) · Also External (External)
Game team action items
Keep network receiving going on a thread separate from the game loop, guarantee a minimum processing rate in the background, turn on the engine’s run-in-background setting (runInBackground in Unity).
External action items
Tell players to turn off the graphics driver’s background frame limit and the PC’s power-saving mode.
Ballpark numbers
In Unity, with runInBackground off, the game loop stops the moment the window loses focus. If receiving happens only in that loop, no packets are processed at all in the meantime.
On the graph
Gap then burst · Client frame interval, packets processed per frame
Where to look
On the same PC, one window in front and the other behind, swapping roles and comparing. Frame intervals of both processes measured with PresentMon, and with game logs, window focus state and packets processed per frame
Confirmed if
Frame intervals grow sharply (with a driver limit, flattening at the interval that matches the configured frame rate) or processing stops only while the window is in the background, and switching windows moves the problem to the other client
Ruled out if
It happens the same way in the foreground window: not background throttling. Always the same client misbehaving regardless of window position: display settings or version difference
Check with
The player’s own environment
Sources (4)
Application.runInBackgroundUnity runInBackground defaults to false, in which case the app pauses in the background
ID pt-asset-lock · Primary owner Game team (Client development)
If two clients write to the same cache folder at the same time or lock its files, one of them can’t load NPC models or textures.
Why Two clients write the cache and patch files in the same install folder at the same time → Effect Loading fails because a file lock failed or a half-written file was read → On screen A name tag with no character model, or a transparent NPC
Right after login or maintenance, While moving or changing zones
Owner
Primary owner Game team (Client development)
Game team action items
Use a separate cache folder per client, retry when a file lock fails, show at least a default model when loading fails.
On the graph
Outliers only · Asset loading failures per client
Where to look
On the player’s PC, Process Monitor filtered to the game install and cache folder paths, showing the file open and write results of both game processes. With client logs, asset loading failures and the file open error code (ERROR_SHARING_VIOLATION)
Confirmed if
Opening the missing model’s file ended in a sharing violation or lock failure while the other client was writing that file. It goes away with only one client running or with separate install and cache folders
Ruled out if
The same model is missing even with only one client running: file corruption or a client version/data mismatch. Files open fine but nothing is drawn: memory/VRAM shortage
Check with
The player’s own environment
Sources (2)
Creating and Opening FilesMicrosoft A file opened without a sharing mode can’t be opened by another process, which gets ERROR_SHARING_VIOLATION
Process MonitorMicrosoft Records file system, registry, and process activity in real time and can filter on any field, such as path
ID pt-vram · Primary owner Game team (Client development) · Also External (External)
When two clients share graphics memory, there’s no room to load newly needed models and textures, and some of them don’t get drawn.
Why Two clients share VRAM and RAM. The OS may also shrink a background window’s graphics memory allowance first → Effect The engine can’t load new models and textures, or keeps evicting and reloading them → On screen NPCs appear late, look blurry, or are invisible, and the game stutters
While moving or changing zones, When crowds gather
Owner
Primary owner Game team (Client development) · Also External (External)
Game team action items
Adjust quality automatically to fit the memory budget, show a fallback model when loading fails.
External action items
Tell players who run two clients at once to lower graphics quality or use low-spec mode, publish recommended VRAM and RAM specs.
On the graph
Hits a ceiling · Dedicated GPU memory usage per process
Where to look
In Task Manager on the player’s PC, the dedicated GPU memory column added to the “Details” tab, with the two clients’ combined usage compared against the graphics card’s VRAM. On the game side, the budget (Budget) and current usage (CurrentUsage) reported by DXGI’s QueryVideoMemoryInfo, logged
Confirmed if
The two clients’ combined usage flattens near VRAM capacity, and model and texture loading failures cluster when current usage goes over budget. It goes away with lower quality or only one client running
Ruled out if
Invisible even with VRAM to spare: simultaneous access to cache or asset files, or a display settings difference
Check with
The player’s own environment
Sources (3)
Residency (Direct3D 12)Microsoft The video memory budget can shrink a lot when you switch to another app, and going over budget causes stalls or failed allocations; outside the foreground, even reservations aren’t guaranteed
GPUs in the task managerMicrosoft Adding columns to Task Manager’s Details tab shows dedicated and shared GPU memory usage per process; dedicated GPU memory is the graphics card’s VRAM
ID pt-display-option · Primary owner Game team (Client development)
If settings such as a visible player limit, hidden NPC name tags or models, or low-spec mode differ between two clients, they see different things.
Why Only one client has a “limit nearby characters shown” setting or low-spec mode on → Effect Distant or low-priority NPCs aren’t drawn (working as intended) → On screen The NPC is missing on one side only
Show when an entity is hidden by a setting, keep settings files separate per client so they don’t get mixed up.
On the graph
Outliers only · Entities drawn on screen per client
Where to look
The two clients’ visible player limit, name tag/model hiding, and low-spec mode settings side by side, with one set to match the other. Also whether the two clients share one settings file and overwrite each other
Confirmed if
With matching settings both screens show the same thing, and the missing NPCs were distant entities beyond the display limit or low-priority entities
Ruled out if
Still missing on one side with identical settings: points to a channel/phase difference or a lost spawn message
ID pt-version · Primary owner Game team (Client development) · Also Game team (Server development)
If the second client is a different install or isn’t fully patched, it doesn’t recognize new NPC IDs the server sends and silently ignores them.
Why An install in a different folder, or a client launched mid-patch → Effect Unknown NPC IDs or model IDs are skipped → On screen Only newly added NPCs are invisible on one side
Primary owner Game team (Client development) · Also Game team (Server development)
Game team action items
Client: send the data version on connect, when an unknown ID arrives log it and show a placeholder. Server: check the data version on connect and, if it differs, refuse the connection and prompt a patch.
On the graph
Outliers only · Unknown IDs received per client version
Where to look
The two clients’ executable paths and the client and data versions shown on screen or in logs. On the game side, the data version sent on connect and how many times unknown NPC or model IDs were received and skipped, logged
Confirmed if
The two clients differ in version or install folder, the invisible NPCs were added in a recent patch, and they show up in the fully patched install
Ruled out if
Same version and install folder but missing on one side only: points to a channel/phase difference or a loading or delivery cause
ID pt-priority · Primary owner Game team (Server development) · Also Game team (Client development)
If the server caps how much it sends per connection and sends the nearest things first, a connection with a low cap gets distant NPCs late or not at all.
Why In crowded places, the server sends in order of importance within each connection’s send cap → Effect A connection with a low bandwidth estimate (e.g., a background window that’s slow to acknowledge) keeps pushing back entities further down the list → On screen Distant NPCs appear late or not at all on one side only
Primary owner Game team (Server development) · Also Game team (Client development)
Game team action items
Server: raise the priority of deferred entities the longer they wait (prevent starvation), guarantee a minimum update interval. Client: send acknowledgments on time even in the background so the bandwidth estimate doesn’t drop.
On the graph
Rises with load · Deferred entities per connection, bytes sent per connection
Where to look
Server-side, per connection: bytes sent per tick, send cap (estimated bandwidth), entities deferred for lack of room, and time since each entity was last sent. In Unreal, Networking Insights shows packet sizes per connection and the replicated objects inside them
Confirmed if
The invisible NPC is an entity deferred for a long time on that connection, that connection’s cap is lower than the others’, and deferred entities grow as it gets more crowded
Ruled out if
Nothing deferred and that NPC was sent on time: a stage after sending (receive buffer, loading, display settings). Every connection is at its cap: a server-wide send volume or AOI design problem
Check with
Game server or client logs and metrics
Sources (3)
Actor Priority in Unreal EngineEpic Games When bandwidth is saturated, actors to replicate are chosen by priority (distance, line of sight, time since last replication); not every actor is replicated every time
State SynchronizationGaffer On Games Priority accumulation: entities that didn’t fit in this packet go first in the next one, and the bandwidth limit is adjusted in real time
Networking Insights in Unreal EngineEpic Games Shows the size of packets sent and received per connection and the replicated objects and properties inside them
ID pt-clock-hold · Primary owner Game team (Client development)
If the client’s estimate of the server time is wrong, it holds back freshly arrived entity data as “still in the future” or discards it as “too old.”
Why One client’s estimate of the server time is far off (measured during loading, or after waking from sleep) → Effect The interpolation reference time and the entity data’s timestamp don’t match → On screen Entities appear late or look frozen
After sitting idle, Right after login or maintenance
Owner
Primary owner Game team (Client development)
Game team action items
Redo time sync periodically and reset immediately when the gap is large, don’t use values measured during loading or right after waking from sleep.
On the graph
Outliers only · Server time estimate error per client
Where to look
Client log of the estimated server time, RTT, when time sync was redone, and how many times entity data was held back or discarded. Try to reproduce right after loading or right after waking from sleep
Confirmed if
Only the affected client’s estimate error exceeds the reset threshold (hardResetThresholdSec in Unity, 0.2 s by default), there are records of entity data held back as future or discarded as past, and redoing time sync fixes it right away
Ruled out if
Estimate error is small but entities still appear late: points to per-connection send budget and priority, or loading
Check with
Game server or client logs and metrics
Sources (2)
NetworkTimeSystem class (Netcode for GameObjects 2.5)Unity If the time gap exceeds hardResetThresholdSec (0.2 s by default), the clock is forced into sync; otherwise adjustmentRatio speeds it up or slows it down a little at a time
ID rt-wireless · Primary owner External (External) · Also Infra team (Server infrastructure), Game team (Server development), Game team (Client development)
Wi-Fi and mobile networks retransmit a few times on the wireless link and drop the packet if that still fails. TCP resends the dropped packet only much later.
Why A weak signal or heavy interference makes wireless transmissions fail several times in a row → Effect Once the wireless device’s retry limit (usually a few to ten-odd attempts) is exceeded, the packet is dropped → On screen Freeze for as long as the TCP retransmission wait, while later packets sit in the receive buffer and then fast-forward
Primary owner External (External) · Also Infra team (Server infrastructure), Game team (Server development), Game team (Client development)
Game team action items
Server: turn on TCP_NODELAY (with Nagle on, RACK has no following packets to use for detecting loss), and while retransmission has the connection blocked, keep only the latest state update pending (cap what queues in the kernel with TCP_NOTSENT_LOWAT). Client: turn on TCP_NODELAY (the client OS recovers losses in the input direction), show network status on screen when losses cluster or ping spikes.
Infra team action items
Speed up loss recovery with RACK-TLP (the server can’t prevent wireless loss; the most it can do is recover faster), confirm the current Linux defaults net.ipv4.tcp_recovery=1 (RACK) and net.ipv4.tcp_early_retrans=3 (TLP) haven’t been changed.
External action items
Tell players to use a wired connection or 5 GHz/6 GHz Wi-Fi, and to move the router or change its channel.
Ballpark numbers
At 1% wireless loss, 1 in every 100 game packets disappears. Receiving 10 a second, that’s a hitch about once every 10 seconds. Without RACK-TLP, each loss freezes the game for an RTO (ping + 200 ms or more).
On the graph
Outliers only · Per-connection retransmission rate, per-connection RTT (ping)
Where to look
From the player’s PC, ping both the router (gateway) and the game server a few hundred times and compare loss and latency spread, then measure again on a wired connection or mobile data. On the server, check that player’s connection in ss -ti for retrans and rtt (mean/deviation)
Confirmed if
Ping to the router already shows loss or erratic latency, and it goes away on a wired connection. From the server, only that player’s connection shows high retrans and RTT deviation
Ruled out if
Clean up to the router with loss starting beyond it: points to the ISP or route (“Bottleneck queue overflow (congestion loss),” “Route change / bad ECMP path”). If several players on the same ISP get worse at the same time, start with the ISP segment
Check with
The player’s own environment
Learn more
Wireless retries add jitter (a few ms per retry), and only packets that exceed the retry limit become losses. So as wireless quality degrades, symptoms grow in this order: “jitter → occasional freezes → frequent freezes.” While a device moves between access points (roaming), it can lose packets in a row for tens of milliseconds to several seconds. Mobile networks retransmit heavily on the radio link to the cell tower, so trouble there more often shows up as latency spikes of hundreds of ms than as loss.
net/wireless/core.cLinux kernel Default retry limits in the Linux wireless stack: 7 for short frames, 4 for long frames (dot11ShortRetryLimit, dot11LongRetryLimit)
Wi-Fi roaming support in Apple devicesApple When a device moves to a new AP, it can’t send data until authentication with the new AP completes, which can take a few seconds with 802.1X
IP SysctlLinux kernel tcp_recovery defaults to 0x1 (RACK), tcp_early_retrans defaults to 3 (TLP on); TCP_NOTSENT_LOWAT and tcp_notsent_lowat cap the amount of data not yet sent
tcp(7) — Linux manual pageLinux man-pages TCP_NODELAY turns off the Nagle algorithm so even small data goes out immediately
ID rt-queue-drop · Primary owner Infra team (Network infrastructure) · Also External (External), Game team (Client development)
When the queue at the narrowest point fills up, such as a home router, a link between ISPs, or a data center uplink, newly arriving packets are dropped.
Why Video, downloads, and other users’ traffic fill up the bottleneck → Effect While the queue is full, newly arriving packets are dropped one after another (tail drop). Packets that get in wait at the back of the full queue → On screen Several packets vanish at once, causing a long freeze then fast-forward; common in the evening
Primary owner Infra team (Network infrastructure) · Also External (External), Game team (Client development)
Game team action items
Show network status on screen when losses cluster or ping spikes (mention that a large transfer on the same connection may be the cause).
Infra team action items
Keep headroom on data center links, check output drop counters on our links and switch ports, route around a congested ISP segment through another link or peering.
External action items
Tell players to use SQM (fq_codel, CAKE) and ECN on their router (so senders slow down before the queue overflows), ask the ISP to add capacity at the bottleneck.
Ballpark numbers
When a queue overflows, a large share of incoming packets can vanish at once over tens of milliseconds. Losing several packets in a row, and even losing the retransmissions, is common, so recovery often has to wait for an RTO.
On the graph
High at certain hours · Retransmission rate, RTT (ping)
Where to look
Server retransmission rate (TcpRetransSegs ÷ TcpOutSegs deltas from nstat run every minute) and per-connection RTT, split by region, ISP, and time of day, alongside output discards (ifOutDiscards) on our links and switch ports. Compare mtr runs to the affected region at peak and off-peak hours
Confirmed if
Retransmission rate rises only at evening peak, and RTT climbs just before the loss (the queue filling up). mtr shows loss and latency growing together from one hop to the end, only at peak hours
Ruled out if
No RTT rise before the loss points to “Policer drops excess traffic.” Similar loss at every hour points to “Physical errors (bad cable, optics, connectors)” or “Route change / bad ECMP path”
RFC 2863: The Interfaces Group MIBIETF ifOutDiscards: outbound packets discarded even though no error was detected, for example to free up buffer space
An Internet-Wide Analysis of Traffic PolicingGoogle Queue overflow raises queuing delay and RTT before the loss, while policing drops the excess with no RTT increase (SIGCOMM 2016)
ID rt-burst · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure)
When a server sends a whole tick of updates for thousands of players in one instant, a switch’s small buffer or a cloud instance’s short-term limit overflows in under 1 ms and some packets are dropped.
Why At the start of each tick, the server sends everyone’s packets all at once → Effect A switch port buffer where traffic from many servers converges (hundreds of KB to a few MB per port) or a cloud instance limit overflows for an instant (average utilization stays low) → On screen Many players teleport or hitch at the same time; averaged metrics don’t reveal the cause
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure)
Game team action items
Spread each tick’s sends across the tick (per-connection pacing does little for thousands of connections that all send at tick start), stagger tick start times across servers, cap the rate of connections that send large data with SO_MAX_PACING_RATE.
Infra team action items
Servers/OS: cap the whole server’s send rate (a shaper in the server OS, Linux tc), smooth out a single connection’s bursts with pacing (Linux fq qdisc, BBR). Network: use deep-buffer switches, check switch port output drop counters at short intervals (average utilization won’t show them).
Ballpark numbers
A 10 Gbps port can send about 1.25 MB in 1 ms. When several servers’ ticks line up and converge on one port, the buffer fills almost instantly.
On the graph
Rises with load · Switch port output drops, retransmission rate
Where to look
Output discards (ifOutDiscards) collected every few seconds on the switch port the server connects to and on the port above it; in the cloud, bw_out_allowance_exceeded and pps_allowance_exceeded from ethtool -S. Line these up against retransmissions collected at the same moments with bcc tcpretrans
Confirmed if
Output discards or allowance overruns grow while per-minute average utilization stays low, and they scale with concurrent users and crowding in one spot. Retransmissions hit many connections on that server at the same moment, with no concentration on particular player IP ranges (ISP or region)
Ruled out if
CRC and input errors rising on the same port point to “Physical errors (bad cable, optics, connectors).” NIC drop counters or softnet dropped rising on the receiving server point to “Packet drops on the receiving host”
Check with
Infra tools (no game code needed)
Learn more
Pacing works per connection. When thousands of connections each send one or two packets at tick start, per-connection pacing does little to spread them out, so the game server has to split up its send timing itself. By contrast, when one connection sends a large amount of data, the NIC cuts tens of KB into packet-sized pieces and sends them back to back (TSO), and pacing spreads that kind of burst out well.
RFC 2863: The Interfaces Group MIBIETF ifOutDiscards: outbound packets discarded even though no error was detected, for example to free up buffer space
ID rt-policer · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)
ISP plans, cloud instance limits, and DDoS protection devices sometimes drop packets over a set rate right away, without queuing them.
Why Momentary send volume exceeds the allowed rate or allowed burst → Effect Packets over the limit are dropped immediately, with no queue (policing) → On screen Each large burst loses several packets, causing a freeze then fast-forward, while the average rate looks below the limit
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)
Game team action items
Spread each tick’s burst across the tick to keep momentary send volume under the allowed burst, and if you hit a packets-per-second limit, combine one tick’s messages into one packet.
Infra team action items
Network: check the device’s policer exceed counters, replace the policer with a shaper, raise the allowed burst. Servers/OS: check cloud limit-exceeded metrics (on AWS, bw_out_allowance_exceeded and pps_allowance_exceeded in ethtool -S), move to a larger instance, pace on the server (Linux fq qdisc).
Ballpark numbers
A shaper (queues packets and delays them) adds latency; a policer (drops them immediately) adds loss. A TCP game connection can freeze for hundreds of ms on a single loss, so when traffic only briefly goes over the limit, the policer usually does more damage.
On the graph
Hits a ceiling · Send volume at short intervals, policer and allowance exceed counters
Where to look
Exceed and drop counters on the device doing the policing; in the cloud, bw_out_allowance_exceeded and pps_allowance_exceeded from ethtool -S. For connections that lost packets, RTT just before the loss from ss -ti rtt or a packet capture
Confirmed if
Exceed counters rise, and send volume at short intervals looks flat, as if cut off at a fixed value. RTT doesn’t rise before the loss, and several packets vanish at once only during large bursts
Ruled out if
RTT rising first, before the loss, points to queue overflow (“Bottleneck queue overflow (congestion loss),” “Send bursts overflow shallow buffers”). Exceed counters unchanged: a different cause
An Internet-Wide Analysis of Traffic PolicingGoogle Policed transfers see loss rates 6 times higher on average, and pacing or shaping can achieve the same goal. Policing drops the excess with no RTT increase, while queue overflow raises RTT before the loss (SIGCOMM 2016)
ID rt-physical · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), External (External)
Damaged cables, dirty fiber connectors, and worn-out optics cause bit errors, and network equipment silently drops the corrupted packets.
Why A bad cable, optic, or connector flips bits → Effect Equipment drops packets whose checksum (CRC) doesn’t match → On screen Only players whose traffic takes that path keep getting short hitches followed by fast-forward, at any time of day
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), External (External)
Infra team action items
CRC errors accumulate on the receiving end of the bad direction, so check both ends. Network: check CRC and input error counters on device ports, check optical power (the switch’s optics info), clean fiber connectors, replace cables and optics. Servers/OS: check rx_crc_errors in ethtool -S on the server (the name varies slightly by driver), check optical power (ethtool -m), replace the server-side cable or NIC.
External action items
If the problem is in the player’s home, tell them to replace the Ethernet cable or router; if it’s on the ISP’s line, ask the ISP to inspect the line.
Ballpark numbers
Even 0.1% loss is one in every 1,000 game packets. With dozens of players on that path, someone hitches every few seconds. Bit errors hit larger packets more often.
On the graph
Outliers only · CRC errors per port, retransmission rate per server and per port
Where to look
CRC counters at both ends of the link: on servers, rx_crc_errors in ethtool -S or crc in ip -s -s link; on switches, the port’s FCS errors (dot3StatsFCSErrors) and input errors (ifInErrors). For fiber links, received optical power from ethtool -m and the switch’s optics info
Confirmed if
CRC errors on one port climb steadily at all hours, and only servers and connections through that port have high retransmission rates. Received optical power is lower than on other links of the same type
Ruled out if
CRC flat while only output discards rise points to queue overflow (“Send bursts overflow shallow buffers,” “Bottleneck queue overflow (congestion loss)”). Late collisions on one side rising together with CRC errors on the other point to “Duplex mismatch”
Check with
Infra tools (no game code needed)
Sources (3)
Interface statisticsLinux kernel rx_crc_errors counts packets the receiving interface flagged with CRC errors; check with ip -s -s link and ethtool -S
ethtool(8) — Linux manual pageethtool -S for NIC and driver statistics, -m for optic module (SFP+, QSFP) EEPROM and optical diagnostics
ID rt-duplex · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)
If one end autonegotiates while the other has speed and duplex hard-set, one side runs half duplex and loses packets to collisions whenever load picks up.
Why Speed and duplex hard-set on only one of the two devices → Effect One side runs full duplex and the other half duplex, causing collisions and late collisions → On screen Fine normally, but as traffic grows, everyone going through that device freezes then fast-forwards
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)
Infra team action items
Set both ends to autonegotiate, or hard-set both ends to the same values. Network: check speed and duplex in the switch port status, and in the port counters check whether late collisions grow on the half-duplex side and CRC errors and runts (frames that are too short) grow on the full-duplex side. Servers/OS: check speed and duplex with ethtool.
Ballpark numbers
Autonegotiation is mandatory on 1 Gbps copper, and half duplex doesn’t exist at all at 10 Gbps and above. So these days it mostly happens on old gear at 100 Mbps or below, on management ports, and on some carrier circuit hand-offs.
On the graph
Rises with load · Port late collisions and CRC errors, retransmission rate
Where to look
Actual speed and duplex on both ends of the link: on servers, ethtool run with just the interface name; on switches, port status or dot3StatsDuplexStatus via SNMP. Late collisions (tx_window_errors on servers, dot3StatsLateCollisions on switches) and CRC errors alongside
Confirmed if
One side reports half duplex and the other full duplex. Every time traffic grows, late collisions rise on the half-duplex side and CRC errors rise on the full-duplex side
Ruled out if
Speed and duplex match on both ends and only CRC errors rise: “Physical errors (bad cable, optics, connectors).” Links at 10 Gbps and above have no half duplex, so rule this cause out for them
Interface statisticsLinux kernel tx_window_errors counts transmissions that failed from late collisions; rx_crc_errors counts packets received with CRC errors
ethtool(8) — Linux manual pageethtool ethtool -s sets speed, duplex, and autonegotiation (speed, duplex, autoneg); given only the interface name, ethtool shows the current settings
ID rt-host-drop · Primary owner Infra team (Server infrastructure)
Packets reach the server but get dropped, because the NIC’s ring buffer (which briefly holds arriving packets) overflows or the kernel cores that handle receive processing are saturated.
Why A surge of players, interrupts piled on one core, CPU steal on a virtual machine, or an overloaded virtual switch → Effect Drops at the ring buffer (rx_missed_errors and similar; the name varies by driver) or at the kernel receive queue (softnet dropped) → On screen When crowds gather, input registers late and hitches hit the whole server at once
Add drop counters such as rx_missed_errors in ethtool -S and dropped in /proc/net/softnet_stat to monitoring, enlarge ring buffers (ethtool -G), spread RSS and interrupts across multiple cores, keep the game thread and receive-processing cores separate, keep CPU headroom, and on virtual machines check CPU steal and virtual switch load.
Ballpark numbers
Packets the server drops on receive are resent by the client. So they barely show up in the server’s retransmission metrics and appear first in drop counters such as rx_missed_errors in ethtool -S (the name varies by driver) and in dropped in /proc/net/softnet_stat.
On the graph
Hits a ceiling · softirq utilization per core, NIC drop counters
Where to look
Drop counters in ethtool -S (rx_missed_errors and similar; on mlx5, rx_out_of_buffer and rx_discards_phy), missed in ip -s -s link, and the 2nd column (dropped) and 3rd column (time_squeeze) of /proc/net/softnet_stat, plus %soft (softirq processing) per core from mpstat -P ALL. On virtual machines, %steal too
Confirmed if
Drop counters or softnet dropped rise when crowds gather, and %soft on the receive-processing cores sits near 100% and can’t go higher. Input slows on every connection to that server at the same time
Ruled out if
Server drop counters flat while retransmissions concentrate on connections from a specific region or ISP: loss on the path. When the path loses packets the server sent, the server’s nstat TcpRetransSegs rises while these counters stay flat
Check with
Infra tools (no game code needed)
Sources (10)
Interface statisticsLinux kernel rx_missed_errors counts packets the host missed because no buffer was available (counted into drop in /proc/net/dev); check with ip -s -s link
Ethtool countersLinux kernel The mlx5 driver’s rx_out_of_buffer (no buffer in the receive queue) and rx_discards_phy (dropped for lack of port buffer)
net/core/net-procfs.cLinux kernel /proc/net/softnet_stat has one line per CPU, in hex; the 2nd column is dropped and the 3rd is time_squeeze
Documentation for /proc/sys/net/Linux kernel netdev_max_backlog: limit of the receive queue that holds packets arriving faster than the kernel can process them
proc_stat(5) — Linux manual pageLinux man-pages steal in /proc/stat: time spent in other operating systems when running in a virtualized environment
mpstat(1) — Linux manual pagesysstat mpstat -P ALL shows per-core utilization; %soft is time spent servicing softirqs, and %steal is time spent waiting while the hypervisor serviced another virtual CPU
net/ipv4/proc.cLinux kernel TcpRetransSegs as shown by nstat (RetransSegs under Tcp)
ID rt-stateful-fw · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development), Game team (Client development)
Firewalls and Linux connection tracking (conntrack, which records passing connections in a table) drop packets when the table is full or when they decide a packet doesn’t match the connection’s state.
Why The connection tracking table is full (table full), or traffic takes a different path each way so only one direction passes through the firewall (asymmetric routing) → Effect The firewall treats the packets as belonging to an “unknown connection” or carrying a “sequence number outside the window” and drops them → On screen A full table blocks new connections; a path mismatch makes only players on that path disconnect after repeated retransmissions
When crowds gather, Right after login or maintenance, Randomly
Owner
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development), Game team (Client development)
Game team action items
Server: in case the table fills up, throttle connection surges with a login queue, reuse connections so you don’t keep opening short-lived ones (including server-to-server calls), proactively close connections whose heartbeats have stopped. Client: when connecting fails or the connection drops, retry at growing, randomized intervals (so players don’t all pile back in at once while the table is full).
Infra team action items
Network: enlarge the firewall’s connection tracking table, exempt game ports from connection tracking, align routing so both directions pass through the same firewall, check the firewall’s TCP window checking settings. Servers/OS: enlarge the Linux table (nf_conntrack_max), exempt game ports from connection tracking (NOTRACK), check the TCP window checking setting (nf_conntrack_tcp_be_liberal), and on AWS also check conntrack_allowance_exceeded.
Ballpark numbers
The default Linux conntrack limit (nf_conntrack_max) ranges from tens of thousands to hundreds of thousands of entries depending on memory. When the current count (nf_conntrack_count) reaches the limit, the log shows “nf_conntrack: table full, dropping packet”.
On the graph
Hits a ceiling · conntrack entry count (nf_conntrack_count), failed new connections
Where to look
On Linux servers, nf_conntrack_count and nf_conntrack_max, “nf_conntrack: table full, dropping packet” in dmesg, and drop and invalid in /proc/net/stat/nf_conntrack (one line per core, in hex). On firewalls, session table usage and drop logs; on AWS, conntrack_allowance_exceeded in ethtool -S
Confirmed if
Entry count flattens at the limit, and table full logs and drop, or conntrack_allowance_exceeded, rise at the same moment. With asymmetric routing, the table has headroom, but invalid and firewall drop logs rise for connections on a specific path
Ruled out if
Entry count far from the limit, with invalid and drop logs flat: a different cause. Table has headroom but the firewall’s CPU or packets per second is maxed out: “Middlebox over capacity (firewall, IPS, DDoS protection)”
Check with
Infra tools (no game code needed)
Sources (5)
Netfilter Conntrack Sysfs variablesLinux kernel nf_conntrack_max defaults to the number of hash buckets (memory ÷ 16384, 1,024–262,144), the current count is nf_conntrack_count, and nf_conntrack_tcp_be_liberal marks only out-of-window RSTs as INVALID
net/netfilter/nf_conntrack_core.cLinux kernel When the table is full, logs “nf_conntrack: table full, dropping packet” and drops the packet (the drop stat increases); packets that don’t match the connection state increase the invalid stat
Amazon EC2 security group connection trackingAWS When an instance exceeds its tracked-connection limit, packets are dropped, visible in conntrack_allowance_exceeded; recommends avoiding asymmetric routing
net/netfilter/nf_conntrack_standalone.cLinux kernel /proc/net/stat/nf_conntrack has one line per core, in hex, with columns such as entries, invalid, insert_failed, drop, and early_drop
ID rt-appliance-pps · Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
Firewalls, intrusion prevention systems (IPS), and DDoS protection devices inspect every packet passing through. The moment traffic exceeds their inspection capacity, they drop the packets they can’t process.
Why Hundreds of thousands or more small game packets per second at peak hours or events, or heavy inspection rules → Effect The device maxes out its CPU or packets-per-second limit and drops packets. False positives block legitimate packets too → On screen Freezes and teleporting hit every server behind that device at once, getting worse only when crowds gather
Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
Game team action items
Share the game’s traffic pattern (ports, packet sizes, packets per second) with the infra team, batch the small messages for one tick and send them together to cut the packet count.
Infra team action items
Watch the device’s CPU, packets per second, and drop counters alongside game metrics, size devices for small packets, exempt game ports from heavy inspection, tune DDoS protection rules to the game’s traffic pattern.
Ballpark numbers
A “10 Gbps” rating on a spec sheet is often based on large 1,500-byte packets. Game packets are around 100 bytes, so the same bandwidth means more than 10 times as many packets, and the packets-per-second limit fills up first even when the link looks idle.
On the graph
Hits a ceiling · Device packets per second and CPU utilization, device drops
Where to look
The device’s CPU, packets per second, and drop counters, plus packet counts on the switch ports in front of and behind it, compared at the same interval. Overlaid on one screen with concurrent users and server retransmission rate
Confirmed if
At peaks and events, the device’s packets per second or CPU stops at one value and can’t go higher, fewer packets leave the device than enter it, and at the same moment retransmission rates rise across every server behind it
Ruled out if
Same packet counts in front of and behind the device and no device drops: a different cause. NIC drop counters or softnet dropped rising on the server: “Packet drops on the receiving host”
ID rt-mtu · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)
If the largest packet size a link along the way can carry shrinks and the “too big” notice (ICMP) is blocked, large packets keep vanishing no matter how many times they’re resent.
Why The maximum size shrinks on a VPN or tunnel segment, and a firewall blocks the “too big” notices → Effect The sender, with no idea why, keeps retransmitting the same large packet, and the RTO doubles each time → On screen Fine normally, but when large data moves (inventory, crowded areas, loading into a zone), everything stops, including the small packets behind it, ending in a disconnect or infinite loading
During specific actions, Right after login or maintenance
Owner
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)
Game team action items
To lower it from the server side, set the socket’s maximum segment size (TCP_MAXSEG); just splitting messages into smaller pieces in game code won’t prevent it (TCP re-packs the outgoing data into MSS-sized segments).
Infra team action items
Network: clamp the MSS on edge devices, allow the “too big” ICMP (type 3 code 4, fragmentation needed) through firewalls and cloud network ACLs. Servers/OS: set the path MTU, make sure the server firewall and cloud security groups don’t block “too big” ICMP either, and as a last safety net set Linux tcp_mtu_probing=1.
Ballpark numbers
Usually 1,500 bytes; around 1,400 through a tunnel. If the same packet is retransmitted 5–6 times, the freeze exceeds 10 seconds.
On the graph
Outliers only · Per-connection RTO and backoff, disconnects per region and ISP
Where to look
Retransmissions on the problem connection from a server-side packet capture or bcc tcpretrans -s (shows sequence numbers), plus mss, pmtu, and backoff for that connection from ss -ti. From the server to that player’s address, a small ping compared with a 1,500-byte ping with DF set (ping -M do -s 1472)
Confirmed if
Full MSS-sized packets are retransmitted over and over with the same sequence number at doubling intervals, while smaller packets get through. No “too big” ICMP arrives (Wireshark filter icmp.type == 3 and icmp.code == 4), and the small ping gets replies while only the large DF ping vanishes without one
Ruled out if
Small packets vanishing too points to loss unrelated to size (“Bottleneck queue overflow (congestion loss),” “Route change / bad ECMP path”). If “too big” ICMP arrives and pmtu in ss -ti drops, path MTU discovery is working properly
Check with
Infra tools (no game code needed)
Learn more
tcp_mtu_probing=1 declares a black hole and lowers the MSS to 1,024 bytes only after retransmission timeouts have gone on for a few seconds (the equivalent of tcp_retries1=3). The connection is frozen until then, so keep it as a last safety net and put MSS clamping, which prevents the problem up front, first.
Sources (15)
RFC 1191: Path MTU discoveryIETF Path MTU discovery: a packet that is too large triggers ICMP “fragmentation needed and DF set” (type 3 code 4)
IP SysctlLinux kernel tcp_mtu_probing: 0 off, 1 only when a black hole is detected, 2 always (starting MSS is tcp_base_mss). tcp_retries1 defaults to 3
net/ipv4/tcp_timer.cLinux kernel When RTO retransmissions go on for tcp_retries1 times, treats it as a detected black hole, turns on MTU probing, and lowers the MSS
iptables-extensions(8) — Linux manual pagenetfilter TCPMSS --clamp-mss-to-pmtu: works around large packets stalling on segments that block ICMP by adjusting the MSS in the SYN
tcp(7) — Linux manual pageLinux man-pages TCP_MAXSEG: maximum segment size for outgoing packets; set before connecting, it also changes the MSS advertised to the peer
MTU considerations | Cloud VPNGoogle Cloud Cloud VPN gateway MTU is 1,460 bytes, and the payload MTU of an IPv4 tunnel is 1,406 bytes (around 1,400 through a tunnel)
ss(8) — Linux manual pageiproute2 mss, pmtu (path MTU), and backoff (how many times the RTO has doubled) in ss -i
ping(8) — Linux manual pageiputils -M do sets DF and won’t send packets larger than the path MTU the kernel knows; -s is the data size (the 8-byte ICMP header comes on top)
ID rt-mapping · Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Network infrastructure), Infra team (Server infrastructure)
If a device along the way deletes the mapping for an idle connection (the entry that records where to forward that connection), the next packet sent can’t be delivered. The connection either retransmits over and over until it disconnects, or the device sends back a reset (RST) and it disconnects right away.
Why A connection with no packets going either way for a while (AFK, lobby) → Effect A home router’s NAT, the ISP’s CGNAT, a firewall, a load balancer, or a cloud security group deletes the idle mapping → On screen When the player moves again, retransmissions go on until a disconnect, or the disconnect is immediate
Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Network infrastructure), Infra team (Server infrastructure)
Game team action items
Client: send heartbeats at no more than half the shortest idle timeout (mappings in players’ routers and ISP CGNAT are reliably refreshed only by outbound packets, and we can’t change their timeouts, so the client sends them), reconnect automatically after a disconnect. Server: answer heartbeats, and close the connection first when none arrive for a set time (shorten the TCP keepalive interval with socket options such as TCP_KEEPIDLE, detect failures quickly with TCP_USER_TIMEOUT), resume sessions with a session token.
Infra team action items
Network: collect the idle timeouts of firewalls and load balancers on the path and share them with the game team, extend them on our own firewalls and load balancers if needed. Servers/OS: check the cloud security group’s connection tracking timeout and share it with the game team.
Ballpark numbers
How long devices keep a TCP mapping varies widely, from a few minutes to several hours. When a cloud security group tracks connections, AWS Nitro v6 instance types delete the tracking entry after 350 seconds by default (other types after 5 days; see “Cloud security group connection tracking expiry”). The Linux TCP keepalive default is “probe after 2 hours idle,” which is later than most devices.
On the graph
Mass disconnect · Disconnects, idle time before disconnect
Where to look
The last few minutes of a dropped connection in a server-side packet capture; for live connections, idle time from lastsnd and lastrcv in ss -ti (ms since the last send and receive). nstat TcpExtTCPAbortOnTimeout (connections abandoned when the timer ran out) alongside
Confirmed if
Each dropped connection had just been idle past a similar value (the idle timeout of a device on the path, e.g., 350 s for security groups on AWS Nitro v6 instances), and from the first packet after the idle period, retransmissions go on with no ACK until the connection gives up, or an RST comes back right away
Ruled out if
Disconnects during play regardless of idle time point to a different cause (“Route change / bad ECMP path,” “Firewall and connection tracking drops”). Rule this cause out for connections that exchange heartbeats at no more than half the shortest idle timeout
Check with
Infra tools (no game code needed)
Sources (8)
RFC 5382: NAT Behavioral Requirements for TCPIETF Recommends a TCP NAT established-connection idle timeout of at least 2 hours 4 minutes (on the premise that devices may delete idle sessions earlier)
Amazon EC2 security group connection trackingAWS The default TCP idle tracking timeout is 350 seconds on Nitro v6 instance types and 432,000 seconds (5 days) on other types; recommends keepalives at intervals shorter than 5 minutes
IP SysctlLinux kernel tcp_keepalive_time defaults to 2 hours
tcp(7) — Linux manual pageLinux man-pages TCP_KEEPIDLE (idle time before keepalive starts), TCP_USER_TIMEOUT (how long to wait for unacknowledged data before closing the connection)
ID rt-path · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)
Packets vanish for a few seconds while an internet route changes, or steadily on connections assigned to a faulty path among several ECMP paths.
Why BGP route recalculation, or faulty equipment or a bad link on one of several paths (ECMP, LAG) → Effect Temporary loss during the route switch, or steady loss only on connections using that path → On screen A sudden freeze of a few seconds then fast-forward, or “it gets better after reconnecting” (assigned to a different path)
Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)
Game team action items
Log per-connection retransmission stats (TCP_INFO) so you can pull the IP, port, and time for affected players, don’t drop connections that stall for a few seconds right away.
Infra team action items
Monitor retransmission rate per region and ISP, check whether reconnecting changes the path, secure links from multiple ISPs, check our ECMP and LAG paths for bad links, run route measurements on the same TCP port as the game (mtr --tcp --port; the path is chosen by address and port, so an ordinary ping can take a different path and look fine).
External action items
Report the bad path to the ISP with route measurements taken on the same TCP port and a comparison from before and after reconnecting.
On the graph
Step change · RTT (ping), retransmission rate per region and ISP
Where to look
Retransmissions grouped per connection with bcc tcpretrans -c to pull the affected players’ addresses and ports; mtr on the same TCP port as the game (mtr -T -P PORT) from the server toward the player and from the player toward the server, compared. Results before and after reconnecting compared too
Confirmed if
From a certain moment, RTT for one region or ISP shifts like a step and a few seconds of loss cluster, or even within one ISP only some connections (address and port combinations) keep retransmitting and get better after reconnecting. A plain ping can look fine while only TCP mtr shows loss
Ruled out if
All connections on that ISP getting worse together at evening peak: “Bottleneck queue overflow (congestion loss).” Only one player affected, with loss already on the ping to their router: “Wireless link loss”
ID rt-spurious-delay · Primary owner External (External) · Also Infra team (Server infrastructure), Game team (Client development)
A packet that isn’t lost, just very late for a moment, still gets retransmitted if the delay is longer than the RTO, because the sender treats it as lost.
Why Bufferbloat, Wi-Fi power saving, mobile radio state changes, or a virtual machine pause cause momentary delays of hundreds of ms → Effect The RTO expires first and the packet is retransmitted; the original arrives soon after (the receiver gets a duplicate) → On screen The freeze and fast-forward come from the latency spike itself. The spurious retransmission barely lengthens the freeze; it only pushes up retransmission metrics, which get mistaken for loss
Primary owner External (External) · Also Infra team (Server infrastructure), Game team (Client development)
Game team action items
On Android 10 and later clients, request low-latency Wi-Fi mode during play (the WIFI_MODE_FULL_LOW_LATENCY Wi-Fi lock, which applies only while the screen is on and the game is in the foreground) to cut latency spikes caused by power saving.
Infra team action items
Avoid burstable instances, don’t set the RTO minimum too low, keep F-RTO and timestamps on (tcp_frto, tcp_timestamps), read retransmission metrics together with nstat TCPSpuriousRTOs and TCPDSACKRecv so they aren’t mistaken for loss.
External action items
To reduce the latency spikes themselves, tell players to use SQM on their router and turn off Wi-Fi power saving.
Ballpark numbers
Linux uses F-RTO to detect spurious RTOs and can undo the cut in sending rate. Check with nstat’s TCPSpuriousRTOs (times an RTO was judged spurious) and TCPDSACKRecv (times the receiver reported “already got this”).
On the graph
Random spikes · RTT (ping), spurious RTOs
Where to look
Deltas of TcpExtTCPTimeouts (RTO expirations), TcpExtTCPSpuriousRTOs, TcpExtTCPDSACKRecv, and TcpExtTCPLostRetransmit together, from nstat run every minute. With a packet capture, the Wireshark filter tcp.analysis.spurious_retransmission
Confirmed if
When RTOs rise, TcpExtTCPSpuriousRTOs or TcpExtTCPDSACKRecv rise with them, and RTT jumps to hundreds of ms at the same moment. The receiver-side capture shows both the original and the retransmission arriving
Ruled out if
TcpExtTCPSpuriousRTOs and DSACK flat while TcpExtTCPLostRetransmit (even the resent packet lost again) rises: real loss. RTT not jumping but DSACK steadily high: “Spurious fast retransmit from reordering”
WifiManagerAndroid (Google) WIFI_MODE_FULL_LOW_LATENCY (API 29, Android 10): a low-latency Wi-Fi lock that applies only while connected to an AP, with the screen on and the app in the foreground
net/ipv4/proc.cLinux kernel Counter names as shown by nstat: TCPTimeouts, TCPSpuriousRTOs, TCPDSACKRecv, TCPLostRetransmit
net/ipv4/tcp_timer.cLinux kernel TCPTimeouts increments when the retransmission timer (RTO) expires
ID rt-reorder · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)
When packets get out of order crossing multiple paths or bundled links, the receiver signals “a packet is missing” with duplicate ACKs, and the sender resends a packet that arrived fine.
Why Devices that split traffic across paths per packet, LAGs (link bundles) that spread traffic per packet, and route changes shuffle packet order → Effect Later packets arrive first and three duplicate ACKs pile up → fast retransmit → On screen Sparse game packets are barely affected. Large updates in crowded areas and patch downloads slow down, with occasional stutter
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)
Infra team action items
Network: switch from per-packet to per-connection load balancing (hash ECMP and LAG on address and port). Servers/OS: use RACK (time-based loss detection that tolerates reordering; when DSACK reveals spurious retransmissions, it automatically widens its reordering allowance), check the reordering degree Linux estimates automatically for each connection (the reordering value in ss -ti, starting from tcp_reordering=3).
On the graph
Always high · Reordering detections, DSACKs received
Where to look
nstat TcpExtTCPSACKReorder and TcpExtTCPTSReorder (reordering detections) and TcpExtTCPDSACKRecv; per connection, reordering (shown when not 3) and reord_seen in ss -ti. In a packet capture, the Wireshark filter tcp.analysis.out_of_order
Confirmed if
Reordering counters and DSACK climb steadily at all hours, and connections through a specific path or device show a reordering value above 3. The receiver-side capture shows later packets arriving first, with the earlier ones following shortly
Ruled out if
Reordering counters flat while TcpExtTCPLostRetransmit rises: real loss. DSACK rising only at moments when RTT jumps: “Spurious retransmission from latency spikes”
IP SysctlLinux kernel tcp_reordering starts at 3 (adjusted automatically per connection up to tcp_max_reordering), RACK setting in tcp_recovery
misc/ss.ciproute2 ss -ti shows reordering:value when a connection’s reordering differs from the default of 3, and reord_seen:count once the connection has seen reordering
SNMP counterLinux kernel TcpExtTCPSACKReorder, TcpExtTCPTSReorder (reordering detected), TcpExtTCPDSACKRecv (DSACKs received), TcpExtTCPLostRetransmit (a retransmitted packet lost again)
include/uapi/linux/tcp.hLinux kernel tcpi_reord_seen in tcp_info: how many times the connection has seen reordering
ID rt-ack-path · Primary owner External (External) · Also Game team (Client development)
Data arrives fine, but if the “got it” ACK is delayed or dropped in a full upload queue, the sender treats the data as lost and retransmits.
Why Video uploads or cloud backups at home saturate the upload → Effect ACKs sit in the router’s queue for hundreds of ms or get dropped when it overflows → On screen Game packets from the server mostly arrive on time. Your inputs, stuck in the same upload queue, go out late, causing input lag and rubber-banding, with occasional spurious retransmissions
Primary owner External (External) · Also Game team (Client development)
Game team action items
Show network status on screen when ping spikes, display a hint to “check for programs that are uploading.”
External action items
Tell players to keep the upload queue short with SQM on their router, prioritize small packets (ACKs), and cap upload speed (video uploads, cloud backups).
Ballpark numbers
A later ACK confirms everything an earlier one did, so losing a few ACKs is usually fine. The real trouble is ACKs delayed in the queue.
On the graph
Outliers only · Per-connection RTT (ping)
Where to look
Ping to the game server from the player’s PC with an upload (video upload, cloud backup) running and stopped, compared. On the server, rtt for that player’s connection in ss -ti
Confirmed if
Ping climbs to hundreds of ms only during the upload, with input lag and rubber-banding, and recovers soon after the upload stops. From the server, that connection’s rtt rises at the same time
Ruled out if
Loss and latency regardless of uploads: “Wireless link loss” or a path-side cause. Only the server-to-player direction slow, with no link to uploads: “Bottleneck queue overflow (congestion loss)”
Check with
The player’s own environment
Sources (4)
RFC 3449: TCP Performance Implications of Network Path AsymmetryIETF On asymmetric links with a narrow upload, delayed or lost ACKs hurt TCP performance; ACKs are cumulative, so a later ACK covers for lost ones; remedies such as ACK-prioritizing scheduling
Smart Queue ManagementBufferbloat.net Keeping router queues short with queue management and shaping
tc-cake(8) — Linux manual pageiproute2 CAKE separates flows and minimizes latency for flows that send sparsely (sparse flows)
ID rt-rto-setting · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Set the RTO minimum too low and even small delays cause spurious retransmissions; leave the default (200 ms) and it’s too long for games, so every loss means a long freeze.
Why RTO minimum lowered sharply for data center use, or the default left as is on internet paths → Effect Too low: retransmission storms on momentary delays. Too high: a long wait on every loss → On screen At the default, each loss means a freeze of hundreds of ms then fast-forward; set too low, freezes get shorter, but spurious retransmissions surge and waste bandwidth
Primary owner Infra team (Server infrastructure) · Also Game team (Server development)
Game team action items
On Linux 6.15 and later, consider lowering the RTO cap for game connections with TCP_RTO_MAX_MS (this also shortens the time until the connection gives up, so set the disconnect detection time with TCP_USER_TIMEOUT alongside it), lower the RTO minimum only on internal server-to-server connections with the TCP_RTO_MIN_US socket option (6.15 and later), consider the TCP_THIN_LINEAR_TIMEOUTS socket option so that consecutive RTOs don’t double on game connections.
Infra team action items
Lower rto_min per route only for internal server-to-server connections, keep the default on internet paths and compensate with RACK-TLP and the thin stream setting (tcp_thin_linear_timeouts).
Ballpark numbers
Linux RTO = round-trip time + max(200 ms, RTT deviation × 4). It doubles on every failure, up to 120 seconds. On Linux 6.15 and later, TCP_RTO_MAX_MS can lower this cap to as little as 1 second.
On the graph
Always high · Per-connection RTO, spurious RTOs
Where to look
The server’s RTO minimum setting (rto_min in ip route show; on Linux 6.11 and later, sysctl net.ipv4.tcp_rto_min_us), rto and rtt in ss -ti, and deltas of nstat TcpExtTCPSpuriousRTOs
Confirmed if
On a server with a lowered minimum, rto on internet connections hugs rtt and TcpExtTCPSpuriousRTOs climbs sharply. At the default, rto on game connections sits 200 ms or more above rtt, and every loss freezes the game for that long
Ruled out if
rto follows the default calculation (about rtt + 200 ms) and spurious RTOs are few, yet freezes are unusually long: points to consecutive losses or the recovery method (“Slow recovery on thin streams,” “Middlebox strips TCP options”)
IP SysctlLinux kernel tcp_rto_min_us defaults to 200000 (the route option rto_min and the socket option TCP_RTO_MIN_US take precedence), tcp_rto_max_ms 1,000–120,000 (default 120,000), tcp_thin_linear_timeouts
ip-route(8) — Linux manual pageiproute2 Per-route rto_min option: the RTO minimum used when communicating with that destination
Thin-streams and TCPLinux kernel TCP_THIN_LINEAR_TIMEOUTS can turn off exponential backoff for thin stream connections only
tcp(7) — Linux manual pageLinux man-pages TCP_USER_TIMEOUT: how long to wait for unacknowledged data before closing the connection
ID rt-thin · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Game team (Client development)
When a game sends small packets sparsely, the RTO fires before “three following packets” can pile up. The same loss causes a much longer freeze than it would on a bulk transfer.
Why Packets go out about 100 ms apart, so only a few packets are ever in flight (not yet ACKed) → Effect Collecting three duplicate ACKs takes over 300 ms, so the RTO (ping + 200 ms) fires first, doubling on consecutive losses → On screen Each loss freezes the game for about 0.3 seconds; if the retransmission is lost too, the freeze lasts close to 1 second, then fast-forward
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Game team (Client development)
Game team action items
Server: turn on TCP_NODELAY (with Nagle on, RACK has no following packets to base its decision on), send real-time packets over UDP with your own retransmission. Client: turn on TCP_NODELAY, send real-time packets the same way as the server (UDP).
Infra team action items
Use RACK-TLP (the default on current Linux), use tcp_thin_linear_timeouts so consecutive RTOs don’t double.
Ballpark numbers
With 100 ms between packets and a 60 ms ping, fast retransmit takes about 360 ms (until three later packets arrive and their acknowledgments come back), while the RTO is about 260 ms. With RACK, the packet is resent right away at about 160 ms, when the acknowledgment for the next packet comes back. Once packets are more than 200 ms apart, RACK is no faster than the RTO either.
On the graph
Gap then burst · Per-connection receive volume, RTO expirations
Where to look
Deltas of nstat TcpExtTCPTimeouts (RTO expirations), TcpExtTCPFastRetrans (fast retransmits), and TcpExtTCPLossProbes and TcpExtTCPLossProbeRecovery (TLP), compared, plus rto and backoff on game connections in ss -ti. The server’s net.ipv4.tcp_recovery, tcp_early_retrans, and tcp_sack values too
Confirmed if
Among retransmissions, RTO expirations outnumber fast retransmits, and game connections often show backoff above 0 (in the middle of an RTO). Received volume sits at 0 during the freeze and then arrives all at once on recovery
Ruled out if
Bulk transfers on the same server stalling just as long: a loss problem unrelated to the connection’s traffic pattern. Concentrated on connections missing SACK or timestamps: “Middlebox strips TCP options”
Check with
Infra tools (no game code needed)
Learn more
Linux once had an option to retransmit on a single duplicate ACK for thin streams (tcp_thin_dupack), but it was removed in 2017, and RACK now fills that role. With Nagle on (TCP_NODELAY off), no new packets go out while the sender waits for the lost packet’s acknowledgment, so RACK has no following packets to base its decision on and the connection ends up waiting for the RTO.
Sources (11)
Thin-streams and TCPLinux kernel Thin streams that send sparsely, like games, don’t trigger fast retransmit well and rely on long timeouts; the threshold is fewer than 4 packets in flight (not yet ACKed)
ID rt-sack-stripped · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)
When some firewalls or accelerators remove or rewrite TCP options, multiple losses get recovered only one per round trip, or the window (how much can be sent at once) shrinks, and everything slows down.
Why A firewall’s “TCP normalization” or an old accelerator strips the SACK, timestamp, and window scale options → Effect With several packets lost, recovery goes one packet per round trip, and the window is capped at 64 KB → On screen Every loss causes a much longer freeze (without SACK, RACK-TLP can’t be used either), then fast-forward when it clears. Bulk transfers such as patches are slow too
Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)
Infra team action items
Network: turn off TCP normalization on the device in question, check the firewall’s sequence number randomization too, compare the options in the SYN with packet captures at both ends. Servers/OS: check in ss -ti whether connections missing sack or wscale concentrate on a specific path (Windows PCs may not use ts depending on their settings, so ts alone missing can be normal), check that the server’s net.ipv4.tcp_sack is 1.
On the graph
Always high · Recoveries started without SACK (TcpExtTCPRenoRecovery)
Where to look
Whether each connection shows sack and wscale in ss -ti, the ratio of nstat TcpExtTCPRenoRecovery (recovery started without SACK) to TcpExtTCPSackRecovery, and TcpExtTCPSACKDiscard (SACK blocks discarded as inconsistent). On a suspect path, SYNs captured at both ends with their options compared (Wireshark tcp.options.sack_perm and similar)
Confirmed if
Only connections through a specific path or device lack sack and wscale, and TcpExtTCPRenoRecovery makes up a large share. The SACK-permitted option present in the SYN as sent is missing from the SYN as received. When sequence number randomization is the cause, the options survive but TcpExtTCPSACKDiscard rises
Ruled out if
sack missing on every connection: check the server’s net.ipv4.tcp_sack value first. Options intact and TcpExtTCPSACKDiscard flat: the slow recovery has another cause (“Slow recovery on thin streams”)
Check with
Infra tools (no game code needed)
Learn more
SACK can break even when the options survive. If a firewall’s sequence number randomization rewrites only the sequence numbers in the header and leaves the numbers inside SACK blocks untouched, the sender discards the inconsistent SACKs. A server where tcp_sack=0 was set during the 2019 SACK security issue and then forgotten ends up the same way.
SNMP counterLinux kernel TcpExtTCPRenoRecovery (recovery started without SACK), TcpExtTCPSackRecovery (recovery started with SACK), TcpExtTCPSACKDiscard (invalid SACK blocks)
ID rt-zero-window · Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Server infrastructure)
When the receiving program doesn’t read its socket in time and the buffer fills up, the sender stops sending and sends only zero window probes. The network itself is fine.
Why A client frame freeze or a blocked server thread keeps the socket from being read → Effect The receive window drops to 0, so the sender stops sending and sends only probes (at growing intervals) → On screen Freeze then fast-forward. A packet capture shows “ZeroWindow” and no loss
Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Server infrastructure)
Game team action items
Start with the side that sent ZeroWindow in the packet capture (the side that can’t read its socket), keep reading network input on a separate thread, size the receive buffer appropriately. Client: fix the causes of frame freezes such as loading and GC. Server: fix whatever blocks the thread that reads the socket.
Infra team action items
Add the server’s nstat TcpExtTCPToZeroWindowAdv (times the server advertised a zero receive window) to monitoring (if it rises, the problem is on the server side, so pass it to server development), provide server-side packet captures.
On the graph
Gap then burst · Per-connection receive volume, zero window count
Where to look
In a packet capture, the side that advertised window 0, found with the Wireshark filter tcp.analysis.zero_window. In the server’s nstat, TcpExtTCPToZeroWindowAdv (the server advertised window 0) and TcpExtTCPWinProbe (probes sent in response to the peer’s window 0) viewed separately, plus Recv-Q on the server socket (bytes in ss the program hasn’t read yet)
Confirmed if
No retransmissions during the freeze, only zero window and probe packets. TcpExtTCPToZeroWindowAdv or server socket Recv-Q rising: the server isn’t reading in time. TcpExtTCPWinProbe rising: the client isn’t reading in time
Ruled out if
No zero window in the capture and the same data being resent: points to loss or spurious retransmission
ID rt-syn · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure), Game team (Client development)
If a connection request is lost because the connection queue (backlog) overflows or a firewall blocks it, the client OS resends it starting 1 second later, at growing intervals.
Why A connection surge right after maintenance overflows the server’s connection queue, or a firewall or DDoS protection drops the SYN → Effect The client OS retransmits the SYN starting 1 second later, at set intervals (older Linux: 1 s → 2 s → 4 s) → On screen After pressing Connect, delays come in whole seconds, such as 1 or 3 seconds; if it keeps failing: can’t connect / infinite loading
Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure), Game team (Client development)
Game team action items
Server: raise the backlog argument to listen (together with somaxconn), make sure the game server calls accept promptly, use a login queue. Client: lengthen connection retry intervals (randomized to spread retries out).
Infra team action items
Servers/OS: confirm server connection queue overflow with nstat TcpExtListenOverflows and TcpExtListenDrops and the “Possible SYN flooding” warning in the logs, raise somaxconn (together with the listen argument), use SYN cookies. Network: relax SYN limits on firewalls and DDoS protection.
Ballpark numbers
On Linux (Android included), the first SYN retransmission comes after 1 second. Older kernels then double the interval each time and resend at 1, 3, 7, 15 seconds …, while 6.5 and later (tcp_syn_linear_timeouts=4) resend five times at 1, 2, 3, 4, and 5 seconds, then double (7, 11, 19 seconds …). Android phones often keep their launch kernel even after OS updates, so behavior can differ between devices on the same Android version. Either way, if every attempt fails, the OS gives up after about 2 minutes. Windows starts at 1 or 3 seconds depending on version and settings and resends 2–4 times, so it gives up after 20–30 seconds (check that PC’s value with Max SYN Retransmissions in netsh int tcp show global).
On the graph
Surge after opening · Connection attempts, connection queue overflows
Where to look
In the server’s nstat, TcpExtListenOverflows and TcpExtListenDrops, plus the “Possible SYN flooding on port” warning in dmesg; in ss -lnt, whether a listening socket’s Recv-Q (connections waiting for accept) reaches its Send-Q (backlog limit). In a server-side capture, whether SYNs arrive and whether SYN-ACKs go back
Confirmed if
During the connection surge right after maintenance, TcpExtListenOverflows rises and Recv-Q sits at Send-Q. The capture shows the same client’s SYN coming back at whole-second intervals with no server response
Ruled out if
SYNs never reach the server and server counters stay flat: a firewall or DDoS protection in front dropped them, so check that device’s SYN limits and drop logs. Server sent SYN-ACKs but connecting is still slow: loss in the return direction
Check with
Infra tools (no game code needed)
Sources (11)
include/net/tcp.hLinux kernel Initial RTO TCP_TIMEOUT_INIT = 1 second (the initial value from RFC 6298)
IP SysctlLinux kernel tcp_syn_retries defaults to 6, tcp_syn_linear_timeouts defaults to 4 (SYN RTO 1, 1, 1, 1, 1, 2, 4 …), last retransmission at 67 seconds and giving up at 131 seconds, somaxconn defaults to 4096, tcp_syncookies defaults to 1
tcp: make the first N SYN RTO backoffs linearLinux kernel Commit that made the first SYN retransmissions use a fixed interval, from Linux 6.5 (the default of 4 follows the macOS and iOS behavior)
Android common kernelsAndroid (Google) Common kernels 5.10 through 6.18 are supported side by side, and a kernel for an earlier platform (e.g., android14-6.1) can be used to launch or upgrade new Android devices
TcpMaxConnectRetransmissionsMicrosoft Older Windows defaults: 2 SYN retransmissions, starting with a 3-second wait and doubling, then waiting double again after the last one before giving up (3+6+12=21 seconds)
listen(2) — Linux manual pageLinux man-pages The backlog argument to listen is truncated to somaxconn (default 4096 since Linux 5.4, 128 before that)
SNMP counterLinux kernel When the accept queue is full, SYNs are dropped and TcpExtListenOverflows and TcpExtListenDrops rise together; TcpExtTCPSynRetrans
net/ipv4/tcp_diag.cLinux kernel For a listening socket, Recv-Q in ss is the number of connections waiting for accept, and Send-Q is the backlog limit
Playbooks
Lag after a patch
When lag reports increase after a specific patch or deployment. Use it when reports like “it’s been weird since this update” pile up, or when a graph steps up at some point and stays there.
Pin down the start time and gather every change around it: Find when reports first spiked and when the graph stepped up, and list every change that went out around that time. Cover client patches, server deployments, configuration changes, DB schema changes (DDL) and restarts, network and firewall work, and infrastructure swaps (instance type, kernel, drivers). If you use your monitoring tool’s annotation feature to mark every deployment with a vertical line on all graphs, this step goes quickly. If a game patch and infrastructure work went out in the same maintenance window, keep both as suspects. Who to call first: the game team and the infra team, for the changes each of them shipped. (Deploys and restarts, Performance changes after OS, kernel, driver, or firmware updates, Schema change (DDL) lock during live service, Query slowdown from a query plan change, Cold cache (right after a restart))
Break down the scope: build, device, server, region: Find the dimension where the problem clusters. Suspect the client first if only players on the new build are affected; client performance or drivers if only certain OSes, graphics cards, or devices are; the server if only certain servers, channels, or zones are; the network path if only certain countries or ISPs are; and shared resources (DB, load balancers, gateways) or the server deployment that just went out if everyone is affected at once. If client telemetry includes the build number, put ping, FPS, frame spikes, and disconnect counts for the old and new builds side by side. If ping is unchanged and only FPS got worse, client performance is a likelier suspect than the network. Who to call first: the game team (client) if it clusters by build or device; for servers or channels, the game team (server) if host metrics look normal, or the infra team (servers/OS) if they don’t; the infra team (network) if it clusters by country or ISP. (Frame time spike, Synchronous loading and shader compilation on the main thread, Out of graphics memory (VRAM), Client crash)
Compare the new and old versions over the same time window: A plain before-and-after comparison mixes in changes from the day of the week, time of day, and events, which muddies the picture. If possible, roll the new version out to a few servers first (canary), and compare tick time p50 and p99, tick overrun count, CPU, memory, and error rate side by side with old-version servers (the control group) over the same time window. If it’s already deployed everywhere, compare against the same day and time last week. A server-wide average hides problems on individual servers or zones, so break it down by server and zone. Who to call first: the game team (server). (Tick overrun, Allocation surge, Memory leak, Broadcast fan-out overload)
Compare the traffic fingerprint before and after: Even without knowing the server code, you can tell from values visible on the network whether the patch changed the shape of the traffic. Compare before and after: packets per second (pps) and bytes per player, average and maximum packet size, connection count, and the size of the send burst that goes out all at once each tick. If UDP packets have started to exceed the path MTU (usually 1,500 bytes), IP fragmentation occurs. Losing a single fragment loses the whole packet, and some NATs and firewalls drop fragments outright. Players whose path crosses a segment with a smaller MTU (a tunnel or VPN) lose only the large packets. If pps went up, check whether you’re hitting the cloud instance’s PPS limit or the throughput limit of firewalls or DDoS protection appliances. Who to call first: if the fingerprint changed, the game team (server), with the evidence attached; if the fingerprint is the same and only loss and retransmissions went up, the infra team (network). (Patch changes the traffic pattern, IP fragmentation of UDP packets, MTU black hole (only large packets keep getting lost), Cloud PPS limit exceeded, Middlebox over capacity (firewall, IPS, DDoS protection), Send bursts overflow shallow buffers)
Compare DB query types and counts before and after: If DB latency went up, start by checking whether query volume (QPS) went up with it. PostgreSQL’s pg_stat_statements and the digest summaries in MySQL Performance Schema group queries that differ only in their values into one entry and track execution count and total time. Comparing the top query lists from before and after the patch reveals new queries, queries whose count multiplied (N+1), and queries that read the whole table without an index (the SUM_NO_INDEX_USED column in MySQL). Who to call first: the game team (server) if QPS or the shape of the queries changed; the infra team (DB: query plans, IOPS, locks) if the queries are the same and only latency went up. (Queries with no index, Login storm and N+1 queries, Query slowdown from a query plan change, Cache stampede)
Use host and server process metrics to find the layer: Use values visible from the OS, without needing the game code, to tell problems inside the server process from problems on the host. If the receive queue (Recv-Q) on the server socket is building up, the server process isn’t reading in time (tick stalls, GC, locks). If one thread alone is at 100%, it’s a single-thread bottleneck. If pause times in the GC log went up, the memory usage pattern changed. Also check whether the build shipped with a higher log level and log writes went up. Conversely, if CPU steal, throttling, or NIC drops went up, look at infrastructure that changed at the same time (instance type, kernel, container limits). Who to call first: the game team (server) for signals inside the process; the infra team (servers/OS) for host signals. (Server GC stop-the-world pause, Single-threaded zone overload (hotspot), Synchronous log writes, Container CPU throttling (CFS quota), CPU steal (virtual machines), Performance changes after OS, kernel, driver, or firmware updates)
Roll back to confirm, then record the result: Revert the most likely change for only some servers or some players (roll back, or turn off a feature flag), or set the configuration back to its previous value, and see whether the symptom goes away with it. If only the reverted side improves, the cause is confirmed. Reverting can itself cause a brief slowdown from restarts and cold caches, so if it isn’t urgent, do it during off-peak hours. Record the result in the incident log along with the cause ID, and add limits on packet size, query count, and tick time to the pre-deployment checklist for the next patch. Who to call first: the team that made the change. (Deploys and restarts, Cold cache (right after a restart))
Launching in a new country or region
When you launch the service in a new country or add a new region or data center. Use it both for pre-launch checks and for sorting out reports like “it’s fine back home, but players in the new country are lagging.”
Measure path quality for each local ISP before launch: For each major ISP (ASN) in the target country, measure the round-trip time (RTT) distribution, jitter, and packet loss to each candidate game server location. A single average hides the differences between ISPs, so look at the median and 95th percentile per ISP, separately for the evening peak and the early morning hours. The public measurement network RIPE Atlas lets you pick countries and ASNs and send ping and traceroute from probes around the world, or you can spin up temporary VMs in the candidate regions and measure from there. Devices along the path sometimes rate-limit ICMP responses, so when possible, also measure with the same protocol and port the game uses. If one ISP’s traffic stands out by passing through distant cities, it’s a peering or routing problem. ISPs choose lower-cost routes even when lower-latency ones exist, so even nearby destinations can end up taking long detours. Who to call first: the infra team (network); external (ISP, IX) if the route problem is on the ISP side. (Propagation delay (physical distance), Detour routing, Peak-hour congestion at peering links, Submarine cable / international link outage)
Compare the measurements with the limits the game design can tolerate: Compare the measured RTT and jitter with the game’s timing windows (reaction times for dodges, parries, and so on), lag compensation limit, interpolation buffer length, and input buffer size. For example, if the parry window is 0.2 s, players on ISPs where round-trip delay plus the interpolation buffer adds up to more than that will be late even when they react in time. Widen lag compensation to make up for it, and now the players on the receiving end start reporting “I got hit behind a wall.” If many ISPs exceed the limits, the infra team should look at placing regions or edge PoPs closer, and the game team should review the timing window, interpolation, and lag compensation values. The white paper’s chapter “Same ping, different feel: netcode models” serves as the reference table. Who to call first: the game team (server and client: design limits) and the infra team (network: region and PoP locations). (Short timing windows eaten up by ping, Hit registration without lag compensation, Too much lag compensation, Missing or too-short interpolation buffer)
Check MTU and whether UDP gets through: Check that the game’s largest packets make it through local networks intact. Measure the path MTU by sending pings of various sizes with the Don’t Fragment (DF) bit set, and look for segments smaller than 1,500 bytes, such as PPPoE, tunnels, and mobile networks. The standard for datagram transports such as UDP (RFC 8899) recommends 1,200 bytes as the base size that can cross most paths on IPv4, so if the game’s largest packet is bigger than that, decide with the game team whether to shrink it or split it up. Also check whether UDP or the game’s ports are blocked or throttled on public Wi-Fi, corporate networks, or certain ISPs, and whether there’s a fallback path (TCP, port 443) for when they are. Who to call first: the infra team (network) and the game team (server: packet size). (MTU mismatch (only large packets vanish), MTU black hole (only large packets keep getting lost), IP fragmentation of UDP packets, Country- or ISP-level UDP restrictions and packet inspection, Public Wi-Fi and corporate network restrictions, ISP throttling and traffic management)
Measure NAT and CGNAT idle timeouts and set the heartbeat interval to match: Measure how long local home routers and mobile networks (CGNAT) keep the mapping for an idle UDP connection before deleting it. For each trial, send one packet from a test device to the server to create the mapping. The device then sends nothing, and the server sends a packet back to the device after a set wait (30 s, 60 s, 120 s …). The wait at which the device stops receiving that packet is the network’s idle timeout. The standard (RFC 4787) says UDP mappings must not expire in less than 2 minutes and recommends a default of 5 minutes or more, but values vary widely between devices, and some delete mappings sooner. Only outbound packets from the device reliably refresh the mapping, so have the client send the heartbeat, and check that its interval is at most half of the shortest of these: the measured value and the idle timeouts of the load balancer and cloud security groups. Who to call first: the game team (client: heartbeat interval; server: timeout values) and the infra team (load balancer and security group settings). (NAT mapping expiry, ISP-shared IP addresses (CGNAT), NAT or load balancer mapping expires mid-connection, Load balancer idle timeout, Cloud security group connection tracking expiry)
Check the external services and security appliances on the local path: Check that local platform login, payment, and identity verification respond at normal speed, that local DNS resolves the login and patch server addresses correctly, and that the CDN serves patches from locations close to that country. Make sure the new country’s IP ranges aren’t caught by country-blocking rules or rate limits in DDoS protection and firewalls. In particular, make sure CGNAT ranges, where many subscribers share one IP, don’t get blocked wholesale. Who to call first: the infra team (security appliances, DNS, CDN) and external parties (platforms, payment providers, ISPs). (External service dependency, DNS failures and delays, DDoS protection detours and false positives, ISP-shared IP addresses (CGNAT))
After launch, break the data down by country and ASN: Tag client IPs in connection logs and load balancer logs with country and ASN, and look at RTT, retransmissions, and disconnect counts and reasons (heartbeat timeout, RST, server kick) by country and ISP. Free databases such as MaxMind GeoLite ASN map IPs to an ASN and organization name. To comply with local privacy rules, store IPs truncated to /24 or reduced to the ASN. Look first at that ISP’s route if problems concentrate in one ASN (infra team, external); at distance and design limits if the whole new country is bad (infra team, game team); and at peering congestion if it only gets worse in the evening. If some players always have high ping, check with the game team (server) whether they’re being assigned to a distant region because of GeoIP errors, VPNs, or region assignment based on the party leader. If synthetic monitoring looks normal and only players are having problems, it points to the player’s environment or the client. (Peak-hour congestion at peering links, Detour routing, Bottleneck queue overflow (congestion loss), Validation false positives concentrated on one ISP, Matchmaking and region assignment errors)
Check how distant players affect everyone else: When more players connect from far away, the damage doesn’t stop at their own screens. A slow player’s inputs arrive in bunches, so on other players’ screens that one character moves in fast-forward, and it trips the server’s speed and cooldown checks, causing rubber-banding or rejected skills. In party mechanics, one slow player’s late reaction can fail the whole party, and in lockstep, everyone waits for the slowest player. Check whether reports of “one character looks off” from existing players went up after the new country opened, and work out input buffers, validation tolerances, and separate matchmaking regions with the game team. Who to call first: the game team (server). (A lagging player moves in bursts on others’ screens, Validation false positives concentrated on one ISP, One lagging party member and boss mechanics, Lockstep waiting on the slowest player)
Real-world incidents
Only postmortems that game studios and infrastructure companies published themselves.
CCP Games 2014: Server overload in EVE Online’s massive HED-GP fleet battle
What happened
During the massive fleet battle in the HED-GP system, covered in a January 2014 retrospective, the server was badly overloaded. Even after Time Dilation (a feature that slows game time under overload) hit its 10% floor and the whole battlefield went into slow motion, load kept piling up. The backlog in processing module deactivations and repeat cycles (Dogma Lateness) peaked at 193 seconds of game time, about 32 minutes in real time. The July 2013 battle in 6VDT, of nearly the same size, peaked at 42 seconds (about 7 minutes in real time).
Cause
CCP cautioned that it couldn’t be certain, because its profiling tools add load of their own and aren’t run in situations like this, and named two likely causes. The first was unprocessed load that kept building up as the battle dragged on. The second was heavier drone use: unique drones deployed during the battle went from 21,123 in 6VDT to 38,852 in HED-GP, 84% more. Telling everyone in view about each player’s actions takes traffic that grows with the square of the player count (O(n²)), and drones generate more messages per attack. The code drones use to pick targets also often scans every attackable target on the same battlefield, so its cost grows close to n².
Lessons
When the processing load in one crowded area exceeds its limit, the whole area goes into slow motion, and the longer the battle runs, the more the backlog grows and the worse the input lag gets. Signals to check: tick time and backlog on the server (node) handling that area, plus player and entity counts. A telltale sign is that other areas stay fine. The primary owner is the game team (server), and the things to fix are how many recipients each action is sent to and the cost of AI target searches. Slowing game time can’t eliminate the overload, but it slows everyone down at the same rate, which keeps a subset of actions from falling behind indefinitely.
Riot Games 2015: League of Legends traffic on roundabout routes, and Riot Direct
What happened
A technical post in which Riot Games explains why the internet is a poor fit for real-time games. Real traffic reported by a League of Legends player should have gone straight from San Francisco to Portland, but it went through Los Angeles, Denver, and Seattle, taking 70 ms for a trip that would take 14 ms on a direct path. Riot explained that when routers overflow and drop packets, other champions appear to jump around the screen and projectiles seem to teleport.
Cause
Riot pointed to routes and routers. Backbone providers and ISPs send traffic along the cheapest path even when a lower-latency path exists, and when the route BGP settles on takes a long detour, traffic also passes through more routers. A router’s processing load depends on the number of packets, whatever their size. Game packets are around 55 bytes, so the same amount of data takes 27 times as many packets as it would in 1,500-byte packets, and fills router input buffers that much faster. According to Riot, many routers drop UDP packets first when they’re overloaded. As a fix, Riot built its own network, Riot Direct, with routers at 10 major internet hubs in the US and direct connections (peering) with as many ISPs as possible. According to Part II, the share of players with a ping under 80 ms rose from 31% to 50% in a little over 9 months, and hit 80% overnight after the game servers moved to Chicago.
Lessons
If only customers of one ISP have unusually high ping, even within the same country, suspect the route. Signals to check: the RTT distribution per ISP (ASN) and the cities that show up as hops in traceroute. The primary owner is the infra team (network), and the fixes are direct peering with ISPs, connecting at IXs (internet exchanges), and choosing server locations. Routing policy on the ISP side has to be worked out with the external party (the ISP). The case also shows that simply moving servers closer to the center of the player base makes a big difference.
Riot Games 2020: Edge host overload on League of Legends servers in Europe and Brazil
What happened
In late February 2020, the League of Legends EUW, EUNE, and BR servers had several outages, and the number of new games starting dropped sharply. Backend services such as matchmaking and game servers all reported healthy, yet almost no traffic was coming in. Riot pushed the tournament mode (Clash) back a week to avoid launching it on clusters that might be unstable. The postmortem doesn’t say how long each outage lasted.
Cause
Three things came together. Requests to one service were malformed, so in certain cases they kept failing and being retried, and request volume exploded. A known compatibility problem between the container system and the OS version was leaking memory inside the OS. The OS upgrade was finished on only about 60% of Riot’s entire container environment and was still in progress on the Europe and Latin America clusters. Edge containers, which receive internet traffic, filter it, and pass it to the backend, were kept apart within a shard (server group), but nothing kept different shards apart, so in every outage edge containers from at least three shards were packed onto a single host. The retry surge landed on that host, and the memory leak brought it to a halt.
Lessons
When every backend service reports “healthy, but no traffic is coming in,” look at what sits in front of them (edge, gateways, load balancers). Signals to check: inbound connection counts skewed toward particular hosts, and the failure and retry rate of specific requests. The primary owner is the game team (server: the malformed request and the retry behavior), and the infra team (servers/OS) handles container placement rules, OS upgrades, and skew alerts. Riot fixed the request code, changed retries so they wouldn’t spike, and put skew alerts in place until it could implement spreading across shards.
Riot Games 2021: League of Legends EUW 5-hour outage: one auxiliary DB halted the whole server
What happened
On January 22, 2021, the League of Legends EUW server didn’t work properly for a little over 5 hours. The metrics for logged-in players and players in a game cut out at the same moment, and between the two restarts, logins went up but almost no games started.
Cause
The primary server of a database behind a non-critical feature had a hardware failure, and that database had no automatic failover to a standby configured. Each database had its own connection pool, but all the pools shared one thread pool; work sent to the failed database never finished and held on to threads, until the whole system ran out of threads. Amid a flood of alerts, the team first suspected a recent malicious network attack and hardware work in another region, so the alert for the failed database wasn’t noticed until about 1 hour later. Because every system ran inside a single JVM, when GC paused the process for several seconds at a time under the reconnect load after the restart, metrics collection also developed large gaps. The login queue also didn’t hold to its configured limit, so players flowed in unevenly.
Lessons
Even one auxiliary database that nobody considered critical can halt everything through a shared resource such as a thread pool. Signals to check: pending requests per database, thread pool utilization, and a ratio of game starts to logins that is far too low. The owners are the game team (server: thread pool isolation, timeouts) and the infra team (DB: automatic failover). When alerts flood in, it’s easy to suspect whatever hit you recently (an attack, for example) first, so rule things out one at a time in the decision order (scope → timing → layer). After a restart, also check that the login queue actually limits inflow as configured.
Roblox 2021: Roblox 73-hour outage: contention in the service discovery (Consul) cluster
What happened
It began on the afternoon of October 28, 2021 (Pacific Time) with high CPU load on one Consul server. At 16:35 the number of players online fell to half of normal, and then the entire service went down. Not until 16:45 on October 31 could all players get back in, 73 hours after the outage began. Roblox said 50 million people use it every day.
Cause
Roblox uses HashiCorp Consul for service discovery (how services find each other’s addresses), health checks, and a key-value (KV) store, and a single Consul cluster was handling several workloads at once. There were two root causes. First, the day before the outage, Consul’s new streaming feature, which had been rolled out gradually over several months, was turned on for the traffic routing service too, and that service’s node count was raised by 50%. Under very heavy read and write load, the feature caused contention on a single shared resource (a Go channel). The contention was even worse on the dual-socket (NUMA) servers with more cores that were swapped in during the outage. Second, free-page list (freelist) management in BoltDB, which Consul uses to store its Raft log, became pathologically slow and wrote 7.8 MB to disk for every append of 16 kB or less. Median KV write latency, normally under 300 ms, rose to 2 s, and zero windows (full TCP buffers) were seen on the slow leader server. Because telemetry depended on Consul, the metrics needed to find the cause disappeared along with it.
Lessons
When a foundational system that many services rely on (service discovery, configuration store, authentication) slows down, every feature stops at once. Signals to check: that system’s write latency, leader changes, and CPU, plus any configuration change made just before the outage. Ownership is shared between the game team (server) and the infra team (servers/OS). Keep monitoring separate from the systems it watches, so you can still see metrics during an outage. During recovery, caches are empty, and letting everyone in at once can knock things over again, so Roblox used DNS to control the share of players let in and raised it about 10% at a time.
Square Enix 2021: FINAL FANTASY XIV expansion launch congestion and login queue errors
What happened
From the start of early access for the Endwalker expansion in December 2021, every World was extremely congested. Login queues grew long, and Error 2002 appeared often when logging in from the character selection screen or while waiting in the queue. Some Worlds and zones also went down (Error 3001), and queues timed out (Error 4004). As of the December 11 notice, on day 8 of early access, the congestion was still going on.
Cause
Error 2002 occurs in two cases. The first is when more than 17,000 players are waiting on a logical data center. This cap exists to keep the login server from going down under an overly long queue, and when it’s hit, the client shuts down completely. On December 7, spare development hardware was added to the lobby servers to raise the cap; this error became less common, but the queues actually got longer. The second case is when a waiting player’s connection is unstable. As waits got longer, brief disconnects caused by packet loss on the internet path or unstable Wi-Fi became more common. The lobby server waits somewhere from tens of seconds to about 1 minute for a reconnect. Players who reconnect in time keep their place in the queue, but anyone who takes longer goes to the back of the line. Square Enix said most reports fell into this case. The semiconductor shortage also meant new Worlds couldn’t be added right away.
Lessons
The longer the queue, the more often a brief connection drop for a waiting player turns into a connection error. Under the same congestion, errors cluster among players on Wi-Fi or unstable connections, so it becomes a problem that hits “only some players.” Signals to check: queue length and wait time, and the share of disconnects that happen while players wait in the queue. The primary owner is the game team (server: the queue cap and the reconnect grace period), and the infra team joins in on adding lobby and World servers. A generous reconnect grace period keeps more of these brief drops from costing players their place in the queue.
Cloudflare 2020: Traffic loss in some cities from a Cloudflare backbone configuration error
What happened
Many games rely on CDN providers for their websites, APIs, and DDoS protection, so this is the kind of infrastructure outage that affects games too. For 27 minutes on July 17, 2020, from 21:12 to 21:39 (UTC), traffic across Cloudflare’s entire network dropped by about 50%. The impact was limited to some city locations (PoPs) in the US, Europe, Russia, and Brazil that were connected to the backbone. Other locations were fine.
Cause
An outage on the Newark–Chicago backbone link congested the Atlanta–Washington link, so an engineer changed a router configuration to take some backbone traffic off Atlanta. The change was supposed to disable an entire policy term, but it disabled only the condition inside it (the prefix-list), so the Atlanta router advertised all its BGP routes across the whole backbone with a higher preference (local-preference 200). Each location gave the routes to its own servers a preference of 100, so traffic from every backbone-connected location was pulled to Atlanta. Atlanta was overloaded, and the affected locations were left with almost no traffic to handle. Service recovered once the Atlanta router was removed from the backbone. Cloudflare stated that the outage was unrelated to any attack or breach.
Lessons
If players in particular cities or regions all hit disconnects or can’t connect / infinite loading at once while everyone else is fine, first suspect a routing configuration change made just before. On the graph, CPU and traffic spike at one location only, while the affected locations actually drop to close to 0. The primary owner is the infra team (network), or external if the outage is on the provider’s side. Cloudflare decided to cap the number of routes each backbone BGP session can accept (maximum-prefix), and adjusted preferences so that one location can’t pull in traffic meant for other locations.
Many games deliver patch files, launchers, and web pages through a CDN, so this is the kind of infrastructure outage that affects games too. Starting at 09:47 (UTC) on June 8, 2021, 85% of Fastly’s network returned errors. Within 49 minutes, 95% of the network was back to normal, and the incident was resolved at 12:35.
Cause
A software deployment that began on May 12 contained a bug that would trigger when a specific customer configuration met specific conditions. On June 8, a customer pushed a valid configuration change that met those conditions. Fastly detected the problem within 1 minute, and recovery began once it identified and disabled the customer configuration that triggered it. Deployment of the bug fix began at 17:25 the same day.
Lessons
Code deployed weeks earlier can still cause a global outage in an instant when it meets a rare condition. On the game side, the signals are HTTP error rates for patch, launcher, and web requests rising in every region at the same time, and the CDN provider’s status page. A telltale sign is that game connections already in progress stay fine if they don’t go through the CDN, and only new connections, patch downloads, and web logins are blocked. The primary owner is external (the CDN provider). The game and infra teams should have a fallback ready, such as a second CDN or a path that fetches directly from the origin server.
Meta 2021: Facebook outage: one backbone command took DNS down with it
What happened
An infrastructure outage whose lessons apply directly to a game company’s own network and DNS. On October 4, 2021, Facebook (now Meta) services were unreachable worldwide. The backbone linking its data centers went down completely, and Facebook’s DNS servers could no longer be found from the internet. The postmortem doesn’t say how long the outage lasted.
Cause
During routine maintenance, a command issued to assess global backbone capacity unintentionally took down every connection in the backbone, and an audit tool meant to block commands like this failed to stop it because of a bug. DNS servers at smaller locations are designed to mark themselves unhealthy and withdraw their BGP advertisements when they can’t talk to the data centers, so the DNS servers became unreachable from the internet even though they were still running. The normal access paths and out-of-band access were both down, and internal tools had lost DNS as well, so engineers had to be sent to the data centers in person, and security procedures slowed that down further. By the time of recovery, power draw at each data center had dropped by tens of MW, and bringing everything back at once could put everything from electrical systems to caches at risk, so load was raised in stages.
Lessons
If can’t connect / infinite loading hits every region and every ISP at the same time, check DNS and BGP routes before the game servers. You can confirm this from outside the company with external DNS lookups and public BGP route data. The primary owner is the infra team (network). Check in advance that your out-of-band access path and the internal tools you’d use in an outage don’t depend on the same DNS and network, and during recovery, raise load in stages so reconnects don’t all arrive at once.
Many games run their servers, login, and data on public clouds, so this is the kind of infrastructure outage that affects games too. At 7:30 AM PST on December 7, 2021, the internal network in the Northern Virginia region (us-east-1) became congested. From 7:33 AM, EC2 API errors and latency rose, making it hard to launch new instances (instance launches recovered at 2:40 PM), followed by console login failures, Route 53 configuration changes being blocked, and delayed and partly lost CloudWatch metrics. The network devices fully recovered at 2:22 PM. EC2 instances that were already running and existing DNS responses were not affected.
Cause
An automated activity to scale capacity for a service in the main network triggered unexpected behavior from a large number of clients in the internal network, causing a surge in connection attempts. The devices linking the internal network to the main network were overwhelmed and communication slowed down. The delays in turn drove more connection attempts and retries, so the congestion persisted. The clients had backoff behavior that spaces out requests during congestion like this, but a latent defect kept it from working properly. Internal monitoring depended on the same network, so the operations team had to respond using logs, without real-time metrics.
Lessons
When retries can’t back off, a brief bout of congestion turns into an outage that lasts hours. From the game’s point of view, game servers that are already running may be fine, but new server capacity (autoscaling), any login, matchmaking, or payment flow that calls cloud APIs, and monitoring can all be blocked at the same time. Signals to check: the cloud provider’s status page, cloud API error rates, and instance launch failures. The primary owner is external (the cloud provider). The game team should give every retry exponential backoff with randomized delays and a retry limit, and the infra team should keep enough spare capacity to ride out blocked scale-out, plus a fallback in another region.
Cloudflare 2025: Cloudflare public DNS 1.1.1.1 outage
What happened
An outage of a public DNS resolver that players set up themselves on their devices or routers. This type of outage blocks every game and service at once, but only for players using that setting. For 62 minutes on July 14, 2025, from 21:52 to 22:54 (UTC), the 1.1.1.1 resolver stopped responding worldwide. Cloudflare said that for many users, this meant being unable to use essentially any internet service. Queries over UDP, TCP, and DNS over TLS were affected, while DNS over HTTPS, which connects by domain name, stayed relatively stable.
Cause
On June 6, while a service topology (the configuration that decides which locations advertise which IP ranges) was being prepared for a different service not yet in use, the 1.1.1.1 resolver’s IP ranges were mistakenly included in it. When that service’s configuration was changed on July 14, the locations advertising the resolver ranges shrank from every location to a single offline one, and the BGP routes were withdrawn worldwide. The change skipped canary deployment and went straight to every data center. Reverting the configuration at 22:20 brought traffic back to about 77%, but in the meantime about 23% of edge servers had lost required IP configuration, and re-applying it meant service wasn’t back to normal until 22:54. Cloudflare said it was an internal configuration error unrelated to any attack or BGP hijack.
Lessons
If the game servers and other players are fine but some players get can’t connect / infinite loading on the login or patch servers, suspect the DNS those players use. A telltale sign is that existing sessions stay connected and only new connections fail. Having those players change their DNS settings or look up the server address directly settles it right away. The primary owner is external (the DNS operator or ISP). If the game team (client) reports name resolution failures separately from other errors, customer support can make the call right away.
AWS 2025: AWS us-east-1 DynamoDB DNS outage and long recovery
What happened
Many games run their servers, login, and data on public clouds, so this is the kind of infrastructure outage that affects games too. From 11:48 PM on October 19, 2025, to 2:20 PM on October 20 (PDT), the Northern Virginia region was affected in three stages. DynamoDB API errors were elevated until 2:40 AM on the 20th; new EC2 instance launches failed from 2:25 AM to 10:36 AM (connection problems on some new instances cleared up at 1:50 PM); and some Network Load Balancers (NLB) saw more connection errors from 5:30 AM to 2:09 PM.
Cause
The automation that manages DynamoDB’s DNS had a latent race condition. Among the processes that apply DNS plans in different Availability Zones (DNS Enactors), one that was running unusually late overwrote a newer plan with an old one. Right after that, another Enactor’s cleanup job deleted that old plan, leaving the DNS record for the regional endpoint (dynamodb.us-east-1.amazonaws.com) empty. The automation couldn’t repair this, so it had to be fixed by hand. EC2’s system for managing physical servers depends on DynamoDB, so in the meantime the lease it kept for each physical server expired. After DynamoDB came back, there were so many physical servers that re-establishing the leases timed out before finishing, and retries piled up again, pushing the system into “congestive collapse.” Network configuration for newly launched instances propagated slowly, so NLB health checks flapped between passing and failing, and even healthy nodes were repeatedly removed from DNS and added back.
Lessons
A DNS record error in one place spreads to the other services that depend on that service, and even after the cause is fixed, backlogged work and flapping health checks stretch recovery out by several more hours. From the game’s side, servers already running may hold up, but new servers can’t launch, so autoscaling stalls, and flapping health checks can make the load balancer pull healthy servers out of rotation. Signals to check: the cloud status page, managed service API error rates, instance launch failures, and the load balancer’s healthy target count. The primary owner is external (the cloud provider). The infra team should cap how many servers can drop out at once on failed health checks and have a fallback in another region ready.
Ping, RTT. The time it takes for a signal you send to reach the server and come back (round trip). The ping a game displays sometimes also includes time spent waiting for server processing.
Latency
Latency. The time a packet takes to get from where it was sent to where it arrives. It often refers to one direction only, so it’s roughly half the ping.
Jitter
Jitter. Variation in packet arrival intervals. Even at the same average ping, high jitter makes the screen stutter.
Packet
Packet. A chunk of data sent over the network in one go. Usually 1,500 bytes at most; game updates are tens to hundreds of bytes.
Packet loss
Packet loss. Packets that are sent but never arrive. In a game that uses TCP, even 1% loss causes a noticeable hitch every few seconds to ten-odd seconds, while a UDP game with interpolation and redundant input sending can sometimes hide a few percent.
Bandwidth
Bandwidth. The maximum amount of data a connection can carry per second (Mbps). It’s a different concept from how quickly data arrives (latency).
Tick
Tick. One step in which the server computes the game state. A 20-tick server computes 20 times per second, every 50 ms.
Tick rate
Tick rate. How many ticks the server runs per second. Higher means faster response, but server cost and bandwidth go up. To save bandwidth, some games send packets less often than the tick rate.
Tick budget
Tick budget. The time limit for finishing one tick. Go over it and the next tick starts late, stretching the tick interval.
FPS
Frames per second. How many times per second the screen is drawn. At 60 FPS, each frame gets 16.7 ms.
Frame time
Frame time. How long it took to draw one frame. Occasional spiky frames matter more to how a game feels than average FPS.
Snapshot
Snapshot. A summary of “the game state right now” that the server sends every tick: position, health, status, and so on. It usually carries only what changed relative to what the receiver already has (delta compression).
Interpolation
Interpolation. Drawing smooth motion between two received snapshots. The trade-off is that it shows a moment slightly in the past.
Interpolation buffer
Interpolation buffer. Time the client deliberately waits before drawing, so it has something to interpolate. It’s the slack that absorbs jitter and one or two lost packets. It’s usually twice the packet interval (100 ms at 20 updates per second), and some games grow it automatically when jitter increases.
Extrapolation
Extrapolation, Dead reckoning. When no new packet has arrived, guessing where an entity is headed from its last velocity and drawing it there. A wrong guess looks like teleporting, so many games extrapolate for only about 0.25 s and then stop (the Source engine default is 0.25 s).
Client-side prediction
Client-side prediction. Moving your own character on screen right away, without waiting for server confirmation.
Server reconciliation
Reconciliation. When the server’s result arrives, comparing it with the prediction and correcting your character’s position. The client starts from the position the server confirmed and replays the inputs the server hasn’t confirmed yet. A large correction looks like rubber-banding.
Lag compensation
Lag compensation. When the server judges an attack, it rewinds to the moment the attacker was seeing and checks whether the attack hit. To keep things fair for the player being hit, the rewind is capped: 0.2–0.25 s is common in competitive shooters, and some games rewind as far as 1 s, like the Source engine default.
Authoritative server
Authoritative server. A design in which only the server makes final decisions. It stops cheating, but every outcome needs a round trip to the server, so prediction and client-side feedback hide the wait.
Lockstep
Deterministic lockstep. Everyone exchanges only inputs and runs the exact same computation on the same turn. Inputs get a fixed delay, and if anyone’s input is late, everyone waits.
Server-side input buffer
Server-side input buffer. A per-player buffer in which the server collects a few inputs and consumes one per tick. Even a player with high jitter looks smooth to others, but that player’s actions are confirmed on the server correspondingly later.
Listen server
Listen server. One player’s PC plays the game and acts as the server at the same time. The host has zero ping, but if the host’s connection or PC is slow, everyone lags.
Phasing
Phasing. Showing different NPCs and terrain in the same place depending on quest progress. If two characters are at different points in a quest, it’s normal for an NPC to be missing for one of them.
Rollback netcode
Rollback netcode (GGPO). Predicting the opponent’s input to keep the game moving, then rewinding to a past frame and recomputing when the real input turns out different. Common in fighting games. Unrelated to a database rollback.
Input buffering
Input buffer, spell queue. Accepting the next input pressed shortly before a cooldown or animation ends, and executing it the moment it ends. This keeps a round trip from sneaking in between combo steps.
Client-side feedback
Client-side feedback. Playing animations, sounds, and effects right away, without waiting for server confirmation. Only results that need confirmation, like damage and rewards, wait for the server’s reply. If the server rejects the action, what was shown has to be reverted.
TCP
Transmission Control Protocol. A protocol that delivers data in order, with nothing missing. It holds later packets back from the game until a lost packet has been received again.
UDP
User Datagram Protocol. A protocol that delivers whatever is sent, with no guarantees. There’s no waiting, but the game has to handle loss and ordering itself.
Reliable UDP
Reliable UDP (KCP, ENet…). An approach that implements just as much retransmission and ordering as needed on top of UDP.
HOL blocking
Head-of-line blocking. One stuck item at the front makes everything behind it wait. This is what causes fast-forward on TCP.
RTO
Retransmission timeout. The retransmission timer: how long TCP waits before deciding a packet is lost and sending it again. On Linux it’s ping + 200 ms or more, doubling after each failure.
Nagle’s algorithm
Nagle’s algorithm. A TCP feature that holds small pieces of data until the acknowledgment (ACK) for earlier data arrives, then sends them together to save packets. Games should usually turn it off.
TCP_NODELAY
TCP_NODELAY. The socket option that turns off Nagle’s algorithm. Small messages go out immediately.
Delayed ACK
Delayed ACK. Sending the acknowledgment of receipt a little later, bundled with other data. Linux typically waits 40 ms (up to 200 ms); older Windows versions wait 200 ms and current ones 40 ms.
Socket buffer
SO_SNDBUF / SO_RCVBUF. The size of the send and receive queues the OS keeps for each socket. Too small and they overflow; too large and stale data piles up and waits.
keepalive
SO_KEEPALIVE. A TCP feature that checks whether an idle connection is still alive. It’s off by default, and even when turned on, the default is to check only after 2 hours.
RST
TCP reset. A TCP signal that kills a connection on the spot. Any data not yet sent is discarded.
Heartbeat
Heartbeat. An “I’m alive” signal the game itself sends periodically. It’s used to detect dead connections and to keep connections open on devices along the path.
Timeout
Timeout. How long to wait for a response before treating it as a failure. Too short causes false alarms; too long delays detection.
NAT
Network Address Translation. A router feature that sends traffic from several devices at home out through one public IP, recording each connection in the NAT table.
CGNAT
Carrier-grade NAT. Large-scale NAT in which an ISP shares one IP address among many subscribers.
MTU
Maximum Transmission Unit. The largest packet that can be sent in one piece. Usually 1,500 bytes, and smaller on VPN and PPPoE segments.
Bufferbloat
Bufferbloat. Network equipment letting its queues grow so large that latency climbs to hundreds of ms.
SQM
Smart Queue Management (fq_codel, CAKE). A router feature that keeps queues short and sends traffic out fairly per flow. The fix for bufferbloat.
QoS
Quality of Service. A feature that gives important traffic priority so it goes out first.
Peering
Peering. The points where ISPs connect their networks to each other. They tend to get congested in the evening.
BGP
Border Gateway Protocol. The protocol ISPs use to tell each other which routes to take across the internet. When routes change, the path and ping change too.
DDoS
Distributed Denial of Service. An attack that floods a service with traffic from many sources to knock it offline.
Scrubbing center
DDoS scrubbing center. A DDoS mitigation provider’s site that receives traffic headed for your servers during an attack, filters out the attack, and forwards only legitimate traffic. A distant site makes the route longer.
Firewall
Firewall. A device or program that lets through only permitted connections. It tracks connections in a session table.
Load balancer
Load balancer. A device that spreads incoming connections across several servers.
Session table
Session table, conntrack. The table a device or OS uses to track current connections. Its size is limited.
Microburst
Microburst. Traffic piling up in very short windows of 1 ms or less, even though the average is low.
NIC
Network Interface Card. A server’s network card.
Ring buffer
Ring buffer. The buffer that holds packets the NIC has received until the CPU picks them up. It cycles through a fixed number of slots, and when every slot is full, new packets are dropped.
Interrupt
Interrupt. A signal a device sends to tell the CPU that there’s work to handle.
RSS
Receive Side Scaling. A NIC feature that spreads incoming packets across several receive queues so multiple CPU cores can process them.
PPS
Packets per second. Packets per second. Game servers often hit this limit before they run out of bandwidth.
Kernel
Kernel. The core of the operating system. It manages networking, memory, and how CPU time is shared out.
backlog
Listen backlog. The queue where new connection requests wait until the server accepts them. When it’s full, Linux silently drops new requests and Windows sends a rejection.
TIME_WAIT
TIME_WAIT. The state in which the side that closed a connection first keeps that port pair reserved for a while (60 seconds on Linux) in case late packets arrive.
CPU steal
Steal time. Time a virtual machine spent waiting for CPU because the physical host was giving it to other virtual machines. Shown as the st value in top.
CPU throttling
CFS throttling. Forcibly pausing a container until the next period once it has used up its CPU quota within a set period (CFS period, usually 100 ms).
File descriptor
File descriptor. The number (fd) attached to each file or connection a process opens. There’s a limit on how many a process can have.
Thread
Thread. A unit of work that runs independently inside a program. Several threads can run at the same time.
Context switching
Context switch. The CPU swapping out the running thread for another one. It has a cost.
Lock
Lock, Mutex. A guard that lets only one thread at a time use shared data.
Deadlock
Deadlock. A state in which threads each wait for a lock the other holds and stay stuck forever.
Thread pool
Thread pool. A set of worker threads created in advance. When all of them are busy, new work waits.
Asynchronous I/O
epoll, IOCP, io_uring. A model in which a thread keeps doing other work while I/O is in progress and gets notified when it completes.
AOI
Area of Interest. The range each player “can see.” Only changes inside it are sent, which cuts bandwidth. To reduce the cost of working out who’s in range, the map is usually divided into a grid and only nearby cells are checked.
Broadcast
Broadcast, fan-out. Sending one change to everyone who can see it. If everyone in a crowd can see everyone else, the amount to send grows with the square of the player count.
GC
Garbage collection. Automatically reclaiming memory that’s no longer in use. Programs can pause during GC.
Heap
Heap. The memory area a program allocates from whenever it needs memory while running.
Memory leak
Memory leak. A bug in which memory that’s no longer needed is never released, so usage keeps growing. It happens even with GC if something still holds a reference to objects the program is done with.
Swap
Swap, paging. Moving part of memory to disk when RAM runs short. Using that memory again is more than 1,000 times slower than RAM.
OOM killer
Out-of-memory killer. A Linux feature that, when memory runs out, picks the process using the most memory and forcibly kills it. In a container it kicks in as soon as the container’s memory limit is hit.
Cache miss
Cache miss. The data isn’t in the cache close to the CPU, so it has to come from slower memory.
IOPS
I/O operations per second. How many reads and writes a disk can handle per second. On cloud disks, the limit depends on what you pay for.
fsync
fsync. A call that waits until data is safely written to disk. Normal writes land in OS memory first and reach the disk later, so they can be lost if the server loses power in between. fsync is safe but slow.
Burst credits
Burst credits. A balance that cloud disks and servers build up so they can briefly run above their baseline performance. When it runs out, performance drops back to baseline.
Index
Index. A lookup structure in a database. Without one, the database has to read the whole table.
Full table scan
Full table scan. A query that checks every row of a table without using an index.
Query plan
Query plan. The database’s plan for executing a query: in what order and with which indexes. Even with unchanged code, the same query can suddenly slow down if the database switches plans.
Transaction
Transaction. A group of DB operations that either all succeed or all fail. Trades must always be handled in a transaction. Rows it modifies stay locked until it finishes, so shorter is better.
Connection pool
Connection pool. A set of DB connections opened in advance. When all of them are in use, new requests wait.
Hot row
Hot row. A single row that many requests try to modify at the same time. A cause of lock contention.
Replication lag
Replication lag. How far a replica database has fallen behind the primary.
Rollback
Rollback. A save is canceled and data returns to its previous state. Players experience it as “my item disappeared.”
Cache
Cache (Redis etc.). A copy of frequently used data kept somewhere fast. It reduces database load.
Checkpoint
Checkpoint. The database periodically writing the changes it has collected in memory out to disk in one batch. Saves and queries can slow down briefly while it happens.
Failover
Failover. Switching over to a standby when the primary server or DB goes down. Writes stop briefly during the switch, and if replication was lagging, the most recent data can be lost.
MVCC
Multi-version concurrency control. A way for a database to keep old row versions for a while so readers and writers don’t block each other. A long-open transaction lets old versions pile up and slows things down.
Cache stampede
Cache stampede. Cache entries empty out all at once and requests flood the origin (DB).
Gateway
Gateway. An intermediate server that accepts client connections and forwards them to the game servers behind it.
Circuit breaker
Circuit breaker. A mechanism that temporarily stops calling a service that keeps failing and returns a failure right away, which prevents cascading failures. After a while it tries a call or two, and if the service has recovered, it lets calls through again.
Cascading failure
Cascading failure. A failure in one place spreading to other services along the call chain.
Autoscaling
Autoscaling. Automatically adding and removing servers based on load. Adding servers takes time.
Watchdog
Watchdog. A timer that watches whether the server has frozen. If the game loop stalls longer than a set time (a few seconds to tens of seconds), it writes a state dump and forcibly kills the server so it restarts.
Utilization
Utilization. The fraction of time a worker (anything that processes requests, such as a CPU core, thread, or DB connection) is busy. Above 80–90%, waiting time climbs steeply.
p99
99th percentile. The value that 99 out of 100 measurements beat, with about 1 in 100 slower. It reflects the lag players feel better than the average does.
V-Sync
Vertical sync. Sending frames in step with the display’s refresh cycle. It eliminates screen tearing but adds input lag, and when FPS drops below the refresh rate, it can bounce between 60 and 30 and stutter.
Variable refresh rate
VRR, G-Sync, FreeSync. The monitor refreshes whenever a frame is ready. It reduces the stutter and input lag caused by V-Sync bouncing between 60 and 30.
Anti-cheat
Anti-cheat. A security module that blocks game hacks. When its periodic scans or its heartbeats to the server fail, it can cause stutter or disconnects.
Overlay
Overlay. A feature of chat, recording, or FPS-counter programs that draws on top of the game screen. It hooks into the game’s rendering and can cause stutter.
Shader compilation
Shader compilation. Converting graphics effect programs into code for the GPU. If it isn’t done ahead of time, the screen hitches the first time an effect appears, and updating the graphics driver invalidates the saved results, so it happens all over again.
Main thread
Main thread, Game thread. The game’s central thread, which runs game logic and prepares each frame in turn. If any one task on it takes long, the screen freezes for that long.
Timer resolution
Timer resolution. The shortest interval at which the OS can wake a sleeping program. The Windows default is 15.6 ms, so unless a program changes it, even “wake me in 1 ms” wakes up late.
Thermal throttling
Thermal throttling. A protective feature that lowers CPU and GPU speed when a device gets hot. On phones it’s common after a few to tens of minutes of play.
VRAM
Video memory. Dedicated memory on the graphics card. Textures and models are loaded here for rendering. When it runs short, data shuttles to and from PC memory over a slower path and the game stutters.
Net graph
Net graph. A dev and debug overlay that shows ping, packet loss, FPS, and tick as live graphs on the game screen. Having it in a lag report video makes finding the cause much easier.
Retransmission rate
Retransmission rate. The share of sent TCP packets that had to be sent again. There’s no official threshold, but a server-wide average below 0.1% is generally healthy, and above 1% many players are likely to feel lag. Also check how many times higher it is than the usual level.
SACK
Selective ACK. A TCP feature in which the receiver reports in detail: “I got this range; only this part is missing.” It can recover several losses at once.
RACK-TLP
Recent ACK, Tail Loss Probe. A TCP feature that detects loss based on time and, when no ACK arrives for a while, resends the last packet to speed up recovery. It’s the default on current Linux and Android. On Windows, TLP and RACK are on by default from Windows 10 (1607) and Server 2016, and the newer RACK that also recovers lost retransmissions arrived with Server 2022. It works only on connections with SACK enabled.
Spurious retransmission
Spurious retransmission. Resending a packet that wasn’t lost: it arrived late or out of order, and the sender mistook it for lost. It wastes bandwidth and needlessly cuts the sending rate.
Zero window
Zero window. The receiver’s buffer is full, so it has told the sender “stop sending for now.” It looks like retransmission, yet the network is fine; the receiving program just didn’t read the data in time.
thin stream
Thin stream. A connection that sends small packets sparsely, like a game. Fast retransmit signals rarely build up, so losses cause long stalls.
Policer
Policer. A rate limiter that drops packets exceeding a set rate right away, without queuing them. A limiter that queues them and releases them slowly is called a shaper.
Pacing
Pacing. Spreading outgoing packets evenly over time so they don’t all go out in one burst. It keeps small buffers from overflowing.
ECN
Explicit Congestion Notification. Under congestion, marking packets with a “congested” flag, without dropping them, so the sender slows down. It signals congestion with no loss. It only works if both endpoints and the equipment on the congested segment all support it.
MSS
Maximum Segment Size. The maximum amount of data TCP puts in one packet. Usually 1,460 bytes; lowering it to fit tunnel segments prevents MTU black holes.
Handover
Handover. A moving phone switching the cell tower it’s connected to.
Percentile
Percentile (p50, p95, p99). The value at a given percentage position when all values are sorted from smallest to largest. p50 is the median; p99 is the value near the slowest 1 in 100. It reveals the spikes an average hides.
Tail latency
Tail latency. Long delays that happen occasionally even though most requests are fast. They barely show up in the average, but they’re what players remember as lag.
Synthetic monitoring
Synthetic monitoring. Measuring path quality with dedicated probes or servers that periodically run ping, traceroute, and similar tests from fixed locations, standing in for real players. RIPE Atlas is the best-known public tool.
Aggregation interval
Aggregation interval. How many seconds or minutes of data one point on a graph combines. The longer the interval, the more short spikes get averaged away.
Postmortem
Postmortem. A write-up after an incident covering what happened, why it happened, and what will change. It is written blamelessly, with the goal of preventing a repeat.
C-state
CPU idle state. Power-saving states a CPU enters when idle. Deeper states save more power but take longer to wake from.
Live migration
Live migration. A cloud provider moving a running virtual machine to another host, for example for host maintenance. The VM can pause briefly during the move.
SNAT
Source NAT. NAT that rewrites the source address of outgoing packets to a public address. Each public address has a limited number of usable ports, and new connections fail once they run out.
NAT gateway
NAT gateway. A cloud device that lets servers on a private network share one public address to reach the internet. It limits concurrent connections per destination.
LEO satellite internet
LEO satellite internet. Internet service through satellites orbiting hundreds to thousands of km up. Latency is far lower than with geostationary satellites, but it can spike when the connection switches to another satellite.
GeoIP
IP geolocation. A database that estimates country, city, and ISP from an IP address. Wrong or outdated entries can get players assigned to a distant server.
TLS certificate
TLS certificate. A digital document proving a server is who it claims to be. It has an expiry date, and once it expires, encrypted connections fail and players can’t connect.
Frame generation
Frame generation. A technology in which the graphics card inserts predicted frames between the frames it actually rendered to raise FPS. Motion looks smoother, but input-to-screen latency can increase.
References
616 sources from 83 publishers: standards documents, official kernel, OS, cloud, engine, and database documentation, research papers, and technical posts by the original developers.
Microsoft 85
_WDF_TIMER_CONFIG (wdftimer.h)Microsoft Standard timer accuracy is the system clock tick interval, 15.6 ms by default; high-resolution timers get 1 ms
/fp (Specify floating-point behavior)Microsoft /fp:fast may reorder or combine floating-point operations and give results different from other /fp settings, and operations fused with FMA can also differ from a separate multiply and add
About Windows Filtering PlatformMicrosoft Packets are allowed or blocked through hooks in the Windows network stack and a filter engine, and third-party security vendors can plug in their own filter modules (callouts)
Acquiring high-resolution time stampsMicrosoft QueryPerformanceCounter (used by Stopwatch) is a clock for elapsed time that isn’t synced to external time; use system time only when UTC time is needed
ASP.NET Core Best PracticesMicrosoft Call data access, I/O, and long-running work asynchronously; synchronous blocking calls lead to thread pool exhaustion and slow responses
closesocket function (winsock.h)Microsoft Turning SO_LINGER on with a time of 0 makes close an abortive close that resets the connection immediately, and unsent data is lost
Collecting User-Mode DumpsMicrosoft Configure Windows Error Reporting (WER) to collect full or mini dumps locally when a user-mode program crashes
CPU AnalysisMicrosoft WPA’s DPC/ISR graph: the duration of each uninterrupted DPC or ISR run and the module containing that function (Module)
CreateMutexW function (synchapi.h)Microsoft If a named mutex already exists, ERROR_ALREADY_EXISTS is returned, which is used to detect duplicate launches and allow only a single instance
Creating and Opening FilesMicrosoft A file opened without a sharing mode can’t be opened by another process, which gets ERROR_SHARING_VIOLATION
Customize the Windows performance power sliderMicrosoft Simulation’s power-saving mode: Windows power modes change power and CPU settings to extend battery life at the cost of performance
Debug ThreadPool StarvationMicrosoft When the pool has no threads left and new work has to wait, responses slow down; blocking code that holds threads is the cause. In dotnet-counters, dotnet.thread_pool.thread.count creeping up while CPU is well below 100% signals exhaustion (dotnet.thread_pool.queue.length is often large too); dotnet-stack shows where threads are waiting
Delivery Optimization referenceMicrosoft Windows Update downloads (Delivery Optimization) adjust dynamically to available bandwidth by default, and caps can be set on background and foreground download bandwidth
Direct3D 12 Return CodesMicrosoft D3D12_ERROR_DRIVER_VERSION_MISMATCH: a PSO cache built with a different driver version can’t be reused (recompiled after a driver update)
DirectStorage is coming to PCMicrosoft Older hard drives read tens of MB per second and NVMe SSDs several GB per second; the asset streaming budget of previous-generation games was around 50 MB per second; open-world games read and discard distant scenery in real time as the player moves
Guidelines for Writing DPC RoutinesMicrosoft All threads on that core stop while a DPC runs, so the guidance is to keep each one under 100 µs
I/O Completion PortsMicrosoft Handle many asynchronous I/Os with a pre-created thread pool and IOCP, and match the number of concurrently running threads to CPU concurrency
Introduction to the page fileMicrosoft The page file is a file on disk used to move rarely used, modified memory pages out of RAM
ipconfigMicrosoft Run with no parameters, it shows each adapter’s IPv4 and IPv6 addresses and default gateway
listen function (winsock2.h)Microsoft On Windows, when the queue is full the client gets a WSAECONNREFUSED error
Logging in C#Microsoft .NET logging methods are synchronous, so with slow storage it recommends writing to fast storage first and moving the logs later
Low Latency Workloads Management and OperationsMicrosoft Dropped Datagrams and Dropped Datagrams/sec in the Microsoft Winsock BSP counter set: UDP datagrams dropped because they arrived faster than the app could process them or the receive socket buffer was too small
Minidump FilesMicrosoft A minidump holds only a useful subset of crash dump information, so it is fast and small
MultitaskingMicrosoft Windows gives each thread a time slice and moves on to the next thread when it’s used up; a time slice is about 20 ms (varies by OS and CPU)
netstatMicrosoft -a shows TCP and UDP ports, -n numeric addresses, -o the process ID (PID), and -p udp UDP only
Priority BoostsMicrosoft The process in the front window (foreground) gets its priority raised to at least that of background processes
Process MonitorMicrosoft Records file system, registry, and process activity in real time and can filter on any field, such as path
Pushing the Limits of Windows: Virtual MemoryMicrosoft On Windows, once the commit limit is reached, allocations that commit memory fail, which can lead to app errors or system failures
Quality of ServiceMicrosoft Programs whose windows can’t be seen or heard get Low QoS and, on battery, are scheduled at the most efficient CPU speed and on efficiency cores
recvfrom function (winsock.h)Microsoft WSAECONNRESET on a UDP socket means a previous send got an ICMP Port Unreachable back
Reduce latency with DXGI 1.3 swap chainsMicrosoft Present blocks until the queue drains, so a rendered frame waits almost one extra frame before it’s displayed; a waitable swap chain reduces this
Request schedulingMicrosoft Orleans grains (actors) use a single-threaded execution model that runs each request to completion one at a time, so state is never modified concurrently; grains waiting on each other’s responses can deadlock
ResidencyMicrosoft Each process has a graphics memory budget; when it’s exceeded, the kernel moves part of the discrete GPU’s heap to system memory (a last resort, so managing the budget is recommended)
Resolve-DnsNameMicrosoft Resolves a name against the DNS server specified with -Server
Results for the Idle Energy Efficiency AssessmentMicrosoft Default system timer resolution is 15.6 ms; the “Platform Timer Resolution” entries in the energy report show which processes changed the timer resolution
Scheduling PrioritiesMicrosoft Among runnable threads, those with the highest priority take turns getting time slices (round robin)
send function (winsock2.h)Microsoft Winsock send also blocks when buffer space runs out, unless the socket is in non-blocking mode
TCP/IP port exhaustion troubleshootingMicrosoft Windows dynamic ports default to 49152–65535, and a closed connection holds its port in TIME_WAIT for 4 minutes by default
TcpMaxConnectRetransmissionsMicrosoft Older Windows defaults: 2 SYN retransmissions, starting with a 3-second wait and doubling, then waiting double again after the last one before giving up (3+6+12=21 seconds)
timeBeginPeriod function (timeapi.h)Microsoft Before Windows 10 2004 it was a global setting; since then it applies only to the requesting process, and Windows 11 doesn’t guarantee high resolution to processes whose windows are hidden or minimized
WDI low latency connection qualityMicrosoft Scanning and roaming move the radio off the connected channel, so low-latency mode limits scanning and time spent off-channel
Well-known EventCounters in .NETMicrosoft Monitor Lock Contention Count (monitor-lock-contention-count): number of times contention occurred when trying to acquire a monitor lock
Windows Firewall RulesMicrosoft Inbound connections are blocked by default, so apps need exception rules, which the app installer usually creates
Windows Performance Monitor Disk Counters ExplainedMicrosoft Avg. Disk sec/Read is the average time one read takes to complete (I/O latency); Current Disk Queue Length is the disk queue length at the moment of measurement
WlanSetInterface function (wlanapi.h)Microsoft Windows API that turns background scanning (wlan_intf_opcode_background_scan_enabled) and media streaming mode on and off
Working SetMicrosoft Touching a page that isn’t in RAM causes a page fault; a hard fault can only be resolved by reading from disk, such as from the page file
Xbox Series X: What’s the Deal with Latency?Microsoft Input lag is the sum of the path controller → console → HDMI → TV; older controllers read and sent input every 8 ms; sending one frame over HDMI takes 16.6 ms at 60 Hz and 8.3 ms at 120 Hz; ALLM switches the TV to game mode automatically
Linux kernel 62
ABI stable symbolsLinux kernel /sys/block/(disk)/queue/rotational: whether the device is rotational or non-rotational
CFS Bandwidth ControlLinux kernel Threads that use up the quota for a period stop until the next period (throttling), default period 100 ms, nr_throttled statistic
Concepts overviewLinux kernel Reclaims page cache backed by files on disk and swappable pages; if that’s still not enough, the OOM killer kills a process
Control Group v2Linux kernel cpu.max takes the form “$MAX $PERIOD” (quota, period), and the default is “max 100000” (100 ms period)
CPU Idle Time ManagementLinux kernel Each power-saving state has a wakeup time (exit latency) and a minimum stay (target residency), and deeper states are chosen based on expected idle time; per-state latency, usage, and time in sysfs; PM QoS (/dev/cpu_dma_latency) and intel_idle.max_cstate limit deep states
CPU Performance ScalingLinux kernel Check and change the governor with scaling_governor; performance requests the highest allowed frequency, powersave the lowest
Documentation for /proc/sys/net/Linux kernel netdev_max_backlog: limit of the receive queue that holds packets arriving faster than the kernel can process them
Documentation for /proc/sys/vm/Linux kernel min_free_kbytes: the minimum free memory (watermark) the kernel keeps in reserve
drivers/idle/intel_idle.c (Linux v6.12)Linux kernel Wakeup times of C-states on Intel server CPUs: Skylake-SP C1 2 µs, C1E 10 µs, C6 133 µs; Ice Lake C6 170 µs; Sapphire Rapids C1 1 µs, C6 290 µs
EEVDF SchedulerLinux kernel Linux began moving from CFS to the EEVDF scheduler in 6.6
Ethtool countersLinux kernel The mlx5 driver’s rx_out_of_buffer (no buffer in the receive queue) and rx_discards_phy (dropped for lack of port buffer)
include/net/sock.h (Linux v6.18)Linux kernel The default socket buffer is defined as 256 packets of 256 bytes including sk_buff overhead (SKB_TRUESIZE(256)×256); even a small frame counts as sk_buff + MTU (about 208 KB is the computed value on x86-64)
include/uapi/linux/tcp.hLinux kernel tcpi_reord_seen in tcp_info: how many times the connection has seen reordering
intel_pstate CPU Performance Scaling DriverLinux kernel The intel_pstate powersave algorithm differs from the generic powersave governor and scales with load (similar to schedutil and ondemand)
Interface statisticsLinux kernel rx_crc_errors: number of packets received with CRC errors; ip -s -s link shows errors by type
IP SysctlLinux kernel tcp_mtu_probing=1 is normally off and turns on TCP path MTU probing only when it detects an ICMP black hole
MDS - Microarchitectural Data SamplingLinux kernel Mitigations flush CPU buffers when returning from the kernel to user space and when entering a VM; the files under /sys/devices/system/cpu/vulnerabilities/ show vulnerability and mitigation status; many CPUs need SMT off for full protection, and turning SMT off can have a large performance impact depending on the workload
mm/oom_kill.c (Linux v6.12)Linux kernel The process using the most memory gets the highest score (oom_score_adj factored in); the kernel logs “Out of memory: Killed process …” when it kills
NAPILinux kernel A large gro_flush_timeout batches more work together but adds latency under low load
net/core/net-procfs.cLinux kernel /proc/net/softnet_stat has one line per CPU, in hex; the 2nd column is dropped and the 3rd is time_squeeze
net/core/sock_reuseport.c (Linux v6.12)Linux kernel Without a BPF program, the packet hash is mapped onto the number of sockets in the group to pick the socket
net/ipv4/proc.c (Linux v6.12)Linux kernel Counter names shown by nstat: RcvbufErrors and SndbufErrors in the Udp group
net/ipv4/tcp_bbr.cLinux kernel BBR sets pacing_rate from its estimated bottleneck bandwidth
net/ipv4/tcp_diag.c (Linux v6.12)Linux kernel Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed
net/ipv4/tcp_input.c (Linux v6.12)Linux kernel Linux RTO = smoothed RTT + RTT variation, and the variation term has a floor of tcp_rto_min (200 ms), so RTO is at least RTT + 200 ms
net/ipv4/tcp_ipv4.cLinux kernel Default initialization: tcp_early_retrans 3, tcp_recovery RACK, tcp_syn_linear_timeouts 4, tcp_base_mss 1,024 (tcp_mtu_probing isn’t set explicitly, so it’s 0)
net/ipv4/tcp_output.cLinux kernel Linux schedules TLP only on connections that use SACK
net/ipv4/tcp_recovery.cLinux kernel RACK reordering window = min(min_RTT/4 × step count, SRTT); retransmissions acknowledged faster than the minimum RTT are excluded from the reference
net/ipv4/tcp_timer.c (Linux v6.12)Linux kernel Each time the retransmission timer expires, TCPTimeouts goes up, backoff increases by one, and RTO doubles (up to the maximum)
net/ipv4/udp.c (Linux v6.12)Linux kernel When the UDP receive queue exceeds the socket buffer size, packets are dropped immediately and RcvbufErrors goes up
net/netfilter/nf_conntrack_core.c (Linux v6.12)Linux kernel When the connection tracking table is full, the kernel logs “nf_conntrack: table full, dropping packet” and drops packets for new connections
net/netfilter/nf_conntrack_standalone.cLinux kernel /proc/net/stat/nf_conntrack has one line per core, in hex, with columns such as entries, invalid, insert_failed, drop, and early_drop
net/sched/sch_generic.c (Linux v6.12)Linux kernel When a transmit queue stalls, the kernel watchdog logs “NETDEV WATCHDOG … transmit queue N timed out” and calls the driver’s reset function
net/wireless/core.cLinux kernel Default retry limits in the Linux wireless stack: 7 for short frames, 4 for long frames (dot11ShortRetryLimit, dot11LongRetryLimit)
Netfilter Conntrack Sysfs variablesLinux kernel The maximum number of entries in Linux’s connection tracking table (nf_conntrack_max) and the default timeouts for each connection state
PSI - Pressure Stall InformationLinux kernel some (share of time some tasks stalled waiting for memory) and full (share of time all tasks stalled) in /proc/pressure/memory
Runtime locking correctness validatorLinux kernel Taking two locks in opposite orders causes circular waiting and a deadlock (lock inversion deadlock); the Linux kernel checks lock ordering and warns in advance
Scaling in the Linux Networking StackLinux kernel RSS (the NIC spreads packets across multiple receive queues) and RPS (the kernel spreads them), giving each queue its own interrupt and spreading those across cores; RSS is recommended when receive interrupt handling is the bottleneck
SNMP counterLinux kernel TcpExtListenOverflows: number of connection requests (SYN) dropped because the accept queue was full; TcpExtListenDrops goes up along with it
Spectre Side ChannelsLinux kernel As a mitigation, branch prediction buffers are flushed on context switches and VM transitions, and stronger mitigations add overhead to every program
tcp: make the first N SYN RTO backoffs linearLinux kernel Commit that made the first SYN retransmissions use a fixed interval, from Linux 6.5 (the default of 4 follows the macOS and iOS behavior)
tcp: use RACK to detect lossesLinux kernel Introduction of RACK and tcp_recovery (Linux 4.4), initially working as a supplement to the existing method
The /proc FilesystemLinux kernel The OOM killer picks the process to kill by a score (badness) based on its share of memory use, adjustable with oom_score_adj
The kernel’s command-line parametersLinux kernel mitigations=: off disables all CPU vulnerability mitigations for more performance but leaves the system exposed, the default auto mitigates with SMT on, auto,nosmt turns SMT off when needed
Thin-streams and TCPLinux kernel TCP_THIN_LINEAR_TIMEOUTS can turn off exponential backoff for thin stream connections only
Transparent Hugepage SupportLinux kernel With defrag=always, a failed THP allocation reclaims and compacts memory on the spot and stalls; with madvise, only regions that asked for it do
What is NUMA?Linux kernel Memory in the same cell is faster and has higher bandwidth; memory in another (remote) cell is slower to access
IETF 59
RFC 1191: Path MTU discoveryIETF Path MTU discovery: a packet that is too large triggers ICMP “fragmentation needed and DF set” (type 3 code 4)
RFC 1812: Requirements for IP Version 4 RoutersIETF Routers must be able to rate-limit ICMP error messages such as Time Exceeded, and may also limit Echo Replies (read mtr and ping results with care)
RFC 2863: The Interfaces Group MIBIETF ifOutDiscards: number of packets discarded even though no error occurred, for reasons such as freeing buffer space
RFC 2923: TCP Problems with Path MTU DiscoveryIETF If a firewall blocks ICMP (Fragmentation Needed), path MTU discovery fails and only large packets keep vanishing (black hole); pings and small messages still work, which makes it hard to diagnose
RFC 3449: TCP Performance Implications of Network Path AsymmetryIETF On asymmetric links with a narrow upload, delayed or lost ACKs hurt TCP performance; ACKs are cumulative, so a later ACK covers for lost ones; remedies such as ACK-prioritizing scheduling
RFC 5382: NAT Behavioral Requirements for TCPIETF Recommends a TCP NAT established-connection idle timeout of at least 2 hours 4 minutes (on the premise that devices may delete idle sessions earlier)
RFC 7871: Client Subnet in DNS QueriesIETF DNS that answers differently by location guesses location from the address of the resolver sending the query, and gives inappropriate answers when the player uses a central resolver far away. EDNS Client Subnet (an optional feature) passes along part of the player’s address
RFC 7999: BLACKHOLE CommunityIETF The BLACKHOLE community, announced over BGP to ask neighboring networks to drop traffic headed to a specific address
RFC 8085: UDP Usage GuidelinesIETF Losing one fragment means the packet can’t be reassembled and is lost entirely; UDP apps should avoid IP fragmentation
RFC 8325: Mapping Diffserv to IEEE 802.11IETF 802.11 CSMA/CA: sends only when the channel is clear; if it’s busy, defers until it clears, then waits an additional random backoff
RFC 8952: Captive Portal ArchitectureIETF Captive portal: a network that restricts access until requirements such as accepting terms or authenticating are met
RFC 9000: QUIC: A UDP-Based Multiplexed and Secure TransportIETF Connection IDs keep a connection alive even when the IP address or port changes (Section 9); a load balancer that distributes by address and port alone may send packets from a changed address to a different server (Section 5.2.3)
Amazon CloudWatch metrics for Amazon EBSAWS VolumeQueueLength (requests waiting to complete), VolumeAvgWriteLatency (1-minute average write latency, Nitro instances)
Amazon CloudWatch metrics for Amazon EC2 Auto ScalingAWS Group metrics are published every minute only when enabled; GroupDesiredCapacity (the count the group tries to maintain), GroupPendingInstances (instances not yet in service), GroupInServiceInstances (instances in service)
Amazon EBS-optimized instance typesAWS Some instances sustain maximum EBS performance for only 30 minutes once every 24 hours, then return to baseline
Amazon EC2 Auto Scaling lifecycle hooksAWS During scale-out and scale-in, instances are held in a wait state to finish setup and cleanup (up to 1 hour by default)
Amazon EC2 instance network bandwidthAWS The “up to N Gbps” of instances with 16 vCPUs or fewer is a burst that spends network I/O credits (usually 5–60 minutes), falling back to baseline bandwidth when the credits run out
Amazon EC2 security group connection trackingAWS Once an instance exceeds the number of connections it can track, packets for new connections are dropped; idle connections can exhaust the tracking table
Amazon GameLift Servers UDP ping beaconsAWS The game client measures latency to a UDP endpoint at each hosting location and uses it for placement and matchmaking; closer to real game traffic than ICMP ping
Avoiding insurmountable queue backlogsAWS Amazon Builders’ Library. Monitor backlog by the age of waiting messages; real-time systems process the newest data first (closer to LIFO) and may drop old messages
CloudWatch metrics for your Application Load BalancerAWS TargetResponseTime (time from the request leaving the load balancer until the target starts responding), HTTPCode_Target_5XX_Count (5xx responses generated by targets), UnHealthyHostCount (number of unhealthy targets)
Create a player latency policyAWS Places the game at the location with the lowest average latency across all players, but players with extreme latency get placed too; example policy widens the ping cap from 50 ms to 100 ms and then 200 ms
Edit attributes for your Application Load BalancerAWS ALB idle timeout default 60 seconds (1–4,000 seconds); the load balancer closes the connection if the client or target connection is silent for that long
Exponential Backoff And JitterAWS Exponential backoff alone still bunches retries together; adding randomness (jitter) reduces contention (the simulation’s retry method)
Failing over a Multi-AZ DB instance for Amazon RDSAWS Multi-AZ failover usually takes 60–120 seconds; connections must be reestablished afterward, and a JVM DNS cache TTL of 60 seconds or less is recommended
FlexMatch rule typesAWS The latency rule (maxLatency) looks at player latency per location; for parties it uses the members’ average by default (partyAggregation avg); queues can place games in regions that don’t meet the latency rule
Flow log recordsAWS srcaddr in a VPC flow log record: for inbound traffic, the sender’s IP address
Health checks for Network Load Balancer target groupsAWS Default health check every 30 seconds, target removed after 2 failures; UDP services are checked with TCP or HTTP health checks, so configuring them to reflect the real service state is recommended
High availability for Amazon AuroraAWS Reads and writes fail during the outage, and recovery usually takes under 60 seconds (often under 30)
How Amazon Route 53 uses EDNS0 to estimate the location of a userAWS If the resolver doesn’t support edns-client-subnet, the player’s location is guessed from the resolver’s address and the answer is based on the resolver’s location (same for geolocation and latency-based routing)
Infrastructure layer attacksAWS Volumetric attacks such as UDP reflection and SYN floods overwhelm network capacity or tie up firewall and load balancer resources
Initialize Amazon EBS volumesAWS Volumes created from snapshots have higher latency and lower performance while blocks are fetched from S3; initialize them in advance by reading every block with dd or fio
NAT gateway basicsAWS 55,000 concurrent connections per IPv4 address to the same destination (destination IP, port, protocol), expandable by attaching up to 8 IPs (public NAT gateways get 2 Elastic IPs by default, more through a quota increase request); bandwidth scales automatically from 5 to 100 Gbps and throughput from 1 million to 10 million packets per second, and packets beyond that limit are dropped
NAT gateway metrics and dimensionsAWS ErrorPortAllocation: number of times a source port couldn’t be allocated (above 0 means too many concurrent connections), ActiveConnectionCount, IdleTimeoutCount (connections cleaned up after 350 seconds idle), PacketsDropCount
Network Load BalancersAWS NLB TCP idle default 350 seconds (60–6,000 seconds); after that it only stops tracking and answers later data with RST; UDP flows are fixed at 120 seconds
Processor state control for Amazon EC2 Linux instancesAWS Only some instance types let the OS control C-states and P-states, which can be changed to reduce latency; the default settings give maximum performance and suit most workloads; Graviton runs at a fixed frequency, so the OS doesn’t control it
Renewal for domains validated by DNSAWS 45 days before expiry, checks whether the certificate is in use by an AWS service and whether the validation CNAME record exists, then renews automatically; if it can’t validate, sends notices 30, 15, 7, 3, and 1 days before expiry
Scheduled events for Amazon EC2 instancesAWS Scheduled event types (system-reboot reboots and moves to a new host; system-maintenance means brief impact from network or power maintenance), notified by email and AWS Health, checked with describe-instance-status, reschedulable for some types
Timeouts, retries, and backoff with jitterAWS Amazon Builders’ Library. Add jitter to all timers, periodic jobs, and delayed work to spread out load that would otherwise land at the same moment; a case where one-minute requests from many servers piled into the first few seconds of every minute
Troubleshoot NAT gatewaysAWS Connections expire after 350 seconds idle and later sends get an RST; keepalives shorter than 350 seconds recommended; when hitting the connection limit, add gateways per availability zone, add IPs, or reduce connections
Working with DB instance read replicasAWS The architecture simulation’s replication lag: read replicas are updated asynchronously, so reads can return stale data
MySQL 32
Configuring Buffer Pool FlushingMySQL When the redo log fills, a sharp checkpoint briefly drops throughput; adaptive flushing spreads the writes out evenly
Deadlock DetectionMySQL At very high concurrency, detection itself can slow things down, so it is sometimes turned off in favor of the lock wait timeout
EXPLAIN Output FormatMySQL type ALL means a full table scan, usually avoided by adding an index
General Thread StatesMySQL Waiting for table metadata lock: thread state while waiting for a metadata lock
How MySQL Uses IndexesMySQL Without an index, the whole table is read starting from the first row, and the bigger the table, the higher the cost
How to Minimize and Handle DeadlocksMySQL Recommends keeping transactions small and short and committing right after related changes to reduce conflicts
InnoDB LockingMySQL When one transaction locks a row (index record), other transactions can’t modify that row and wait
InnoDB Multi-VersioningMySQL While a transaction that can see old versions remains, update undo logs can’t be discarded and the rollback segment grows; recommends committing often, even for read-only transactions
InnoDB Standard Monitor and Lock Monitor OutputMySQL LATEST DETECTED DEADLOCK: the two transactions in the most recent deadlock, the locks they held and waited for, and which one was rolled back
InnoDB Startup Options and System VariablesMySQL With detection on (the default), InnoDB detects deadlocks immediately and rolls one back; innodb_lock_wait_timeout defaults to 50 seconds
Online DDL Performance and ConcurrencyMySQL Even online DDL briefly needs an exclusive metadata lock to finish; it waits if there is a long transaction, and the waiting lock request blocks every transaction behind it
Purge ConfigurationMySQL Purge cleans up the list of undo logs from committed transactions (history list); the backlog shows as History list length in the TRANSACTIONS section of SHOW ENGINE INNODB STATUS
Replica Server Options and VariablesMySQL replica_parallel_workers lets multiple threads apply transactions in parallel (default 4; 0 means a single thread applies them in order)
Saving and Restoring the Buffer Pool StateMySQL Saves the list of recently used pages (25% by default) at shutdown and reads them back at startup to shorten warm-up after a restart; both are on by default
Semisynchronous ReplicationMySQL With asynchronous replication, committed transactions may be missing from the replica when the primary dies; semi-synchronous replication narrows this by waiting for one replica to acknowledge receipt, at the cost of higher latency
Server Status VariablesMySQL Innodb_row_lock_waits and Innodb_row_lock_time give the count and total time of row lock waits; Innodb_row_lock_current_waits gives how many are waiting right now
Server System VariablesMySQL lock_wait_timeout: metadata lock wait limit, default 31,536,000 seconds (1 year)
SHOW REPLICA STATUS StatementMySQL Seconds_Behind_Source: time elapsed since the event the replica is currently applying was written on the primary (replication lag)
Statement Summary TablesMySQL events_statements_summary_by_digest: SUM_NO_INDEX_USED (times executed without an index) and SUM_ROWS_EXAMINED per normalized query
AI.NavMesh.pathfindingIterationsPerFrameUnity Pathfinding processes only a set number of nodes per frame, spreading the work over several frames, so the game stays smooth even with long paths or many requests at once
Application.runInBackgroundUnity Defaults to false, so the game pauses in the background; Android pauses in the background regardless of the setting, and iOS ignores it
Authority (Netcode for GameObjects 2.5)Unity In a distributed authority model, each game instance (client) takes authority over some network objects and simulates them
Entity struct (Entities 1.3)Unity An Entity consists of an Index and a generation number (Version), which tells whether a reused Index is still valid
Garbage collection modesUnity Incremental GC is the default and collects across several frames; with it off, the main thread stops while the whole heap is scanned, which can take up to hundreds of ms
Handling variation in timeUnity Maximum Allowed Timestep defaults to 1/3 s (0.3333333): even if the game stops for 1 s, game time advances only 0.333 s. The cap prevents the vicious cycle where catch-up steps slow things down even more
Introduction to level of detailUnity Without LOD, even objects that look tiny on screen are drawn at full complexity; LOD cuts the rendering cost
Introduction to prediction (Netcode for Entities 6.5)Unity Client and server predict with the same simulation code; when the result differs from the server state (misprediction), the client rolls back and resimulates, and the correction becomes visible
Physics (Netcode for Entities 6.5)Unity Lag compensation: the server finds the collision world the client was seeing at that tick and decides whether the shot hit
Profiler markers referenceUnity GC.Collect: time program code is paused during garbage collection (under 1 ms to hundreds of ms); GC.Alloc: managed heap allocations
Shader loadingUnity The first time a shader variant is used, the graphics driver has to build it for the GPU, which can cause a noticeable freeze; once built, it’s cached and doesn’t freeze again
Texture and mesh loadingUnity Synchronous upload reads and uploads in one frame on the main thread and causes a visible freeze; asynchronous upload streams over several frames
Asynchronous Commit (PostgreSQL Documentation)PostgreSQL Batching writes and flushing them late raises throughput, but the most recent transactions can be lost in a failure (the same tradeoff)
CREATE INDEX (PostgreSQL Documentation)PostgreSQL Building with CONCURRENTLY creates the index without blocking writes; a regular build blocks writes until it finishes
Number Of Database ConnectionsPostgreSQL Once DB resources are fully used, adding connections actually lowers throughput; matching active connections to resources and queueing the rest gives better latency and throughput
Reliability (PostgreSQL Documentation)PostgreSQL Ordinary SATA disks and many SSDs have write caches that are lost on power failure; durable writes need a cache with battery or power-loss protection
Routine Vacuuming (PostgreSQL Documentation)PostgreSQL Old row versions can’t be removed while other transactions can still see them; long-open transactions must be ended or their sessions terminated
WAL Configuration (PostgreSQL Documentation)PostgreSQL By default, a checkpoint runs every 5 minutes or every 1 GB of WAL (max_wal_size) and is expensive because it writes all dirty pages. checkpoint_completion_target spreads the writes out to avoid I/O bursts. If checkpoints come closer together than checkpoint_warning, the log suggests raising max_wal_size
A Multifaceted Look at Starlink Performance (WWW 2024)ACM Starlink reassigns routes every 15 seconds at the same moment worldwide; at these boundaries latency and throughput fluctuate and brief outages under 1 second occur (unrelated to switching between satellites); terminal↔satellite↔ground station latency about 40 ms
Capacity of Ad Hoc Wireless Networks (MobiCom 2001)ACM When 802.11 relays over several wireless hops, a node can’t send while it’s receiving and adjacent hops interfere with each other, so the throughput of a chain of relays can drop to 1/3 in theory (about 1/7 in simulation)
Data Center TCP (DCTCP) (SIGCOMM 2010)ACM Commodity switches have shallow buffers (48 ports share 4 MB, and one port can use up to about 700 KB); loss occurs when many flows converge on one port for a brief moment
The Internet at the Speed of Light (HotNets 2014)ACM Actual router paths are about 1.5 times the straight fiber distance (median), and packets between two nearby points sometimes travel around the far side of the globe (hairpinning)
The QUIC Transport Protocol: Design and Internet-Scale Deployment (SIGCOMM 2017)ACM 2016: 4.4% of clients couldn’t use QUIC over UDP because UDP/QUIC was blocked or the path MTU was small (mostly behind corporate firewalls; no ISP-wide blocking observed); 0.3% were on networks that appeared to throttle UDP (higher loss at peak hours; down from 1% in 2015 after requests to ISPs); a case where a firewall let only the first few packets through after a 1-bit header change and blocked the rest, defeating the TCP fallback logic
AActor::SetLifeSpanEpic Games Setting a lifespan on an actor destroys it automatically when it expires
Actor Priority in Unreal EngineEpic Games When bandwidth runs short, not every actor is replicated every time; actors are prioritized by distance from the viewer and time since their last replication
Actor Relevancy in Unreal EngineEpic Games The server replicates only relevant actors to each connection and doesn’t send irrelevant ones
Actor Ticking in Unreal EngineEpic Games Actors and components tick once every frame unless given their own interval, and ticking can be turned off when not needed
Detailed Actor Replication Flow in Unreal EngineEpic Games NetUpdateFrequency sets the update rate per actor; actors are sent in priority order, and once the connection is saturated, the rest wait for the next tick
Introduction to Iris in Unreal EngineEpic Games Holds the state to replicate as a single quantized copy to cut expensive work, and shares that work across connections
Networking Insights in Unreal EngineEpic Games Shows the size of packets sent and received per connection and the replicated objects and properties inside them
Networking Overview for Unreal EngineEpic Games A listen server’s host has an advantage over other clients and carries a heavy load, running both the server and rendering
PSO Precaching for Unreal EngineEpic Games With r.PSOPrecache.Validation on, stat PSOPrecache shows missed PSO statistics and the log records “PSO PRECACHING MISS”; runtime PSO creation over 20 ms (default) counts as a hitch
Replication Graph in Unreal EngineEpic Games Games with many players and replicated objects (such as MMORPGs) need to group them by location and send only what each player needs, or the server CPU becomes the bottleneck
Stat Commands in Unreal EngineEpic Games stat GC (garbage collection statistics), stat Hitches (logs frames that exceed t.HitchFrameTimeThreshold)
Texture Streaming Overview for Unreal EngineEpic Games The streamer raises and lowers texture resolution (mips) to match the camera view, does most of the work on async worker threads, and loads the mips visible on screen first
Using Gameplay Abilities in Unreal EngineEpic Games Local Predicted runs immediately on press and the server makes the final decision; Server Initiated has no prediction, so the player sees the delay
Using Network Emulation in Unreal EngineEpic Games Test with minimum and maximum latency and a packet loss percentage on server and client; set from the console, e.g., NetEmulation.PktLag
Using the Anti-Cheat InterfacesEpic Games If the server doesn’t receive the client’s anti-cheat message within the set time (RegisterTimeout), it kicks the client for an authentication timeout (a client frozen by loading is a common cause); if a recent module update is the problem, roll back to the previous module
Android (Google) 17
Android common kernelsAndroid (Google) Common kernels 5.10 through 6.18 are supported side by side, and a kernel for an earlier platform (e.g., android14-6.1) can be used to launch or upgrade new Android devices
ApplicationExitInfoAndroid (Google) REASON_LOW_MEMORY: the system’s low memory killer terminated the app process (devices that don’t support it report REASON_SIGNALED with SIGKILL)
Cached apps freezerAndroid (Google) Android 14 and later freezes cached app processes after 10 s; once frozen, all threads stop
CrashesAndroid (Google) A crash is when an app exits unexpectedly because of an unhandled exception or signal (SIGSEGV and so on); tallied in Android vitals in Play Console
Frame Pacing libraryAndroid (Google) On a 60 Hz screen, the previous frame is shown again when there’s no new frame; example of a 30 FPS game whose frame times become uneven, such as 49, 16, and 33 ms
Memory allocation among processesAndroid (Google) Android holds out by compressing memory into zRAM; when that’s not enough, the low memory killer terminates processes, and a foreground app being killed looks like a crash
Network security configurationAndroid (Google) With certificate pinning, you must include backup keys to prepare for key rotation or CA changes; otherwise connections break until the app is updated
Optimize network accessAndroid (Google) Radio state transition delay and tail time vary with the radio technology (3G, LTE, 5G) and carrier settings; 3G example: low power → full power about 1.5 s, idle → full power over 2 s
Read network stateAndroid (Google) When the default network changes, new connections go over the new network and connections on the old network are eventually forced closed; registerDefaultNetworkCallback detects the switch
Security with network protocolsAndroid (Google) If the server omits the intermediate certificate, Android apps fail with SSLHandshakeException, while desktop browsers may fill it in from cached intermediates and show no error; check the chain the server sends with openssl s_client
Slow renderingAndroid (Google) To hit 60 FPS, a frame must render within 16 ms; late frames get skipped and show up as stutter (jank)
Slow Sessions (games only)Android (Google) Android vitals counts a game frame as slow when it takes longer than 50 ms (20 FPS) or 34 ms (30 FPS)
TelephonyDisplayInfoAndroid (Google) OVERRIDE_NETWORK_TYPE_NR_NSA: the network indicator shown when the device is on LTE and can use, or is using, dual connectivity (EN-DC) with 5G (NR)
Thermal APIAndroid (Google) Devices can sustain high performance only for a limited time before heat forces throttling; recommends watching thermal status and lowering the load ahead of time
Wi-Fi low-latency modeAndroid (Google) Low-latency mode turns off Wi-Fi power saving; how scanning and roaming settings are optimized depends on the device maker’s implementation
WifiManagerAndroid (Google) WIFI_MODE_FULL_LOW_LATENCY (API 29, Android 10): a low-latency Wi-Fi lock that applies only while connected to an AP, with the screen on and the app in the foreground
Window.setPreferMinimalPostProcessingAndroid (Google) Latency-sensitive windows such as games request minimal video processing from the display; over HDMI, the ALLM and Game Content Type signals switch the TV to low-latency mode
clock_gettime(2) — Linux manual pageLinux man-pages CLOCK_MONOTONIC is not affected by discontinuous jumps in the system clock and never goes backward
connect(2) — Linux manual pageLinux man-pages EADDRNOTAVAIL: every port in the ephemeral port range is in use, so the connection can’t be opened
core(5) — Linux manual pageLinux man-pages RLIMIT_CORE caps core file size, coredump_filter selects which memory regions to include, and core dumps can be piped to a program for separate handling
fsync(2) — Linux manual pageLinux man-pages fsync flushes modified data all the way to the disk (including the disk cache) and blocks until the device reports completion
getrlimit(2) — Linux manual pageLinux man-pages RLIMIT_NOFILE: the limit on how many fds a process can open; going over it returns EMFILE
listen(2) — Linux manual pageLinux man-pages A listen backlog larger than somaxconn is silently truncated; somaxconn defaults to 4,096 (since 5.4; 128 before); when the queue is full, requests may be ignored and left to client retries
mallopt(3) — Linux manual pageLinux man-pages To reduce thread contention, glibc malloc creates arenas up to a multiple of the CPU count, and more arenas mean more memory use (limit with M_ARENA_MAX, or set the MALLOC_ARENA_MAX environment variable)
send(2) — Linux manual pageLinux man-pages send() blocks when the send buffer has no room, and returns immediately with EAGAIN in non-blocking mode
socket(7) — Linux manual pageLinux man-pages SO_RCVBUF is the maximum socket receive buffer size; the default comes from rmem_default and the maximum from rmem_max (Android also runs the Linux kernel)
tcp(7) — Linux manual pageLinux man-pages 9 probes 75 s apart after 7,200 s of idle time (about 11 more minutes), applies only to sockets with SO_KEEPALIVE on, TCP_KEEPIDLE and TCP_USER_TIMEOUT
Azure network round-trip latency statisticsMicrosoft Azure Measured median round-trip times from Seoul (Korea Central): Tokyo 30 ms, Singapore 68 ms, US West 124–136 ms, Europe 234–244 ms
Bulkhead PatternMicrosoft Azure With separate connection and thread pools for each called service, a failure in one service blocks only its own pool
Chatty I/O antipatternMicrosoft Azure Many small I/O requests add up to latency that badly hurts responsiveness; recommends fewer, larger requests
Circuit Breaker PatternMicrosoft Azure Requests blocked until their timeout hold threads and DB connections and make unrelated features fail; once failures pile up within a set time, calls are rejected immediately
Configure load balancer TCP reset and idle timeoutMicrosoft Azure Azure Load Balancer idle timeout default 4 minutes (4–100 minutes), no guarantee the session is kept beyond that, TCP reset is optional
Disk metricsMicrosoft Azure Disk and VM burst credit usage (5-minute intervals), such as Data Disk Used Burst IO Credits Percentage
Maintenance and updatesMicrosoft Azure Maintenance without a reboot almost always pauses for under 10 seconds, rarely (no more than once every 18 months for general-purpose sizes) for about 30 seconds, and live migration usually 5 seconds or less; the clock syncs automatically after the pause; long-lived TCP connections may drop, or recovery may take longer as peers retransmit data sent to the paused VM with exponential backoff; load balancer health checks mark the VM unhealthy within about 10 seconds; confirm with Microsoft.Compute/virtualMachines/liveMigration/action in the Activity Log and VmAvailabilityMetric dropping to 0 during the pause; pick when maintenance applies with Maintenance Configuration
Managed disk burstingMicrosoft Azure Premium SSD P20 and smaller use credit-based bursting; a full credit balance gives 30 minutes at max burst speed
Metrics and alerts for Azure NAT GatewayMicrosoft Azure SNAT Connection Count filtered by the Failed state above 0 suggests SNAT port exhaustion; Dropped Packets
Scheduled Events for Linux VMs in AzureMicrosoft Azure Freeze (a pause of a few seconds; CPU and network may stop) is announced at least 15 minutes ahead; for host hardware failures, recovery starts right away with no notice period
Source Network Address Translation (SNAT) with Azure NAT GatewayMicrosoft Azure 64,512 SNAT ports per public IP (up to 16 IPs); each connection to the same destination needs a different port; closed ports go through a cooldown before reuse for the same destination
Oracle 9
Available CollectorsOracle ZGC gives up a little throughput to keep maximum pauses under 1 ms, independent of heap size
Garbage Collector ImplementationOracle Minor GC when the young generation fills; some surviving objects move to the old generation, and when the old generation fills, the whole heap is collected (much slower than a minor GC); -Xlog:gc writes one line per GC
Garbage-First (G1) Garbage CollectorOracle G1 default pause target 200 ms (MaxGCPauseMillis); if memory runs out during collection, it falls back to a full GC that stops and compacts the whole heap
Garbage-First Garbage Collector TuningOracle By default (GCTimeRatio=12), G1 sizes the heap to keep GC time at about 8% of total time or less; full GCs caused by excessive heap occupancy show up in the log as Pause Full (G1 Compaction Pause)
The java CommandOracle Table mapping old GC log options to -Xlog: -XX:+PrintGCDetails becomes -Xlog:gc*
The Parallel CollectorOracle Parallel GC throws OutOfMemoryError if it spends more than 98% of total time in GC and recovers less than 2% of the heap
The Z Garbage CollectorOracle ZGC does its expensive work concurrently and doesn’t pause for more than 1 ms, but if reclamation falls behind, the application can stall waiting for GC (the GC simulation’s concurrent mode)
Troubleshoot Memory LeaksOracle Suspect a leak when execution gets gradually slower; memory eventually runs out and the program terminates abnormally. The key data for leak analysis is a heap dump
Redis 9
Diagnosing latency issuesRedis One thread processes requests in turn, so a slow command blocks everything behind it; use SCAN in place of KEYS; fork measured at about 9–13 ms per GB on physical servers and modern VMs; THP causes latency and memory spikes from copying after fork; mass expiry in the same second causes stalls
INFORedis keyspace_hits and keyspace_misses (successful and failed key lookups), expired_keys (number of expired keys), uptime_in_seconds (time since startup)
KEYSRedis Use with extreme care in production; can ruin performance on large databases (40 ms for 1 million keys on an entry-level laptop)
Redis CLIRedis --bigkeys: scans the keyspace to find large keys
Redis latency monitoringRedis latency-monitor-threshold defaults to 0 (off); LATENCY LATEST and LATENCY DOCTOR; records latency per event such as fork and expire-cycle
Redis persistenceRedis Taking RDB snapshots every few minutes means accepting the loss of the last few minutes of data on an abnormal shutdown
SLOWLOGRedis Slow command log that records commands exceeding slowlog-log-slower-than; execution time excludes I/O with the client
UNLINKRedis Asynchronous deletion: unlinks the key immediately and reclaims its memory in another thread
Cloudflare 1.1.1.1 Incident on July 14, 2025Cloudflare When a public DNS resolver stopped for 62 minutes, users who could no longer resolve names effectively lost access to all internet services
How "expensive" is crypto anyway?Cloudflare BoringSSL measurements: AES-128-GCM about 3.7 GB per second (varies widely with record size); per core per second, 1,120 RSA 2048 signatures, 18,477 ECDSA P-256 signatures, and 9,394 P-256 ECDHE operations; on Cloudflare edge servers, the TLS library used about 1.8% of CPU
How to receive a million packets per secondCloudflare Measurements where a receive queue served by a single core topped out at about 350,000–430,000 packets per second; a case where the NIC hashed UDP by IP address only and everything piled into one queue
Maximum transmission unit and maximum segment sizeCloudflare Inbound traffic is delivered after filtering over a GRE tunnel (MTU 1,476) while outbound responses go straight to the internet (DSR); limiting TCP MSS to 1,436 or less is recommended, and without it large packets are dropped or fragmented
Q1 2024 Internet disruption summaryCloudflare West African cable cuts (March 14) were repaired 3–6 weeks later, with traffic moved to other cables in the meantime
Q2 2024 Internet disruption summaryCloudflare Red Sea cables damaged in February 2024 were still under repair in July (conflict zone); the EASSy and Seacom cuts in May were repaired in 19 days
Why does one NGINX worker take all the load?Cloudflare SO_REUSEPORT splits queues per worker with a simple hash, so if one worker is blocked, every connection queued to it stalls
Gaffer On Games 7
Deterministic LockstepGaffer On Games Frame n can be computed only after all its inputs arrive, so a late input means waiting; a small playout delay buffer for absorbing jitter causes hitches
Floating Point DeterminismGaffer On Games The same floating-point code can give different results depending on compiler, CPU architecture, and debug or release build; includes a case where AMD and Intel CPUs returned slightly different values from transcendental functions
Snapshot CompressionGaffer On Games Changes must be built only against a baseline the other side has acknowledged (acked), and the initial state is sent separately
Snapshot InterpolationGaffer On Games Drawing received snapshots immediately stutters because of jitter; holding them briefly in an interpolation buffer before drawing makes motion smooth
State SynchronizationGaffer On Games Sending state along with inputs keeps both sides in sync without perfect determinism
UDP vs. TCPGaffer On Games UDP doesn’t guarantee delivery or order, so the application must detect and resend lost packets itself
IP addresses and portsGoogle Cloud 64,512 ports each for TCP and UDP per NAT IP; default minimum ports per VM is 64 (static allocation) or 32 (dynamic allocation); the number of ports reserved for a VM caps its concurrent connections to the same destination; ports of closed connections can’t be used during TIME_WAIT
Live migration process during maintenance eventsGoogle Cloud Live migration pauses are usually much shorter than 1 second; the system clock jumps forward by up to 5 seconds during the pause; disk, CPU, memory, and network performance drop briefly during the move; VMs that don’t live-migrate are terminated for maintenance (bare metal instances don’t support live migration)
Logs and metricsGoogle Cloud dropped_sent_packets_count with reason OUT_OF_RESOURCES: packets dropped for lack of NAT IPs or ports
MTU considerations | Cloud VPNGoogle Cloud Cloud VPN gateway MTU is 1,460 bytes, and the payload MTU of an IPv4 tunnel is 1,406 bytes (around 1,400 through a tunnel)
Query metadata server for maintenance event noticesGoogle Cloud The maintenance-event metadata value changes 60 seconds before live migration (when the VM is set to live-migrate and the value was queried at least once since the last maintenance)
ss(8) — Linux manual pageiproute2 In -o, timer:(on,…) is the retransmission timer; in -i, backoff is the number of times the retransmission wait has doubled
tc-cake(8) — Linux manual pageiproute2 CAKE separates flows and minimizes latency for flows that send sparsely (sparse flows)
tc-fq(8) — Linux manual pageiproute2 The fq qdisc paces each socket (connection), and SO_MAX_PACING_RATE sets a per-connection maximum rate
tc-netem(8) — Linux manual pageiproute2 A test tool that adds delay and jitter (delay TIME JITTER) and loss (loss random PERCENT) to outgoing packets to mimic a real network
Apple 6
Extending your app’s background execution timeApple When the app goes to the background, applicationDidEnterBackground gets 5 s before the app is suspended; if it needs more, it requests time with beginBackgroundTask (time left in backgroundTimeRemaining)
Recommended settings for Wi-Fi routers and access pointsApple Other routers and devices on the same channel are sources of interference; 20 MHz channel width is recommended on 2.4 GHz; interference is less of a concern on 5 GHz and 6 GHz
thermalStateApple The current thermal level reported by iOS; as the level rises, the app should reduce its resource use
Wi-Fi roaming support in Apple devicesApple When a device moves to a new AP, it can’t send data until authentication with the new AP completes, which can take a few seconds with 802.1X
Bufferbloat.net 6
CakeBufferbloat.net CAKE: router SQM that combines a shaper with fq_codel-style queue management
IntroductionBufferbloat.net When network equipment such as a router buffers too much data, latency spikes sharply (bufferbloat)
Setting up SQM for CeroWrt 3.10Bufferbloat.net Set SQM to 95% of the measured speed (85% if based on the advertised speed) to move the bottleneck from the ISP’s equipment into the router; that’s what makes it work
Tests for BufferbloatBufferbloat.net If ping rises while a speed test saturates the connection with ping running, it’s bufferbloat
What Can I Do About Bufferbloat?Bufferbloat.net Use a router that supports SQM such as cake or fq_codel, and tune the SQM rate while measuring latency under load
Google 6
An Internet-Wide Analysis of Traffic PolicingGoogle Queue overflow raises queuing delay and RTT before the loss, while policing drops the excess with no RTT increase (SIGCOMM 2016)
Load Balancing in the DatacenterGoogle Plain round robin lets CPU usage differ by up to 2× between tasks; weighted distribution where backends report their load in responses and health checks; a lame duck state in which a backend asks not to be sent new requests
The Tail at ScaleGoogle Occasional long delays (tail latency) increasingly dominate how the whole service feels as scale grows
Microsoft SQL Server 6
Deadlocks guideMicrosoft SQL Server Deadlock checks run every 5 seconds by default, dropping to as little as 100 ms when deadlocks are frequent; the system_health session, on by default, collects xml_deadlock_report; the victim gets error 1205
Monitor performance by using the Query StoreMicrosoft SQL Server Plans change with statistics, schema, and index changes, and the plan cache keeps only the latest plan; Query Store’s plan forcing pins a good plan; the Regressed Queries view compares slowed-down queries and their plans
Query Processing Architecture GuideMicrosoft SQL Server Parameter sniffing: the query plan is built for the parameter values passed at compile or recompile time
Transaction Locking and Row Versioning GuideMicrosoft SQL Server Lock escalation when one statement holds 5,000 or more locks on one table (or index); recorded with the lock_escalation extended event
Troubleshoot a full transaction log (SQL Server Error 9002)Microsoft SQL Server When the log is full, the DB is read-only and can’t be modified; missed log backups, replication lag, and long transactions are common causes that block log truncation; see what is blocking it in log_reuse_wait_desc of sys.databases
Wireshark 6
7.5. TCP AnalysisWireshark TCP ZeroWindow: a packet in which the receiver advertises window 0, telling the sender to stop sending
8.7. Packet LengthsWireshark Splits captured packets into length ranges and shows the count, average, minimum, and maximum for each
8.8. The “I/O Graphs” WindowWireshark Graphs the packet count and bytes matching a display filter per time interval
.NET runtime metrics.NET dotnet.monitor.lock_contentions since .NET 9: number of times contention occurred when trying to acquire a monitor lock since the process started
Background garbage collection.NET Background GC applies only to gen2 collections; gen0 and gen1 collections (foreground GC) stop all managed threads
Debug a memory leak in .NET.NET Even with GC, holding references to objects no longer needed is a leak, leading to performance degradation and OutOfMemoryException. Check memory trends and analyze dumps
dotnet-counters diagnostic tool.NET .NET 9 and later show System.Runtime meters (dotnet.gc.pause.time and others); .NET 8 and earlier show the older EventCounters (% Time in GC since last GC and others)
Efficient Querying.NET ORM lazy loading creates the N+1 problem, sending one more query per item and badly hurting performance; recommends loading in one batch (eager loading)
JEP 271: Unified GC LoggingOpenJDK JDK 9 reimplemented GC logging on unified logging (-Xlog); -Xlog:gc writes one line per GC, like the old -XX:+PrintGC
JEP 439: Generational ZGCOpenJDK ZGC pauses are 1 ms or less regardless of heap size; G1 pauses range from a few ms to a few seconds. Risk of allocation stalls when allocation outpaces reclamation
coredumpctl(1) — Linux manual pagesystemd list queries core dumps saved by systemd-coredump, showing crash time, PID, and the signal that caused the crash
systemd.service(5) — Linux manual pagesystemd Restart=on-failure automatically restarts the service after an abnormal exit, a kill by signal (including core dumps), or a watchdog timeout; recommended for long-running services
ITU-T G.114: One-way transmission timeITU Planning value for optical fiber propagation delay: 5 µs/km (about 200,000 km per second, 10 ms round trip per 1,000 km)
iostat(1) — Linux manual pagesysstat -x: w_await (average time per write request, including time waiting in the queue), aqu-sz (average queue length, formerly avgqu-sz)
mpstat(1) — Linux manual pagesysstat %soft: share of CPU time spent handling software interrupts; per core with -P ALL
Source SDK 2013: player.cppValve A per-tick command processing budget that accumulates (up to sv_maxusrcmdprocessticks, 24 ticks) lets bunched-up commands through; a developer comment says stricter restrictions caused stutter even for legitimate players
Steam Overlay (Steamworks Documentation)Valve The Steam overlay automatically hooks into games launched through Steam, and because of how it does that, it can expose memory errors in the game’s use of the rendering API and cause crashes
AMD FSR Frame GenerationAMD Frame generation is recommended at 60 FPS or higher before interpolation (below 30 FPS should be avoided); AMD Radeon Anti-Lag 2 reduces system latency by keeping CPU and GPU work in step
HED-GP Technical Retrospective: What a HED-acheCCP Games Under overload, EVE Online slows game time with Time Dilation down to a floor of 10% (10 times slower); node CPU normally stays below 80%
Introducing Time Dilation (TiDi)CCP Games A design that slows the game clock when the server is overloaded (Time Dilation) so everything runs slower
Time Dilation – How’s That Going?CCP Games EVE Online’s Time Dilation works per node, so even distant star systems on the same node slow down; big battles run on reinforced nodes that host only 4 star systems
chrony 3
chrony – Frequently Asked Questionschrony Recommended: allow steps only a few times right after startup, as in makestep 1 3; a VM that was paused and resumed can wake up with the wrong time
chrony.conf(5)chrony logchange: clock adjustments larger than this value (default 1 s) are written to syslog
chronyc(1)chrony chronyc tracking fields: System time (difference between NTP time and the system clock), Last offset (offset estimated at the last update), Ref time (when the last measurement from the time source was applied)
Go 3
A Guide to the Go Garbage CollectorGo Go’s GC runs mostly concurrently with only short stop-the-world pauses; under heavy allocation, goroutines take on GC work (assist), which adds latency
runtime packageGo GODEBUG=gctrace=1: one line per GC with wall-clock time per phase, heap size at GC start and end, and the heap goal
IEEE 3
Amdahl's Law in the Multicore EraIEEE IEEE Computer 2008 paper (authors’ copy). If the fraction that can’t be parallelized is 1−f, the speedup can never exceed 1/(1−f) no matter how many cores you add (Amdahl’s law)
Liveness, Readiness, and Startup ProbesKubernetes A liveness probe catches a deadlock, where the app is running but can’t make progress, and restarts the container
CUDA C++ Best Practices GuideNVIDIA Graphics memory bandwidth (V100: 898 GB/s) is far higher than PCIe 3.0 x16 (16 GB/s), so it recommends minimizing transfers to and from system memory
perf-trace(1) — Linux manual pageperf -p traces the system calls of a running process; --duration shows only calls that took longer than the given ms
Riot Games 3
Peeking into VALORANT's NetcodeRiot Games When the guesses that fill in late or missing data are wrong, the client drifts from the server and characters jump or slide into place when corrected
Peeking into VALORANT's NetcodeRiot Games The server rewinds to the game state the player saw at the moment of the shot to check the hit; the client sends the simulation time it was seeing
VALORANT's 128-Tick ServersRiot Games A 128-tick server must finish each frame within 7.8125 ms; server frame time is measured per subsystem and the budget is split among them
RIPE NCC 3
BGPlay (RIPEstat Data API)RIPE NCC Shows the BGP routes for an address prefix at the start time, the BGP updates observed during the period, and the ASes on the path
Probe Selection (RIPE Atlas REST API)RIPE NCC Select RIPE Atlas probes by country, region, ASN, or address prefix and run ping and traceroute from them
RIPE Atlas documentationRIPE NCC A public synthetic monitoring tool that runs ping and traceroute from probes around the world
APNIC 2
BGP updates in 2024APNIC Daily average time for unstable routes to settle again: 25–35 seconds (IPv4), 40–50 seconds (IPv6)
IPv6 Performance – RevisitedAPNIC Comparison of IPv6 and IPv4 round-trip times for the same dual-stack users: some access networks handle IPv6 packets completely differently, producing groups within one ISP where IPv6 is 15, 25, or 75 ms slower
Istio 2
Istio Standard MetricsIstio istio_request_duration_milliseconds (distribution of HTTP/gRPC request duration); the reporter label separates the sending (source) and receiving (destination) proxy
Performance and ScalabilityIstio In sidecar mode, a request passes through the sender’s sidecar proxy and then the receiver’s; every added feature lengthens the processing path inside the proxy, and telemetry collection adds queueing time to the next request
Let's Encrypt 2
Decreasing Certificate Lifetimes to 45 DaysLet's Encrypt Default lifetime cut to 64 days in February 2027 and 45 days in February 2028; a fixed 60-day renewal interval will no longer be enough, so renewal at about two-thirds of the lifetime is recommended
FAQLet's Encrypt Default certificate lifetime of 90 days, renewal every 60 days recommended
Geolocation accuracyMaxMind About 99.8% at the country level, about 66% at the US city level (within 50 km); with a VPN it gives the VPN server’s location in place of the end user’s; mobile network IPs are used across wide areas, so fine-grained location is unknown; the database needs continuous updates; corrections can be requested
numastat(8) — Linux manual pagenumactl numa_miss (allocated on a node other than the intended one) and other_node (allocated on this node by a process running on another node) counters; -p shows a process’s memory per node
OpenSSL 2
openssl-s_clientOpenSSL -showcerts: shows the certificates the server sent, in the order it sent them (not a validated chain)
openssl-x509OpenSSL -enddate: prints the certificate’s expiry date (notAfter); -checkend: checks whether it expires within the given number of seconds
SK텔레콤 2
SKT, 고객 선택권과 혜택 강화한 신규 5G 요금제 출시SK텔레콤 Examples of speed control after a 5G plan’s base data runs out: up to 400 kbps, 1 Mbps, 3 Mbps
SKT, 요금제 개편SK텔레콤 Service continues at up to 400 kbps after the included data runs out (Korea’s nationwide “safety net data” program)
Solidigm 2
D3-S4520 SSDSolidigm Server SATA SSD 4 KB random read/write up to 92K/48K IOPS
Solidigm™ D7-P5520 and D7-P5620 Product BriefSolidigm 99.99th percentile latency (four-nines latency) of 130 µs for server NVMe SSDs: basis for a single SSD read taking around 100 µs
Response to Congestion (as of Dec. 11)Square Enix Once the queue passed 17,000 players per logical data center, new entries were refused so the login servers wouldn’t go down (Error 2002); a player disconnected while waiting got tens of seconds to 1 minute from the lobby server to reconnect and resume mid-queue, and went to the back of the line after that
Scaling Memcache at Facebook (NSDI '13)USENIX When a hot key is invalidated, many reads rush to the DB (thundering herd); prevented with leases (only one client refreshes) and by returning stale values; clusters with an empty cache are warmed up separately
util-linux 2
ionice(1) — Linux manual pageutil-linux Jobs in the idle class get disk I/O only when no other program is using the disk
lsblk(8) — Linux manual pageutil-linux -o selects output columns; device topology columns include ROTA (whether the device is rotational)
Amazon Builders' Library 1
Using load shedding to avoid overloadAmazon Builders' Library Load shedding: refusing excess requests early so the server keeps processing the requests it can handle
Apache Software Foundation 1
Asynchronous loggersApache Software Foundation Asynchronous logging absorbs short bursts in a queue, but if output stays slow, the queue fills and logging drops to the speed of the slowest output, or logs are dropped (Discard) depending on policy
GGPO Rollback Networking SDKGGPO The game predicts the opponent’s input and moves ahead; if the real input differs, it recomputes from the point where they diverged up to the present
GNU Project 1
Threads (Debugging with GDB)GNU Project thread apply all runs the same command (bt: print the call stack) on every thread
HDMI Licensing Administrator 1
Auto Low Latency Mode (ALLM)HDMI Licensing Administrator ALLM lets a device switch the display to low-latency mode (often called game mode) automatically; in low-latency mode, the TV stops some video processing to reduce delay
id Software 1
Quake III Arena source: code/server/sv_snapshot.cid Software Delta compression uses the snapshot the client acknowledged as its baseline, and a full snapshot is sent when the baseline gets too old
iputils 1
ping(8) — Linux manual pageiputils -M do sets the DF flag and refuses packets larger than the path MTU; -s sets the data size (default 56 bytes, plus the 8-byte ICMP header)
IRTF 1
RFC 9505: A Survey of Worldwide Censorship TechniquesIRTF Inspection equipment in the network can pick out and block TCP and UDP flows by address, port, and protocol (blocking of UDP endpoints has been observed with QUIC); allowing only approved protocols leads to overblocking, and throttling of specific traffic is also used
jemalloc 1
jemalloc memory allocatorjemalloc A general-purpose malloc implementation that emphasizes fragmentation avoidance and scalable concurrency
Juniper Networks 1
Flow-Based SessionsJuniper Networks The simulation’s corporate firewall: SRX firewall default session timeouts of 1,800 seconds (30 minutes) for TCP and 60 seconds for UDP
Lua.org 1
Lua 5.4 Reference ManualLua.org Incremental mode splits collection into small steps interleaved with execution (large steps make it stop-the-world); a major collection in generational mode is a stop-the-world pass over every object; collectgarbage("count") returns the total memory Lua is using (KB)
ntpd - Network Time Protocol (NTP) daemonNetwork Time Foundation Offsets above the 128 ms step threshold are corrected in one step and smaller ones gradually, at 0.5 ms per second, so correcting 1 second takes 2,000 s (about 33 minutes)
OpenWrt 1
SQM (Smart Queue Management)OpenWrt Enter 90% of the measured download and upload speeds; cake is the recommended queue discipline (fq_codel if the CPU is weak)
procps-ng 1
vmstat(8) — Linux manual pageprocps-ng The cs (context switches per second) and r (processes running or waiting to run) fields
Red Hat 1
Chapter 2. Getting started with TuneDRed Hat The latency-performance profile turns off power-saving features, sets the governor to performance, and uses PM QoS to allow only shallow C-states; check the current profile with tuned-adm active
Seagate 1
Exos X18 Data SheetSeagate 4K random reads on a 7,200 rpm server HDD: 170 IOPS (QD16)
Starlink 1
Improving Starlink’s LatencyStarlink US peak-hour median 48.5 ms→33 ms, slowest 1% (p99) over 150 ms→under 65 ms (2024); one satellite hop 1.8–3.6 ms; routing over laser links adds latency, and so does the distance from the ground station to the internet access point (PoP)
VLDB Endowment 1
Optimal Probabilistic Cache Stampede PreventionVLDB Endowment When a popular item expires, many requests regenerate it at once (cache stampede); prevented by probabilistically refreshing early, before expiry
과학기술정보통신부 1
5G 통신서비스 품질평가 결과 발표과학기술정보통신부 As announced in 2020, 5G in Korea is offered in NSA mode, with the move to SA still planned