Interactive site: https://jungrok5.github.io/mmo-lag-anatomy/en/ Cause pages: https://jungrok5.github.io/mmo-lag-anatomy/en/c/[cause ID].html # Game Lag White Paper knowledge base Auto-generated (2026-10-03, `node tools/export.cjs`). Edit src/js/ and regenerate; do not edit this file by hand. 228 causes, 135 glossary terms. Causes are referenced by **ID**. Append `#c-ID` to the site URL to go to that cause card. ## Owner codes | Code | Team | Owner | Scope | |---|---|---|---| | cli | Game team | Client development | Game client code: frames, GC, loading, interpolation, extrapolation, prediction, and the client’s network handling (including sending heartbeats and auto-reconnect) | | srv | Game team | Server development | Game server code: ticks, threads, locks, netcode design, connection handling (accept loop, listen arguments), heartbeat replies and dead-connection cleanup, socket options, query and transaction design | | net | Infra team | Network infrastructure | Circuits and data center network gear (switches, routers, firewalls, load balancers, DDoS protection); cloud network ACLs, VPC routing, and load balancers; ISPs and peering | | sys | Infra team | Server infrastructure | Server hardware and cloud instances (including security groups and connection tracking), OS and kernel settings, NICs, deployment and monitoring environments | | dba | Infra team | DB infrastructure | DB servers and storage; DB configuration, replication, and backups; cache servers | | ext | External | External | Players’ PCs and home networks, ISP segments (outside our contracts), cloud providers. We can’t fix these directly, so we respond with player guidance, requests to the provider, or workarounds | For each cause, the “primary owner” is the team that removes the root cause, and “also” lists teams that have real work to do. ## Symptoms - **Stutter** (`stutter`, also called: choppy, hitching, feels like frame drops): Movement isn’t smooth: it keeps pausing briefly and moving again. If ping looks normal, the likely cause is frame timing on your PC (client or OS); if ping jumps around, jitter on Wi-Fi or the connection is more likely. Keep in mind that most in-game ping readouts are measured inside the game loop, which runs once per frame, so a frame spike can make the ping number spike too. - **Teleporting** (`teleport`, also called: warping, skipping, freeze then jump): A character moves to a distant spot in one step, with no movement in between. It usually means packets stopped arriving for a while. Check for packet loss, brief connection drops, server stalls, and extrapolation failures. If one player jumps around while everyone else is fine, suspect that player’s connection first. - **Rubber-banding** (`rubber`, also called: snapping back, getting pulled back, position rollback): Your character runs forward, then gets dragged back to where it just was. Your screen (prediction) and the server’s call disagree. Your input never reached the server (packet loss), the server’s movement validation cut it off, or the two sides compute movement differently. - **Fast-forward** (`burst`, also called: everything at once, speed-up, catch-up): A frozen screen starts moving again, and the backlog of movement, hits, and damage plays out all at once at high speed. Packets piled up somewhere and were released all at once. Typical causes are waiting on a TCP retransmission, the server catching up, and a processing backlog on the client. - **Slow motion** (`slowmo`, also called: world slows down, everything sluggish): Everything moves slowly. Skill casts and monster movement look stretched out. Depending on the server design, speed can stay normal while it shows up as stutter or teleporting. The server can’t finish its ticks on time. The network is fine, so ping measured outside the game stays the same; in-game ping can rise slightly if it includes time waiting for server processing. Check for player surges, AOI calculation, broadcast, and memory shortage. - **Input lag** (`delay`, also called: delayed response, sluggish, mushy controls): It takes a while from pressing a button to seeing the result. The screen itself can still be smooth. Round-trip time (ping) is long, or a queue is building up somewhere. Check distance, router queues, Nagle (a TCP feature that collects small packets and sends them together), and server queues. If it always feels sluggish even with low ping, look at your PC (V-Sync, low FPS) or at a design that waits for server confirmation on every action (see the netcode models chapter). - **Freeze** (`freeze`, also called: frozen, hang, not responding): Everything on screen stops for a moment (0.5 s to a few seconds), then moves again. The whole server stopped (GC, deadlock, blocking call), the connection dropped briefly, or your PC froze. - **Dropped action / rollback** (`dropped`, also called: skill didn’t go off, item reverted, failed trade): Something you definitely did never happened, or its result gets reversed much later. The request was lost (packet loss, queue overflow), the server ruled differently from what your screen showed (timing difference, rejection after client-side feedback), or saving failed partway (DB lock or outage, server crash). - **Disconnect** (`disconnect`, also called: kicked out, connection lost, “Disconnected from server”): The connection drops mid-game and you land back at the login screen or a reconnect dialog. Not a single packet arrived within the timeout. Check for long connection drops, idle timeouts, server crashes or restarts, and a server or PC that stalled longer than the timeout (long loading). If the game just closed with no message, look at a forced client shutdown (crash, out of memory) before the connection. - **Can’t connect / infinite loading** (`noconnect`, also called: can’t log in, stuck on the loading screen): You can’t get into the game, or you’re stuck on a loading or entry screen. Whatever accepts new connections (the server’s connection queue, firewall, login server, DB) is full. It’s especially common right after maintenance. - **Invisible / ghost entities** (`invisible`, also called: missing NPCs, invisible characters, dead monsters still standing): NPCs, monsters, or players that should be there are missing only on your screen, or entities that are already gone remain only on your screen. This usually has little to do with speed: one packet went missing, or drawing the entity failed. Check for channel or phasing differences, lost spawn and despawn messages, messages discarded during loading, and asset loading failures. The decisive clue is whether it shows up after you leave view range and come back. ## The four factors - **Latency** (`lat`, Latency): Distance, queues, and processing time make every packet arrive late by a steady amount. How games cope: Prediction and client-side feedback show your own actions right away, and the server rewinds to the past to judge hits fairly (lag compensation). - **Jitter** (`jit`, Jitter): The average looks fine, but some packets arrive early and others late. Wi-Fi, congested links, and busy CPUs cause it. How games cope: The game holds a few packets in the interpolation buffer and plays them out at a steady rate. Jitter larger than the buffer can’t be hidden. - **Packet loss** (`loss`, Packet loss): Overflowing queues, radio interference, and faulty equipment drop packets. A connection that cuts out for a moment is loss too, just many packets in a row. How games cope: UDP games fill the gaps with interpolation and extrapolation, and send your inputs redundantly so losing one or two costs nothing. TCP holds every later packet back from the game until the lost one is received again. - **Stall** (`stall`, Stall): Server ticks run late or stop (GC, locks, blocking calls, overload), or frames on your PC stop. It happens even when the network is fine. How games cope: The game works off the backlog in a rush, skips it, or lets time run slow. ## Causes ### L1 Client game process (causes: 16) #### cg-hitch · Frame time spike · Frame hitch One frame takes several times longer than usual to compute, so the screen freezes for a moment. - Why → Effect → On screen: A burst of skill effects, a mass spawn, or a full UI refresh all land in one frame → The frame can’t finish within 16.7 ms and takes 50–300 ms → The screen hitches, then everyone moves at once on the next frame - Symptoms: Stutter, Freeze / Factors: Stall - Who: Just me / When: When crowds gather, During specific actions, Randomly - Primary owner: Game team (Client development) - Game team action items: Split heavy work across several frames, find spiking frames with a profiler, cap the number of effects. - Ballpark numbers: At 60 FPS, one frame is 16.7 ms. A single frame over 50 ms is often enough to make players feel “it stuttered.” - On the graph: Random spikes (Frame time) - Where to look: Record frame time (FrameTime) and the time the CPU and GPU spent on each frame (CPUBusy, GPUBusy) with PresentMon during play. On mobile, the slow sessions and slow rendering metrics in Android vitals - Confirmed if: Frame time, normally around 16.7 ms, jumps above 50 ms at the same moments as skill effects, mass spawns, or full UI refreshes, while ping stays the same - Ruled out if: Frame time steady but other characters hitch: points to the network side (e.g., “Missing or too-short interpolation buffer”). Spikes at regular intervals: check “Client garbage collection” first - Check with: The player’s own environment - Sources: - [Slow rendering](https://developer.android.com/topic/performance/vitals/render) · Android (Google) · To hit 60 FPS, a frame must render within 16 ms; late frames get skipped and show up as stutter (jank) - [Slow Sessions (games only)](https://developer.android.com/topic/performance/issues/slow-session) · Android (Google) · Android vitals counts a game frame as slow when it takes longer than 50 ms (20 FPS) or 34 ms (30 FPS) - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · FrameTime (CPU time between frames), CPUBusy and GPUBusy (time the CPU and GPU spent producing that frame) #### cg-gc · Client garbage collection · Client GC (Unity C#, Unreal, Lua) The whole game freezes while it reclaims memory that was used and thrown away (garbage). The telltale sign is stutter at regular intervals. - Why → Effect → On screen: Temporary strings, arrays, and lists are created and thrown away every frame → Once garbage piles up, GC pauses the main thread to reclaim it → Regular stutter every few seconds to tens of seconds - Symptoms: Stutter, Freeze / Factors: Stall - Who: Just me / When: At regular intervals, When crowds gather - Primary owner: Game team (Client development) - Game team action items: Reduce allocations (avoid string concatenation, LINQ, and lambda captures), use object pools, keep incremental GC on (the default since Unity 2020), run GC ahead of time when a pause does no harm, such as on a loading screen. - Ballpark numbers: Usually a few ms to 100 ms per pause, longer on low-end phones or in games that use a lot of memory (about 150–170 ms in the simulation). Unity’s GC scans the whole heap every time it runs, so the more memory the game is using, the longer it takes. - On the graph: Periodic spikes (Frame time, GC run times) - Where to look: In a development build, check the GC.Collect and GC.Alloc markers in the Unity Profiler. In Unreal, stat GC and stat Hitches (logs frames longer than the threshold set by t.HitchFrameTimeThreshold) - Confirmed if: Every spiking frame contains a GC.Collect section about as long as the spike, at regular intervals of a few seconds to tens of seconds. GC.Alloc per frame rises in crowded places - Ruled out if: No GC section in the spiking frames: “Synchronous loading and shader compilation on the main thread” or “Rendering load from large crowds” - Check with: Game server or client logs and metrics - Learn more: Common in clients that use C#, such as Unity. The usual culprit is code that builds combat log strings, damage numbers, and UI text from scratch every frame. If it only stutters in crowded places, some code is producing garbage in proportion to the player count. Incremental GC collects a little at a time in each frame (3 ms by default in Unity), but if the game makes garbage faster than it collects, it still ends up pausing all at once. Unreal Engine also has its own GC that cleans up unused game objects. It depends on the engine version and settings, but with default settings it runs about once a minute and can cause a hitch every minute. Clients that write game rules in a scripting language such as Lua also run that script’s GC separately. - Sources: - [Garbage collection modes](https://docs.unity3d.com/Manual/performance-incremental-garbage-collection.html) · Unity · Incremental GC is the default and collects across several frames; with it off, the main thread stops while the whole heap is scanned, which can take up to hundreds of ms - [Scripting.GarbageCollector.CollectIncremental](https://docs.unity3d.com/ScriptReference/Scripting.GarbageCollector.CollectIncremental.html) · Unity · The default target for the time incremental GC spends per run (time slice) is 3 ms (incrementalTimeSliceNanoseconds) - [Garbage Collection Settings in the Unreal Engine Project Settings](https://dev.epicgames.com/documentation/en-us/unreal-engine/garbage-collection-settings-in-the-unreal-engine-project-settings) · Epic Games · The setting that makes Unreal’s GC run at a fixed interval (Time Between Purging Pending Kill Objects, in seconds). This page doesn’t give the default value - [Profiler markers reference](https://docs.unity3d.com/Manual/profiler-markers.html) · Unity · GC.Collect: time program code is paused during garbage collection (under 1 ms to hundreds of ms); GC.Alloc: managed heap allocations - [Stat Commands in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/stat-commands-in-unreal-engine) · Epic Games · stat GC (garbage collection statistics), stat Hitches (logs frames that exceed t.HitchFrameTimeThreshold) #### cg-sync-load · Synchronous loading and shader compilation on the main thread · Synchronous asset load, shader compile The game freezes to read files and build shaders right before it draws an area, monster, or effect for the first time. - Why → Effect → On screen: Entering a new area, or a skill, piece of gear, or monster appearing for the first time → The main thread waits for file reads and shader compilation → A 0.1–1 s freeze the first time only; fine from the second time on - Symptoms: Freeze, Stutter / Factors: Stall - Who: Just me / When: While moving or changing zones, During specific actions - Primary owner: Game team (Client development) - Game team action items: Load asynchronously, preload (prewarm), precompile shaders on the loading screen or at first launch, make use of loading screens. - Ballpark numbers: Compiling one shader takes tens of ms, sometimes over 100 ms. Loading a texture takes tens to hundreds of ms depending on storage speed. - On the graph: Surge after opening (Frame spike count (right after a patch or driver update)) - Where to look: Record the same route twice with PresentMon and compare the first visit with the second. In a development build, turn on r.PSOPrecache.Validation in Unreal and check stat PSOPrecache and “PSO PRECACHING MISS” in the log; in Unity, check the loading and shader sections of spiking frames in the Profiler’s Timeline view - Confirmed if: 0.1–1 s spikes only at places visited or skills used for the first time, gone on the second visit. Reports surge right after a patch or graphics driver update, then taper off - Ruled out if: Spikes every time at the same place: not a shader cache problem. Repeats on every move only on PCs with slow storage: “Slow storage delays asset streaming.” Dedicated GPU memory full: “Out of graphics memory (VRAM)” - Check with: Game server or client logs and metrics - Learn more: On PC, the graphics driver stores each shader it builds in a shader cache and reuses it. So right after a graphics driver update or a game patch, that cache is invalidated, and even players who were fine stutter again for a while. The classic report is “after the patch, it hitches everywhere I go for the first time.” - Sources: - [Shader loading](https://docs.unity3d.com/Manual/shader-loading.html) · Unity · The first time a shader variant is used, the graphics driver has to build it for the GPU, which can cause a noticeable freeze; once built, it’s cached and doesn’t freeze again - [Optimizing Rendering With PSO Caches in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/optimizing-rendering-with-pso-caches-in-unreal-engine) · Epic Games · Creating a pipeline state object (PSO) at the moment it’s needed can take over 100 ms, so PSOs should be created in advance - [Direct3D 12 Return Codes](https://learn.microsoft.com/en-us/windows/win32/direct3d12/d3d12-graphics-reference-returnvalues) · Microsoft · D3D12_ERROR_DRIVER_VERSION_MISMATCH: a PSO cache built with a different driver version can’t be reused (recompiled after a driver update) - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · Records per-frame time as FrameTime (CPU time between frames) - [PSO Precaching for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/pso-precaching-for-unreal-engine) · Epic Games · With r.PSOPrecache.Validation on, stat PSOPrecache shows missed PSO statistics and the log records “PSO PRECACHING MISS”; runtime PSO creation over 20 ms (default) counts as a hitch #### cg-asset-stream · Slow storage delays asset streaming · Slow storage stalls asset streaming On slow storage such as an HDD, reading open-world textures and models can’t keep up with movement, so objects appear late or the game stutters while it waits for reads. - Why → Effect → On screen: Moving fast on a mount or by teleport, or entering a crowded area, suddenly calls for many new textures and models → Slow storage such as an HDD can’t read at the needed speed, so read requests pile up, and some loads make the main thread wait until they finish → Textures stay blurry for a while, buildings and characters pop in late, and the game stutters or freezes while it waits on reads - Symptoms: Invisible / ghost entities, Stutter, Freeze / Factors: Stall - Who: Just me / When: While moving or changing zones, When crowds gather - Primary owner: Game team (Client development) / Also: External (External) - Game team action items: Use asynchronous streaming so the main thread never waits on reads, prefetch based on movement direction and speed, show low-resolution textures (mipmaps) and simple models first and swap them in later, briefly cap movement speed or use a loading screen when streaming falls behind during fast travel, state in the minimum and recommended specs whether an SSD is required. - External action items: Tell players to install the game on an SSD and to check whether downloads or antivirus scans are running on the same disk. - Ballpark numbers: According to Microsoft, older hard drives read tens of MB per second and NVMe SSDs read several GB per second, while previous-generation games used around 50 MB per second for streaming. A game built around SSD speeds that reads far more than that can easily fall behind your movement when it runs from an HDD. - On the graph: Outliers only (Frame spike count (by storage type), disk read latency) - Where to look: While moving fast, record PhysicalDisk\Avg. Disk sec/Read (average time per read) and Current Disk Queue Length in Windows Performance Monitor together with PresentMon frame time. In a development build, check Unreal’s stat Streaming and stat AsyncLoad, or AssetBundle.asset/allAssets warnings in the Unity Profiler (the result was requested before loading finished, so the main thread waits) - Confirmed if: Disk read latency and queue length shoot up when moving fast, and at the same moments frames spike or textures and objects load late. Gone when the same scene runs from an SSD - Ruled out if: Disk idle but textures blurry: “Out of graphics memory (VRAM).” Fine from the second visit to the same place: “Synchronous loading and shader compilation on the main thread” - Check with: The player’s own environment - Learn more: If it freezes only the first time and is fine afterward, it’s closer to “Synchronous loading and shader compilation on the main thread.” If it repeats on every move only on PCs with slow storage, it’s this cause. “Out of graphics memory (VRAM),” where textures get dropped and reloaded because graphics memory runs short, also causes blurry textures, so check disk read waits and graphics memory usage together. Antivirus real-time scanning can also step in every time the game opens a file and slow reads down further. - Sources: - [DirectStorage is coming to PC](https://devblogs.microsoft.com/directx/directstorage-is-coming-to-pc/) · Microsoft · Older hard drives read tens of MB per second and NVMe SSDs several GB per second; the asset streaming budget of previous-generation games was around 50 MB per second; open-world games read and discard distant scenery in real time as the player moves - [Texture and mesh loading](https://docs.unity3d.com/Manual/LoadingTextureandMeshData.html) · Unity · Synchronous upload reads and uploads in one frame on the main thread and causes a visible freeze; asynchronous upload streams over several frames - [Texture Streaming Overview for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/texture-streaming-overview-for-unreal-engine) · Epic Games · The streamer raises and lowers texture resolution (mips) to match the camera view, does most of the work on async worker threads, and loads the mips visible on screen first - [Profiler markers reference](https://docs.unity3d.com/Manual/profiler-markers.html) · Unity · AssetBundle.asset/allAssets warning: the result was requested before loading finished, so the main thread stops and waits - [Stat Commands in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/stat-commands-in-unreal-engine) · Epic Games · stat Streaming (memory use and count of streaming textures), stat AsyncLoad (async loading statistics) - [Windows Performance Monitor Disk Counters Explained](https://learn.microsoft.com/en-us/archive/blogs/askcore/windows-performance-monitor-disk-counters-explained) · Microsoft · Avg. Disk sec/Read is the average time one read takes to complete (I/O latency); Current Disk Queue Length is the disk queue length at the moment of measurement - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · Records per-frame time as FrameTime (CPU time between frames) - [About regular quick and full scans with Microsoft Defender Antivirus](https://learn.microsoft.com/en-us/defender-endpoint/schedule-antivirus-scans) · Microsoft · Real-time protection scans a file every time it’s opened or closed #### cg-crowd · Rendering load from large crowds · Render/animation cost of crowds When hundreds of players fill one screen, as in a siege or a world boss fight, the cost of drawing them is more than the device can handle. - Why → Effect → On screen: Hundreds of players and effects overlap on one screen → Animation, shadow, name tag, and effect costs grow with the player count → FPS drops 60 → 15: all movement stutters, and input lag sets in - Symptoms: Stutter, Input lag / Factors: Stall - Who: Specific zone/channel, Just me / When: When crowds gather - Primary owner: Game team (Client development) - Game team action items: Simplify by distance (LOD), cap the number of players shown, offer a simplified effects option, lower the animation update rate. - Ballpark numbers: Even at 0.02–0.1 ms per character, 300 characters add up to 6–30 ms. The frame budget at 60 FPS is 16.7 ms, so this alone uses over 1/3 of it, and at the high end exceeds it. - On the graph: Rises with load (Frame time, characters on screen) - Where to look: Compare PresentMon frame time and CPUBusy/GPUBusy before and after a siege or world boss. In a development build, Unreal’s stat Unit (game thread, rendering thread, and GPU time) - Confirmed if: Frame time climbs as more players come on screen and improves right away when a cap on displayed players or the simplified effects option is turned on - Ruled out if: Spikes regardless of player count: “Frame time spike” or “Client garbage collection.” FPS fine but only other players’ movement lags behind: “Packet processing bottleneck on the main thread” - Check with: The player’s own environment - Sources: - [Introduction to level of detail](https://docs.unity3d.com/Manual/LevelOfDetail.html) · Unity · Without LOD, even objects that look tiny on screen are drawn at full complexity; LOD cuts the rendering cost - [Animation Budget Allocator in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/animation-budget-allocator-in-unreal-engine) · Epic Games · Dynamically reduces animation updates (ticks) for skeletal meshes to keep animation time within budget - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · FrameTime, CPUBusy, and GPUBusy show whether the CPU or the GPU is holding back the frame - [Stat Commands in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/stat-commands-in-unreal-engine) · Epic Games · stat Unit: total frame time plus game thread, rendering thread, and GPU time #### cg-net-mainthread · Packet processing bottleneck on the main thread · Network processing on the main thread If the client processes only a fixed amount of received packets per frame, a flood of packets keeps getting pushed to the next frame. - Why → Effect → On screen: Thousands of updates per second arrive in crowded places → The main thread hits its per-frame processing limit and can’t read them all → Other players’ movements show up later and later, then all at once - Symptoms: Fast-forward, Input lag / Factors: Stall, Latency - Who: Specific zone/channel, Just me / When: When crowds gather - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: receive and parse on a separate thread, merge stale position updates for the same entity and apply only the latest. Server: in crowded places, send updates for distant characters less often to cut traffic. - Ballpark numbers: Once unprocessed packets pile up, it takes only a few seconds to fall a full second behind. - On the graph: Rises with load (Unprocessed received packets, receive-to-apply delay) - Where to look: Log how many packets the client leaves unprocessed each frame and the delay from packet arrival to when the game applies it, and view them alongside the number of nearby players - Confirmed if: In crowded places the leftover packet count and apply delay keep growing, while ping and the server’s send interval stay normal - Ruled out if: No apply delay but the packets themselves arrive late: the network path. Frame time rises sharply: “Rendering load from large crowds” - Check with: Game server or client logs and metrics - Sources: - [Actor Priority in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-priority-in-unreal-engine) · Epic Games · When bandwidth runs short, not every actor is replicated every time; actors are prioritized by distance from the viewer and time since their last replication - [Replication Graph in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/replication-graph-in-unreal-engine) · Epic Games · Games with many players and replicated objects (such as MMORPGs) need to group them by location and send only what each player needs, or the server CPU becomes the bottleneck #### cg-no-buffer · Missing or too-short interpolation buffer · Missing/short interpolation buffer If the client draws server packets the moment they arrive, jitter (variation in packet arrival times) shows up directly on screen. - Why → Effect → On screen: Received positions are drawn immediately, or the buffer is shorter than the jitter → Characters stop for as long as a packet is late, then jump when delayed packets arrive together → Other characters move in fits and starts - Symptoms: Stutter / Factors: Jitter - Who: Just me / When: Always - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: add an interpolation buffer, adjust its length automatically to connection quality. Server: apply lag compensation (rewind-based hit registration) so hits still register correctly when players aim at positions a buffer’s length in the past. - Ballpark numbers: The usual buffer is about twice the server’s send interval (100 ms when receiving 20 updates a second). - On the graph: Always high (Packet arrival intervals, times the interpolation buffer ran empty) - Where to look: On the client, record the distribution of server packet arrival intervals and the number of frames that froze or fell back to extrapolation because there was no next snapshot to interpolate to - Confirmed if: Variation in arrival intervals often exceeds the interpolation buffer length, and each time the buffer runs empty and other characters hitch. A longer buffer reduces it - Ruled out if: Hitches even with enough buffer: check whether the server’s send interval itself is irregular (tick delays) - Check with: Game server or client logs and metrics - Learn more: A longer buffer makes movement smoother, but you see opponents that much further in the past. That’s why hit registration comes paired with lag compensation, where the server rewinds to “the past that player was seeing” to check the hit. - Sources: - [Interpolation and extrapolation (Netcode for Entities 6.5)](https://docs.unity3d.com/Packages/com.unity.netcode@6.5/manual/interpolation.html) · Unity · Buffered interpolation deliberately renders late to wait for late packets; a bigger buffer is more accurate but adds that much latency - [Struct ClientTickRate (Netcode for Entities 6.5)](https://docs.unity3d.com/Packages/com.unity.netcode@6.5/api/Unity.NetCode.ClientTickRate.html) · Unity · Default interpolation buffer InterpolationTimeNetTicks = 2 (two server sends’ worth) - [Physics (Netcode for Entities 6.5)](https://docs.unity3d.com/Packages/com.unity.netcode@6.5/manual/physics.html) · Unity · Lag compensation: the server finds the collision world the client was seeing at that tick and decides whether the shot hit #### cg-extrap · Excessive extrapolation (dead reckoning) · Over-extrapolation / dead reckoning While no packets arrive, the client keeps showing characters moving at their last velocity, then snaps them back when it turns out to be wrong. - Why → Effect → On screen: Packets stop arriving, so the character keeps moving in its last direction and speed → In reality, the other player stopped or changed direction → The other character runs on for a while, then snaps to its real position or passes through walls. With erratic packet arrival intervals, it keeps overshooting and getting pulled back, so it looks like it’s shaking - Symptoms: Teleporting, Stutter / Factors: Packet loss, Jitter - Who: Just me / When: Randomly - Primary owner: Game team (Client development) - Game team action items: Cap extrapolation time (e.g., 200–250 ms), converge smoothly when wrong. - Ballpark numbers: At 6 m/s, being off by just 300 ms puts a character 1.8 m out of place. - On the graph: Random spikes (Extrapolation time, position correction distance) - Where to look: Record how long other characters are drawn by extrapolation and how far their position is corrected when a new packet arrives - Confirmed if: Every time packets stop, extrapolation time grows with no cap, then the correction distance grows to several meters - Ruled out if: Teleporting even though extrapolation runs only briefly: packet loss or latency itself is high, so look at the connection or route - Check with: Game server or client logs and metrics - Sources: - [Interpolation and extrapolation (Netcode for Entities 6.5)](https://docs.unity3d.com/Packages/com.unity.netcode@6.5/manual/interpolation.html) · Unity · Extrapolation that keeps moving in the same direction and speed when the next snapshot is late is often wrong, so it gets a cap (Unity default 20 ticks, about 1/3 s at 60 Hz) - [Peeking into VALORANT's Netcode](https://technology.riotgames.com/news/peeking-valorants-netcode) · Riot Games · When the guesses that fill in late or missing data are wrong, the client drifts from the server and characters jump or slide into place when corrected #### cg-predict · Client-side prediction mismatch · Prediction mismatch / reconciliation Your client shows your character moving before the server confirms it, but if the server calculates something different, your character gets pulled back. - Why → Effect → On screen: The client moves the character before the server confirms (prediction) → The server calculates collisions, movement speed, or buffs differently, or never receives the command → When the confirmation arrives, your character gets pulled back - Symptoms: Rubber-banding / Factors: Packet loss, Latency - Who: Just me / When: While moving or changing zones, Randomly - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: use the same movement code as the server, send inputs redundantly, smooth out corrections. Server: use the same movement code as the client, filter duplicate inputs by input number so each one is processed only once. - Ballpark numbers: The pull-back distance is “time out of sync × movement speed”. Losing just a few commands means 1–3 m. - On the graph: Random spikes (Server corrections (mispredictions)) - Where to look: Record how often and how far the server corrects position. In Unreal, count the server’s ClientAdjustPosition corrections; in Unity Netcode for Entities, count rollbacks and resimulations caused by mispredictions - Confirmed if: Corrections cluster at the times of rubber-banding reports, and correction distance is repeatedly large with specific buffs, terrain, or movement skills - Ruled out if: Corrections cluster only when loss is high: input packet loss on the connection. Only other characters look pulled back with no corrections: “Excessive extrapolation (dead reckoning)” - Check with: Game server or client logs and metrics - Sources: - [Introduction to prediction (Netcode for Entities 6.5)](https://docs.unity3d.com/Packages/com.unity.netcode@6.5/manual/intro-to-prediction.html) · Unity · Client and server predict with the same simulation code; when the result differs from the server state (misprediction), the client rolls back and resimulates, and the correction becomes visible - [Use the command stream to handle user inputs (Netcode for Entities 6.5)](https://docs.unity3d.com/Packages/com.unity.netcode@6.5/manual/command-stream.html) · Unity · Sends the inputs of the previous few ticks again with the latest input to cover packet loss - [Understanding Networked Movement in the Character Movement Component for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/understanding-networked-movement-in-the-character-movement-component-for-unreal-engine) · Epic Games · The client moves first and the server replays the same move; if the positions differ, the server corrects with ClientAdjustPosition and the client reapplies its saved moves #### cg-fixed-step · Fixed-timestep catch-up spiral · Fixed-timestep catch-up / spiral of death After one stall, the game runs its backlog of calculations all at once, and that extra work puts it behind again. - Why → Effect → On screen: The game simulation runs at a fixed interval and stalls once → The backlog of steps is computed in a single frame → A chain of long frames causes spikes, or the cap kicks in and the world slows down - Symptoms: Stutter, Fast-forward, Slow motion / Factors: Stall - Who: Just me / When: Randomly, When crowds gather - Primary owner: Game team (Client development) - Game team action items: Cap catch-up per frame, handle the leftover time with interpolation. - On the graph: Random spikes (Frame time, fixed steps per frame) - Where to look: In a development build profiler, check how many fixed steps ran per frame (in Unity, the number of FixedUpdate-phase markers such as FixedBehaviourUpdate) together with frame time - Confirmed if: One long frame is followed by a chain of long frames that each run several steps; once the cap (Unity’s Maximum Allowed Timestep) is reached, game time runs slower than real time - Ruled out if: A single long frame that doesn’t repeat: “Frame time spike” or “Client garbage collection” - Check with: Game server or client logs and metrics - Learn more: Unity’s physics (FixedUpdate) is the classic fixed step (0.02 s by default, 50 times a second). Maximum Allowed Timestep in the Time settings (the most time the game catches up in one frame, about 0.33 s by default) is the catch-up cap. If a frame runs longer than that, the excess time is dropped, and the game clock falls behind real time by that much. - Sources: - [Set fixed timestep to optimize physics simulation frequency](https://docs.unity3d.com/Manual/physics-optimization-cpu-frequency.html) · Unity · Fixed Timestep defaults to 0.02 s (50 times a second); when a frame runs long, several physics steps run in that frame and the load grows - [Handling variation in time](https://docs.unity3d.com/Manual/time-handling-variations.html) · Unity · Maximum Allowed Timestep defaults to 1/3 s (0.3333333): even if the game stops for 1 s, game time advances only 0.333 s. The cap prevents the vicious cycle where catch-up steps slow things down even more - [Profiler markers reference](https://docs.unity3d.com/Manual/profiler-markers.html) · Unity · FixedBehaviourUpdate: the section where MonoBehaviour.FixedUpdate runs; physics markers are called in the FixedUpdate phase #### cg-clock · Clock sync error · Clock sync error If the client’s estimate of the server time is wrong, interpolation timing and cooldown checks drift out of step with the server. - Why → Effect → On screen: The client syncs to server time only once when connecting and never adjusts as ping changes → The point to interpolate to and the time a cooldown ends drift away from the server’s → Opponents occasionally hitch; skills get rejected even after the cooldown has ended - Symptoms: Stutter, Dropped action / rollback / Factors: Latency - Who: Just me / When: The longer it runs, Randomly - Primary owner: Game team (Client development) - Game team action items: Sync time periodically (measure round-trip time and correct), adjust gradually with no sudden jumps, switch elapsed-time measurement from the PC’s clock to a monotonic clock. - On the graph: Slow climb (Error in the estimated server time) - Where to look: Periodically log the difference between the client’s estimated server time and the server time (tick number) the server puts in its packets - Confirmed if: The error grows the longer the session runs, or jumps all at once when the PC clock gets corrected, and reports of rejected skills and hitches rise around then - Ruled out if: Error stays small but skills still get rejected: points to server-side validation or latency - Check with: Game server or client logs and metrics - Learn more: If elapsed time is measured with the PC’s date and time (wall clock), game time jumps the moment Windows corrects the clock against internet time or the user changes the clock. Measure elapsed time with a monotonic clock, which never goes backward (Stopwatch and so on). - Sources: - [RFC 5905: Network Time Protocol Version 4: Protocol and Algorithms Specification](https://www.rfc-editor.org/rfc/rfc5905) · IETF · A method that computes round-trip delay and clock offset from the four timestamps of a request and its response - [Acquiring high-resolution time stamps](https://learn.microsoft.com/en-us/windows/win32/sysinfo/acquiring-high-resolution-time-stamps) · Microsoft · QueryPerformanceCounter (used by Stopwatch) is a clock for elapsed time that isn’t synced to external time; use system time only when UTC time is needed - [Time synchronization (Netcode for Entities 6.5)](https://docs.unity3d.com/Packages/com.unity.netcode@6.5/manual/time-synchronization.html) · Unity · Estimates server time from round-trip time and converges by adjusting the clock rate slightly, with no large time jumps #### cg-float-time · Float time precision loss · Float time precision loss on long sessions If the game keeps its clock in a low-precision floating-point type (float), the longer it runs, the worse its time resolution (the smallest time difference it can tell apart) becomes, and movement and effects start to shake. - Why → Effect → On screen: Time elapsed since launch is accumulated in a float or passed to shaders as is → The longer the game runs, the larger the smallest difference a float can represent → Only clients left running for days see characters, animations, and scrolling effects shake; a restart fixes it - Symptoms: Stutter / Factors: Jitter - Who: Just me / When: The longer it runs - Primary owner: Game team (Client development) - Game team action items: Store elapsed time as a double (64-bit) or an integer, wrap the time passed to shaders back around at a fixed interval, run automated tests that last several days. - Ballpark numbers: A 32-bit float has just over 7 significant digits, so after a day of running (about 86,400 s) the time resolution is about 8 ms, roughly half a frame at 60 FPS (16.7 ms), and after a week about 60 ms, more than a whole frame. - On the graph: Slow climb (Shaking reports by how long the client has been running) - Where to look: Collect how long the client has been running with each shaking report and compare before and after a restart. On the development side, a test that sets the game start time several days back and runs from there - Confirmed if: Shaking only on clients left running for days, gone after a restart, and worse the longer the client has been running - Ruled out if: Shaking right after launch: “Missing or too-short interpolation buffer” or “Timer resolution” - Check with: The player’s own environment - Learn more: This shows up especially in mobile MMOs where players leave auto-hunting on for days without closing the game. Unity’s Time.time is also a float, so Unity provides a separate double version, Time.timeAsDouble, and recommends it. - Sources: - [Time.timeAsDouble](https://docs.unity3d.com/ScriptReference/Time-timeAsDouble.html) · Unity · The double version of Time.time; more precise than float the longer the game runs, so it’s recommended in most cases - [Floating-point numeric types (C# reference)](https://learn.microsoft.com/en-us/dotnet/csharp/language-reference/builtin-types/floating-point-numeric-types) · Microsoft · float precision about 6–9 digits, double about 15–17 digits #### cg-vsync · V-Sync and the render queue · V-Sync, render queue Input is delayed while several finished frames wait in a queue to be sent out in step with the monitor’s refresh. - Why → Effect → On screen: The graphics driver queues 1–3 frames ahead → Input takes that much longer to show up on screen → Ping is low, but controls feel heavy and sluggish - Symptoms: Input lag, Stutter / Factors: Latency - Who: Just me / When: Always - Primary owner: Game team (Client development) / Also: External (External) - Game team action items: Support a low-latency mode, shorten the frame queue, offer a frame cap option set slightly below the refresh rate, turn on frame pacing on phones. - External action items: Tell players to pair a variable refresh rate monitor with a frame cap slightly below the refresh rate and to turn on the graphics driver’s low-latency mode. - Ballpark numbers: At 60 Hz, each frame is 16.7 ms. When the CPU runs ahead of the GPU or the display refresh and the three-frame queue (the DirectX 11 default) fills up, 50 ms is added. With double-buffered V-Sync, a frame that took 17 ms waits for the next refresh (33.3 ms), and the previous frame is shown once more in the meantime. - On the graph: Always high (Input-to-display latency) - Where to look: Compare PresentMon’s MsClickToPhotonLatency and MsAllInputToPhotonLatency (from mouse or keyboard input to output on screen) and DisplayLatency while switching V-Sync, low-latency mode, and frame caps. MsPCLatency (from when the PC receives input to when it sends the frame to the display) is recorded only if the game emits PC Latency events - Confirmed if: This latency grows by one or two frames (tens of ms) with V-Sync on or with no frame cap, and shrinks with low-latency mode or a frame cap slightly below the refresh rate. Ping unchanged - Ruled out if: Controls still feel late while latency inside the PC is low: “Display, input device, and frame generation latency.” High ping: the network side - Check with: The player’s own environment - Learn more: V-Sync (vertical sync) is a setting that sends out a new frame only at the moment the monitor refreshes the screen. Screen tearing goes away, but input is delayed by the wait for that moment, and when FPS drops below 60, it bounces between 60 and 30 and stutters. A variable refresh rate monitor shortens this wait by refreshing when a frame is ready. The same thing happens on phones. If a 30 FPS game can’t pace its frames evenly on a 60 Hz screen, the average is 30 FPS, but individual frames stay on screen for uneven times such as 49, 16, and 33 ms, which shows up as stutter (an example from the Android developer docs). Android’s frame pacing library (which evens out the intervals between frames) or the equivalent engine option reduces it. - Sources: - [IDXGIDevice1::SetMaximumFrameLatency](https://learn.microsoft.com/en-us/windows/win32/api/dxgi/nf-dxgi-idxgidevice1-setmaximumframelatency) · Microsoft · The default number of frames the driver can queue is 3 (range 1–16) - [Reduce latency with DXGI 1.3 swap chains](https://learn.microsoft.com/en-us/windows/uwp/gaming/reduce-latency-with-dxgi-1-3-swap-chains) · Microsoft · Present blocks until the queue drains, so a rendered frame waits almost one extra frame before it’s displayed; a waitable swap chain reduces this - [Frame Pacing library](https://developer.android.com/games/sdk/frame-pacing) · Android (Google) · On a 60 Hz screen, the previous frame is shown again when there’s no new frame; example of a 30 FPS game whose frame times become uneven, such as 49, 16, and 33 ms - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · MsPCLatency (from when the PC receives input to when it sends the frame to the display), MsClickToPhotonLatency (mouse click to screen), MsAllInputToPhotonLatency (keyboard or mouse input to screen), DisplayLatency (frame submission to output to the monitor) - [PresentMon Console Application (README-ConsoleApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-ConsoleApplication.md) · Intel · MsPCLatency is recorded only if the app emits PC Latency events (--track_pc_latency); MsAllInputToPhotonLatency is based on keyboard and mouse input #### cg-leak · Client memory leak · Client memory leak The longer the game stays open, the more memory it uses; it gets slower and slower until the game is eventually force-closed. - Why → Effect → On screen: Textures, UI, and effects aren’t released when moving between areas → GC runs more often, and the OS runs short of memory and starts swapping → After hours of play it stutters more and more, then gets force-closed (looks like a disconnect to the player) - Symptoms: Stutter, Disconnect / Factors: Stall - Who: Just me / When: The longer it runs - Primary owner: Game team (Client development) - Game team action items: Measure memory usage at area transitions, find and fix textures, UI, and effects that never get released, run long automated tests (soak tests). - On the graph: Slow climb (Game process memory) - Where to look: Record Process(game)\Private Bytes in Performance Monitor for a few hours. On mobile, the exit reason in Android’s ApplicationExitInfo (REASON_LOW_MEMORY) and iOS jetsam reports - Confirmed if: Memory goes up every time the player moves between areas and never comes back down, and stutter and forced shutdowns increase the longer the game runs - Ruled out if: Memory stays flat, and the only thing that grows with uptime is shaking: “Float time precision loss” - Check with: The player’s own environment - Learn more: Phones mostly hold out by compressing memory. If memory still runs short, the OS shuts the game down on the spot (players see it as getting kicked out). Devices with less RAM get shut down first. - Sources: - [Memory allocation among processes](https://developer.android.com/topic/performance/memory-management) · Android (Google) · Android holds out by compressing memory into zRAM; when that’s not enough, the low memory killer terminates processes, and a foreground app being killed looks like a crash - [Identifying high-memory use with jetsam event reports](https://developer.apple.com/documentation/xcode/identifying-high-memory-use-with-jetsam-event-reports) · Apple · iOS force-quits apps (jetsam) when memory pressure doesn’t ease, and any app that exceeds its per-app memory limit becomes a target - [Find user-mode memory leaks with Performance Monitor (PerfMon)](https://learn.microsoft.com/en-us/windows-hardware/drivers/debugger/using-performance-monitor-to-find-a-user-mode-memory-leak) · Microsoft · Record Process > Private Bytes (private memory allocated by the process) and Virtual Bytes over a long period; if they only ever grow, it’s a leak - [ApplicationExitInfo](https://developer.android.com/reference/android/app/ApplicationExitInfo) · Android (Google) · REASON_LOW_MEMORY: the system’s low memory killer terminated the app process (devices that don’t support it report REASON_SIGNALED with SIGKILL) #### cg-crash · Client crash · Client crash An unhandled error closes the game. To the player it looks like a disconnect, but the server is fine. - Why → Effect → On screen: Null reference, out of memory, graphics driver error → The game process is forcibly terminated → Reports of “I got kicked out” while everyone else is fine at the same moment - Symptoms: Disconnect / Factors: Stall - Who: Just me / When: During specific actions, Randomly - Primary owner: Game team (Client development) / Also: External (External) - Game team action items: Collect crash reports, break down stats by device and driver, fix the most frequent errors first. - External action items: If crashes cluster on a specific graphics driver version, tell players to update their driver. - On the graph: Outliers only (Crash count (by device, graphics driver, and build)) - Where to look: Crash reports and the Android vitals crash rate by device, driver, and build. On players’ PCs, Event ID 1000 in the Event Viewer Application log (faulting module name) and “Display driver stopped responding and has recovered” entries - Confirmed if: A crash record exists at the reported disconnect time, and other players on the same server are fine at the same moment. Crashes cluster on specific devices, driver versions, or modules - Ruled out if: Connection dropped with no crash record: “NAT mapping expiry” or the connection side - Check with: Game server or client logs and metrics - Sources: - [Crashes](https://developer.android.com/topic/performance/vitals/crash) · Android (Google) · A crash is when an app exits unexpectedly because of an unhandled exception or signal (SIGSEGV and so on); tallied in Android vitals in Play Console - [WDDM Support for Timeout Detection and Recovery (TDR)](https://learn.microsoft.com/en-us/windows-hardware/drivers/display/timeout-detection-and-recovery) · Microsoft · If the GPU can’t finish its work within 2 s (default), Windows resets the graphics driver and the GPU - [The application or service crashing behavior troubleshooting guidance](https://learn.microsoft.com/en-us/troubleshoot/windows-server/performance/troubleshoot-application-service-crashing-behavior) · Microsoft · Event ID 1000 in the Application log is the actual crash record and includes the faulting application and the faulting module (Faulting module name) #### cg-anticheat · Anti-cheat scans · Anti-cheat scan and heartbeat The anti-cheat module that runs alongside the game to block cheats scans the system periodically. If a scan is heavy, or the heartbeat (a periodic keepalive signal) to the anti-cheat server is late, the game stutters or disconnects. - Why → Effect → On screen: The anti-cheat module periodically scans game memory, running programs, and drivers → The game thread stalls during the scan, or the heartbeat doesn’t go out on time → Hitches at regular intervals; in bad cases, a disconnect with a security error message - Symptoms: Stutter, Freeze, Disconnect / Factors: Stall - Who: Just me / When: At regular intervals, Right after login or maintenance, Randomly - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: run heavy scans off the game thread in small pieces, compare stutter and kick stats by anti-cheat module version (if they spike right after an update, report it to the anti-cheat vendor). Server: tolerate a heartbeat that’s late once or twice. - Ballpark numbers: A light scan usually takes under 1 ms, but a heavy scan running on the game thread can take tens to hundreds of ms at a time, depending on the implementation. - On the graph: Periodic spikes (Frame time, anti-cheat kicks) - Where to look: Measure the spacing between spikes in PresentMon frame time, and tally the anti-cheat kick reasons the server received (for EOS, AuthenticationFailed / Authentication Timed Out and others in ClientActionReason) by anti-cheat module version and hardware - Confirmed if: Brief pauses repeat at a regular interval regardless of what’s happening in the game, and right after an anti-cheat update, stutter and authentication-timeout kicks rise on specific hardware - Ruled out if: Same interval on all hardware regardless of anti-cheat version: “Client garbage collection” or “Background processes taking up CPU” - Check with: Game server or client logs and metrics - Learn more: Anti-cheat modules sit deep in the OS as drivers, so they sometimes conflict with antivirus software, overlays, and other games’ anti-cheat modules. If stutter and kick reports surge on specific hardware right after an anti-cheat update, suspect this first. - Sources: - [Using the Anti-Cheat Interfaces](https://dev.epicgames.com/docs/epic-online-services/trust-and-safety/anti-cheat-interfaces/using-anti-cheat) · Epic Games · If the server doesn’t receive the client’s anti-cheat message within the set time (RegisterTimeout), it kicks the client for an authentication timeout (a client frozen by loading is a common cause); if a recent module update is the problem, roll back to the previous module - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · Records per-frame time as FrameTime (CPU time between frames) ### L2 Client OS and device (causes: 15) #### co-background · Background processes taking up CPU · Background CPU contention When an antivirus scan, Windows Update, streaming software, or a browser video takes over CPU cores, the game thread has to wait for CPU time. - Why → Effect → On screen: Other programs hold CPU cores for a long time → The game thread waits for CPU time → Frames come late, and received packets are processed late too - Symptoms: Stutter, Fast-forward / Factors: Stall - Who: Just me / When: Randomly, At regular intervals - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Adjust game thread priority; record whole-PC CPU usage in the logs taken when stutter happens, to tell whether another program is to blame. - External action items: Tell players to turn on Windows Game Mode and to close unneeded programs while playing (antivirus scans, Windows Update, streaming software, browser videos). - Ballpark numbers: Windows usually hands out a core for a few ms to tens of ms at a time. Getting passed over by the scheduler just once is enough to lose a frame. - On the graph: Random spikes (Whole-PC CPU usage, frame time) - Where to look: Record the CPU column in Task Manager’s Processes tab and Processor Information(_Total)\% Processor Time in Performance Monitor together with PresentMon frame time. If antivirus is suspected, record with New-MpPerformanceRecording and use Get-MpPerformanceReport to find the files and processes with the longest scan times - Confirmed if: Another program’s CPU use (antivirus scan, update, streaming software) spikes at the stutter times, or files in the game folder rank near the top for scan time. Closing that program or adding an exclusion makes it go away - Ruled out if: The whole screen hitches and audio crackles even though CPU usage is low: “NIC power saving and driver issues” (DPC latency) - Check with: The player’s own environment - Learn more: Windows gives the program in the front window (foreground) a little extra priority, but when there’s more work than there are cores, the game waits too. Antivirus software gets in the way through “real-time protection” more often than through CPU use. It scans every time the game opens a file, so the freezes while assets load get longer. - Sources: - [Multitasking](https://learn.microsoft.com/en-us/windows/win32/procthread/multitasking) · Microsoft · Windows gives each thread a time slice and moves on to the next thread when it’s used up; a time slice is about 20 ms (varies by OS and CPU) - [Priority Boosts](https://learn.microsoft.com/en-us/windows/win32/procthread/priority-boosts) · Microsoft · The process in the front window (foreground) gets its priority raised to at least that of background processes - [About regular quick and full scans with Microsoft Defender Antivirus](https://learn.microsoft.com/en-us/defender-endpoint/schedule-antivirus-scans) · Microsoft · Real-time protection scans every time a file is opened or closed and every time a folder is opened - [Network-Related Performance Counters](https://learn.microsoft.com/en-us/windows-server/networking/technologies/network-subsystem/net-sub-performance-counters) · Microsoft · Processor Information: % Processor Time counter - [Performance analyzer for Microsoft Defender Antivirus](https://learn.microsoft.com/en-us/defender-endpoint/tune-performance-defender-antivirus) · Microsoft · Record with New-MpPerformanceRecording and use Get-MpPerformanceReport to see the top files, paths, and processes that affected scan time - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · Records per-frame time as FrameTime (CPU time between frames) #### co-power · Power saving and thermal throttling · Power saving, thermal throttling Laptop battery mode, phone power-saving mode, and device heat slow down the CPU and GPU. With heat, the telltale sign is that the game runs fine at first and slows down only after a while. - Why → Effect → On screen: Battery or power-saving mode is on, or the device gets hot → CPU and GPU clocks drop by 30–50%, depending on the device → FPS drops and the game stutters, right from the start with power saving, or after a few minutes to about 20 minutes of play with heat - Symptoms: Stutter, Input lag / Factors: Stall - Who: Just me / When: The longer it runs, Always - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Adjust graphics options automatically, manage heat with a frame cap, lower options ahead of time based on the thermal level the OS reports (iOS thermalState, Android thermal status API), mark the executable so laptops with two graphics chips use the discrete GPU (export NvOptimusEnablement and AmdPowerXpressRequestHighPerformance). - External action items: Tell players to turn off power-saving mode and to keep laptops plugged in; for reports like “it’s a good laptop but FPS is low,” check which graphics chip the game runs on and tell players to assign the game to the high-performance GPU in Windows graphics settings. - On the graph: Slow climb (FPS, CPU/GPU clocks) - Where to look: Record PresentMon’s CPUFrequency, GPUFrequency, CPUTemperature, and GPUTemperature with frame time for 20–30 minutes, and check which graphics chip the game runs on with the GPU engine column in Task Manager’s Processes tab. On mobile, record Android’s thermal API (getThermalHeadroom, thermal status) and iOS thermalState together with FPS - Confirmed if: FPS drops from the point where the temperature rises and clocks fall, or clocks are low only in battery or power-saving mode. Or the game is running on integrated graphics - Ruled out if: FPS drops while clocks and temperature hold steady: “Background processes taking up CPU” or “Client memory leak” - Check with: The player’s own environment - Learn more: Laptops with two graphics chips sometimes run games on the slower integrated graphics to save power. For a report like “it’s a good laptop but FPS is low,” first check which graphics chip the game is running on. - Sources: - [Thermal API](https://developer.android.com/games/optimize/adpf/thermal) · Android (Google) · Devices can sustain high performance only for a limited time before heat forces throttling; recommends watching thermal status and lowering the load ahead of time - [thermalState](https://developer.apple.com/documentation/foundation/processinfo/thermalstate-swift.property) · Apple · The current thermal level reported by iOS; as the level rises, the app should reduce its resource use - [Selecting the Best Graphics Device to Run a 3D Intensive Application](https://gpuopen.com/learn/amdpowerxpressrequesthighperformance/) · AMD · On laptops with two graphics chips, running on integrated graphics can turn a 60 FPS game into a 30 FPS one; exporting AmdPowerXpressRequestHighPerformance selects the discrete GPU - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · Records CPUFrequency and GPUFrequency (clocks) and CPUTemperature and GPUTemperature (temperatures) per frame - [GPUs in the task manager](https://devblogs.microsoft.com/directx/gpus-in-the-task-manager/) · Microsoft · Task Manager has columns showing per-process GPU usage and which GPU and engine that usage belongs to #### co-timer · Timer resolution · Timer resolution (Windows 15.6ms) Windows’ default timer ticks every 15.6 ms, so “sleep for just 1 ms” actually lasts until the next timer tick, up to 15.6 ms. - Why → Effect → On screen: Frame limiting and packet sending are implemented with Sleep (a short wait) → The OS wakes the thread only every 15.6 ms → Frame intervals and input send intervals become uneven - Symptoms: Stutter / Factors: Jitter - Who: Just me / When: Always - Primary owner: Game team (Client development) - Game team action items: Use high-resolution timers, replace Sleep-based waits with event- or V-Sync-based pacing. - Ballpark numbers: 15.6 ms steps can’t hit a 16.7 ms interval, so frame intervals alternate between 15.6 ms and 31.2 ms. - On the graph: Always high (Frame interval distribution) - Where to look: Look at the distribution of PresentMon’s MsBetweenPresents (frame interval) and the “Platform Timer Resolution” entries in the powercfg /energy report (processes that changed the timer resolution) - Confirmed if: Frame intervals cluster at multiples of 15.6 ms, such as 15.6 ms and 31.2 ms, and the game doesn’t request a higher timer resolution - Ruled out if: Intervals spread out evenly: more likely “Background processes taking up CPU” or frame load than the timer - Check with: The player’s own environment - Learn more: In older versions of Windows, when one program set the timer to 1 ms, it applied to every program. That’s where the saying “leave a browser open and your game runs smoother” came from. Since Windows 10 version 2004, the change applies only to the program that requested it, and Windows 11 may ignore requests from windows that are minimized or fully hidden and not playing sound. - Sources: - [_WDF_TIMER_CONFIG (wdftimer.h)](https://learn.microsoft.com/en-us/windows-hardware/drivers/ddi/wdftimer/ns-wdftimer-_wdf_timer_config) · Microsoft · Standard timer accuracy is the system clock tick interval, 15.6 ms by default; high-resolution timers get 1 ms - [timeBeginPeriod function (timeapi.h)](https://learn.microsoft.com/en-us/windows/win32/api/timeapi/nf-timeapi-timebeginperiod) · Microsoft · Before Windows 10 2004 it was a global setting; since then it applies only to the requesting process, and Windows 11 doesn’t guarantee high resolution to processes whose windows are hidden or minimized - [CreateWaitableTimerExW function (synchapi.h)](https://learn.microsoft.com/en-us/windows/win32/api/synchapi/nf-synchapi-createwaitabletimerexw) · Microsoft · High-resolution waitable timer flag CREATE_WAITABLE_TIMER_HIGH_RESOLUTION - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · MsBetweenPresents: time between this Present() call and the previous one (ms) - [Results for the Idle Energy Efficiency Assessment](https://learn.microsoft.com/en-us/windows-hardware/test/assessments/results-for-the-idle-energy-efficiency-assessment) · Microsoft · Default system timer resolution is 15.6 ms; the “Platform Timer Resolution” entries in the energy report show which processes changed the timer resolution - [Powercfg command-line options](https://learn.microsoft.com/en-us/windows-hardware/design/device-experiences/powercfg-command-line-options) · Microsoft · powercfg /energy: analyzes the system and generates an energy report (HTML) #### co-mobile-bg · Mobile app sent to the background · App suspended in background If you leave the game for a moment to check a notification, the OS suspends the app a few seconds later, and meanwhile the server disconnects you. - Why → Effect → On screen: The player leaves the game to read a message or take a call → The game engine pauses gameplay, and the OS soon suspends the app and its networking → Already disconnected on return, so the game reconnects - Symptoms: Disconnect / Factors: Stall, Packet loss - Who: Just me / When: After sitting idle, During specific actions - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: on return, reconnect automatically right away with a session token (resume without logging in again) without waiting on the dead connection, then fetch the latest state in one go to catch up. Server: when heartbeats stop, clean up the connection but keep the character session for a short grace period (don’t kick it right away), and resume it by session token if the player reconnects within that window. - Ballpark numbers: Game engines usually pause the moment the app goes to the background. iOS suspends the app within a few seconds, or usually within tens of seconds even with extra time granted, and Android 14 and later freezes apps that leave the screen after about 10 seconds. - On the graph: Mass disconnect (Disconnects (heartbeat timeouts), app suspend records) - Where to look: Match the app suspend and resume times in the client log (OnApplicationPause in Unity) against the server’s disconnect reasons and times by session ID. On Android, also check the process exit reasons recorded in ApplicationExitInfo (REASON_LOW_MEMORY and so on) - Confirmed if: The client went into suspend just before the server’s heartbeat-timeout disconnect and reconnected right after resuming - Ruled out if: Disconnected while the app was in the foreground: “NAT mapping expiry,” “ISP-shared IP addresses (CGNAT),” or “Wi-Fi ↔ LTE/5G switching” - Check with: Game server or client logs and metrics - Learn more: When memory runs short, phones sometimes kill a backgrounded game outright. That’s why the game starts over from scratch after a trip to the camera or a payment or authentication app. It’s more common on low-end devices. - Sources: - [Extending your app’s background execution time](https://developer.apple.com/documentation/uikit/extending-your-app-s-background-execution-time) · Apple · When the app goes to the background, applicationDidEnterBackground gets 5 s before the app is suspended; if it needs more, it requests time with beginBackgroundTask (time left in backgroundTimeRemaining) - [Cached apps freezer](https://source.android.com/docs/core/perf/cached-apps-freezer) · Android (Google) · Android 14 and later freezes cached app processes after 10 s; once frozen, all threads stop - [Application.runInBackground](https://docs.unity3d.com/ScriptReference/Application-runInBackground.html) · Unity · Defaults to false, so the game pauses in the background; Android pauses in the background regardless of the setting, and iOS ignores it - [MonoBehaviour.OnApplicationPause(bool)](https://docs.unity3d.com/ScriptReference/MonoBehaviour.OnApplicationPause.html) · Unity · Sends OnApplicationPause(true/false) to all MonoBehaviours when the app is paused or resumed - [ApplicationExitInfo](https://developer.android.com/reference/android/app/ApplicationExitInfo) · Android (Google) · REASON_LOW_MEMORY: the system’s low memory killer terminated the app process (devices that don’t support it report REASON_SIGNALED with SIGKILL) #### co-netswitch · Wi-Fi ↔ LTE/5G switching · Network switch changes IP When you walk out of the house and your phone drops Wi-Fi for LTE or 5G, your IP address changes and the existing connection stops working. - Why → Effect → On screen: The Wi-Fi signal weakens and the phone switches to the mobile network → Your IP address changes, so the connection made from the old address can’t carry any more data → A brief freeze, then a disconnect or a reconnect - Symptoms: Freeze, Disconnect / Factors: Packet loss - Who: Just me / When: While moving or changing zones - Primary owner: Game team (Server development) / Also: Game team (Client development), Infra team (Network infrastructure) - Game team action items: Server: resume the same player by session token even when the address changes, clean up the old address’s connection right away, consider a protocol that survives address changes (such as QUIC connection migration). Client: when a network switch is detected, reconnect right away with the session token without waiting for a heartbeat timeout. - Infra team action items: When using QUIC connection migration, configure the load balancer to choose the server by connection ID (choosing by address and port sends packets from a changed address to a different server). - On the graph: Mass disconnect (Disconnects and reconnects, IP changes on reconnect) - Where to look: Look in the server connection log for the same session token reconnecting from a different IP, and match the times against the client’s default network change callback (registerDefaultNetworkCallback) - Confirmed if: Right after the disconnect, the reconnecting IP moves from the Wi-Fi (home connection) range to a mobile carrier range or the other way around, and a network switch callback arrives just before - Ruled out if: Disconnected while the IP stayed the same: “Cell tower handover (while moving)” or “Weak mobile signal and dead zones” - Check with: Game server or client logs and metrics - Sources: - [Read network state](https://developer.android.com/training/basics/network-ops/reading-network-state) · Android (Google) · When the default network changes, new connections go over the new network and connections on the old network are eventually forced closed; registerDefaultNetworkCallback detects the switch - [RFC 9000: QUIC: A UDP-Based Multiplexed and Secure Transport](https://www.rfc-editor.org/rfc/rfc9000) · IETF · Connection IDs keep a connection alive even when the IP address or port changes (Section 9); a load balancer that distributes by address and port alone may send packets from a changed address to a different server (Section 5.2.3) - [RFC 9293: Transmission Control Protocol (TCP)](https://www.rfc-editor.org/rfc/rfc9293) · IETF · A TCP connection is identified by the pair of sockets (address and port) at its two ends #### co-security · Packet inspection by security software · Antivirus / firewall inspection When antivirus software or a firewall inspects every packet, latency goes up, and in bad cases it mistakes the game for an attack and blocks it. - Why → Effect → On screen: Security software inspects every packet sent and received, one by one → Each packet picks up delay, and packets get dropped when inspection falls behind → Ping spikes irregularly, or connections get blocked - Symptoms: Stutter, Can’t connect / infinite loading / Factors: Jitter, Packet loss - Who: Just me / When: Always, Right after login or maintenance - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Maintain a security software compatibility list, register a Windows Firewall exception for the game at install time. - External action items: Tell players to add the game as an exception in their security software; if it mistakes the game for an attack, ask the security vendor to fix the false positive. - Ballpark numbers: When everything works normally, packet inspection usually takes under 1 ms. The trouble starts when the inspection module falls behind or has a bug, or when it mistakes game traffic for an attack. - On the graph: Outliers only (RTT, connection failures (per player)) - Where to look: Compare after briefly turning off the security software or adding the game as an exception. On Windows, turning on Audit Filtering Platform Connection and Audit Filtering Platform Packet Drop in the audit policy logs 5157 (connection blocked) and 5152 (packet blocked) in the Security log, and Performance Monitor’s WFPv4\Packets Discarded/sec shows the number of discarded packets - Confirmed if: Block records show up for connections or packets to the game server’s address, or ping spikes and connection failures go away with the security software off - Ruled out if: Other devices in the same household behave the same way regardless of security software: the router or connection side - Check with: The player’s own environment - Sources: - [About Windows Filtering Platform](https://learn.microsoft.com/en-us/windows/win32/fwp/about-windows-filtering-platform) · Microsoft · Packets are allowed or blocked through hooks in the Windows network stack and a filter engine, and third-party security vendors can plug in their own filter modules (callouts) - [Windows Firewall Rules](https://learn.microsoft.com/en-us/windows/security/operating-system-security/network-security/windows-firewall/rules) · Microsoft · Inbound connections are blocked by default, so apps need exception rules, which the app installer usually creates - [Address false positives/negatives in Microsoft Defender for Endpoint](https://learn.microsoft.com/en-us/defender-endpoint/defender-endpoint-false-positives-negatives) · Microsoft · How to set an exclusion and submit a file to Microsoft for analysis when a legitimate program is mistaken for a threat (false positive) - [5157(F): The Windows Filtering Platform has blocked a connection.](https://learn.microsoft.com/en-us/previous-versions/windows/it-pro/windows-10/security/threat-protection/auditing/event-5157) · Microsoft · Event 5157: Windows Filtering Platform blocked a connection (Audit Filtering Platform Connection) - [Audit Filtering Platform Packet Drop](https://learn.microsoft.com/en-us/previous-versions/windows/it-pro/windows-10/security/threat-protection/auditing/audit-filtering-platform-packet-drop) · Microsoft · Event 5152: Windows Filtering Platform blocked a packet - [Network-Related Performance Counters](https://learn.microsoft.com/en-us/windows-server/networking/technologies/network-subsystem/net-sub-performance-counters) · Microsoft · WFPv4/WFPv6: Packets Discarded/sec counter #### co-rcvbuf · Receive buffer overflow · Socket receive buffer overflow If the game is busy and pulls packets out of the socket (the network send/receive interface the OS provides) late, the OS buffer overflows. - Why → Effect → On screen: Frames fall behind and the game reads the socket late → The OS receive buffer fills up: UDP packets get dropped, and TCP shrinks the receive window so the sender stops sending → Teleporting (UDP) or fast-forward (TCP) - Symptoms: Teleporting, Fast-forward / Factors: Packet loss, Stall - Who: Just me / When: When crowds gather - Primary owner: Game team (Client development) - Game team action items: Use a dedicated receive thread, tune the buffer size (SO_RCVBUF). - Ballpark numbers: The default receive buffer is tens to hundreds of KB, depending on the OS and settings. Updates in crowded places can reach hundreds of KB per second. - On the graph: Rises with load (UDP receive buffer drops, frame time) - Where to look: Record Microsoft Winsock BSP\Dropped Datagrams (UDP datagrams dropped for lack of socket receive buffer space) and UDPv4\Datagrams Received Errors in Windows Performance Monitor along with frame time; in the game, count gaps in the sequence numbers of received packets - Confirmed if: Dropped Datagrams rises in crowded places or right after a long frame, and gaps appear in the game’s sequence numbers at the same moment. No loss on the connection at the same time - Ruled out if: Dropped Datagrams flat but sequence numbers still go missing: loss along the path - Check with: The player’s own environment - Sources: - [socket(7) — Linux manual page](https://man7.org/linux/man-pages/man7/socket.7.html) · Linux man-pages · SO_RCVBUF is the maximum socket receive buffer size; the default comes from rmem_default and the maximum from rmem_max (Android also runs the Linux kernel) - [SOL_SOCKET Socket Options (Winsock2.h)](https://learn.microsoft.com/en-us/windows/win32/winsock/sol-socket-socket-options) · Microsoft · Windows SO_RCVBUF: buffer space reserved for receiving on each socket - [RFC 9293: Transmission Control Protocol (TCP)](https://www.rfc-editor.org/rfc/rfc9293) · IETF · The TCP window field is the number of bytes the receiver can still accept; at 0, the sender waits, sending only zero window probes - [Low Latency Workloads Management and Operations](https://learn.microsoft.com/en-us/previous-versions/windows/it-pro/windows-server-2012-R2-and-2012/hh997022(v=ws.11)) · Microsoft · Dropped Datagrams and Dropped Datagrams/sec in the Microsoft Winsock BSP counter set: UDP datagrams dropped because they arrived faster than the app could process them or the receive socket buffer was too small - [Network-Related Performance Counters](https://learn.microsoft.com/en-us/windows-server/networking/technologies/network-subsystem/net-sub-performance-counters) · Microsoft · UDPv4/UDPv6: Datagrams Received Errors; Microsoft Winsock BSP: Dropped Datagrams counter #### co-swap · Client low on memory and swapping · Paging / swap on client With dozens of browser tabs open alongside the game, the OS moves part of the game’s memory out to disk. - Why → Effect → On screen: The PC runs low on RAM overall → The OS moves game memory that isn’t in use right now to disk → The moment that memory is used again, the game freezes for tens to hundreds of ms, depending on storage - Symptoms: Freeze, Stutter / Factors: Stall - Who: Just me / When: While moving or changing zones, Randomly - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Reduce memory usage, show a warning when free memory is low. - External action items: Tell players the minimum specs and to close other programs (such as browser tabs) while playing. - On the graph: Random spikes (Hard page faults, memory usage) - Where to look: Record Memory\Pages Input/sec (pages read from disk to resolve hard page faults) in Performance Monitor and memory usage and commit on Task Manager’s Performance tab, together with frame time - Confirmed if: Pages Input/sec spikes at the moment of each freeze and memory is nearly full. Closing the browser and other programs makes it go away - Ruled out if: Memory has headroom and Pages Input/sec stays quiet: “Synchronous loading and shader compilation on the main thread” or “Slow storage delays asset streaming” - Check with: The player’s own environment - Sources: - [Introduction to the page file](https://learn.microsoft.com/en-us/troubleshoot/windows-client/performance/introduction-to-the-page-file) · Microsoft · The page file is a file on disk used to move rarely used, modified memory pages out of RAM - [Working Set](https://learn.microsoft.com/en-us/windows/win32/memory/working-set) · Microsoft · Touching a page that isn’t in RAM causes a page fault; a hard fault can only be resolved by reading from disk, such as from the page file - [Chapter 12 - Detecting Memory Bottlenecks](https://learn.microsoft.com/en-us/previous-versions/cc749872(v=technet.10)) · Microsoft · Memory\Pages Input/sec: pages read from disk to resolve page faults (hard page faults) #### co-vram · Out of graphics memory (VRAM) · VRAM over-commit When the graphics settings need more memory than the graphics card has, the OS moves textures out to system memory and brings them back, and the game stutters. - Why → Effect → On screen: High texture settings plus all the gear and effects in a crowded place fill up graphics card memory → The OS moves textures that aren’t in use right now to system memory, then brings them back over the slow PCIe bus when needed → A hitch every time a new scene or character comes into view; textures stay blurry for a while - Symptoms: Stutter, Freeze / Factors: Stall - Who: Just me / When: When crowds gather, While moving or changing zones - Primary owner: Game team (Client development) / Also: External (External) - Game team action items: Set default options to match graphics card memory size, lower texture quality automatically when the memory budget is exceeded, simplify character textures in crowded places. - External action items: Tell players to lower texture settings, and to lower them further when running two clients. - Ballpark numbers: Graphics card memory reads hundreds of GB per second, but the PCIe bus to system memory carries roughly 16–64 GB per second depending on the generation, more than ten times slower. - On the graph: Hits a ceiling (Dedicated GPU memory, shared GPU memory) - Where to look: Watch the Dedicated GPU memory and Shared GPU memory graphs under GPU on Task Manager’s Performance tab (per-process columns can also be added on the Details tab) alongside PresentMon frame time - Confirmed if: Hitches are frequent while dedicated GPU memory sits flat at its limit and shared GPU memory grows, and they go away when texture settings are lowered - Ruled out if: Dedicated memory has headroom: “Slow storage delays asset streaming” or “Synchronous loading and shader compilation on the main thread” - Check with: The player’s own environment - Learn more: If “Dedicated GPU memory” under GPU in Windows Task Manager is full and “Shared GPU memory” keeps growing, this is what’s happening. Running two clients on the same PC fills it faster (see “Streaming failure from memory or VRAM shortage”). - Sources: - [Residency](https://learn.microsoft.com/en-us/windows/win32/direct3d12/residency) · Microsoft · Each process has a graphics memory budget; when it’s exceeded, the kernel moves part of the discrete GPU’s heap to system memory (a last resort, so managing the budget is recommended) - [GPUs in the task manager](https://devblogs.microsoft.com/directx/gpus-in-the-task-manager/) · Microsoft · In Task Manager, dedicated GPU memory is the graphics card’s VRAM, and shared GPU memory is system memory used by both the GPU and the CPU - [CUDA C++ Best Practices Guide](https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html) · NVIDIA · Graphics memory bandwidth (V100: 898 GB/s) is far higher than PCIe 3.0 x16 (16 GB/s), so it recommends minimizing transfers to and from system memory - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · Records per-frame time as FrameTime (CPU time between frames) #### co-wifi-scan · Wi-Fi background scanning · Periodic Wi-Fi background scan Communication pauses briefly while the OS periodically hops across channels to look for nearby Wi-Fi networks. - Why → Effect → On screen: The OS or driver searches for nearby Wi-Fi networks on a fixed schedule → Sending and receiving pause briefly during the scan → Ping spikes at exactly regular intervals (e.g., every 60 s) - Symptoms: Stutter, Teleporting / Factors: Jitter - Who: Just me / When: At regular intervals - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Request a mode that reduces wireless scanning during play (on Android, the low-latency Wi-Fi mode WIFI_MODE_FULL_LOW_LATENCY; on Windows, the media streaming mode of WlanSetInterface; it may have no effect on some devices and drivers). - External action items: Tell players to use a wired connection, adjust location services and automatic Wi-Fi scanning settings, and update wireless drivers. - Ballpark numbers: Usually tens to hundreds of ms each time. If the spikes are suspiciously regular, suspect this first. - On the graph: Periodic spikes (RTT to the router) - Where to look: During play, ping the router address (Default Gateway in ipconfig) with ping /t for a few minutes and measure the interval between spikes. Repeat the same measurement over a wired connection - Confirmed if: Ping to the router spikes by tens to hundreds of ms at exactly regular intervals (e.g., 60 s) and goes away on a wired connection - Ruled out if: Irregular spike intervals: “Wi-Fi interference and weak signal.” Fine up to the router but spiking beyond it: the connection or ISP segment - Check with: The player’s own environment - Sources: - [WDI low latency connection quality](https://learn.microsoft.com/en-us/windows-hardware/drivers/network/wdi-low-latency-connection-quality) · Microsoft · Scanning and roaming move the radio off the connected channel, so low-latency mode limits scanning and time spent off-channel - [WlanSetInterface function (wlanapi.h)](https://learn.microsoft.com/en-us/windows/win32/api/wlanapi/nf-wlanapi-wlansetinterface) · Microsoft · Windows API that turns background scanning (wlan_intf_opcode_background_scan_enabled) and media streaming mode on and off - [Wi-Fi low-latency mode](https://source.android.com/docs/core/connect/wifi-low-latency) · Android (Google) · Low-latency mode turns off Wi-Fi power saving; how scanning and roaming settings are optimized depends on the device maker’s implementation - [ping](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/ping) · Microsoft · /t: keeps sending echo requests until stopped - [ipconfig](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/ipconfig) · Microsoft · Run with no parameters, it shows each adapter’s IPv4 and IPv6 addresses and default gateway #### co-driver · NIC power saving and driver issues · NIC power saving, driver bugs When a network card or Wi-Fi chip enters a power-saving state between packets, it takes time to wake back up. - Why → Effect → On screen: Network device power saving is on, or the driver is outdated → Wake-up delay, occasional device restarts → Irregular delays, occasional freezes lasting several seconds - Symptoms: Stutter, Freeze / Factors: Jitter, Packet loss - Who: Just me / When: After sitting idle, Randomly - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: On Android clients, request the low-latency Wi-Fi mode (WIFI_MODE_FULL_LOW_LATENCY) during play to turn off Wi-Fi power saving. - External action items: Tell players to update network drivers and turn off network device power saving in Device Manager; if the whole screen hitches and audio crackles, have them find the driver at fault with LatencyMon. - On the graph: Random spikes (DPC/ISR time, RTT to the router) - Where to look: Record with Windows Performance Recorder (WPR), find long-running drivers (Module column) in the DPC/ISR graph in Windows Performance Analyzer (WPA), and check the network adapter’s power management (power saving) settings in Device Manager - Confirmed if: At the hitch times, network driver DPCs and ISRs run for several ms at a stretch, or turning off power saving makes the irregular delays go away - Ruled out if: DPCs are short and turning off power saving changes nothing: “Wi-Fi interference and weak signal” or “Wi-Fi background scanning” - Check with: The player’s own environment - Learn more: When a driver holds the CPU for a long time handling interrupts (Windows calls this DPC latency), the game thread can’t use that core either. The whole screen then hitches and audio crackles even though CPU usage is low. A tool such as LatencyMon can find which driver is at fault; Wi-Fi and Ethernet drivers are common culprits. - Sources: - [Introduction to NDIS Selective Suspend](https://learn.microsoft.com/en-us/windows-hardware/drivers/network/ndis-selective-suspend) · Microsoft · Windows can put an idle network adapter into a low-power state (selective suspend) - [Guidelines for Writing DPC Routines](https://learn.microsoft.com/en-us/windows-hardware/drivers/kernel/guidelines-for-writing-dpc-routines) · Microsoft · All threads on that core stop while a DPC runs, so the guidance is to keep each one under 100 µs - [Wi-Fi low-latency mode](https://source.android.com/docs/core/connect/wifi-low-latency) · Android (Google) · In Android’s low-latency Wi-Fi mode, the framework explicitly turns off Wi-Fi power saving - [CPU Analysis](https://learn.microsoft.com/en-us/windows-hardware/test/wpt/cpu-analysis) · Microsoft · WPA’s DPC/ISR graph: the duration of each uninterrupted DPC or ISR run and the module containing that function (Module) #### co-other-apps · Other apps on the same device using up bandwidth · Other apps saturating the link When cloud sync, a large download, or a game patch runs on the same PC, game packets have to wait in a queue. - Why → Effect → On screen: Another app maxes out the upload or download → Game packets pile up in the queues on the PC and the router → Ping shoots up, input lag, fast-forward - Symptoms: Input lag, Fast-forward / Factors: Latency, Jitter - Who: Just me, Same household / When: Randomly - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Make our own launcher and patcher pause or throttle background downloads during play. - External action items: Tell players to cap download speeds and turn off automatic updates while playing. - On the graph: Rises with load (RTT, PC traffic sent and received) - Where to look: Record Network Interface\Bytes Sent/sec and Bytes Received/sec in Performance Monitor together with ping. Same approach as a bufferbloat test, where you deliberately start a large transfer with ping running - Confirmed if: Ping rises by tens to hundreds of ms while downloads or uploads run close to the connection speed, and drops back as soon as the transfer stops - Ruled out if: Ping rises while PC traffic is low: “Bufferbloat (router queue)” from another device in the same household, or the ISP segment - Check with: The player’s own environment - Sources: - [Introduction](https://www.bufferbloat.net/projects/bloat/wiki/Introduction/) · Bufferbloat.net · When network equipment such as a router buffers too much data, latency spikes sharply (bufferbloat) - [Delivery Optimization reference](https://learn.microsoft.com/en-us/windows/deployment/do/waas-delivery-optimization-reference) · Microsoft · Windows Update downloads (Delivery Optimization) adjust dynamically to available bandwidth by default, and caps can be set on background and foreground download bandwidth - [Network-Related Performance Counters](https://learn.microsoft.com/en-us/windows-server/networking/technologies/network-subsystem/net-sub-performance-counters) · Microsoft · Network Interface: Bytes Received/sec and Bytes Sent/sec counters - [Tests for Bufferbloat](https://www.bufferbloat.net/projects/bloat/wiki/Tests_for_Bufferbloat/) · Bufferbloat.net · If ping rises while a speed test saturates the connection with ping running, it’s bufferbloat #### co-unfocused · Throttling when the window is minimized or unfocused · Minimized / unfocused window throttling When you switch to another window or minimize the game, the game and Windows slow it down to save power. When you come back, the backlog of packets floods in, or you’ve already been disconnected. - Why → Effect → On screen: Switching to another window with Alt+Tab, or minimizing the game → While it isn’t visible, the game lowers FPS sharply or pauses, and Windows also lowers the priority of programs that aren’t visible → Fast-forward the moment you return; a disconnect if the game stayed minimized for a long time - Symptoms: Fast-forward, Stutter, Disconnect / Factors: Stall - Who: Just me / When: During specific actions, After sitting idle - Primary owner: Game team (Client development) - Game team action items: Keep receiving packets and sending heartbeats on a separate thread even when the window is hidden, check the engine’s “Run in background” setting, catch up to the latest state in one go on return. - Ballpark numbers: If FPS drops to 5–10 while the window is hidden, each frame takes 100–200 ms. A game that processes packets once per frame reads them that much later. - On the graph: Gap then burst (Frame interval (before and after switching windows), packets processed) - Where to look: With PresentMon running, try Alt+Tab and minimizing, and look at frame intervals while the window is hidden. Log window focus changes in the game log and match them against disconnect reasons - Confirmed if: Frame intervals stretch past 100 ms or recording stops while the window is hidden, and the moment you return, the game processes the packet backlog all at once and fast-forwards. Left minimized for long, it disconnects on a heartbeat timeout - Ruled out if: Same behavior with the window in front: “Background processes taking up CPU” or the network side - Check with: The player’s own environment - Learn more: Windows 11 doesn’t guarantee a 1 ms timer for programs whose windows are minimized or fully hidden and not playing sound. On a laptop running on battery, it slows such programs to the most power-efficient speed, and on CPUs with mixed core types it may run them on the slower efficiency cores. If only the background one of two clients on the same PC misbehaves, also see “Background window throttling.” - Sources: - [Quality of Service](https://learn.microsoft.com/en-us/windows/win32/procthread/quality-of-service) · Microsoft · Programs whose windows can’t be seen or heard get Low QoS and, on battery, are scheduled at the most efficient CPU speed and on efficiency cores - [timeBeginPeriod function (timeapi.h)](https://learn.microsoft.com/en-us/windows/win32/api/timeapi/nf-timeapi-timebeginperiod) · Microsoft · Windows 11 doesn’t guarantee a higher-than-default timer resolution to processes whose windows are hidden or minimized - [Application.runInBackground](https://docs.unity3d.com/ScriptReference/Application-runInBackground.html) · Unity · Unity defaults to false, so the game loop stops when the window goes to the background - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · MsBetweenPresents: time between this Present() call and the previous one (ms) #### co-overlay · Overlay software interference · Overlays and screen hooks Chat apps, launchers, recording tools, and FPS counters hook into the game’s rendering to draw their own UI on top of the game screen (hooking). That adds work to every frame and sometimes clashes with the game, causing hitches or crashes. - Why → Effect → On screen: Overlays from chat apps, game launchers, graphics card tools, or recording software are turned on → Every time a frame goes out to the screen, the overlay steps in and draws its own UI on top → Frames get slightly later, and when a notification pops up the game hitches, shows graphics glitches, or crashes (looks like a disconnect to the player) - Symptoms: Stutter, Freeze, Disconnect / Factors: Stall - Who: Just me / When: Always, Randomly - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Collect the list of running overlays along with crash reports and stutter logs. - External action items: When reports come in, tell players to turn off all overlays and test again. - On the graph: Outliers only (Frame time and crash count (players with overlays on)) - Where to look: Turn off all overlays and compare PresentMon frame time in the same scene; for crashes, check the faulting module name (Faulting module name) in Event Viewer Event ID 1000 - Confirmed if: Hitches and graphics glitches go away with overlays off, or the faulting module in the crash is an overlay program’s DLL - Ruled out if: Same with every overlay off: the graphics driver or “Client crash” - Check with: The player’s own environment - Learn more: When only certain players stutter or crash and their specs don’t explain it, suspect a conflict between an overlay and the anti-cheat module first. - Sources: - [Steam Overlay (Steamworks Documentation)](https://partner.steamgames.com/doc/features/overlay) · Valve · The Steam overlay automatically hooks into games launched through Steam, and because of how it does that, it can expose memory errors in the game’s use of the rendering API and cause crashes - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · Records per-frame time as FrameTime (CPU time between frames) - [The application or service crashing behavior troubleshooting guidance](https://learn.microsoft.com/en-us/troubleshoot/windows-server/performance/troubleshoot-application-service-crashing-behavior) · Microsoft · Event ID 1000 in the Application log includes the faulting module name (Faulting module name); a Windows module sometimes shows up as the faulting module because of corruption caused by another module #### co-display-input · Display, input device, and frame generation latency · Display, input device and frame generation latency If ping is normal but controls feel heavy, a TV’s video processing, a wireless controller, or frame generation may be adding delay between your input and the screen. - Why → Effect → On screen: The TV’s game mode is off, a Bluetooth or wireless controller is in use, or frame generation (DLSS or FSR frame generation) is on → The TV sends frames out late while it processes the picture, wireless input arrives late by its polling interval plus any interference, and frame generation waits for the next frame to create an in-between frame → Ping and FPS numbers look good, but there’s a delay between pressing a button and seeing the result on screen: input lag - Symptoms: Input lag / Factors: Latency - Who: Just me / When: Always - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Make frame generation optional and warn that turning it on can increase input lag, integrate the GPU vendor’s low-latency feature (NVIDIA Reflex, AMD Anti-Lag 2) when frame generation is used, show the PC-side input-to-display latency in the game, request the TV’s low-latency mode (ALLM) with Window.setPreferMinimalPostProcessing(true) in Android TV and set-top box builds. - External action items: Tell players to turn on game mode (ALLM) on their TV or monitor, use a wired controller and turn off frame generation for competitive play, keep Bluetooth devices close, and use 5 GHz Wi-Fi. - Ballpark numbers: Just sending one frame takes 16.7 ms on a 60 Hz screen and 8.3 ms at 120 Hz. Older Xbox controllers read and sent input every 8 ms. The delay added by a TV’s video processing varies by device, so no single number fits; game mode is the setting that cuts down this processing. AMD recommends using frame generation at a frame rate of 60 FPS or higher before generation. - On the graph: Always high (Input-to-display latency) - Where to look: Compare PresentMon’s MsAllInputToPhotonLatency (from keyboard or mouse input to output on screen) with frame generation on and off, and use FrameType (recorded only when the driver or SDK reports it) to see whether generated in-between frames are mixed in. This value doesn’t include the controller’s wireless link or the processing inside the TV, so compare those parts by switching to TV game mode or a wired controller - Confirmed if: Ping is normal, but input-to-display latency drops with frame generation off, or the perceived delay goes away with TV game mode or a wired controller - Ruled out if: No change after switching all of these, and ping is high or spiking: the network side. PC-side latency is high because of V-Sync or the frame queue: “V-Sync and the render queue” - Check with: The player’s own environment - Learn more: Network latency shows up in ping, but this latency doesn’t. That’s why it’s the first thing to check in a “low ping but still lagging” report. Frame generation roughly doubles the FPS number on screen, but to create an in-between frame it has to wait for the next real frame, so the time until your input shows up on screen gets longer (AMD states that latency increases by design). Bluetooth devices use the same 2.4 GHz band as Wi-Fi, so interference can make input drop out or jump. For V-Sync and the render queue, which add latency inside the PC, see “V-Sync and the render queue.” - Sources: - [Auto Low Latency Mode (ALLM)](https://www.hdmi.org/spec21sub/autolowlatencymode) · HDMI Licensing Administrator · ALLM lets a device switch the display to low-latency mode (often called game mode) automatically; in low-latency mode, the TV stops some video processing to reduce delay - [Xbox Series X: What’s the Deal with Latency?](https://news.xbox.com/en-us/2020/03/16/xbox-series-x-latency/) · Microsoft · Input lag is the sum of the path controller → console → HDMI → TV; older controllers read and sent input every 8 ms; sending one frame over HDMI takes 16.6 ms at 60 Hz and 8.3 ms at 120 Hz; ALLM switches the TV to game mode automatically - [AMD FSR 3 game integrations out now + more details for developers](https://gpuopen.com/news/fsr3-in-games-technical-details/) · AMD · Frame interpolation increases latency by design; recommended at a frame rate of 60 or higher before interpolation; outputs up to 120 FPS from 60 FPS input - [AMD FSR Frame Generation](https://gpuopen.com/amd-fsr-framegeneration/) · AMD · Frame generation is recommended at 60 FPS or higher before interpolation (below 30 FPS should be avoided); AMD Radeon Anti-Lag 2 reduces system latency by keeping CPU and GPU work in step - [NVIDIA DLSS](https://developer.nvidia.com/rtx/dlss) · NVIDIA · DLSS Frame Generation is designed to keep responsiveness when paired with NVIDIA Reflex (a low-latency feature) - [Resolve Wi-Fi and Bluetooth issues caused by wireless interference](https://support.apple.com/en-us/102319) · Apple · Wireless interference causes dropouts and poor performance on Wi-Fi and Bluetooth devices, and Bluetooth and Wi-Fi use the same 2.4 GHz band - [PresentMon Capture Application (README-CaptureApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-CaptureApplication.md) · Intel · MsAllInputToPhotonLatency (input-to-display latency), DisplayLatency (frame submission to output to the monitor), FrameType (tells frames rendered by the app apart from frames interpolated by the driver or SDK) - [PresentMon Console Application (README-ConsoleApplication.md)](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README-ConsoleApplication.md) · Intel · MsAllInputToPhotonLatency is based on keyboard and mouse input; FrameType is recorded only if the app or driver emits Intel-PresentMon events (--track_frame_type) - [Window.setPreferMinimalPostProcessing](https://developer.android.com/reference/android/view/Window#setPreferMinimalPostProcessing(boolean)) · Android (Google) · Latency-sensitive windows such as games request minimal video processing from the display; over HDMI, the ALLM and Game Content Type signals switch the TV to low-latency mode ### L3 Home network (causes: 10) #### hn-wifi · Wi-Fi interference and weak signal · Wi-Fi interference, weak signal With a weak signal or interference, packets get resent several times over the wireless link, so they arrive unevenly. - Why → Effect → On screen: Walls, distance, microwaves, Bluetooth, and neighbors’ routers degrade the radio signal → Transmissions fail on the wireless link → resent several times → Packets arrive unevenly (jitter), so characters move in fits and starts; with heavy loss, they teleport - Symptoms: Stutter, Teleporting, Rubber-banding / Factors: Jitter, Packet loss - Who: Just me, Same household / When: Randomly, Always - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Adjust interpolation buffer length automatically to connection quality, show network status on screen when jitter or loss is high. - External action items: Tell players to use a wired connection, switch to 5 GHz or 6 GHz, and move the router. - Ballpark numbers: Each retransmission adds about 1–4 ms. With a weak signal, the radio resends several times at a low rate and also waits for the channel to clear, so latency can spike by 50–200 ms. The trap is that average ping looks fine. - On the graph: Random spikes (RTT to the router) - Where to look: Ping the router address (Default Gateway in ipconfig) with ping /t for a few minutes, and check your router’s signal strength and channel with netsh wlan show networks mode=bssid. Compare over a wired connection from the same spot - Confirmed if: Even ping to the router spikes irregularly by tens to hundreds of ms with occasional loss, and signal strength is low. Gone on a wired connection or close to the router - Ruled out if: Steady up to the router but spiking beyond it: the connection or ISP segment. Spikes only at a fixed interval: “Wi-Fi background scanning” - Check with: The player’s own environment - Learn more: In mesh Wi-Fi, when the routers (nodes) link to each other wirelessly (wireless backhaul), a relaying node can’t send while it’s receiving and shares transmit opportunities with the hops before and after it on the same channel, so throughput can drop and latency can rise when the network is busy. Products with a dedicated wireless backhaul band may suffer less, and wiring the nodes together (Ethernet) takes this hop off the air. Powerline adapters (PLC) also check that the medium is free before sending (CSMA/CA), like Wi-Fi, and their quality changes constantly with electrical noise from appliances and with appliances switching on and off, which can cause retransmissions and jitter. - Real incidents: ffxiv-2021 - Sources: - [Resolve Wi-Fi and Bluetooth issues caused by wireless interference](https://support.apple.com/en-us/102319) · Apple · Interference sources such as microwaves and cordless phones; Wi-Fi and Bluetooth share the 2.4 GHz band; recommends moving to 5 GHz and choosing a channel with less interference - [RFC 8325: Mapping Diffserv to IEEE 802.11](https://www.rfc-editor.org/rfc/rfc8325) · IETF · 802.11 CSMA/CA: sends only when the channel is clear; if it’s busy, defers until it clears, then waits an additional random backoff - [Ending the Anomaly: Achieving Low Latency and Airtime Fairness in WiFi (USENIX ATC 2017)](https://www.usenix.org/system/files/conference/atc17/atc17-hoiland-jorgensen.pdf) · USENIX · The default queue on a loaded Wi-Fi network adds hundreds of ms of latency; devices connected at low rates (weak signal) exceed a 200 ms median even with FQ-CoDel - [ping](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/ping) · Microsoft · /t: keeps sending echo requests until stopped - [ipconfig](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/ipconfig) · Microsoft · Run with no parameters, it shows each adapter’s IPv4 and IPv6 addresses and default gateway - [Wireless network connectivity issues troubleshooting](https://learn.microsoft.com/en-us/troubleshoot/windows-client/networking/wireless-network-connectivity-issues-troubleshooting) · Microsoft · netsh wlan show networks mode=bssid: shows the BSSID, signal strength, channel, and radio type of every visible Wi-Fi network - [Capacity of Ad Hoc Wireless Networks (MobiCom 2001)](https://pdos.csail.mit.edu/papers/grid:mobicom01/paper.pdf) · ACM · When 802.11 relays over several wireless hops, a node can’t send while it’s receiving and adjacent hops interfere with each other, so the throughput of a chain of relays can drop to 1/3 in theory (about 1/7 in simulation) - [Electri-Fi Your Data: Measuring and Combining Power-Line Communications with WiFi (IMC 2015)](https://conferences.sigcomm.org/imc/2015/papers/p325.pdf) · ACM · Commercial powerline equipment (IEEE 1901, HomePlug AV) sends with CSMA/CA much like Wi-Fi, and short-term unfairness can increase jitter; channel quality changes with electrical noise from appliances and with appliances switching on and off (over minutes to hours) #### hn-channel · Congested Wi-Fi channel · Crowded Wi-Fi channel Where there are dozens of routers, as in an apartment building, they share the same channel and have to wait for a chance to transmit. - Why → Effect → On screen: Dozens of routers use the same 2.4 GHz channel → Before transmitting, a device waits until other devices finish and the channel clears → In the evening, when people get home, jitter (variation in packet arrival times) rises and the game stutters - Symptoms: Stutter, Input lag / Factors: Jitter, Latency - Who: Same household / When: Evening peak hours - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Lengthen the interpolation buffer automatically when jitter rises. - External action items: Tell players to use 5 GHz or 6 GHz, a less crowded channel, or a wired connection. - On the graph: High at certain hours (RTT and jitter to the router) - Where to look: Check the channels and signal strength of nearby Wi-Fi networks with netsh wlan show networks mode=bssid, and compare ping to the router in the evening and during the day - Confirmed if: Many nearby routers show up on the same 2.4 GHz channel, and jitter to the router grows only in the evening. Moving to 5 GHz or 6 GHz or to a less crowded channel reduces it - Ruled out if: Spikes regardless of time of day: “Wi-Fi interference and weak signal.” Fine up to the router but bad beyond it in the evening: “Peak-hour congestion at peering links” - Check with: The player’s own environment - Sources: - [Recommended settings for Wi-Fi routers and access points](https://support.apple.com/en-us/102766) · Apple · Other routers and devices on the same channel are sources of interference; 20 MHz channel width is recommended on 2.4 GHz; interference is less of a concern on 5 GHz and 6 GHz - [RFC 8325: Mapping Diffserv to IEEE 802.11](https://www.rfc-editor.org/rfc/rfc8325) · IETF · 802.11 defers transmission while the channel is busy and sends after a random backoff once it clears (CSMA/CA) - [Wireless network connectivity issues troubleshooting](https://learn.microsoft.com/en-us/troubleshoot/windows-client/networking/wireless-network-connectivity-issues-troubleshooting) · Microsoft · netsh wlan show networks mode=bssid: shows the BSSID, signal strength, channel, and radio type of every visible Wi-Fi network - [ping](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/ping) · Microsoft · /t: keeps sending echo requests until stopped #### hn-bufferbloat · Bufferbloat (router queue) · Bufferbloat When someone in the household uploads a video or downloads a large file, hundreds of ms worth of packets pile up in the router’s queue, and game packets wait behind them. - Why → Effect → On screen: The connection fills up with a family member’s video upload or cloud backup, your own live stream, or a large download → The router or modem holds the overflow of packets in a large queue → Game packets wait at the back of the queue too, and ping shoots up to hundreds of ms - Symptoms: Input lag, Fast-forward, Teleporting / Factors: Latency, Jitter - Who: Same household, Just me / When: Randomly, Evening peak hours - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Show network status on screen when ping suddenly jumps to hundreds of ms (mention a possible large transfer on the same connection). - External action items: Tell players to use a router with SQM (fq_codel, CAKE) or QoS, set the SQM speed to 90–95% of the connection speed (so the queue forms inside the router and SQM can take effect), and cap upload speeds. - Ballpark numbers: On a connection with 10 Mbps upload, a 1 MB buffer lets the queue grow to 800 ms. - On the graph: Rises with load (RTT, upload and download usage on the connection) - Where to look: With ping running, saturate the connection with a speed test, or use a web test that measures latency under load (see Bufferbloat.net). Check alongside the upload and download usage shown on the router’s admin page - Confirmed if: Ping climbs to hundreds of ms while an upload or download saturates the connection and recovers when the transfer ends (suspect it when latency under load exceeds 50 ms). Gone with SQM on - Ruled out if: Ping spikes even when the connection is idle: “Wi-Fi interference and weak signal” or “Poor line quality” - Check with: The player’s own environment - Learn more: The upload side clogs especially easily, because cable and mobile connections often have far less upload bandwidth than download. In homes with plenty of fiber bandwidth, the Wi-Fi link becomes the bottleneck, and the same thing happens in the router’s wireless queue. Game packets are small and use almost no bandwidth, but they still have to wait in the queue. If only the upload direction is clogged, only your own input is late, while other players’ movement looks fine. On phones, photo backups and app updates on the same phone fill the queues in the phone’s modem and at the cell tower, with the same result. - Sources: - [Setting up SQM for CeroWrt 3.10](https://www.bufferbloat.net/projects/cerowrt/wiki/Setting_up_SQM_for_CeroWrt_310/) · Bufferbloat.net · Set SQM to 95% of the measured speed (85% if based on the advertised speed) to move the bottleneck from the ISP’s equipment into the router; that’s what makes it work - [SQM (Smart Queue Management)](https://openwrt.org/docs/guide-user/network/traffic-shaping/sqm) · OpenWrt · Enter 90% of the measured download and upload speeds; cake is the recommended queue discipline (fq_codel if the CPU is weak) - [Ending the Anomaly: Achieving Low Latency and Airtime Fairness in WiFi (USENIX ATC 2017)](https://www.usenix.org/system/files/conference/atc17/atc17-hoiland-jorgensen.pdf) · USENIX · When the Wi-Fi link is saturated, the router’s wireless queue adds hundreds of ms of latency; fixing the wireless queue cuts latency under load to about a tenth - [Tests for Bufferbloat](https://www.bufferbloat.net/projects/bloat/wiki/Tests_for_Bufferbloat/) · Bufferbloat.net · If ping rises while a speed test saturates the connection with ping running, it’s bufferbloat; a fix is recommended when latency under load exceeds 50 ms (or the grade is below B) #### hn-nat · NAT mapping expiry · NAT mapping timeout Routers remove idle connections that haven’t carried packets for a while from their NAT table. It’s a common reason the connection drops the moment you move after standing still. - Why → Effect → On screen: The router records the “inside device ↔ outside server” connection in its NAT table (address translation table) → If no packets pass for a while, the entry is deleted (for UDP, often after 30–120 s) → Server packets can no longer get into the home, so the connection drops - Symptoms: Disconnect / Factors: Packet loss - Who: Just me, Same household / When: After sitting idle - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: send heartbeats at no more than half the shortest idle timeout (UDP mappings are reliably refreshed only by packets going out of the home, so the client sends them), reconnect automatically after a disconnect. Server: answer heartbeats, clean up the connection first when none arrive for a set time, and when a deleted mapping changes the outside address and port, confirm it’s the same player with the session token (an ID issued when the player connects) and resume the session. - On the graph: Mass disconnect (Disconnects (heartbeat timeouts), idle time before disconnect) - Where to look: Collect the server’s disconnect reasons and the time elapsed since the last packet on that connection before it dropped (idle time), and look at the distribution. To test, increase the UDP packet interval to 30 s, 60 s, and 120 s and find the interval at which responses stop - Confirmed if: Only idle connections drop, and idle times cluster just past a specific value such as 30–120 s. A heartbeat interval shorter than that makes it go away - Ruled out if: Drops even while moving: the connection or route side. Clustered at a short value only on a specific mobile carrier: “ISP-shared IP addresses (CGNAT)” - Check with: Game server or client logs and metrics - Sources: - [RFC 4787: Network Address Translation (NAT) Behavioral Requirements for Unicast UDP](https://www.rfc-editor.org/rfc/rfc4787) · IETF · The UDP mapping timer must not expire in under 2 minutes, and a default of 5 minutes or more is recommended; refreshing on outbound packets is required, refreshing on inbound packets is optional - [An Experimental Study of Home Gateway Characteristics (IMC 2010)](https://conferences.sigcomm.org/imc/2010/papers/p260.pdf) · ACM · Measurements of 34 home routers: UDP mappings lasted 30–691 s with a 90 s median, over half under 2 minutes; TCP median about 60 minutes - [RFC 9308: Applicability of the QUIC Transport Protocol](https://www.rfc-editor.org/rfc/rfc9308) · IETF · On an internet path with NATs, keep-alives about every 30 s are reasonable; more frequent ones waste traffic and power #### hn-router · Underpowered or overheating router · Router CPU / session table exhaustion When dozens of devices and thousands of connections pile onto a cheap router, the router itself can’t keep up. - Why → Effect → On screen: Dozens of devices, plus P2P and torrent clients opening thousands of connections → The router’s CPU and session table are saturated → Delayed and lost packets, failed new connections - Symptoms: Stutter, Can’t connect / infinite loading, Disconnect / Factors: Packet loss, Jitter - Who: Same household / When: The longer it runs, Randomly - Primary owner: External (External) - External action items: Tell players to reboot the router (temporary fix), replace the router, and clean up programs that open many connections (P2P, torrents). - On the graph: Hits a ceiling (Router CPU and connection count, RTT to the router) - Where to look: Check CPU usage, connection (session) count, and number of connected devices on the router’s admin page (if the router supports it), and compare ping to the router itself before and after a reboot - Confirmed if: When connection counts are high, even ping to the router spikes or loses packets and new connections fail. After a reboot it’s fine for a while, then gets worse again - Ruled out if: Fine up to the router but bad beyond it: the connection or ISP segment - Check with: The player’s own environment - Sources: - [An Experimental Study of Home Gateway Characteristics (IMC 2010)](https://conferences.sigcomm.org/imc/2010/papers/p260.pdf) · ACM · Home routers allow 16 to about 1,024 TCP connections to a single server port (median 135), and low-end devices may manage only a few Mbps of throughput - [Netfilter Conntrack Sysfs variables](https://docs.kernel.org/networking/nf_conntrack-sysctl.html) · Linux kernel · The maximum number of entries in Linux’s connection tracking table (nf_conntrack_max) and the default timeouts for each connection state #### hn-handover · Cell tower handover (while moving) · Cellular handover When you travel by bus or subway, the connection drops out while your phone switches cell towers. - Why → Effect → On screen: The phone switches to a different cell tower while on the move → Usually a gap of tens of ms, but if the signal is bad and the switch fails, it can drop out for hundreds of ms to several seconds → A freeze, then teleporting; if it lasts long, a disconnect - Symptoms: Freeze, Teleporting, Disconnect / Factors: Packet loss - Who: Just me / When: While moving or changing zones - Primary owner: External (External) / Also: Game team (Client development), Game team (Server development) - Game team action items: Client: use timeouts that tolerate short dropouts, reconnect quickly. Server: use timeouts that don’t kick players right away after a few seconds of silence, resume the same session on reconnect. - External action items: Tell players that disconnects while traveling (bus, subway) are caused by cell tower switching. - On the graph: Gap then burst (Packets received, RTT) - Where to look: Check whether the disconnect report came from someone traveling (bus, subway), and look at the receive gap times in the client log together with changes in network type and signal - Confirmed if: Only while traveling, receiving stops for hundreds of ms to several seconds and then packets arrive in a rush; it doesn’t reproduce when standing still - Ruled out if: Same when standing still: “Weak mobile signal and dead zones” or “Frequent 5G↔LTE switching (at 5G coverage edges)” - Check with: The player’s own environment - Sources: - [Report ITU-R M.2134: Requirements related to technical performance for IMT-Advanced radio interface(s)](https://www.itu.int/pub/R-REP-M.2134-2008) · ITU · Required limit on the time no data can be exchanged during a handover: 27.5 ms on the same frequency, 40–60 ms on a different frequency - [Understanding Operational 5G: A First Measurement Study on Its Coverage, Performance and Energy Consumption (SIGCOMM 2020)](https://xyzhang.ucsd.edu/papers/DXu_SIGCOMM20_5Gmeasure.pdf) · ACM · Handover delays measured on commercial networks: 4G↔4G average 30 ms, between 5G (NSA) cells average 108 ms #### hn-rrc · RRC state transition delay (mobile radio power saving) · Radio state promotion (RRC) When a phone has no traffic for a while, it drops its radio connection to a low-power state, and the next packet is delayed while it powers back up. - Why → Effect → On screen: After a short period with no traffic, the phone puts its radio connection into a power-saving state → To send the next packet, the connection has to be brought back up → Only the first action after standing idle is noticeably slow - Symptoms: Input lag / Factors: Latency - Who: Just me / When: After sitting idle - Primary owner: Game team (Client development) - Game team action items: Keep the radio active with light periodic traffic (at a cost in battery life). - Ballpark numbers: LTE usually drops to power saving after about 10 seconds with no traffic, and coming back takes tens to hundreds of ms (measured example: about 0.3–0.6 s). On 3G it’s over 1 second. - On the graph: Outliers only (RTT of the first request after idle (mobile)) - Where to look: Break down in-game RTT by the gap since the previous traffic. On mobile, compare the RTT of the first packet sent after more than 10 s of idle with that of packets sent back to back - Confirmed if: On a mobile network, only the first packet after idle is hundreds of ms late, and packets sent right after it are normal. No difference on Wi-Fi - Ruled out if: Late even when sent back to back: the signal, connection, or route side - Check with: Game server or client logs and metrics - Sources: - [An In-depth Study of LTE: Effect of Network Protocol and Application Behavior on Performance (SIGCOMM 2013)](https://conferences.sigcomm.org/sigcomm/2013/papers/sigcomm/p363.pdf) · ACM · On the measured LTE network: power-saving transition timer (tail) 10 s, median delay to come back from power saving 435 ms (25th–75th percentile: 319–558 ms); 3G about 1.5–2 s - [Optimize network access](https://developer.android.com/develop/connectivity/network-ops/network-access-optimization) · Android (Google) · Radio state transition delay and tail time vary with the radio technology (3G, LTE, 5G) and carrier settings; 3G example: low power → full power about 1.5 s, idle → full power over 2 s - [Report ITU-R M.2134: Requirements related to technical performance for IMT-Advanced radio interface(s)](https://www.itu.int/pub/R-REP-M.2134-2008) · ITU · Required control-plane latency from idle to active state: under 100 ms (excluding paging and the wired segment) #### hn-weak-cell · Weak mobile signal and dead zones · Weak cellular signal In elevators, basements, and deep inside buildings, retransmissions increase, speed drops, and eventually the connection drops. - Why → Effect → On screen: Moving into a place with weak signal → More radio retransmissions, lower speed, momentary dropouts → Jitter and loss cause stutter and teleporting, and eventually a disconnect - Symptoms: Stutter, Teleporting, Disconnect / Factors: Jitter, Packet loss - Who: Just me / When: While moving or changing zones - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Polish the reconnect flow, show network quality. - External action items: Tell players the problem happens in places with weak signal (elevators, basements, deep inside buildings). - On the graph: Outliers only (RTT and loss (per mobile player)) - Where to look: Check where the player was when the disconnect was reported (elevator, basement, inside a building) and the phone’s signal indicator, and compare by repeating the same actions where the signal is good - Confirmed if: RTT and loss rise and the connection drops only where the signal is weak, and it goes away after moving to a place with good signal - Ruled out if: Same even with good signal: the ISP segment or the server side - Check with: The player’s own environment - Sources: - [An In-depth Study of LTE: Effect of Network Protocol and Application Behavior on Performance (SIGCOMM 2013)](https://conferences.sigcomm.org/sigcomm/2013/papers/sigcomm/p363.pdf) · ACM · LTE hides radio link loss with physical and MAC layer retransmissions; available bandwidth swings widely from second to second with signal strength and other factors #### hn-5g-flip · Frequent 5G↔LTE switching (at 5G coverage edges) · 5G NSA / LTE switching Inside buildings with weak 5G signal or at the edge of 5G coverage, the phone switches between 5G and LTE often, and each switch causes a ping spike or a brief dropout. - Why → Effect → On screen: In a place where the 5G signal comes and goes (inside a building, at the edge of 5G coverage) → The phone keeps switching between 5G and LTE, with a short gap each time → Ping spikes at random even when standing still, with occasional freezes and teleporting - Symptoms: Stutter, Teleporting, Freeze / Factors: Jitter, Packet loss - Who: Just me / When: Randomly, While moving or changing zones - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Lengthen the interpolation buffer automatically when jitter rises; record network type changes (5G, LTE) in the logs taken when lag happens, to tell causes apart. - External action items: Tell players to switch to LTE-preferred mode in settings and compare, recommend using Wi-Fi. - Ballpark numbers: Each switch takes tens to hundreds of ms. Most 5G in Korea runs bundled with LTE (NSA), so the 5G part easily connects and drops over and over. - On the graph: Random spikes (RTT, network type changes (5G/LTE)) - Where to look: Switch the phone to LTE-preferred mode and compare from the same spot. It’s more conclusive if the client logs changes in the network indicator from Android’s TelephonyDisplayInfo (OVERRIDE_NETWORK_TYPE_NR_NSA and so on) together with RTT - Confirmed if: RTT spikes line up with the times the 5G↔LTE indicator changes, and the spikes disappear in LTE-preferred mode - Ruled out if: Spikes even though the network indicator doesn’t change: “Weak mobile signal and dead zones” or the connection side - Check with: The player’s own environment - Sources: - [Understanding Operational 5G: A First Measurement Study on Its Coverage, Performance and Energy Consumption (SIGCOMM 2020)](https://xyzhang.ucsd.edu/papers/DXu_SIGCOMM20_5Gmeasure.pdf) · ACM · NSA 5G leaves control to LTE, so when changing 5G cells the phone drops 5G, goes through LTE, and reattaches: average 108 ms (4G→5G 80 ms); TCP throughput falls 73–83% right after a handover involving 5G - [5G 통신서비스 품질평가 결과 발표](https://www.korea.kr/briefing/policyBriefingView.do?newsId=156404679) · 과학기술정보통신부 · As announced in 2020, 5G in Korea is offered in NSA mode, with the move to SA still planned - [TelephonyDisplayInfo](https://developer.android.com/reference/android/telephony/TelephonyDisplayInfo) · Android (Google) · OVERRIDE_NETWORK_TYPE_NR_NSA: the network indicator shown when the device is on LTE and can use, or is using, dual connectivity (EN-DC) with 5G (NR) #### hn-captive · Public Wi-Fi and corporate network restrictions · Captive portal, restrictive network A café Wi-Fi login page or a corporate firewall blocks the game’s connection. - Why → Effect → On screen: The login page hasn’t been completed yet, or a firewall blocks the game’s ports or UDP → Connection attempts are blocked outright, or only some traffic gets through → Can’t connect, or login works but entering the game fails - Symptoms: Can’t connect / infinite loading / Factors: Packet loss - Who: Just me / When: Right after login or maintenance - Primary owner: External (External) / Also: Game team (Client development), Game team (Server development) - Game team action items: Client: tell players why the connection is blocked (login page not completed, UDP blocked, and so on), switch to a fallback path automatically when UDP is blocked. Server: provide a fallback path such as TCP 443. - External action items: Tell players to complete the login page first on public Wi-Fi and to use a different network where traffic is restricted, such as a corporate network. - On the graph: Outliers only (Connection failures (by network)) - Where to look: Have the failing player try connecting from another network such as mobile data, and check the server connection log for whether the first UDP packet arrived and whether the TCP 443 fallback path connects - Confirmed if: Fails only on specific Wi-Fi (café, office) and connects right away on other networks. The login page hasn’t been completed, or only UDP fails to reach the server - Ruled out if: Fails on every network: the account, the server, or “DNS failures and delays.” Fails across an entire country or ISP: “Country- or ISP-level UDP restrictions and packet inspection” - Check with: The player’s own environment - Sources: - [RFC 8952: Captive Portal Architecture](https://www.rfc-editor.org/rfc/rfc8952) · IETF · Captive portal: a network that restricts access until requirements such as accepting terms or authenticating are met - [RFC 9308: Applicability of the QUIC Transport Protocol](https://www.rfc-editor.org/rfc/rfc9308) · IETF · Measurement studies show 3–5% of networks block UDP entirely, so UDP-based apps need a TCP (TLS) fallback path ### L4 Internet path (causes: 14) #### isp-distance · Propagation delay (physical distance) · Propagation delay Even light travels only about 200,000 km per second in optical fiber. A distant server is slow no matter how good it is. - Why → Effect → On screen: The server is far away (an overseas server, another continent) → Round-trip time grows with distance (at least 10 ms per 1,000 km) → Constant input lag on every action and a disadvantage in hit registration - Symptoms: Input lag / Factors: Latency - Who: Specific region/ISP / When: Always - Primary owner: Infra team (Server infrastructure) / Also: Infra team (Network infrastructure), Game team (Server development) - Game team action items: Mitigate only (code can’t change physics), offer region selection so players pick a nearby server, use lag compensation (rewind) to reduce the hit-registration disadvantage. - Infra team action items: Servers/OS: put regional servers where most players are. Network: put edge locations (PoPs) close to players, choose links and routes with fewer detours. - Ballpark numbers: Seoul–Tokyo about 30 ms, Seoul–Singapore about 75 ms, Seoul–US West Coast about 140 ms, Seoul–Europe about 230–270 ms (round trip, over real routes). Few major cables run along the direct line to Europe, so traffic goes around through Southeast Asia and Suez or through the US, and latency is much higher than the distance alone suggests. - On the graph: Always high (RTT (by country/region)) - Where to look: Tag client IPs with a country and look at the RTT distribution per country. Run ping and traceroute to the server from a cloud-region VM in that area or from RIPE Atlas probes (selected by country or ASN) - Confirmed if: RTT from distant countries is always high regardless of time of day, and close to both the minimum delay computed from distance (10 ms round trip per 1,000 km) and public latency statistics - Ruled out if: Far above what distance explains: points to “Detour routing.” Rises only in the evening: “Peak-hour congestion at peering links” - Check with: Infra tools (no game code needed) - Real incidents: riot-direct-2015 - Sources: - [ITU-T G.114: One-way transmission time](https://www.itu.int/rec/T-REC-G.114-200305-I/en) · ITU · Planning value for optical fiber propagation delay: 5 µs/km (about 200,000 km per second, 10 ms round trip per 1,000 km) - [Azure network round-trip latency statistics](https://learn.microsoft.com/en-us/azure/networking/azure-network-latency) · Microsoft Azure · Measured median round-trip times from Seoul (Korea Central): Tokyo 30 ms, Singapore 68 ms, US West 124–136 ms, Europe 234–244 ms - [AAE-1 & SMW5 cable cuts impact millions of users across multiple countries](https://blog.cloudflare.com/aae-1-smw5-cable-cuts/) · Cloudflare · Traffic between Europe and Asia mostly runs over submarine cables through Egypt (Suez) - [Probe Selection (RIPE Atlas REST API)](https://atlas.ripe.net/docs/apis/rest-api-manual/measurements/creating-measurements/probe-selection/) · RIPE NCC · Select RIPE Atlas probes by country, region, ASN, or address prefix and run ping and traceroute from them #### isp-satellite · Satellite internet (LEO/GEO) · Satellite internet (LEO, GEO) Satellite signals have to travel to space and back. With geostationary satellites the round trip alone exceeds 0.5 seconds. Low Earth orbit satellites such as Starlink are usually fast, but latency fluctuates and the link can drop briefly at the moment routes are reassigned. - Why → Effect → On screen: Connecting from home, a ship, or a plane over GEO or LEO satellite internet, or over in-flight Wi-Fi that uses satellites → GEO satellites sit at about 36,000 km, so the distance itself is long. LEO systems reassign the terminal–satellite–ground station path at short intervals, with a brief burst of delay and loss at each reassignment → GEO: heavy input lag on every action. LEO: fine most of the time, then stutter and teleporting at regular intervals - Symptoms: Input lag, Stutter, Teleporting / Factors: Latency, Jitter, Packet loss - Who: Just me, Same household, Specific region/ISP / When: Always, At regular intervals - Primary owner: External (External) / Also: Game team (Client development), Game team (Server development) - Game team action items: Client: lengthen the interpolation buffer automatically to match jitter, send inputs redundantly to survive brief loss, show connection quality. Server: account for satellite latency when setting timing windows and lag compensation limits, use timeouts that don’t kick players over gaps of about 1 second. - External action items: Tell players that satellite internet can have high latency or periodic spikes, and advise a wired terrestrial connection for competitive content where possible. - Ballpark numbers: At geostationary orbit (36,000 km altitude) the signal takes 260 ms one way just to cross space, so the round trip exceeds 520 ms (ITU-T G.114). For LEO Starlink, official data (15-second averages) put the US peak-hour median at 33 ms, with even the worst 1% (p99) under 65 ms (2024). Measurement studies found that latency shifts each time routes are reassigned every 15 seconds, along with brief outages under 1 second. In 2018 in-flight internet measurements, satellite-based connections averaged 750 ms round trip. - On the graph: Outliers only (RTT/jitter (by satellite ISP ASN)) - Where to look: Check whether the client IP’s ASN belongs to a satellite internet provider, and plot the RTT distribution and time series of that provider’s players separately. Ping the server for a few minutes straight from RIPE Atlas probes in that ASN, or have the player leave ping running and measure the interval between spikes - Confirmed if: GEO providers: RTT always above 500 ms. LEO providers: normally tens of ms, with RTT shifting or brief drops about every 15 seconds - Ruled out if: Not a satellite provider but RTT always high: “Propagation delay (physical distance)” or “Detour routing.” Irregular spikes: points to Wi-Fi or mobile signal - Check with: Infra tools (no game code needed) - Learn more: LEO satellites are close (one Starlink hop takes 1.8–3.6 ms), so typical latency can be similar to a terrestrial connection. If the point where the ground station hands traffic to the internet (PoP) is far from the game server, though, the path gets longer by that much, and routing through laser links between satellites adds more latency. Measurement studies attribute the 15-second fluctuation to route reassignment that happens at the same moment worldwide; it is unrelated to switching between satellites. In-flight Wi-Fi latency varies widely by technology (satellite or ground-based cell towers), and systems that use geostationary satellites have the same long round trip described above. - Sources: - [ITU-T G.114: One-way transmission time](https://www.itu.int/rec/T-REC-G.114-200305-I/en) · ITU · One-way propagation delay planning values for satellite links: 12 ms at 400 km altitude, 110 ms at 14,000 km, 260 ms at 36,000 km (geostationary) - [Improving Starlink’s Latency](https://starlink.com/public-files/StarlinkLatency.pdf) · Starlink · US peak-hour median 48.5 ms→33 ms, slowest 1% (p99) over 150 ms→under 65 ms (2024); one satellite hop 1.8–3.6 ms; routing over laser links adds latency, and so does the distance from the ground station to the internet access point (PoP) - [A Multifaceted Look at Starlink Performance (WWW 2024)](https://www.nitindermohan.com/documents/2024/pubs/starlinkWWW2024.pdf) · ACM · Starlink reassigns routes every 15 seconds at the same moment worldwide; at these boundaries latency and throughput fluctuate and brief outages under 1 second occur (unrelated to switching between satellites); terminal↔satellite↔ground station latency about 40 ms - [Mile High WiFi: A First Look at In-Flight Internet Connectivity (WWW 2018)](https://aqualab.cs.northwestern.edu/publication/2018/jrula-www18/) · ACM · 45 hours of in-flight internet measurements: average round-trip latency 200 ms for ground-based (cell tower) systems and 750 ms for satellite systems, median loss rate 7% for satellite systems - [Probe Selection (RIPE Atlas REST API)](https://atlas.ripe.net/docs/apis/rest-api-manual/measurements/creating-measurements/probe-selection/) · RIPE NCC · Select RIPE Atlas probes by country, region, ASN, or address prefix and run ping and traceroute from them #### isp-routing · Detour routing · Suboptimal routing Because of interconnection agreements between ISPs, traffic to even a nearby server can take a long way around. - Why → Effect → On screen: Your ISP and the server’s ISP aren’t directly connected → Traffic passes through another country or city, adding distance and hops → Only players on certain ISPs have unusually high ping - Symptoms: Input lag / Factors: Latency - Who: Specific region/ISP / When: Always - Primary owner: Infra team (Network infrastructure) / Also: External (External) - Infra team action items: Connect to multiple ISPs (multihoming), monitor ping per ISP to find the ones taking detours, negotiate route changes with ISPs. - External action items: Ask the ISP to adjust its routing. - Ballpark numbers: Even within one country, ping can differ two- or threefold depending on the route. - On the graph: Always high (RTT (by ISP/ASN)) - Where to look: Compare RTT by ISP (ASN), and use traceroute or mtr from RIPE Atlas probes on the slow ISP, or from players, to see which countries and cities the route passes through. Measure IPv4 and IPv6 separately (mtr -4, -6) - Confirmed if: Within the same region, one ISP is always higher and its route passes through another country or a distant city. Or only one address family (IPv4 or IPv6) is high - Ruled out if: All ISPs similarly high: “Propagation delay (physical distance).” High only in the evening: “Peak-hour congestion at peering links” - Check with: Infra tools (no game code needed) - Learn more: IPv4 and IPv6 routes are chosen separately, so for the same server one of them can take a long detour and be slow (a 2016 APNIC measurement found separate groups of users within a single ISP for whom IPv6 was 15, 25, or 75 ms slower than IPv4). Apps that use Happy Eyeballs (RFC 8305) go with whichever of IPv6 and IPv4 connects first. They try IPv6 first, and if it connects within the recommended 250 ms, they never try IPv4. So they tend to end up on IPv6 even when that path is a little slower. If ping is high only on a particular ISP, measure IPv4 and IPv6 separately. - Real incidents: riot-direct-2015 - Sources: - [Quantifying the Causes of Path Inflation (SIGCOMM 2003)](https://conferences.sigcomm.org/sigcomm/2003/papers/p113-spring.pdf) · ACM · Analysis of 65 ISPs: peering policies between ISPs and interdomain routing make paths much longer - [The Internet at the Speed of Light (HotNets 2014)](https://conferences.sigcomm.org/hotnets/2014/papers/hotnets-XIII-final111.pdf) · ACM · Actual router paths are about 1.5 times the straight fiber distance (median), and packets between two nearby points sometimes travel around the far side of the globe (hairpinning) - [Probe Selection (RIPE Atlas REST API)](https://atlas.ripe.net/docs/apis/rest-api-manual/measurements/creating-measurements/probe-selection/) · RIPE NCC · Select RIPE Atlas probes by country, region, ASN, or address prefix and run ping and traceroute from them - [mtr(8) manual page source](https://raw.githubusercontent.com/traviscross/mtr/master/man/mtr.8.in) · mtr · -4 and -6 measure the route over IPv4 only or IPv6 only - [RFC 8305: Happy Eyeballs Version 2: Better Connectivity Using Concurrency](https://www.rfc-editor.org/rfc/rfc8305) · IETF · An address or address family (IPv4/IPv6) can be blocked, broken, or slow depending on the network; try IPv6 first and wait a recommended 250 ms before the next connection attempt - [IPv6 Performance – Revisited](https://blog.apnic.net/2016/08/22/ipv6-performance-revisited/) · APNIC · Comparison of IPv6 and IPv4 round-trip times for the same dual-stack users: some access networks handle IPv6 packets completely differently, producing groups within one ISP where IPv6 is 15, 25, or 75 ms slower #### isp-peak · Peak-hour congestion at peering links · Peak-hour congestion at peering Around 9–11 PM, video traffic surges and the links between ISPs (peering) tend to get congested. - Why → Effect → On screen: Evening streaming and downloads pile up → Queues build and packets drop on peering links → Players on certain ISPs get stutter and teleporting only in the evening - Symptoms: Stutter, Teleporting, Rubber-banding / Factors: Jitter, Packet loss, Latency - Who: Specific region/ISP / When: Evening peak hours - Primary owner: Infra team (Network infrastructure) / Also: External (External) - Infra team action items: Add more direct connections with the affected ISP, route around congested paths, monitor evening loss and ping per ISP. - External action items: Ask the ISP to add capacity on the peering link. - On the graph: High at certain hours (RTT/loss (by ISP)) - Where to look: Plot RTT and loss per ISP (ASN) by time of day, and get mtr runs in the evening and during the day from RIPE Atlas probes on that ISP or from players to find the hop where loss starts - Confirmed if: Only one ISP sees RTT and loss rise around 9–11 PM every evening, and in mtr the loss and delay persist from the inter-ISP link all the way to the destination - Ruled out if: All ISPs rise together: points to our own links or servers. Only one household is bad in the evening: “Congested Wi-Fi channel” - Check with: Infra tools (no game code needed) - Sources: - [Inferring Persistent Interdomain Congestion (SIGCOMM 2018)](https://www.caida.org/catalog/papers/2018_inferring_persistent_interdomain_congestion/inferring_persistent_interdomain_congestion.pdf) · ACM · Recurring congestion on some inter-ISP links, with latency rising at peak hours every day and loss rates also rising during congested periods - [Probe Selection (RIPE Atlas REST API)](https://atlas.ripe.net/docs/apis/rest-api-manual/measurements/creating-measurements/probe-selection/) · RIPE NCC · Select RIPE Atlas probes by country, region, ASN, or address prefix and run ping and traceroute from them #### isp-cable · Submarine cable / international link outage · Submarine cable fault When a submarine cable is cut, traffic takes long detours for weeks (sometimes months) until it is repaired, and the remaining links get congested. - Why → Effect → On screen: Cable cut or equipment failure → Traffic crowds onto long detour routes and the remaining links → Ping surges and packet loss for overseas players that last days to weeks - Symptoms: Input lag, Teleporting / Factors: Latency, Packet loss - Who: Specific region/ISP / When: Always - Primary owner: External (External) / Also: Infra team (Network infrastructure) - Infra team action items: Secure links on other routes, move traffic onto them during an outage. - External action items: Tell overseas players the cause and the expected recovery time, ask the link provider for its repair schedule. - On the graph: Step change (RTT (by overseas country)) - Where to look: Find when RTT and loss rose on the per-country graph, match that time against Cloudflare Radar’s internet outage summaries and submarine cable operators’ notices, and check with traceroute whether the route now goes around another continent - Confirmed if: From a certain moment, RTT for a specific overseas region steps up and stays there for days to weeks, with cable outage reports from the same period. The route has switched to an unusually long detour - Ruled out if: Back to normal within a few days with no outage reports: “BGP route changes and convergence” or a problem in the ISP’s segment - Check with: Infra tools (no game code needed) - Sources: - [AAE-1 & SMW5 cable cuts impact millions of users across multiple countries](https://blog.cloudflare.com/aae-1-smw5-cable-cuts/) · Cloudflare · Repairing a submarine cable requires sending a repair ship and usually takes days to weeks (38 days in the Tonga case); cuts increase latency and loss on Europe–Asia routes - [Q2 2024 Internet disruption summary](https://blog.cloudflare.com/q2-2024-internet-disruption-summary/) · Cloudflare · Red Sea cables damaged in February 2024 were still under repair in July (conflict zone); the EASSy and Seacom cuts in May were repaired in 19 days - [Q1 2024 Internet disruption summary](https://blog.cloudflare.com/q1-2024-internet-disruption-summary/) · Cloudflare · West African cable cuts (March 14) were repaired 3–6 weeks later, with traffic moved to other cables in the meantime #### isp-bgp · BGP route changes and convergence · Route change / BGP convergence When internet routing information changes, packets are lost for the few seconds to tens of seconds (rarely a few minutes) it takes to converge again. - Why → Effect → On screen: Routing information changes somewhere in an ISP’s network → For a few seconds to tens of seconds, packets vanish or switch to a new route → A sudden freeze of a few seconds, then ping settles at a different value (e.g., 40 → 70 ms) - Symptoms: Freeze, Teleporting / Factors: Packet loss, Latency - Who: Specific region/ISP / When: Randomly - Primary owner: Infra team (Network infrastructure) / Also: Game team (Server development), External (External) - Game team action items: Use timeouts that survive brief outages (don’t drop a connection right away just because it went silent for a few seconds). - Infra team action items: Monitor routes (watch for route and ping changes on our IP prefixes), detect failures on our own links within 1 second with BFD and fail over (the default BGP hold time is 90–180 seconds), move traffic to another link if the route switches to a long path and doesn’t come back. - External action items: Ask the ISP to investigate segments of its network where routes change often. - On the graph: Step change (RTT, traceroute path) - Where to look: Compare traceroute and mtr paths from before and after the moment RTT changed, and check the BGP route change history for our prefix in RIPEstat BGPlay - Confirmed if: A freeze of a few seconds, then RTT moves to a different level, with BGP updates and AS path changes at the same time - Ruled out if: No route changes on record but high only in the evening: “Peak-hour congestion at peering links.” Only some connections bad: “One faulty ECMP path” - Check with: Infra tools (no game code needed) - Real incidents: cloudflare-2020, meta-2021, cloudflare-dns-2025 - Sources: - [RFC 4271: A Border Gateway Protocol 4 (BGP-4)](https://www.rfc-editor.org/rfc/rfc4271) · IETF · Recommended default BGP hold time is 90 seconds (the session is torn down if no message arrives from the peer within that time) - [BGP updates in 2024](https://blog.apnic.net/2025/01/07/bgp-updates-in-2024/) · APNIC · Daily average time for unstable routes to settle again: 25–35 seconds (IPv4), 40–50 seconds (IPv6) - [Delayed Internet Routing Convergence (SIGCOMM 2000)](https://conferences.sigcomm.org/sigcomm/2000/conf/paper/sigcomm2000-5-2.pdf) · ACM · Convergence after a route failure can take up to several minutes, with higher loss and latency in the meantime (measured in 2000) - [BGPlay (RIPEstat Data API)](https://stat.ripe.net/docs/data-api/api-endpoints/bgplay) · RIPE NCC · Shows the BGP routes for an address prefix at the start time, the BGP updates observed during the period, and the ASes on the path #### isp-ecmp · One faulty ECMP path · ECMP / link bundle member fault ISPs and data centers keep several paths to the same destination and send each connection down one of them. If a single path fails, only the players assigned to it keep lagging. - Why → Effect → On screen: One link or device in a bundle of links is faulty or congested → The path is chosen from the address and port combination (hash), so only connections assigned to that path see loss and delay → Same region and ISP, but only some players keep teleporting. Reconnecting sometimes fixes it - Symptoms: Teleporting, Rubber-banding, Stutter / Factors: Packet loss, Jitter - Who: Just me, Specific region/ISP / When: Always - Primary owner: Infra team (Network infrastructure) / Also: Game team (Server development), External (External) - Game team action items: Record per-connection loss and retransmission stats so you can pull the IP, port, and time for affected players (for TCP, the retransmission count from TCP_INFO; for UDP, compute it from missing packet sequence numbers). - Infra team action items: Collect affected players’ IPs, ports, and times and pass them to the ISP or data center, monitor loss per path, measure paths with the same protocol and port as the game (mtr --tcp or --udp with --port), remove the faulty link or device from the bundle if the path runs over our equipment. - External action items: Ask the ISP to check and replace the faulty path, tell players they can work around it for now by reconnecting (when reconnecting changes the port). - Ballpark numbers: With 4 paths, only about a quarter of players are affected. A ping test may take a different path from the game and come back perfectly fine. - On the graph: Outliers only (Per-connection loss/retransmissions (by IP/port)) - Where to look: Split per-connection loss and retransmissions by source IP and port. Run mtr in UDP mode (-u) against the game port (-P) with a fixed source port (-L), and repeat several times with different source ports. With -P and no -L, the source port changes on every probe and several paths get mixed together - Confirmed if: Within the same region and ISP, only certain source port (or address) combinations keep losing packets, and the problem goes away when a reconnect changes the port - Ruled out if: Bad no matter which port: congestion or failure across a whole segment - Check with: Infra tools (no game code needed) - Learn more: To keep a connection’s packets in order, network devices (ECMP, LAG) pin each connection to one path using a value computed from its addresses and ports (or only its addresses, depending on device settings). Where only addresses are used, reconnecting lands on the same path and doesn’t help. So when reports like “ping is fine but the game lags” and “reconnecting fixed it” come in together, suspect this cause. - Sources: - [RFC 7424: Mechanisms for Optimizing Link Aggregation Group (LAG) and Equal-Cost Multipath (ECMP) Component Link Utilization in Networks](https://www.rfc-editor.org/rfc/rfc7424) · IETF · LAG and ECMP pick one link per flow from a hash of header fields to keep packets in order (many-to-one mapping of flows to links) - [RFC 2991: Multipath Issues in Unicast and Multicast Next-Hop Selection](https://www.rfc-editor.org/rfc/rfc2991) · IETF · What defines a flow varies by implementation (destination address only, address pair, or including ports); with multiple paths, ping and traceroute results are hard to trust - [mtr(8) manual page source](https://raw.githubusercontent.com/traviscross/mtr/master/man/mtr.8.in) · mtr · The -u (UDP), -P (destination port), and -L (UDP source port) options; with only -P, the probe sequence number goes into the source port, so it changes on every probe #### isp-shaping · ISP throttling and traffic management · Traffic shaping, data caps When you go over your data allowance, or on plans that manage certain kinds of traffic, packets get delayed or dropped. - Why → Effect → On screen: Speed throttled after the plan’s data runs out, or certain traffic restricted → Packets wait in a queue or get dropped → Lag after a certain amount of usage, especially on mobile - Symptoms: Input lag, Teleporting / Factors: Latency, Packet loss - Who: Just me, Specific region/ISP / When: Always, Evening peak hours - Primary owner: External (External) / Also: Game team (Server development), Infra team (Network infrastructure) - Game team action items: Reduce game traffic (compression, send only what’s needed). - Infra team action items: If game traffic is delayed or dropped only on a particular ISP, gather evidence and escalate to that ISP. - External action items: Tell players to check whether their plan’s data has run out or is being throttled and whether other apps on the same phone are using data, ask the ISP whether it restricts game traffic. - Ballpark numbers: Korean mobile plans usually throttle to 1–5 Mbps once the data allowance runs out, and cheap plans to hundreds of kbps. The game itself uses little data, but when other apps on the same phone use data, a queue forms in front of the throttling equipment. - On the graph: Hits a ceiling (Throughput, RTT) - Where to look: Have the player check remaining data and throttling status in the carrier’s app and run a speed test to see the top speed. On the server side, compare loss and RTT per ISP - Confirmed if: Throughput stops rising at one value such as 1–5 Mbps or a few hundred kbps, and from then on RTT and loss grow whenever other apps on the same phone use data. Goes away after topping up data or switching to Wi-Fi - Ruled out if: No throttling but one ISP is bad: “Peak-hour congestion at peering links” or “Detour routing” - Check with: The player’s own environment - Sources: - [SKT, 고객 선택권과 혜택 강화한 신규 5G 요금제 출시](https://news.sktelecom.com/180213) · SK텔레콤 · Examples of speed control after a 5G plan’s base data runs out: up to 400 kbps, 1 Mbps, 3 Mbps - [SKT, 요금제 개편](https://news.sktelecom.com/225487) · SK텔레콤 · Service continues at up to 400 kbps after the included data runs out (Korea’s nationwide “safety net data” program) #### isp-udp-block · Country- or ISP-level UDP restrictions and packet inspection · UDP blocking, throttling and inspection by networks Some networks block specific UDP addresses and ports or throttle UDP, and packet inspection equipment filters out protocols it doesn’t recognize. Games that communicate over UDP can’t connect on those networks or disconnect often. - Why → Effect → On screen: Connecting from an ISP network that throttles UDP, or from a network with country- or ISP-level traffic inspection (censorship) equipment → Specific UDP addresses and ports are blocked, UDP is throttled at busy hours, ports or protocols not on an allowlist are filtered out, or the first few packets get through before the flow is blocked → Only players in certain countries or on certain ISPs can’t connect or get infinite loading, disconnect soon after connecting, or teleport from packet loss at busy hours - Symptoms: Can’t connect / infinite loading, Disconnect, Teleporting / Factors: Packet loss - Who: Specific region/ISP / When: Right after login or maintenance, Always, Evening peak hours - Primary owner: External (External) / Also: Game team (Client development), Game team (Server development), Infra team (Network infrastructure) - Game team action items: Client: fall back automatically to a TCP/TLS 443 path if UDP doesn’t connect within a few seconds, also detect connections that drop soon after succeeding and retry them on the fallback path, log which path was used. Server: accept the same game protocol over TCP 443 (TLS) as well, adjust timeouts because the fallback path can add latency. - Infra team action items: Before launching in a new country, measure UDP reachability and peak-hour loss on local ISP networks, place relays or gateways that accept the TCP 443 fallback close to the region, monitor UDP and TCP connection success rates per country and ASN, gather evidence and escalate for ISPs confirmed to throttle UDP. - External action items: Ask the ISP or agency about its UDP restriction criteria and any relief, tell players to try connecting from another network to compare. - Ballpark numbers: Measurements cited by an IETF document show that 3–5% of networks block UDP entirely. When Google reviewed its 2016 QUIC (UDP-based) usage, 4.4% of clients couldn’t use UDP/QUIC because it was blocked or the path MTU was too small, mostly behind corporate firewalls, and it saw no case of an entire ISP blocking it. Another 0.3% were on networks where loss rose sharply at peak hours, apparently from UDP throttling; Google brought that down from 1% in 2015 by working with the ISPs. - On the graph: Outliers only (UDP connection success rate (by country/ASN)) - Where to look: Split UDP connection success rate and TCP 443 fallback success rate by country and ASN. From a cloud VM or a player’s PC on that ISP’s network, test connections to the game’s UDP port and to TCP 443 separately, and compare mtr -u -P (game port) with mtr -T -P 443 to see at which hop responses disappear - Confirmed if: Only in a specific country or ASN, UDP gets no first response or drops within a few seconds, while TCP 443 works from the same place. With throttling, UDP loss rises clearly only at peak hours and TCP is less affected - Ruled out if: TCP fails too: points to a path outage, IP blocking, or “DNS failures and delays.” Same in every country: our own server or firewall configuration. Loss only during traffic bursts, for UDP and TCP alike: “Policer drops excess traffic” - Check with: Infra tools (no game code needed) - Learn more: According to an IRTF survey document, packet inspection equipment may pick out UDP flows by address, port, and protocol and block them, or block everything except the protocols it allows (an allowlist). If the equipment decides based on only a few fields of the packet, even a small protocol change can get traffic blocked. In the early days of QUIC, one firewall let the first few packets through after a single header bit changed and then blocked the rest, so the client’s logic for falling back to TCP never kicked in. When launching in a new country, this can surface as reports that “it works fine in Korea, but players on some ISPs in that country can’t connect.” If it’s blocked only on one location’s network, such as a café or an office, see the “Public Wi-Fi and corporate network restrictions” entry. - Sources: - [RFC 9308: Applicability of the QUIC Transport Protocol](https://www.rfc-editor.org/rfc/rfc9308) · IETF · Measurement studies show 3–5% of networks block UDP entirely, so UDP-based apps must accept connection failures or provide a TCP (TLS) fallback; firewalls may block ports that aren’t tied to a registered service - [The QUIC Transport Protocol: Design and Internet-Scale Deployment (SIGCOMM 2017)](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/46403.pdf) · ACM · 2016: 4.4% of clients couldn’t use QUIC over UDP because UDP/QUIC was blocked or the path MTU was small (mostly behind corporate firewalls; no ISP-wide blocking observed); 0.3% were on networks that appeared to throttle UDP (higher loss at peak hours; down from 1% in 2015 after requests to ISPs); a case where a firewall let only the first few packets through after a 1-bit header change and blocked the rest, defeating the TCP fallback logic - [RFC 9505: A Survey of Worldwide Censorship Techniques](https://www.rfc-editor.org/rfc/rfc9505) · IRTF · Inspection equipment in the network can pick out and block TCP and UDP flows by address, port, and protocol (blocking of UDP endpoints has been observed with QUIC); allowing only approved protocols leads to overblocking, and throttling of specific traffic is also used - [mtr(8) manual page source](https://raw.githubusercontent.com/traviscross/mtr/master/man/mtr.8.in) · mtr · Sends UDP with -u or TCP SYN with -T and sets the destination port with -P, so the route is measured with the same protocol and port as the game #### isp-line · Poor line quality · Faulty last-mile line / modem Loose connectors, old wiring, or a faulty modem cause steady packet loss and periodic line drops. - Why → Effect → On screen: Damaged cable, poor contact, faulty modem or optical network terminal → Packets dropped from bit errors; now and then the line drops for a few seconds to about a minute while it reconnects → Steady low-level packet loss, occasional freezes of a few seconds or disconnects - Symptoms: Teleporting, Freeze, Disconnect / Factors: Packet loss - Who: Same household / When: Randomly - Primary owner: External (External) - External action items: Tell players to check whether other games and video calls also drop, and if so, to ask their ISP for a line check. - On the graph: Random spikes (Loss rate, line reconnect log) - Where to look: Measure loss up to the ISP’s first hop for a few minutes with pathping (or mtr), and check reconnect times in the internet (WAN) connection log on the router’s admin page - Confirmed if: Steady loss from the ISP’s first hop even when the line is idle, and reconnect times in the router log line up with freezes and disconnects. Other games and video calls drop at the same time - Ruled out if: Loss starts on the wireless hop to the router: “Wi-Fi interference and weak signal.” Loss starts far into the ISP network: the ISP’s route - Check with: The player’s own environment - Sources: - [RFC 3635: Definitions of Managed Objects for the Ethernet-like Interface Types](https://www.rfc-editor.org/rfc/rfc3635) · IETF · Frames that fail the frame check (FCS) are counted as FCS errors (dot3StatsFCSErrors) and included in input errors (ifInErrors) - [Understanding and Mitigating Packet Corruption in Data Center Networks (SIGCOMM 2017)](https://www.microsoft.com/en-us/research/wp-content/uploads/2017/06/CorrOpt_SIGCOMM2017.pdf) · ACM · The main causes of packet corruption are bad optics, damaged fiber, dirty connectors, and poor installation; corruption loss stays steady regardless of utilization - [pathping](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/pathping) · Microsoft · Pings each hop for a set time and calculates loss per router and link to show where loss occurs #### isp-dns · DNS failures and delays · DNS failure / slowness If DNS, which turns server names into addresses, is slow or fails, the game can’t find its login or patch servers. - Why → Effect → On screen: ISP DNS outage or misconfiguration → The login or patch server address can’t be resolved → A long wait after pressing Connect, or no connection at all. Players already connected are fine - Symptoms: Can’t connect / infinite loading / Factors: Latency, Packet loss - Who: Specific region/ISP, Just me / When: Right after login or maintenance - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Cache addresses (remember the last server address that connected successfully), set up multiple DNS resolvers (query another one if one fails). - External action items: Tell players to try switching to another DNS resolver, such as a public DNS. - On the graph: Outliers only (Login failures (by ISP), DNS lookup time) - Where to look: Query the login server name against the ISP’s DNS and a public DNS separately with Resolve-DnsName -Server (or nslookup), and compare response times and results - Confirmed if: Only the ISP’s DNS fails to respond or takes a long time, and switching to a public DNS connects right away. Players already connected are fine - Ruled out if: Every DNS returns the address right away but connecting still fails: points to the route, a firewall, or the server - Check with: The player’s own environment - Real incidents: meta-2021, cloudflare-dns-2025, aws-2025 - Sources: - [Cloudflare 1.1.1.1 Incident on July 14, 2025](https://blog.cloudflare.com/cloudflare-1-1-1-1-incident-on-july-14-2025/) · Cloudflare · When a public DNS resolver stopped for 62 minutes, users who could no longer resolve names effectively lost access to all internet services - [RFC 8767: Serving Stale Data to Improve DNS Resiliency](https://www.rfc-editor.org/rfc/rfc8767) · IETF · Riding out outages by continuing to use expired cache records when authoritative servers can’t be reached (serve-stale) - [Resolve-DnsName](https://learn.microsoft.com/en-us/powershell/module/dnsclient/resolve-dnsname) · Microsoft · Resolves a name against the DNS server specified with -Server - [nslookup](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/nslookup) · Microsoft · Command that queries a DNS server for a name directly #### isp-ddos-path · Shared links saturated by DDoS · DDoS saturating shared links Massive attacks aimed at the game company, or at someone else on the same network, fill up shared links. - Why → Effect → On screen: A flood of attack traffic → Legitimate traffic on the same links gets delayed and dropped too → Many players teleport, disconnect, or can’t connect at the same time - Symptoms: Teleporting, Disconnect, Can’t connect / infinite loading / Factors: Packet loss, Latency - Who: Whole server, Specific region/ISP / When: Randomly, When crowds gather - Primary owner: Infra team (Network infrastructure) / Also: External (External) - Infra team action items: Use a DDoS protection service, reroute traffic during attacks, hide server addresses (keep servers behind protection equipment and don’t expose their real addresses). - External action items: If the attack targets someone else on the same network, ask the ISP to block it upstream. - On the graph: Hits a ceiling (Link inbound traffic (bps/pps), interface drops) - Where to look: Inbound traffic and dropped packet counts on our links and equipment, plus the DDoS protection service’s attack detection log, lined up with the times when disconnects cluster - Confirmed if: Inbound traffic flattens out at link capacity and drops rise, while players across many regions and ISPs teleport or disconnect at the same moment - Ruled out if: Links have headroom but only some ISPs are bad: congestion or routing problems in the ISP segment - Check with: Infra tools (no game code needed) - Sources: - [Infrastructure layer attacks](https://docs.aws.amazon.com/whitepapers/latest/aws-best-practices-ddos-resiliency/infrastructure-layer-attacks.html) · AWS · Volumetric attacks such as UDP reflection and SYN floods overwhelm network capacity or tie up firewall and load balancer resources - [Obfuscating AWS resources (BP1, BP4, BP5)](https://docs.aws.amazon.com/whitepapers/latest/aws-best-practices-ddos-resiliency/obfuscating-aws-resources-bp1-bp4-bp5.html) · AWS · Put edge services such as CloudFront or a load balancer in front of origin servers to reduce direct exposure - [RFC 7999: BLACKHOLE Community](https://www.rfc-editor.org/rfc/rfc7999) · IETF · The BLACKHOLE community, announced over BGP to ask neighboring networks to drop traffic headed to a specific address #### isp-cgnat · ISP-shared IP addresses (CGNAT) · Carrier-grade NAT Mobile networks and some ISPs have many subscribers share one IP address, and they delete the mappings of idle connections after a short time. - Why → Effect → On screen: ISP equipment manages the session table for huge numbers of subscribers → Session table limits, short idle timeouts → Disconnects after sitting idle, and false positives that block everyone sharing the same IP at once - Symptoms: Disconnect, Can’t connect / infinite loading / Factors: Packet loss - Who: Specific region/ISP / When: After sitting idle - Primary owner: Game team (Client development) / Also: Game team (Server development), Infra team (Network infrastructure) - Game team action items: Client: send heartbeats at no more than half the shortest idle timeout (which can be about 30 seconds on mobile networks), send them from the client because ISP CGNAT mappings are reliably refreshed only by outgoing packets, reconnect automatically on disconnect. Server: respond to heartbeats and close the connection proactively if none arrive for a set time, use a session token to keep the same player when a changed mapping gives them a new address and port, be careful with IP-based blocking because many people can share one IP (decide together with account and device signals). - Infra team action items: Adjust per-IP connection limits and new-connections-per-second limits on firewalls and DDoS protection equipment to allow for ISP-shared IPs (raise the thresholds for mobile carrier ranges or exempt them). - Ballpark numbers: On mobile networks the UDP idle timeout can be as short as about 30 seconds. - On the graph: Mass disconnect (Disconnects (heartbeat timeouts), idle time before disconnect (by ISP)) - Where to look: In the connection logs, check how many accounts connect from one IP at the same time and their ISP (ASN), and collect per ISP the idle time of connections that dropped after idling. On the player side, check the internet (WAN) address on the router’s admin page - Confirmed if: Multiple accounts connect from one IP in mobile carrier ranges, and idle times before disconnect cluster at short values around 30–60 seconds. The router’s WAN address is in 100.64.0.0/10 (shared address space for carrier NAT) or differs from the address the server sees - Ruled out if: Concentrated among home router users regardless of ISP: “NAT mapping expiry” - Check with: Game server or client logs and metrics - Sources: - [A Multi-perspective Analysis of Carrier-Grade NAT Deployment (IMC 2016)](https://www.icir.org/vern/papers/cgn-imc16.pdf) · ACM · Measured NAT UDP mapping lifetimes of 10–200 seconds, 74% at 1 minute or less; CGN medians of 65 seconds on mobile networks and 35 seconds on wired networks - [RFC 6888: Common Requirements for Carrier-Grade NATs (CGNs)](https://www.rfc-editor.org/rfc/rfc6888) · IETF · CGNs must support limits on external ports per subscriber and on the rate of creating new mappings - [RFC 6269: Issues with IP Address Sharing](https://www.rfc-editor.org/rfc/rfc6269) · IETF · When many share an address, IP-based blocking (penalty box) also blocks other subscribers on the same address - [RFC 6598: IANA-Reserved IPv4 Prefix for Shared Address Space](https://www.rfc-editor.org/rfc/rfc6598) · IETF · 100.64.0.0/10 is the shared address space used between carrier NAT (CGN) equipment and subscribers’ routers #### isp-vpn · Routing through a VPN or game booster · VPN / game accelerator detour With a VPN or game booster on, packets go through that company’s relay servers. If the relay is far away or busy, the connection can actually get slower. - Why → Effect → On screen: The VPN or booster sends every game packet through its relay servers → Distance to the relay and its congestion add up, and tunnel headers shrink the MTU (the largest packet size that can be sent at once) → Higher ping and packet loss; can’t connect when the relay address gets blocked along with everyone else using it - Symptoms: Input lag, Teleporting, Can’t connect / infinite loading / Factors: Latency, Packet loss - Who: Just me / When: Always, Right after login or maintenance - Primary owner: External (External) / Also: Infra team (Network infrastructure), Game team (Server development) - Game team action items: Keep UDP packets at 1,200 bytes or less (so they don’t fragment even when tunnel headers shrink the MTU), account for shared VPN and booster relay addresses in IP-based blocking and decide together with account and device signals. - Infra team action items: If you have many overseas players, set up your own locations (PoPs) near them, check the routes of ISPs with a cluster of reports saying “turning on the booster made it better.” - External action items: Tell players to turn off the VPN or booster and compare. - Ballpark numbers: A nearby relay adds a few ms; a detour through another country adds tens of ms to over 100 ms. - On the graph: Outliers only (RTT (per player), network operator of the client IP) - Where to look: Check whether the client IP’s ASN belongs to a VPN, booster, or hosting provider, and have the player turn off the VPN or booster and compare ping and traceroute - Confirmed if: RTT and loss rise, or connections get blocked, only with the VPN or booster on, and traceroute shows hops through the relay server - Ruled out if: Same with it on or off: the line or the ISP segment. Better with it on: a problem in the original ISP route (“Detour routing,” “Peak-hour congestion at peering links”) - Check with: The player’s own environment - Learn more: Conversely, when the ISP’s route is bad, a booster can take a better route and lower ping. Reports that “turning on the booster made it better” are a clue to an ISP route problem such as detour routing or evening congestion. - Sources: - [RFC 4459: MTU and Fragmentation Issues with In-the-Network Tunneling](https://www.rfc-editor.org/rfc/rfc4459) · IETF · Fragmentation and path MTU problems that occur when encapsulation headers of tunnels inside the network reduce the size that can be sent - [RFC 8899: Packetization Layer Path MTU Discovery for Datagram Transports](https://www.rfc-editor.org/rfc/rfc8899) · IETF · Recommends 1,200 bytes as the default safe size (BASE_PLPMTU) for datagram transports such as UDP - [Azure network round-trip latency statistics](https://learn.microsoft.com/en-us/azure/networking/azure-network-latency) · Microsoft Azure · Round-trip latency added when going through a location (PoP) in another country: Seoul–Tokyo 30 ms, Seoul–Hong Kong 39 ms, Seoul–Singapore 68 ms ### L5 Data center network equipment (causes: 11) #### dc-firewall · Firewall session table full · Firewall session table exhaustion A firewall tracks every connection it lets through by recording it in a session table. Once the table is full, it can’t accept new connections. - Why → Effect → On screen: A connection surge or an attack pushes the session count to its limit → No free entry to record a new connection, so it’s refused → Players trying to get in can’t connect or get infinite loading, and some existing connections disconnect too - Symptoms: Can’t connect / infinite loading, Disconnect / Factors: Packet loss - Who: Whole server / When: Right after login or maintenance, When crowds gather - Primary owner: Infra team (Network infrastructure) / Also: Game team (Server development), Game team (Client development) - Game team action items: Server: smooth out connection surges with a login queue, reuse connections to avoid opening short ones over and over, proactively close connections whose heartbeats have stopped (so dead connections don’t hold session table entries for long). Client: send heartbeats at no more than half the shortest idle timeout, reconnect automatically on disconnect with growing, randomized retry intervals (so everyone doesn’t pile back in at once). - Infra team action items: Enlarge the session table, clean up short-lived connections quickly (shorten the timeout for closed sessions), tell the game team whenever you shorten the idle session timeout so they can match the heartbeat interval, block attacks, alert on session table utilization. - On the graph: Hits a ceiling (Firewall session count, new connection failures) - Where to look: Graph the firewall’s concurrent session count together with its session limit, and search the device log for packets dropped because a session couldn’t be created. On a Linux firewall, compare nf_conntrack_count with nf_conntrack_max and check dmesg for “nf_conntrack: table full, dropping packet”; on an AWS instance, check conntrack_allowance_exceeded in ethtool -S - Confirmed if: New connection failures rise from the moment the session count flattens at the limit, along with session creation failure logs or drop counters - Ruled out if: Session count well below the limit but connections fail: “Connection queue (listen backlog) overflow” or the login server. Only idle connections drop: “Cloud security group connection tracking expiry” - Check with: Infra tools (no game code needed) - Sources: - [Netfilter Conntrack Sysfs variables](https://docs.kernel.org/networking/nf_conntrack-sysctl.html) · Linux kernel · Maximum entries in the connection tracking table (nf_conntrack_max), how long closing connections are kept (TIME_WAIT and FIN_WAIT default 120 seconds), established TCP default 5 days, current entry count (nf_conntrack_count) - [Amazon EC2 security group connection tracking](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/security-group-connection-tracking.html) · AWS · Once an instance exceeds the number of connections it can track, packets for new connections are dropped; idle connections can exhaust the tracking table - [Infrastructure layer attacks](https://docs.aws.amazon.com/whitepapers/latest/aws-best-practices-ddos-resiliency/infrastructure-layer-attacks.html) · AWS · Attacks such as SYN floods tie up server, firewall, and load balancer resources - [net/netfilter/nf_conntrack_core.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/netfilter/nf_conntrack_core.c?h=v6.12) · Linux kernel · When the connection tracking table is full, the kernel logs “nf_conntrack: table full, dropping packet” and drops packets for new connections - [Monitor network performance for ENA settings on your EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-network-performance-ena.html) · AWS · conntrack_allowance_exceeded: number of packets dropped because the instance exceeded its connection tracking allowance, visible with ethtool -S #### dc-ddos · DDoS protection detours and false positives · DDoS scrubbing latency, false positives Diverting traffic to a scrubbing center to stop attacks makes the route longer, and legitimate players are sometimes mistaken for attackers and blocked. - Why → Effect → On screen: After an attack is detected (or all the time), inbound traffic is diverted to a scrubbing center → The route gets longer, and some legitimate packets are flagged as attack traffic → Ping rises for everyone; players in certain regions or on certain ISPs can’t connect - Symptoms: Input lag, Can’t connect / infinite loading, Teleporting / Factors: Latency, Packet loss - Who: Whole server, Specific region/ISP / When: When crowds gather, Randomly - Primary owner: Infra team (Network infrastructure) / Also: Game team (Server development) - Game team action items: Document the game’s traffic pattern (ports, packet sizes, packets per second) and share it with the infra team, keep UDP packets at 1,200 bytes or less. - Infra team action items: Write protection rules that fit the game’s traffic pattern, use regional scrubbing locations, reduce TCP packet size on tunnel segments (MSS clamping), check for false positives with connection failure rates per region and ISP. - Ballpark numbers: A scrubbing location in the same country adds a few ms; going through a location in another country adds 30–100 ms or more. Usually only inbound traffic takes the detour, and the server’s responses go straight out. If the filtered traffic comes back through a tunnel, the largest packet size that can be sent at once (MTU) also shrinks, which can lead to a problem where only large packets vanish. - On the graph: Step change (RTT (ping), connection failure rate by region/ISP) - Where to look: Put the protection device’s or service’s diversion (scrubbing) start and end records and its block logs on the same timeline as the RTT graph and the connection failure rates per region and ISP. From the affected region, check with mtr or traceroute whether a scrubbing location shows up in the path - Confirmed if: RTT steps up when diversion turns on, stays there, and comes back down when it turns off. Or legitimate player addresses show up in the block log, and only that region or ISP sees its connection failure rate rise - Ruled out if: RTT rises at times with no diversion or block records: “Detour routing” or “BGP route changes and convergence.” Only large packets vanish: “MTU mismatch (only large packets vanish)” - Check with: Infra tools (no game code needed) - Sources: - [Maximum transmission unit and maximum segment size](https://developers.cloudflare.com/magic-transit/reference/mtu-mss/) · Cloudflare · Inbound traffic is delivered after filtering over a GRE tunnel (MTU 1,476) while outbound responses go straight to the internet (DSR); limiting TCP MSS to 1,436 or less is recommended, and without it large packets are dropped or fragmented - [Azure network round-trip latency statistics](https://learn.microsoft.com/en-us/azure/networking/azure-network-latency) · Microsoft Azure · Round-trip latency by location (PoP): Seoul–Busan area 8 ms, Seoul–Tokyo 30 ms, Seoul–Singapore 68 ms #### dc-lb-idle · Load balancer idle timeout · Load balancer idle timeout A load balancer deletes idle connections after a set time. The game assumes the connection is still alive, and then the player gets disconnected. - Why → Effect → On screen: The player sends no packets for a while (chat window open, away from keyboard) → The load balancer cleans up the idle connection (common defaults are 60–350 seconds) → Disconnect the moment the player moves again - Symptoms: Disconnect / Factors: Packet loss - Who: Just me, Whole server / When: After sitting idle - Primary owner: Infra team (Network infrastructure) / Also: Game team (Client development), Game team (Server development) - Game team action items: Client: send heartbeats at no more than half the shortest idle timeout (30 seconds or less behind a 60-second ALB), reconnect automatically on disconnect. Server: respond to heartbeats and close the connection proactively if none arrive for a set time, resume the session with a session token. - Infra team action items: Check the idle timeout of every load balancer on the path, share the values with the game team, and raise them if needed. - Ballpark numbers: Defaults are 60 seconds for AWS ALB, 350 seconds for TCP and 120 seconds for UDP on NLB, and 4 minutes for TCP on Azure Load Balancer. The ALB and NLB TCP values can be changed, but the NLB UDP value of 120 seconds can’t. When the time runs out, ALB also closes the server-side connection, while NLB deletes it silently, so the server often never finds out. - On the graph: Mass disconnect (Disconnects, idle time before disconnect) - Where to look: Check the idle timeout setting of each load balancer on the path, and collect for each dropped connection the time from its last packet to the disconnect. On AWS NLB, also check TCP_ELB_Reset_Count in CloudWatch (number of RSTs sent by the load balancer) - Confirmed if: Idle times of dropped connections cluster just past the setting (60 s for ALB, 350 s for NLB TCP, and so on), and it reproduces when you sit still longer than that and then move. On NLB, TCP_ELB_Reset_Count rises at those times - Ruled out if: Disconnects regardless of idle time: not this cause. Clusters near 350 seconds on a server that clients reach directly without a load balancer: “Cloud security group connection tracking expiry.” On the player’s home router: “NAT mapping expiry” - Check with: Infra tools (no game code needed) - Sources: - [Edit attributes for your Application Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/application/edit-load-balancer-attributes.html) · AWS · ALB idle timeout default 60 seconds (1–4,000 seconds); the load balancer closes the connection if the client or target connection is silent for that long - [Network Load Balancers](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/network-load-balancers.html) · AWS · NLB TCP idle default 350 seconds (60–6,000 seconds); after that it only stops tracking and answers later data with RST; UDP flows are fixed at 120 seconds - [Configure load balancer TCP reset and idle timeout](https://learn.microsoft.com/en-us/azure/load-balancer/load-balancer-tcp-idle-timeout) · Microsoft Azure · Azure Load Balancer idle timeout default 4 minutes (4–100 minutes), no guarantee the session is kept beyond that, TCP reset is optional - [CloudWatch metrics for your Network Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/load-balancer-cloudwatch-metrics.html) · AWS · TCP_ELB_Reset_Count: number of RST packets generated by the load balancer #### dc-cloud-conntrack · Cloud security group connection tracking expiry · Cloud security group connection tracking timeout The firewall attached to a cloud server (security group) also tracks connections, and tracking entries for idle connections expire after a set time. Even on servers that clients reach directly without a load balancer, players who sat idle can get disconnected. - Why → Effect → On screen: The security group is set up so that it tracks game connections (only certain addresses allowed, restricted outbound rules, traffic through an NLB, and so on) → The tracking entry for a connection that sat idle for a while expires, and the security group silently drops packets that arrive after that → After being away, the player moves again, gets no response, then disconnects. The server program doesn’t notice for a long time - Symptoms: Disconnect / Factors: Packet loss - Who: Just me, Whole server / When: After sitting idle - Primary owner: Infra team (Server infrastructure) / Also: Game team (Client development), Game team (Server development) - Game team action items: Client: send heartbeats at no more than half the shortest idle timeout (175 seconds or less for 350 seconds on TCP, 90 seconds or less for 180 seconds on UDP streams), reconnect automatically on disconnect. Server: respond to heartbeats and close the connection proactively if none arrive for a set time, resume the session with a session token. - Infra team action items: Check the instance’s connection tracking timeout (TcpEstablishedTimeout) and raise it if needed (UDP can’t be raised because 180 seconds is already the maximum), review a security group setup that creates no tracking (game ports open to all addresses, all outbound allowed; connections through an NLB are still tracked), run idle tests when moving to a new instance generation. - Ballpark numbers: On AWS, Nitro v6 instance types delete tracking entries for idle TCP connections after 350 seconds by default (5 days on other types). For UDP, the defaults are 180 seconds for flows with several request/response exchanges (streams) and 30 seconds for flows that went only one way or had a single request and response. - On the graph: Mass disconnect (Disconnects, idle time before disconnect) - Where to look: Check the instance’s connection tracking timeout setting and the security group rules (whether the setup creates tracking), and collect the idle times of dropped connections. Right after a disconnect, use ss -tnoi on the server to see whether the connection stays ESTABLISHED with the retransmission timer (timer:(on,…)) running and backoff growing - Confirmed if: Idle times of dropped connections cluster just past 350 s for TCP, 180 s for UDP streams, or 30 s for one-way UDP, and the server-side socket stays ESTABLISHED without noticing the disconnect (if the server has data to send, it just keeps retransmitting) - Ruled out if: The security group setup doesn’t track (game ports open to all addresses, all outbound allowed, no NLB in the path): not this cause. Traffic goes through an NLB: compare the values with “Load balancer idle timeout” - Check with: Infra tools (no game code needed) - Sources: - [Amazon EC2 security group connection tracking](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/security-group-connection-tracking.html) · AWS · TCP idle tracking default 350 seconds (Nitro v6; 432,000 seconds = 5 days on others), UDP one-way 30 seconds and stream 180 seconds (180 max); rules that allow all addresses aren’t tracked; connections through an NLB are always tracked - [Update the TCP idle timeout for your Network Load Balancer listener](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/update-idle-timeout.html) · AWS · If the NLB idle timeout is longer than the target instance’s connection tracking timeout, the instance side silently discards the connection state first - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · In -o, timer:(on,…) is the retransmission timer; in -i, backoff is the number of times the retransmission wait has doubled #### dc-nat-gateway · Cloud NAT gateway connection and port limits · Cloud NAT gateway connection / port limits When servers in a private subnet connect out (platform authentication, payments, external APIs), a NAT gateway rewrites their address and port. If concurrent connections to the same destination exceed the gateway’s port limit, new connections fail. - Why → Effect → On screen: Servers open many short connections to the same external address, such as platform authentication or payments, or keep connections open for a long time → The NAT gateway can’t allocate any more source ports for that destination, so new connections fail → The game itself is fine, but only features that call external services, such as login, payments, and reward delivery, fail or slow down (can’t connect / infinite loading, dropped action / rollback) - Symptoms: Can’t connect / infinite loading, Dropped action / rollback / Factors: Packet loss, Latency - Who: One feature only, Whole server / When: Right after login or maintenance, Evening peak hours, When crowds gather - Primary owner: Infra team (Network infrastructure) / Also: Game team (Server development) - Game team action items: Reuse connections to external APIs (HTTP keep-alive, connection pools) and don’t open a new connection per request, send keepalives on idle pooled connections more often than the NAT idle timeout (350 seconds on AWS) or close them first, retry failures with growing, randomized intervals, record failure rates and latency per external call. - Infra team action items: Add IP addresses to the NAT gateway (an AWS public NAT gateway takes only 2 Elastic IPs by default, so request a quota increase for more), split gateways per availability zone and subnet, alert on port allocation failure metrics (AWS ErrorPortAllocation, Failed in Azure SNAT Connection Count, OUT_OF_RESOURCES in Google Cloud dropped_sent_packets_count), raise the minimum ports per VM or use dynamic port allocation on Google Cloud NAT. - Ballpark numbers: An AWS NAT gateway can open up to 55,000 concurrent connections to the same destination (IP, port, protocol) per IP address, and you can attach up to 8 IPs to raise that. It deletes connections that stay silent for 350 seconds and answers later packets on them with RST. Azure NAT Gateway has 64,512 SNAT ports per public IP (up to 16 IPs). Google Cloud NAT divides 64,512 ports per NAT IP among VMs, and the default minimum is 64 ports per VM (static allocation), so with default settings a single VM is usually limited to 64 concurrent connections to the same destination. - On the graph: Hits a ceiling (NAT gateway concurrent connections, port allocation failures) - Where to look: Line up the NAT gateway metrics ErrorPortAllocation, ActiveConnectionCount, and PacketsDropCount in AWS CloudWatch (on Azure, SNAT Connection Count filtered by the Failed state and Dropped Packets; on Google Cloud, dropped_sent_packets_count with reason OUT_OF_RESOURCES) against the times the game server’s external calls failed - Confirmed if: ErrorPortAllocation (Failed SNAT Connection Count on Azure, OUT_OF_RESOURCES drops on Google Cloud) goes above 0 when external calls fail, and the failures concentrate on calls to one or two heavily used destinations such as authentication or payment servers - Ruled out if: Port allocation failures at 0, but the game server’s connect fails with EADDRNOTAVAIL and TIME_WAIT is close to the size of the ephemeral port range: “Ephemeral port exhaustion on server-to-server connections.” Connections succeed but responses are slow: “External service dependency” - Check with: Infra tools (no game code needed) - Learn more: “Ephemeral port exhaustion on server-to-server connections” is about one server running out of ephemeral ports. This limit sits on the NAT gateway and is shared by all the servers behind it (Google Cloud NAT divides it per VM). If only external calls fail while the servers still have plenty of room in TIME_WAIT and the ephemeral port range, this is the cause. Ports from closed connections also aren’t reused for the same destination right away (Azure applies a cooldown; Google Cloud blocks them during TIME_WAIT), so the more you repeat short connections, the sooner you hit the limit. - Sources: - [NAT gateway basics](https://docs.aws.amazon.com/vpc/latest/userguide/nat-gateway-basics.html) · AWS · 55,000 concurrent connections per IPv4 address to the same destination (destination IP, port, protocol), expandable by attaching up to 8 IPs (public NAT gateways get 2 Elastic IPs by default, more through a quota increase request); bandwidth scales automatically from 5 to 100 Gbps and throughput from 1 million to 10 million packets per second, and packets beyond that limit are dropped - [NAT gateway metrics and dimensions](https://docs.aws.amazon.com/vpc/latest/userguide/metrics-dimensions-nat-gateway.html) · AWS · ErrorPortAllocation: number of times a source port couldn’t be allocated (above 0 means too many concurrent connections), ActiveConnectionCount, IdleTimeoutCount (connections cleaned up after 350 seconds idle), PacketsDropCount - [Troubleshoot NAT gateways](https://docs.aws.amazon.com/vpc/latest/userguide/nat-gateway-troubleshooting.html) · AWS · Connections expire after 350 seconds idle and later sends get an RST; keepalives shorter than 350 seconds recommended; when hitting the connection limit, add gateways per availability zone, add IPs, or reduce connections - [Source Network Address Translation (SNAT) with Azure NAT Gateway](https://learn.microsoft.com/en-us/azure/nat-gateway/nat-gateway-snat) · Microsoft Azure · 64,512 SNAT ports per public IP (up to 16 IPs); each connection to the same destination needs a different port; closed ports go through a cooldown before reuse for the same destination - [Metrics and alerts for Azure NAT Gateway](https://learn.microsoft.com/en-us/azure/nat-gateway/nat-metrics) · Microsoft Azure · SNAT Connection Count filtered by the Failed state above 0 suggests SNAT port exhaustion; Dropped Packets - [IP addresses and ports](https://cloud.google.com/nat/docs/ports-and-addresses) · Google Cloud · 64,512 ports each for TCP and UDP per NAT IP; default minimum ports per VM is 64 (static allocation) or 32 (dynamic allocation); the number of ports reserved for a VM caps its concurrent connections to the same destination; ports of closed connections can’t be used during TIME_WAIT - [Logs and metrics](https://cloud.google.com/nat/docs/monitoring) · Google Cloud · dropped_sent_packets_count with reason OUT_OF_RESOURCES: packets dropped for lack of NAT IPs or ports #### dc-lb-imbalance · Load balancer skew and misjudged health checks · LB imbalance, bad health checks Connections pile onto one server, or players keep getting sent to a server that’s already dead. - Why → Effect → On screen: The distribution rule is a poor fit, or the health check can’t see the real state → One server alone is overloaded, or players try to connect to a dead server → Only some channels or some players get slow motion, can’t connect, or get infinite loading - Symptoms: Slow motion, Can’t connect / infinite loading / Factors: Stall, Packet loss - Who: Specific zone/channel / When: Right after login or maintenance, When crowds gather - Primary owner: Infra team (Network infrastructure) / Also: Game team (Server development) - Game team action items: Implement a health check that answers the load balancer’s probes based on the real game state (tick progress, DB connections), report server load along with it. - Infra team action items: Switch to health checks that verify real game responses, distribute by server load, monitor differences in connection counts between servers. - On the graph: Outliers only (Connections/CPU utilization per server) - Where to look: Overlay connection counts (ss -s) and CPU utilization of each server behind the load balancer on one graph, and compare the load balancer’s target health status (HealthyHostCount and UnHealthyHostCount in CloudWatch on AWS) with the game servers’ actual state - Confirmed if: Only one or two servers have far higher connections and CPU than the rest, or a server whose tick has stopped stays “healthy” and keeps taking new connections - Ruled out if: Connection counts even across servers but one channel is slow: load inside that channel (“Single-threaded zone overload (hotspot)”) - Check with: Infra tools (no game code needed) - Real incidents: aws-2025 - Sources: - [Load Balancing in the Datacenter](https://sre.google/sre-book/load-balancing-datacenter/) · Google · Plain round robin lets CPU usage differ by up to 2× between tasks; weighted distribution where backends report their load in responses and health checks; a lame duck state in which a backend asks not to be sent new requests - [Health checks for Network Load Balancer target groups](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/target-group-health-checks.html) · AWS · Default health check every 30 seconds, target removed after 2 failures; UDP services are checked with TCP or HTTP health checks, so configuring them to reflect the real service state is recommended - [CloudWatch metrics for your Network Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/load-balancer-cloudwatch-metrics.html) · AWS · HealthyHostCount and UnHealthyHostCount: number of targets judged healthy and unhealthy #### dc-microburst · Switch microbursts · Switch microburst drops When several servers send packets to thousands of players at the same instant, the small buffer on the switch port where that traffic converges overflows in less than 1 ms. - Why → Effect → On screen: A world boss spawn or massive skills, or ticks on several servers lining up so they all send at once → Buffers where several ports feed into one, or where a fast port feeds a slower one (hundreds of KB to a few MB per port), fill up in an instant → Some packets dropped; many players teleport or have skills fail to go off at the same moment - Symptoms: Teleporting, Dropped action / rollback / Factors: Packet loss - Who: Specific zone/channel / When: When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Network infrastructure), Infra team (Server infrastructure) - Game team action items: Spread sends evenly across the tick (pacing), offset each server’s tick start time slightly. - Infra team action items: Network: use switches with bigger buffers, spread traffic (place servers across several switches and ports), watch drop counters per switch port. Servers/OS: cap each server’s total send rate (Linux tc shaper). - Ballpark numbers: A 10 Gbps port can send about 1.25 MB in 1 ms. If traffic from two ports converges on one port at once, 1.25 MB piles up every 1 ms. Even at 10% average utilization over 1 second, the port can overflow at the 1 ms scale. - On the graph: Rises with load (Switch port output drops) - Where to look: Collect output drop counters (ifOutDiscards, or output drops depending on the device) at the shortest interval you can on the switch ports the servers connect to and on the ports where their traffic converges, and match them against boss spawns and big battles. 1-second or 1-minute average utilization graphs won’t show it - Confirmed if: Average utilization is low, but output drops rise every time players crowd into one place, and at those moments many players report teleporting and skills not going off - Ruled out if: Drops rise steadily during hours of high average utilization: “Data center link saturation.” Input errors (CRC) rise: “Bad cables and port errors” - Check with: Infra tools (no game code needed) - Sources: - [High-Resolution Measurement of Data Center Microbursts (IMC 2017)](https://conferences.sigcomm.org/imc/2017/papers/imc17-final60.pdf) · ACM · Over 70% of data center bursts end within tens of µs; bursts drop packets even on ports with about 9% average utilization - [Data Center TCP (DCTCP) (SIGCOMM 2010)](https://conferences.sigcomm.org/sigcomm/2010/papers/sigcomm/p63.pdf) · ACM · Commodity switches have shallow buffers (48 ports share 4 MB, and one port can use up to about 700 KB); loss occurs when many flows converge on one port for a brief moment - [RFC 2863: The Interfaces Group MIB](https://www.rfc-editor.org/rfc/rfc2863) · IETF · ifOutDiscards: number of packets discarded even though no error occurred, for reasons such as freeing buffer space #### dc-uplink · Data center link saturation · Uplink saturation When patch distribution, log shipping, or backups share a link with the game, the link fills up. - Why → Effect → On screen: Bulk transfers take over the same link → Link queues and loss grow → Higher ping and teleporting across the whole server - Symptoms: Input lag, Teleporting / Factors: Latency, Packet loss - Who: Whole server / When: At regular intervals, When crowds gather - Primary owner: Infra team (Network infrastructure) / Also: Infra team (Server infrastructure) - Infra team action items: Network: prioritize game traffic (QoS), separate links for bulk transfers, alert on link utilization. Servers/OS: rate-limit backups, log shipping, and deployments, and run them during quiet hours. - On the graph: Hits a ceiling (Link utilization, RTT (ping)) - Where to look: Put the utilization of the data center link (uplink) interface (computed from SNMP ifHCInOctets and ifHCOutOctets) and its output drops (ifOutDiscards) on the same timeline as the backup, deployment, and log shipping schedules - Confirmed if: RTT and drops rise across the whole server when link utilization flattens at the bandwidth limit, and those times overlap with bulk transfer jobs - Ruled out if: Per-minute utilization well below the limit but drops still present: “Switch microbursts” - Check with: Infra tools (no game code needed) - Sources: - [RFC 4594: Configuration Guidelines for DiffServ Service Classes](https://www.rfc-editor.org/rfc/rfc4594) · IETF · Handle real-time interactive traffic such as games and bulk transfers such as backups in separate service classes - [RFC 7567: IETF Recommendations Regarding Active Queue Management](https://www.rfc-editor.org/rfc/rfc7567) · IETF · When more traffic arrives at a device than it can send out, queues build, and excessive queuing is a major cause of latency - [RFC 2863: The Interfaces Group MIB](https://www.rfc-editor.org/rfc/rfc2863) · IETF · ifHCInOctets and ifHCOutOctets: bytes received and sent on the interface (64-bit); ifOutDiscards: packets discarded without being sent #### dc-failover · Network equipment failover · Network device failover When a router or firewall fails and traffic switches to the standby unit (failover), everyone freezes for a few seconds. - Why → Effect → On screen: Switchover to standby equipment because of a failure or maintenance → The switchover takes a few seconds, and connections reset if session state isn’t synced → Every player on the server freezes at once; mass disconnects - Symptoms: Freeze, Disconnect / Factors: Packet loss - Who: Whole server / When: Randomly - Primary owner: Infra team (Network infrastructure) / Also: Game team (Server development), Game team (Client development) - Game team action items: Server: use timeouts that survive brief outages (a few seconds), let players resume their session with a session token when they reconnect after a disconnect. Client: reconnect automatically on disconnect (randomize retry intervals so everyone doesn’t pile in at once). - Infra team action items: Use redundancy that shares connection state, detect failures within 1 second with BFD, test failover regularly. - Ballpark numbers: About 1–3 seconds if the equipment detects the failure immediately. Without fast failure detection (BFD), relying only on default BGP timers, the route can be down for 90–180 seconds before neighboring equipment notices. - On the graph: Mass disconnect (Connections, total server traffic in/out) - Where to look: Router and firewall event logs (VRRP role changes, BFD and BGP sessions going down, failover records) next to total server connections and traffic at the same time - Confirmed if: At the failover time in the device log, traffic for every server behind that device drops to 0 for a few seconds, or connection counts fall together - Ruled out if: Only one server’s connections drop: “Server crash” or “NIC driver and firmware problems.” Device logs clean and the frozen server is a single cloud VM: “Cloud host maintenance and live migration” - Check with: Infra tools (no game code needed) - Sources: - [RFC 5880: Bidirectional Forwarding Detection (BFD)](https://www.rfc-editor.org/rfc/rfc5880) · IETF · Routing protocols’ Hello mechanisms take 1 second or more to detect a failure, so BFD was created to detect failures faster - [RFC 7938: Use of BGP for Routing in Large-Scale Data Centers](https://www.rfc-editor.org/rfc/rfc7938) · IETF · Relying only on BGP keepalives makes convergence slow; tearing down the session as soon as the link goes down detects failures within ms and reconverges - [RFC 5798: Virtual Router Redundancy Protocol (VRRP) Version 3 for IPv4 and IPv6](https://www.rfc-editor.org/rfc/rfc5798) · IETF · VRRP advertisements default to 1 second, and the backup takes over the role when advertisements stop for more than about 3 intervals (just over 3 seconds with default settings) #### dc-bad-cable · Bad cables and port errors · Bad cable / optics (CRC errors) Bad optics or a bad cable corrupt a steady share of the packets that pass through that path. - Why → Effect → On screen: Bit errors from bad optics or cables → The equipment silently drops corrupted packets → Only some servers or players using that path teleport or rubber-band from steady packet loss - Symptoms: Teleporting, Rubber-banding / Factors: Packet loss - Who: Specific zone/channel / When: Always - Primary owner: Infra team (Network infrastructure) - Infra team action items: Monitor and alert on port error (CRC) counters, replace parts such as optics and cables, take the problem link out and route around it until it’s replaced. - On the graph: Outliers only (CRC errors per port, loss rate per server/path) - Where to look: Check the CRC counters on both ends of the link. On switches, the port’s FCS errors (dot3StatsFCSErrors) and input errors (ifInErrors); on servers, crc under RX errors in ip -s -s link (kernel stat rx_crc_errors) - Confirmed if: CRC errors on one port keep rising regardless of traffic volume or time of day, and only servers and players passing through that port see loss - Ruled out if: No CRC errors and only output drops rising: congestion (“Switch microbursts,” “Data center link saturation”) - Check with: Infra tools (no game code needed) - Sources: - [Understanding and Mitigating Packet Corruption in Data Center Networks (SIGCOMM 2017)](https://www.microsoft.com/en-us/research/wp-content/uploads/2017/06/CorrOpt_SIGCOMM2017.pdf) · ACM · Analysis of 350,000 data center links: corruption comes from bad optics, damaged fiber, and dirty connectors; corruption rates stay steady regardless of utilization; problem links are taken out for repair while keeping enough paths in service - [RFC 3635: Definitions of Managed Objects for the Ethernet-like Interface Types](https://www.rfc-editor.org/rfc/rfc3635) · IETF · Frame check (FCS) failure counter (dot3StatsFCSErrors); these errors are included in input errors (ifInErrors) - [Interface statistics](https://docs.kernel.org/networking/statistics.html) · Linux kernel · rx_crc_errors: number of packets received with CRC errors; ip -s -s link shows errors by type #### dc-mtu · MTU mismatch (only large packets vanish) · MTU black hole If the MTU (the largest size that can be sent at once) shrinks somewhere along the path and the “packet too big” messages are blocked, only large packets keep vanishing. - Why → Effect → On screen: The MTU shrinks on a tunnel or VPN segment → A firewall blocks the “packet too big” messages (ICMP), so the sender never finds out → Freezes and then disconnects only when opening large screens such as the inventory or character list - Symptoms: Freeze, Disconnect, Can’t connect / infinite loading / Factors: Packet loss - Who: Specific region/ISP, Just me / When: During specific actions, Right after login or maintenance - Primary owner: Infra team (Network infrastructure) / Also: Infra team (Server infrastructure), Game team (Server development) - Game team action items: To lower it directly on the server side, set the socket’s maximum segment size with TCP_MAXSEG (splitting messages into smaller pieces in game code alone won’t prevent it), keep UDP packets at 1,200 bytes or less. - Infra team action items: Network: reduce TCP packet size on tunnel segments (MSS clamping), allow “packet too big” messages (ICMP) through firewalls and cloud network ACLs. Servers/OS: allow “packet too big” messages (ICMP) in server firewalls and cloud security groups too, turn on MTU probing in the server kernel (tcp_mtu_probing=1) as a last safety net that only kicks in after a freeze of a few seconds. - Ballpark numbers: Usually 1,500 bytes, shrinking to around 1,400 through a tunnel. - On the graph: Outliers only (Disconnects by region/ISP, failed large responses) - Where to look: From the affected player’s PC, ping the server with the don’t-fragment flag (DF) set, varying the size. On Windows, ping /f /l 1472 SERVER_IP; on Linux, ping -M do -s 1472 SERVER_IP (1,472 is the 1,500 MTU minus the 20-byte IP header and the 8-byte ICMP header). Lower the size step by step to find the largest size that gets through, and check whether the server-side security groups and firewalls allow ICMP “packet too big” messages (Fragmentation Needed) - Confirmed if: Small pings get through but the 1,472-byte DF ping fails (no reply, or an error saying fragmentation is needed), and the largest size that passes is small, around 1,400. Players in the same region freeze only when opening large screens - Ruled out if: The 1,472-byte DF ping also gets through: not a path MTU problem. Even small pings fail: ICMP itself is blocked, so this method can’t tell - Check with: The player’s own environment - Sources: - [RFC 2923: TCP Problems with Path MTU Discovery](https://www.rfc-editor.org/rfc/rfc2923) · IETF · If a firewall blocks ICMP (Fragmentation Needed), path MTU discovery fails and only large packets keep vanishing (black hole); pings and small messages still work, which makes it hard to diagnose - [Maximum transmission unit and maximum segment size](https://developers.cloudflare.com/magic-transit/reference/mtu-mss/) · Cloudflare · Internet path MTU 1,500, 1,476 through a GRE tunnel; limiting TCP MSS to 1,436 or less is recommended - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_mtu_probing=1 is normally off and turns on TCP path MTU probing only when it detects an ICMP black hole - [ping(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ping.8.html) · iputils · -M do sets the DF flag and refuses packets larger than the path MTU; -s sets the data size (default 56 bytes, plus the 8-byte ICMP header) - [ping](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/ping) · Microsoft · /f sets the DF flag and is used to find path MTU problems; /l sets the data size ### L6 Server network card (causes: 9) #### nic-irq · NIC interrupts concentrated on one core · Single-queue NIC / no RSS If the NIC sends every packet-arrival interrupt to a single CPU core, that core becomes the bottleneck. - Why → Effect → On screen: A single receive queue, or RSS (which spreads packets across cores) turned off → One core hits 100% and can’t pull packets off in time → Packet loss and latency across the whole server when players crowd in (teleporting, input lag) - Symptoms: Teleporting, Rubber-banding, Input lag / Factors: Packet loss, Latency - Who: Whole server / When: When crowds gather - Primary owner: Infra team (Server infrastructure) - Infra team action items: Configure RSS (spreading by the NIC) and RPS (spreading by the kernel), spread interrupts across several cores, make UDP queue selection include ports (rx-flow-hash udp4 sdfn in ethtool -N), keep interrupt-handling cores separate from the game tick thread’s cores, watch %soft per core. - Ballpark numbers: One core can push roughly hundreds of thousands of packets per second through the kernel, depending on packet size and settings. If per-core utilization shows receive processing (%soft in mpstat) piled onto a single core, this is what’s happening. - On the graph: Hits a ceiling (%soft per core, packets received per second) - Where to look: Check %soft (share of time spent on software interrupts) per core with mpstat -P ALL 1, which core each NIC queue’s interrupts go to in /proc/interrupts, the number of queues with ethtool -l, and packets per queue with ethtool -S (names vary by driver) - Confirmed if: One core’s %soft sits near 100% while the rest are idle, and interrupts and packets pile into one queue. From then on, packets received per second can’t climb any higher - Ruled out if: %soft spread evenly across cores: not this cause. CPU idle but loss present: “Cloud PPS limit exceeded” or “Ring buffer too small” - Check with: Infra tools (no game code needed) - Learn more: Even with multiple queues, if most traffic comes from a handful of addresses, such as gateways or proxies, it all lands in one queue. For UDP, some NICs pick the queue from addresses only by default, and traffic spreads evenly only after you change that to include ports. - Sources: - [Scaling in the Linux Networking Stack](https://docs.kernel.org/networking/scaling.html) · Linux kernel · RSS (the NIC spreads packets across multiple receive queues) and RPS (the kernel spreads them), giving each queue its own interrupt and spreading those across cores; RSS is recommended when receive interrupt handling is the bottleneck - [How to receive a million packets per second](https://blog.cloudflare.com/how-to-receive-a-million-packets/) · Cloudflare · Measurements where a receive queue served by a single core topped out at about 350,000–430,000 packets per second; a case where the NIC hashed UDP by IP address only and everything piled into one queue - [ethtool(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ethtool.8.html) · ethtool · The ethtool -N rx-flow-hash udp4 option that adds ports (f and n) to the UDP hash - [mpstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/mpstat.1.html) · sysstat · %soft: share of CPU time spent handling software interrupts; per core with -P ALL #### nic-ring · Ring buffer too small · RX ring buffer overflow If the NIC’s ring buffer, which briefly holds incoming packets, is small, a sudden burst overflows it and packets get dropped. - Why → Effect → On screen: The ring buffer is left at its small default (256–2,048 slots depending on the driver) → During a burst, the buffer overflows before the CPU can pull packets off → Loss only at burst moments (teleporting, skills not going off). No trace in the game server logs - Symptoms: Teleporting, Dropped action / rollback / Factors: Packet loss - Who: Whole server / When: When crowds gather - Primary owner: Infra team (Server infrastructure) - Infra team action items: Enlarge the ring buffer (ethtool -G), watch drop counters (such as rx_missed_errors in ethtool -S; names vary by driver). - Ballpark numbers: At 1 million packets per second, 1,024 slots fill in about 1 ms. If the CPU is late even once in that window, the buffer overflows. Most NICs can be raised to several thousand slots. - On the graph: Random spikes (NIC receive drop counters) - Where to look: Collect the receive drop counters in ethtool -S (rx_missed_errors, rx_fifo_errors, and so on; names vary by driver) and missed in ip -s -s link at short intervals, and check the current and maximum ring size with ethtool -g - Confirmed if: Drop counters rise at burst moments, and the current ring size is far below the maximum. Enlarging the ring reduces the drops - Ruled out if: Drop counters flat but loss present: the next stage in the kernel (“Kernel socket buffers too small”) or the network path. One core’s %soft at 100%: “NIC interrupts concentrated on one core” - Check with: Infra tools (no game code needed) - Sources: - [Linux Base Driver for Intel(R) Ethernet Network Connection (e1000)](https://docs.kernel.org/networking/device_drivers/ethernet/intel/e1000.html) · Linux kernel · e1000 receive descriptors (ring slots) default to 256 and can be raised to 4,096 - [drivers/net/ethernet/intel/ice/ice.h (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/drivers/net/ethernet/intel/ice/ice.h?h=v6.12) · Linux kernel · ice driver receive descriptors default to 2,048, maximum 8,160 - [Interface statistics](https://docs.kernel.org/networking/statistics.html) · Linux kernel · Packets the device drops for lack of buffers are counted in rx_missed_errors; driver-specific stats are visible with ethtool -S - [ethtool(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ethtool.8.html) · ethtool · -g shows ring size (current and maximum), -G changes it, -S shows driver-specific stats #### nic-coalesce · Excessive interrupt coalescing · Interrupt coalescing When the NIC collects packets and notifies the CPU once per batch to reduce CPU load, packets arrive later by the time spent collecting. - Why → Effect → On screen: The NIC collects packets for a set time or count before raising an interrupt → Packets wait while the batch fills → A small rise in latency. Usually tiny, but ms-scale if overdone - Symptoms: Input lag / Factors: Latency - Who: Whole server / When: Always - Primary owner: Infra team (Server infrastructure) - Infra team action items: Use adaptive coalescing, tune the values for game servers (ethtool -C). - Ballpark numbers: Typically tens to hundreds of µs. That’s negligible for most games, but aggressive settings can push it into the ms range. - On the graph: Always high (Round-trip time within the same data center) - Where to look: Check the current coalescing settings (adaptive-rx, rx-usecs, rx-frames) with ethtool -c, and compare ping round-trip time to another server in the same data center before and after changing them - Confirmed if: rx-usecs is set high (hundreds of µs or more), and lowering it cuts round-trip time within the data center by about the same amount - Ruled out if: Round-trip time unchanged after lowering it: not this cause - Check with: Infra tools (no game code needed) - Sources: - [Linux Driver for Intel(R) Ethernet Network Connection (e1000e)](https://docs.kernel.org/networking/device_drivers/ethernet/intel/e1000e.html) · Linux kernel · Default is adaptive interrupt moderation at 4,000–20,000 interrupts per second (50–250 µs apart); fewer interrupts save CPU but add latency - [Linux Base Driver for Intel(R) Ethernet Network Connection (e1000)](https://docs.kernel.org/networking/device_drivers/ethernet/intel/e1000.html) · Linux kernel · RxIntDelay can delay receive interrupts in 1.024 µs units up to 65,535 (about 67 ms), and larger values increase receive latency - [ethtool(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ethtool.8.html) · ethtool · The adaptive-rx, rx-usecs, and rx-frames settings of ethtool -C #### nic-cloud-pps · Cloud PPS limit exceeded · Cloud PPS / bandwidth allowance Each cloud instance type has limits on packets per second and bandwidth, and traffic over them is silently dropped. - Why → Effect → On screen: Rising CCU pushes packets per second over the instance limit → The cloud network drops the excess → Teleporting and skills not going off from unexplained packet loss. Server CPU has headroom - Symptoms: Teleporting, Dropped action / rollback / Factors: Packet loss - Who: Whole server / When: When crowds gather, Evening peak hours - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development), External (External) - Game team action items: Combine packets (one tick’s messages in one packet), avoid sending tiny packets frequently. - Infra team action items: Check and alert on limit-exceeded counters (pps_allowance_exceeded, conntrack_allowance_exceeded, and so on for AWS), move to a larger instance, avoid connection tracking limits with a security group setup that creates no tracking. - External action items: Ask the cloud provider for packets-per-second and connection tracking limits per instance type. - Ballpark numbers: Limits differ by instance size, and the packets-per-second limit is often not published. The “up to 10 Gbps” on small instances is a burst speed available only while credits last (usually 5–60 minutes); the normal baseline speed is much lower. - On the graph: Hits a ceiling (Packets per second, allowance exceeded counters) - Where to look: Collect the ENA counters pps_allowance_exceeded, bw_in_allowance_exceeded, bw_out_allowance_exceeded, and conntrack_allowance_exceeded from ethtool -S at short intervals and view them alongside packets per second. You can also publish these counters with the CloudWatch agent and set alarms on them - Confirmed if: Allowance exceeded counters rise at the times of loss, and packets per second stops climbing at a fixed value. Server CPU has headroom - Ruled out if: Exceeded counters flat: not this cause. One core’s %soft at 100%: “NIC interrupts concentrated on one core” - Check with: Infra tools (no game code needed) - Learn more: conntrack_allowance_exceeded means the connection tracking table was full and new connections were dropped. If the table has room and only idle connections drop because their tracking expired, see the “Cloud security group connection tracking expiry” entry. - Sources: - [Monitor network performance for ENA settings on your EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-network-performance-ena.html) · AWS · Each instance has bandwidth, PPS, and connection tracking limits, and traffic over them is queued and then dropped; pps_allowance_exceeded and conntrack_allowance_exceeded counters - [Amazon EC2 instance network bandwidth](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-instance-network-bandwidth.html) · AWS · The “up to N Gbps” of instances with 16 vCPUs or fewer is a burst that spends network I/O credits (usually 5–60 minutes), falling back to baseline bandwidth when the credits run out #### nic-saturate · NIC bandwidth saturation · NIC bandwidth saturation Running a 1 Gbps or 10 Gbps card at its limit makes the transmit queue grow until packets get dropped. - Why → Effect → On screen: More broadcasts push traffic to the card’s limit → The transmit queue grows, and packets are dropped when it overflows → Latency and loss across the whole server (input lag, teleporting) - Symptoms: Input lag, Teleporting / Factors: Latency, Packet loss - Who: Whole server / When: When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Reduce traffic (area of interest, compression, send only what changed). - Infra team action items: Upgrade the card (a faster NIC, or a larger instance in the cloud), alert on NIC utilization. - On the graph: Hits a ceiling (NIC transmit volume, transmit drops) - Where to look: Compare txkB/s and %ifutil (utilization relative to interface speed) from sar -n DEV 1 with the NIC speed or instance bandwidth, alongside TX dropped from ip -s link - Confirmed if: Transmit volume flattens near the NIC or instance bandwidth, and from then on transmit drops and server-wide latency rise - Ruled out if: Bandwidth has headroom: not this cause. Many small packets and loss: “Cloud PPS limit exceeded” - Check with: Infra tools (no game code needed) - Sources: - [Interface statistics](https://docs.kernel.org/networking/statistics.html) · Linux kernel · tx_dropped: number of packets dropped during transmission for lack of resources - [Amazon EC2 instance network bandwidth](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-instance-network-bandwidth.html) · AWS · The bandwidth an instance can use is set by its vCPU count (instance size) - [sar(1) — Linux manual page](https://man7.org/linux/man-pages/man1/sar.1.html) · sysstat · rxkB/s, txkB/s, and %ifutil (utilization relative to interface speed) from -n DEV #### nic-noisy · Virtualization overhead and noisy neighbors · Noisy neighbors in virtualization When other virtual machines on the same physical server use a lot of network or CPU, your server’s processing gets delayed at irregular times. - Why → Effect → On screen: Other VMs on the same physical server use a lot of resources → Packet processing on your VM is delayed at irregular times → Occasional jitter (variation in packet arrival times) with no obvious cause, showing up as stutter - Symptoms: Stutter / Factors: Jitter - Who: Whole server / When: Randomly - Primary owner: Infra team (Server infrastructure) / Also: External (External) - Infra team action items: Use dedicated hosts or instances with guaranteed performance, stop and restart instances with persistent jitter to move them to another host. - External action items: Report the problem host to the cloud provider. - On the graph: Random spikes (Round-trip jitter within the data center, %steal) - Where to look: Ping another server in the same data center continuously to record round-trip jitter, and compare it, along with %steal from mpstat, against other instances with the same configuration - Confirmed if: Only this instance shows irregular spikes in round-trip jitter or %steal, while others with the same configuration stay quiet. Stopping and restarting it to move it to another host makes the problem go away - Ruled out if: All instances with the same configuration spike the same way: not a host problem. Look at game server load or the network path - Check with: Infra tools (no game code needed) - Sources: - [How EC2 instance stop and start works](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/how-ec2-instance-stop-start-works.html) · AWS · Stopping and starting an instance usually moves it to a new host (except dedicated hosts) - [mpstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/mpstat.1.html) · sysstat · %steal: share of time this virtual CPU was forced to wait while the hypervisor ran other virtual CPUs #### nic-host-maintenance · Cloud host maintenance and live migration · Cloud host maintenance / live migration When a cloud provider performs maintenance on a physical server (host), it moves VMs to another host (live migration) or pauses them briefly. The whole server freezes during that time, and if the pause is long, connections drop. - Why → Effect → On screen: The provider moves the VM to another host, or pauses it briefly, for host maintenance or a predicted failure → During the move, CPU, memory, and network slow down, and at the end the VM stops completely for a moment (from under 1 second to around 30 seconds, depending on the provider and method) → Everyone on the server freezes at once and then sees fast-forward and teleporting; if the freeze outlasts the timeout, mass disconnects - Symptoms: Freeze, Fast-forward, Teleporting, Disconnect / Factors: Stall, Packet loss - Who: Whole server / When: Randomly - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development), External (External) - Game team action items: Use timeouts that survive pauses of a few seconds, cap how many ticks the server catches up after a pause, compute elapsed time with a monotonic clock, have a procedure that saves progress and moves players to another server when a maintenance notice arrives. - Infra team action items: Subscribe to and alert on maintenance notices (Google Cloud maintenance-event, AWS scheduled events and AWS Health, Azure Scheduled Events), replace servers ahead of time during low-traffic hours when a notice arrives, reschedule maintenance where the provider allows it (Azure Maintenance Configuration, AWS scheduled events depending on type), compare maintenance records with incident records. - External action items: Ask the cloud provider about maintenance schedules and impact, report instances that keep pausing. - Ballpark numbers: Google Compute Engine says live migration pauses are usually much shorter than 1 second, and the system clock can jump forward by up to 5 seconds during the pause. The maintenance-event metadata value changes 60 seconds before the move (if you have queried it at least once beforehand). On Azure, maintenance that doesn’t need a reboot almost always pauses the VM for under 10 seconds, and rarely (no more than once every 18 months for general-purpose sizes) for about 30 seconds; live migration usually takes no more than 5 seconds. Azure Scheduled Events gives notice of these pauses (Freeze) at least 15 minutes ahead. If host hardware fails suddenly, though, recovery starts right away with no notice. - On the graph: Gap then burst (Server packets sent/received, tick interval) - Where to look: Match the freeze time against the provider’s records. Google Cloud: compute.instances.migrateOnHostMaintenance in the audit logs; AWS: scheduled events in describe-instance-status and AWS Health; Azure: Microsoft.Compute/virtualMachines/liveMigration/action in the Activity Log and the time the VM availability metric (VmAvailabilityMetric) dropped to 0. Inside the server, check whether metrics and logs have a gap during the freeze and whether the clock jumped right after (time sync logs) - Confirmed if: The time the whole server froze overlaps with a maintenance or migration time in the provider’s records, and every metric and log inside the server is blank for those few seconds - Ruled out if: Not in the provider’s records and short freezes recur often: “CPU steal (virtual machines).” NIC reset entries in the kernel log: “NIC driver and firmware problems” - Check with: Infra tools (no game code needed) - Learn more: AWS gives notice through scheduled events. system-reboot means the instance will be rebooted and moved to a new host; system-maintenance means network or power maintenance may affect it briefly. Even if the pause lasts only a few seconds, clients that got no ACK for packets sent to the server during that time keep doubling their retransmission wait, so TCP connections can stay stalled for longer after the pause ends (“TCP RTO and exponential backoff”). When the VM wakes up, its clock can jump and lead to a “System clock jump (NTP step),” and failed load balancer health checks may take the server out of rotation for a while. Instances that can’t be moved (such as Google Cloud bare metal instances) are stopped or restarted during maintenance. - Sources: - [Live migration process during maintenance events](https://cloud.google.com/compute/docs/instances/live-migration-process) · Google Cloud · Live migration pauses are usually much shorter than 1 second; the system clock jumps forward by up to 5 seconds during the pause; disk, CPU, memory, and network performance drop briefly during the move; VMs that don’t live-migrate are terminated for maintenance (bare metal instances don’t support live migration) - [Query metadata server for maintenance event notices](https://cloud.google.com/compute/docs/metadata/getting-live-migration-notice) · Google Cloud · The maintenance-event metadata value changes 60 seconds before live migration (when the VM is set to live-migrate and the value was queried at least once since the last maintenance) - [Monitor and plan for a host maintenance event](https://cloud.google.com/compute/docs/instances/monitor-plan-host-maintenance-event) · Google Cloud · Maintenance leaves a compute.instances.migrateOnHostMaintenance system event in the audit logs - [Scheduled events for Amazon EC2 instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-instances-status-check_sched.html) · AWS · Scheduled event types (system-reboot reboots and moves to a new host; system-maintenance means brief impact from network or power maintenance), notified by email and AWS Health, checked with describe-instance-status, reschedulable for some types - [Maintenance and updates](https://learn.microsoft.com/en-us/azure/virtual-machines/maintenance-and-updates) · Microsoft Azure · Maintenance without a reboot almost always pauses for under 10 seconds, rarely (no more than once every 18 months for general-purpose sizes) for about 30 seconds, and live migration usually 5 seconds or less; the clock syncs automatically after the pause; long-lived TCP connections may drop, or recovery may take longer as peers retransmit data sent to the paused VM with exponential backoff; load balancer health checks mark the VM unhealthy within about 10 seconds; confirm with Microsoft.Compute/virtualMachines/liveMigration/action in the Activity Log and VmAvailabilityMetric dropping to 0 during the pause; pick when maintenance applies with Maintenance Configuration - [Scheduled Events for Linux VMs in Azure](https://learn.microsoft.com/en-us/azure/virtual-machines/linux/scheduled-events) · Microsoft Azure · Freeze (a pause of a few seconds; CPU and network may stop) is announced at least 15 minutes ahead; for host hardware failures, recovery starts right away with no notice period #### nic-reset · NIC driver and firmware problems · NIC hang / reset When a driver bug or a malfunctioning feature hangs the card, all traffic in and out stops while it restarts. - Why → Effect → On screen: Driver bug, malfunctioning offload feature → The NIC hangs and restarts (a few seconds) → Everyone on that server freezes together and then teleports or disconnects - Symptoms: Freeze, Disconnect / Factors: Packet loss - Who: Whole server / When: Randomly, The longer it runs - Primary owner: Infra team (Server infrastructure) - Infra team action items: Check and alert on “transmit queue … timed out” and “Link is Down” entries in the kernel log, update drivers and firmware, turn off the problem feature (such as an offload). - On the graph: Gap then burst (Server packets sent/received) - Where to look: Search the kernel log with dmesg for “NETDEV WATCHDOG … transmit queue N timed out”, driver resets, and “Link is Down” or “Link is Up” entries, and check the server’s packets sent and received at those times - Confirmed if: At the time of the freeze, the kernel log has a transmit queue timeout or link down/up entries, and packets sent and received drop to 0 for those few seconds - Ruled out if: Kernel log clean and the switch-side port fine: a stall in the game server process (“Server GC stop-the-world pause,” “Deadlock”) or “Network equipment failover” - Check with: Infra tools (no game code needed) - Sources: - [net/sched/sch_generic.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/sched/sch_generic.c?h=v6.12) · Linux kernel · When a transmit queue stalls, the kernel watchdog logs “NETDEV WATCHDOG … transmit queue N timed out” and calls the driver’s reset function - [drivers/net/ethernet/intel/ixgbe/ixgbe_main.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/drivers/net/ethernet/intel/ixgbe/ixgbe_main.c?h=v6.12) · Linux kernel · The ixgbe driver resets the adapter on a transmit hang and logs “NIC Link is Down” when the link drops - [ethtool(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ethtool.8.html) · ethtool · The ethtool -K option for turning offload features off one at a time #### nic-offload · GRO/LRO batching delay · GRO/LRO batching GRO and LRO bundle several packets into one to reduce CPU load. Depending on settings, a small game packet may wait briefly for the next packet to bundle with. - Why → Effect → On screen: The NIC and kernel bundle arriving packets together for processing → With hardware aggregation (LRO) or a batching wait time setting turned on, packets wait briefly for the next one → A small rise in latency (usually tens of µs or less) - Symptoms: Input lag / Factors: Latency - Who: Whole server / When: Always - Primary owner: Infra team (Server infrastructure) - Infra team action items: Tune for game traffic (turn off LRO, check the batching wait time setting), check it after other causes since the effect is usually small. - On the graph: Always high (Round-trip time within the same data center) - Where to look: Check lro and gro status with ethtool -k and the device’s gro_flush_timeout sysfs setting, and compare round-trip time for small packets within the data center before and after changing them - Confirmed if: LRO is on or gro_flush_timeout is above 0, and turning it off or setting it to 0 reduces round-trip time for small packets - Ruled out if: Difference stays within a few µs after the change: not this cause - Check with: Infra tools (no game code needed) - Sources: - [NAPI](https://docs.kernel.org/networking/napi.html) · Linux kernel · A large gro_flush_timeout batches more work together but adds latency under low load - [Linux Base Driver for the Intel(R) Ethernet 10 Gigabit PCI Express Adapters (ixgbe)](https://docs.kernel.org/networking/device_drivers/ethernet/intel/ixgbe.html) · Linux kernel · GRO saves CPU by merging incoming traffic into large chunks; it evolved from LRO - [ethtool(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ethtool.8.html) · ethtool · The gro and lro on|off settings of ethtool -K ### L7 Server OS (kernel) (causes: 14) #### so-backlog · Connection queue (listen backlog) overflow · Listen backlog / SYN queue overflow When tens of thousands of players connect at once right after maintenance, the kernel’s connection queue (listen backlog) overflows and connection attempts are dropped. - Why → Effect → On screen: As maintenance ends, connections pour in faster than the game server can accept them → The kernel’s connection queue (listen backlog: the smaller of the value the server code passes to listen and the kernel cap) fills up → Connection attempts are dropped and retried again and again: can’t connect / infinite loading - Symptoms: Can’t connect / infinite loading / Factors: Packet loss - Who: Whole server / When: Right after login or maintenance - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure), Game team (Client development) - Game team action items: Server: raise the listen value passed in code, keep the thread that accepts connections from stalling on other work, add a login queue system. Client: lengthen the retry interval (randomized to spread retries out). - Infra team action items: Raise the kernel’s somaxconn (it only helps when the listen value in the server code goes up too), keep SYN cookies on, monitor the overflow count (TcpExtListenOverflows in nstat). - Ballpark numbers: The Linux kernel cap (somaxconn) defaults to 4,096 since 5.4 (128 before that), but if the server code passes a smaller value to listen, that value is the limit. When the queue is full, Linux silently drops connection requests without returning an error. The client OS resends a few times, starting after 1 second, so players just see a long loading screen with no “Connection failed” message. A Windows server sends back a refusal, so the client sees “Connection failed” right away. - On the graph: Surge after opening (Connection queue overflows (ListenOverflows), connection attempts) - Where to look: Increase in TcpExtListenOverflows and TcpExtListenDrops from nstat -az; with ss -ltn, the listening socket’s Recv-Q (connections waiting for accept) compared with its Send-Q (the backlog limit) - Confirmed if: ListenOverflows rises when connections pile in, and the listening socket’s Recv-Q sits at its Send-Q value - Ruled out if: ListenOverflows unchanged: not this cause. Connection established but loading never finishes: “Login storm and N+1 queries.” Blocked at an exact player count: “File descriptor limit” - Check with: Infra tools (no game code needed) - Sources: - [listen(2) — Linux manual page](https://man7.org/linux/man-pages/man2/listen.2.html) · Linux man-pages · A listen backlog larger than somaxconn is silently truncated; somaxconn defaults to 4,096 (since 5.4; 128 before); when the queue is full, requests may be ignored and left to client retries - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_syn_retries: the connection request (SYN) is resent several times, with a 1 s wait before the first retransmission; tcp_abort_on_overflow off by default (no refusal sent on overflow); tcp_syncookies on by default - [listen function (winsock2.h)](https://learn.microsoft.com/en-us/windows/win32/api/winsock2/nf-winsock2-listen) · Microsoft · On Windows, when the queue is full the client gets a WSAECONNREFUSED error - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · TcpExtListenOverflows: number of connection requests (SYN) dropped because the accept queue was full; TcpExtListenDrops goes up along with it - [net/ipv4/tcp_diag.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_diag.c?h=v6.12) · Linux kernel · Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed #### so-fd · File descriptor limit · File descriptor limit (ulimit) Every connection needs a file descriptor (fd: the number the OS gives an open file or socket), and the number of fds one process can open is capped. - Why → Effect → On screen: Concurrent users reach the process’s file descriptor limit → The server can’t accept new connections (Too many open files). Opening log files and DB connections fails too → Past an exact player count, nobody gets in: can’t connect / infinite loading - Symptoms: Can’t connect / infinite loading / Factors: Packet loss - Who: Whole server / When: Right after login or maintenance, When crowds gather - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: Always close the socket when a connection ends (to prevent fd leaks), when accept fails with EMFILE (out of fds), pause accepting briefly or accept with a spare fd kept in reserve and close it right away (so the server doesn’t burn CPU handling the same new-connection notification over and over). - Infra team action items: Check ulimit and the service settings (LimitNOFILE in systemd), alert when usage nears the limit. - Ballpark numbers: On Linux, the limit is still often 1,024 unless the service is configured otherwise. Game servers usually raise it to tens of thousands or hundreds of thousands. Windows has no default limit this low. - On the graph: Hits a ceiling (Open fds of the process, concurrent users) - Where to look: fd-nr (open file descriptors) of the game server process from pidstat -v, the open-files limit in /proc/PID/limits, and accept failures (EMFILE, Too many open files) in the server log - Confirmed if: The fd count flattens at the limit, and from that moment accept fails with EMFILE - Ruled out if: fd count well below the limit: not this cause. Connection requests dropped in the kernel: “Connection queue (listen backlog) overflow.” Connection tracking involved: “Server conntrack table full” - Check with: Infra tools (no game code needed) - Learn more: Connections that weren’t accepted stay in the kernel’s connection queue (listen backlog), so depending on the code, the server may keep getting “new connection” notifications and waste CPU on them. - Sources: - [systemd-system.conf(5) — Linux manual page](https://man7.org/linux/man-pages/man5/systemd-system.conf.5.html) · systemd · DefaultLimitNOFILE for services defaults to 1024:524288 (soft limit 1,024) - [accept(2) — Linux manual page](https://man7.org/linux/man-pages/man2/accept.2.html) · Linux man-pages · When a process hits its fd limit, accept fails with EMFILE - [Maximum Number of Sockets Supported](https://learn.microsoft.com/en-us/windows/win32/winsock/maximum-number-of-sockets-supported-2) · Microsoft · Windows Winsock limits the number of sockets only by available memory - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · fd-nr in -v: number of file descriptors the process has open - [proc_pid_limits(5) — Linux manual page](https://man7.org/linux/man-pages/man5/proc_pid_limits.5.html) · Linux man-pages · /proc/PID/limits lists the soft and hard values of each per-process resource limit #### so-sockbuf · Kernel socket buffers too small · Small socket buffers With small send and receive buffers, a burst of traffic makes the kernel drop packets arriving over UDP, and TCP sends block because the buffer has no room left. - Why → Effect → On screen: SO_SNDBUF and SO_RCVBUF left at their defaults or set too small → During a burst, or while the receiving thread pauses briefly, the UDP receive buffer overflows and drops packets; TCP waits because the send buffer has no room → Teleporting (UDP loss) or fast-forward (TCP waiting) - Symptoms: Teleporting, Fast-forward / Factors: Packet loss, Stall - Who: Whole server / When: When crowds gather - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: Set buffer sizes (SO_SNDBUF, SO_RCVBUF) in code to match the traffic, remember that setting a TCP buffer size explicitly turns off Linux autotuning, keep sizes moderate because oversized buffers let stale data pile up and add delay, keep the receiving thread from stalling. - Infra team action items: Tune the kernel caps (rmem_max, wmem_max; buffer sizes set in code can’t exceed them either) and the default (rmem_default), monitor the buffer overflow counter (RcvbufErrors). - Ballpark numbers: The default UDP receive buffer on Linux is about 208 KB. Even a small packet takes up far more kernel memory than its actual size, so tens to hundreds of packets fill it. On a server receiving 100,000 packets per second, it overflows if the receiving thread stalls for just a few ms. - On the graph: Random spikes (UDP receive buffer overflows (UdpRcvbufErrors)) - Where to look: Increase in UdpRcvbufErrors from nstat -az, and skmem in ss -uamn (rb is the receive buffer size, d is packets dropped because they couldn’t be queued to the socket); for TCP, whether send-queue memory (w) in the skmem of ss -tm has reached the send buffer size (tb) - Confirmed if: UdpRcvbufErrors (or the socket’s d) rises during bursts or when the receiving thread stalls, with rb near the default (about 208 KB). For TCP, w stays pinned at tb and send blocks - Ruled out if: Counters unchanged but loss present: the NIC stage (“Ring buffer too small”) or the network path - Check with: Infra tools (no game code needed) - Sources: - [socket(7) — Linux manual page](https://man7.org/linux/man-pages/man7/socket.7.html) · Linux man-pages · SO_RCVBUF and SO_SNDBUF default to rmem_default and wmem_default and are capped at rmem_max and wmem_max; the kernel doubles the value you set - [include/net/sock.h (Linux v6.18)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/net/sock.h?h=v6.18) · Linux kernel · The default socket buffer is defined as 256 packets of 256 bytes including sk_buff overhead (SKB_TRUESIZE(256)×256); even a small frame counts as sk_buff + MTU (about 208 KB is the computed value on x86-64) - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_rmem, tcp_wmem: setting SO_RCVBUF or SO_SNDBUF directly turns off autotuning for that socket - [net/ipv4/udp.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/udp.c?h=v6.12) · Linux kernel · When the UDP receive queue exceeds the socket buffer size, packets are dropped immediately and RcvbufErrors goes up - [net/ipv4/proc.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/proc.c?h=v6.12) · Linux kernel · Counter names shown by nstat: RcvbufErrors and SndbufErrors in the Udp group - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · skmem in -m: rb receive buffer size, tb send buffer size, w send-queue memory, d packets dropped before reaching the socket #### so-context · Too many threads and context switching · Thread oversubscription, context switching Running far more threads than there are cores makes the OS spend CPU just switching between them. - Why → Effect → On screen: Hundreds to thousands of threads, for example one thread per connection → Higher context-switching cost (swapping out the running thread) and more cache misses → CPU is busy but throughput is low and ticks are uneven: stutter, slow motion - Symptoms: Stutter, Slow motion / Factors: Stall, Jitter - Who: Whole server / When: When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Match thread count to core count, use asynchronous I/O (epoll, IOCP). - Infra team action items: Monitor context switches and runnable threads (cs and r in vmstat). - Ballpark numbers: One context switch costs a few µs, and more once you add the cache misses that follow. - On the graph: Rises with load (Context switches per second, runnable threads) - Where to look: cs (context switches per second) and r (running or waiting for CPU) from vmstat 1 compared with the core count; voluntary (cswch/s) and involuntary (nvcswch/s) context switches per game server thread from pidstat -w -t - Confirmed if: As concurrent users grow, r climbs far above the core count, cs spikes with it, and hundreds of threads show many involuntary context switches - Ruled out if: r stays at or below the core count: not this cause. Mostly voluntary switches: threads are waiting on locks or I/O (“Lock contention,” “Blocking I/O design”) - Check with: Infra tools (no game code needed) - Sources: - [Quantifying The Cost of Context Switch (ExpCS 2007)](https://www.usenix.org/legacy/events/expcs07/papers/2-li.pdf) · ACM · Direct cost of a context switch about 3.8 µs; indirect cost including cache effects ranges from a few µs to over 1,000 µs (in the measured setup) - [vmstat(8) — Linux manual page](https://man7.org/linux/man-pages/man8/vmstat.8.html) · procps-ng · The cs (context switches per second) and r (processes running or waiting to run) fields - [I/O Completion Ports](https://learn.microsoft.com/en-us/windows/win32/fileio/i-o-completion-ports) · Microsoft · Handle many asynchronous I/Os with a pre-created thread pool and IOCP, and match the number of concurrently running threads to CPU concurrency - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · In -w, cswch/s counts voluntary context switches (the task stopped on its own to wait for a resource) and nvcswch/s counts involuntary ones (forced out after using up its time slice); -t shows them per thread #### so-steal · CPU steal (virtual machines) · CPU steal time While the physical server (hypervisor) briefly gives a virtual machine’s CPU time to another VM (CPU steal), the game server stalls. - Why → Effect → On screen: Other VMs on the same host use a lot of CPU → The game server’s VM loses its turn on the CPU for a few ms to tens of ms at a time → Unexplained tick-time spikes: stutter, freeze - Symptoms: Stutter, Freeze / Factors: Stall - Who: Whole server / When: Randomly - Primary owner: Infra team (Server infrastructure) / Also: External (External) - Infra team action items: Monitor steal (st in top and vmstat), use dedicated cores or hosts, avoid burstable instances that slow down when CPU credits run out, stop and start instances with persistently high steal to move them to another host. - External action items: Report hosts with persistently high steal to the cloud provider. - On the graph: Random spikes (%steal, server tick time) - Where to look: %steal from mpstat -P ALL 1 on the same time axis as server tick time - Confirmed if: %steal spikes whenever the tick spikes, and drops after a stop and start moves the instance to another host - Ruled out if: %steal near 0 while ticks still spike: a cause inside the game server (“Server GC stop-the-world pause,” “Lock contention”). In a container: “Container CPU throttling (CFS quota)” - Check with: Infra tools (no game code needed) - Sources: - [proc_stat(5) — Linux manual page](https://man7.org/linux/man-pages/man5/proc_stat.5.html) · Linux man-pages · steal: time lost while other operating systems ran in a virtualized environment - [Standard mode for burstable performance instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/burstable-performance-instances-standard-mode.html) · AWS · Burstable instances spend credits to run above baseline performance, and when the credits run out, CPU utilization drops to the baseline level - [How EC2 instance stop and start works](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/how-ec2-instance-stop-start-works.html) · AWS · Stopping and starting an instance usually moves it to a new host - [mpstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/mpstat.1.html) · sysstat · %steal: share of time this virtual CPU was forced to wait while the hypervisor ran other virtual CPUs #### so-cpu-quota · Container CPU throttling (CFS quota) · Container CPU throttling (CFS quota) With a CPU limit on a container, the moment it uses up its quota within a set period (usually 100 ms), it is forced to stop for the rest of that period (throttling). - Why → Effect → On screen: A CPU limit is set on the game server container (in Kubernetes, for example) → A burst of tick work uses up the quota, and the server stops for tens of ms until the next period → Average CPU is low, yet ticks spike periodically: stutter, slow motion - Symptoms: Stutter, Slow motion / Factors: Stall, Jitter - Who: Whole server / When: When crowds gather, Randomly - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: Match the worker thread count to the CPU limit (so the runtime doesn’t create as many threads as the host has cores). - Infra team action items: Set a generous CPU limit or remove it and assign dedicated cores, monitor the throttle count (nr_throttled). - Ballpark numbers: On a server limited to 2 cores, 8 threads working at once use up the quota for a 100 ms period in 25 ms and then stop for 75 ms. - On the graph: Rises with load (Throttle count (nr_throttled), server tick time) - Where to look: Increase in nr_throttled and throttled_usec (nr_throttled and throttled_time on cgroup v1) in the container cgroup’s cpu.stat, alongside server tick time - Confirmed if: Average CPU utilization stays below the limit, yet nr_throttled and throttled_usec keep rising, lining up with tick spikes - Ruled out if: nr_throttled not rising: not this cause. The VM itself being held back: “CPU steal (virtual machines)” - Check with: Infra tools (no game code needed) - Sources: - [CFS Bandwidth Control](https://docs.kernel.org/scheduler/sched-bwc.html) · Linux kernel · Threads that use up the quota for a period stop until the next period (throttling), default period 100 ms, nr_throttled statistic - [Control Group v2](https://docs.kernel.org/admin-guide/cgroup-v2.html) · Linux kernel · cpu.max takes the form “$MAX $PERIOD” (quota, period), and the default is “max 100000” (100 ms period) - [Resource Management for Pods and Containers](https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/) · Kubernetes · A container’s CPU limit is a hard limit that the kernel enforces with CPU throttling #### so-cstate · Latency spikes from server power management (C-states, frequency scaling) · CPU power management latency (C-states, frequency scaling) Idle CPU cores drop into deep power-saving states (C-states) and lower their frequency to save power. Waking up and raising the frequency when a packet or timer arrives takes time, which adds delay to handling small packets. - Why → Effect → On screen: The OS frequency scaling policy (governor) or the BIOS power settings allow deep C-states and low frequencies → An idle core is up to hundreds of µs late every time it wakes from a deep power-saving state, and a frequency pinned low slows the tick computation itself → Usually hard to notice, but with many server-to-server calls it adds up to input lag that gets worse when the server is quiet. With the frequency pinned low, ticks fall behind when crowds gather: slow motion - Symptoms: Input lag, Slow motion / Factors: Latency, Jitter, Stall - Who: Whole server / When: Always, Randomly - Primary owner: Infra team (Server infrastructure) - Infra team action items: Set the BIOS power settings to performance, set the OS governor to performance (scaling_governor in cpufreq), limit deep C-states on latency-sensitive servers (tuned latency-performance profile, /dev/cpu_dma_latency in PM QoS, the intel_idle.max_cstate kernel parameter), and after the change compare round-trip time within the data center, tick-time jitter, and power usage. - Ballpark numbers: According to the intel_idle driver tables in Linux 6.12, Intel server CPUs take 1–2 µs to wake from the shallow C1 state and 133 µs (Skylake-SP) to 290 µs (Sapphire Rapids) from the deep C6 state. One wakeup is small, but when a request passes through several servers, the delays add up. The kernel picks deeper states the longer it expects to stay idle, so this shows up more on quiet servers where packets arrive only now and then. The generic cpufreq powersave governor pins the frequency at the lowest allowed value (the intel_pstate algorithm of the same name adjusts it to the load). - On the graph: Always high (Round-trip time within the same data center, core frequency) - Where to look: Per-core C-state residency and actual frequency from cpupower monitor; for each state under /sys/devices/system/cpu/cpu0/cpuidle/, its name, latency (µs to wake up), and usage; scaling_governor in cpufreq; and the current profile from tuned-adm active - Confirmed if: Idle cores sit in the deepest C-state for long stretches or the frequency is pinned near the minimum, and switching to the performance governor and shallow C-states reduces the round-trip time and jitter of small requests - Ruled out if: Difference after the change stays within tens of µs: safe to ignore this cause. Spikes in the ms range: “CPU steal (virtual machines)” or another layer - Check with: Infra tools (no game code needed) - Learn more: For bare-metal servers in a data center, check the BIOS (firmware) power settings together with the OS settings. In the cloud, only some instance types let the OS change C-states and frequency, and AWS defaults to maximum performance, so most instances can be left as they are. The tuned latency-performance profile on Red Hat-based systems sets the governor to performance and uses PM QoS to allow only shallow C-states. Turning off power saving raises power consumption, so apply it only to latency-sensitive servers. - Sources: - [CPU Idle Time Management](https://docs.kernel.org/admin-guide/pm/cpuidle.html) · Linux kernel · Each power-saving state has a wakeup time (exit latency) and a minimum stay (target residency), and deeper states are chosen based on expected idle time; per-state latency, usage, and time in sysfs; PM QoS (/dev/cpu_dma_latency) and intel_idle.max_cstate limit deep states - [drivers/idle/intel_idle.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/drivers/idle/intel_idle.c?h=v6.12) · Linux kernel · Wakeup times of C-states on Intel server CPUs: Skylake-SP C1 2 µs, C1E 10 µs, C6 133 µs; Ice Lake C6 170 µs; Sapphire Rapids C1 1 µs, C6 290 µs - [CPU Performance Scaling](https://docs.kernel.org/admin-guide/pm/cpufreq.html) · Linux kernel · Check and change the governor with scaling_governor; performance requests the highest allowed frequency, powersave the lowest - [intel_pstate CPU Performance Scaling Driver](https://docs.kernel.org/admin-guide/pm/intel_pstate.html) · Linux kernel · The intel_pstate powersave algorithm differs from the generic powersave governor and scales with load (similar to schedutil and ondemand) - [Chapter 2. Getting started with TuneD](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/9/html/monitoring_and_managing_system_status_and_performance/getting-started-with-tuned_monitoring-and-managing-system-status-and-performance) · Red Hat · The latency-performance profile turns off power-saving features, sets the governor to performance, and uses PM QoS to allow only shallow C-states; check the current profile with tuned-adm active - [tools/power/cpupower/man/cpupower-monitor.1 (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/tools/power/cpupower/man/cpupower-monitor.1?h=v6.12) · Linux kernel · cpupower monitor: per-core frequency and power-saving state statistics - [Processor state control for Amazon EC2 Linux instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/processor_state_control.html) · AWS · Only some instance types let the OS control C-states and P-states, which can be changed to reduce latency; the default settings give maximum performance and suit most workloads; Graviton runs at a fixed frequency, so the OS doesn’t control it #### so-oom · OOM killer · Out-of-memory killer When memory runs out, Linux picks the process using the most memory and kills it. Usually that’s the game server. - Why → Effect → On screen: Memory runs out from a leak or a surge in usage, or the container hits its memory limit → The kernel kills the game server process → Everyone on that server disconnects at once, and recent progress may be rolled back - Symptoms: Disconnect, Dropped action / rollback / Factors: Stall - Who: Whole server / When: The longer it runs, When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Fix leaks, set a memory usage ceiling with a procedure that saves and shuts down cleanly as usage approaches it. - Infra team action items: Set memory alerts, size the container memory limit to actual usage, adjust which process gets killed first (oom_score_adj). - Ballpark numbers: The kernel log (dmesg) records “Out of memory: Killed process”, and Kubernetes shows OOMKilled. Windows has no OOM killer; there the server usually dies with an error when a memory allocation fails. - On the graph: Mass disconnect (Connection count, memory usage) - Where to look: “Out of memory: Killed process” entries in dmesg, OOMKilled in the pod status on Kubernetes, or an increase in oom_kill in memory.events on cgroup v2, lined up with the time connections dropped - Confirmed if: At the moment connections dropped all at once, a record shows the game server process being killed, and memory usage had been climbing to the limit right before - Ruled out if: No OOM record but the process died: check the crash log and core dump (see “Server crash”) - Check with: Infra tools (no game code needed) - Sources: - [mm/oom_kill.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/mm/oom_kill.c?h=v6.12) · Linux kernel · The process using the most memory gets the highest score (oom_score_adj factored in); the kernel logs “Out of memory: Killed process …” when it kills - [Assign Memory Resources to Containers and Pods](https://kubernetes.io/docs/tasks/configure-pod-container/assign-memory-resource/) · Kubernetes · A container that keeps using memory beyond its limit is terminated, and its status shows OOMKilled - [Pushing the Limits of Windows: Virtual Memory](https://learn.microsoft.com/en-us/archive/blogs/markrussinovich/pushing-the-limits-of-windows-virtual-memory) · Microsoft · On Windows, once the commit limit is reached, allocations that commit memory fail, which can lead to app errors or system failures - [Control Group v2](https://docs.kernel.org/admin-guide/cgroup-v2.html) · Linux kernel · oom_kill in memory.events: number of processes in this cgroup killed by the OOM killer #### so-reclaim · Stalls from memory reclaim and compaction · Memory compaction / reclaim stalls (THP) The process stalls while the OS compacts memory to build huge pages or reclaims memory to free it up. - Why → Effect → On screen: Free memory runs low, or transparent huge pages (THP) trigger memory compaction → The thread that asked for memory waits until reclaim or compaction finishes → Irregular server stalls (a few ms to hundreds of ms) - Symptoms: Freeze, Stutter / Factors: Stall - Who: Whole server / When: The longer it runs, Randomly - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: Cut down large memory allocations at runtime (allocate up front at startup and reuse). - Infra team action items: Configure huge pages (THP) so only regions that ask for them use them (madvise), raise the free-memory threshold (vm.min_free_kbytes and so on). - On the graph: Random spikes (Server tick time, memory PSI) - Where to look: some and full in /proc/pressure/memory (share of time stalled waiting for memory) and the increase in compact_stall from /proc/vmstat, alongside server tick time; the /sys/kernel/mm/transparent_hugepage/defrag setting - Confirmed if: Memory PSI rises and compact_stall increases when ticks spike. defrag is set to always - Ruled out if: PSI and compact_stall unchanged: not this cause. Swap usage rising: “Swap” - Check with: Infra tools (no game code needed) - Sources: - [Transparent Hugepage Support](https://docs.kernel.org/admin-guide/mm/transhuge.html) · Linux kernel · With defrag=always, a failed THP allocation reclaims and compacts memory on the spot and stalls; with madvise, only regions that asked for it do - [Documentation for /proc/sys/vm/](https://docs.kernel.org/admin-guide/sysctl/vm.html) · Linux kernel · min_free_kbytes: the minimum free memory (watermark) the kernel keeps in reserve - [PSI - Pressure Stall Information](https://docs.kernel.org/accounting/psi.html) · Linux kernel · some (share of time some tasks stalled waiting for memory) and full (share of time all tasks stalled) in /proc/pressure/memory #### so-timejump · System clock jump (NTP step) · Wall-clock jump (NTP step) When the server clock is moved forward or back by several seconds in one step, timers that depend on the system clock fire all at once or stop. - Why → Effect → On screen: Time sync moves the clock by a large amount in one step → Timers fire in a batch or stop, and timeouts are misjudged → Buff and cooldown glitches, mass disconnects, fast-forward - Symptoms: Fast-forward, Disconnect, Dropped action / rollback / Factors: Stall - Who: Whole server / When: Randomly - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Compute elapsed time, timeouts, and cooldowns with a monotonic clock that never jumps or goes backward, use the wall clock only for display and logging. - Infra team action items: Adjust the clock gradually (chrony’s makestep only right after startup), monitor time sync status (clock offset). - Ballpark numbers: ntpd steps the clock in one go when the offset exceeds 0.128 s; below that, it slews gradually at a rate that takes a little over 30 minutes to remove a 1-second offset. With the recommended setting (makestep), chrony, now widely used, steps the clock only a few times right after startup, then slews gradually. The clock also jumps when a VM pauses briefly and resumes. - On the graph: Random spikes (Timer firings and disconnects, clock adjustment log) - Where to look: Entries in the time sync service’s log where the clock was stepped by a large amount, lined up with when the problems occurred. With chrony, any adjustment larger than logchange (default 1 s) is written to syslog - Confirmed if: Clock adjustments appear in the log at the times of buff and cooldown glitches, mass disconnects, and fast-forward, and the size of the adjustment matches the size of the glitch - Ruled out if: No clock adjustments logged: not this cause. On a VM, also check for a pause and resume (“Cloud host maintenance and live migration”) - Check with: Infra tools (no game code needed) - Sources: - [ntpd - Network Time Protocol (NTP) daemon](https://www.ntp.org/documentation/4.2.8-series/ntpd/) · Network Time Foundation · Offsets above the 128 ms step threshold are corrected in one step and smaller ones gradually, at 0.5 ms per second, so correcting 1 second takes 2,000 s (about 33 minutes) - [chrony – Frequently Asked Questions](https://chrony-project.org/faq.html) · chrony · Recommended: allow steps only a few times right after startup, as in makestep 1 3; a VM that was paused and resumed can wake up with the wrong time - [clock_gettime(2) — Linux manual page](https://man7.org/linux/man-pages/man2/clock_gettime.2.html) · Linux man-pages · CLOCK_MONOTONIC is not affected by discontinuous jumps in the system clock and never goes backward - [chrony.conf(5)](https://chrony-project.org/doc/4.6/chrony.conf.html) · chrony · logchange: clock adjustments larger than this value (default 1 s) are written to syslog #### so-cron · Scheduled jobs · Cron jobs (log rotation, backup, scans) Log compression, backups, and security scans that run at the same time every day take up CPU and disk. - Why → Effect → On screen: An OS job runs at a scheduled time → It shares CPU and disk with the game server → Stutter and slow motion at a fixed time, such as 4 a.m. every day - Symptoms: Stutter, Slow motion / Factors: Stall - Who: Whole server / When: At regular intervals - Primary owner: Infra team (Server infrastructure) - Infra team action items: Stagger job times, lower their priority (nice, ionice), move them off the game server (run them on a separate server). - On the graph: Periodic spikes (CPU utilization, disk queue, server tick time) - Where to look: Run times of scheduled jobs collected from crontab and systemctl list-timers, and which processes use CPU and disk when ticks spike, from pidstat -u -d - Confirmed if: Ticks spike at the same time every day (or every hour), and at that time a scheduled job process takes up CPU and disk - Ruled out if: Spikes don’t line up with a fixed time of day: not this cause. Spikes every few seconds or minutes: “Server GC stop-the-world pause,” “Timers firing all at once” - Check with: Infra tools (no game code needed) - Sources: - [ionice(1) — Linux manual page](https://man7.org/linux/man-pages/man1/ionice.1.html) · util-linux · Jobs in the idle class get disk I/O only when no other program is using the disk - [systemd.timer(5) — Linux manual page](https://man7.org/linux/man-pages/man5/systemd.timer.5.html) · systemd · RandomizedDelaySec delays scheduled jobs by a random amount to reduce load pileups - [systemctl(1) — Linux manual page](https://man7.org/linux/man-pages/man1/systemctl.1.html) · systemd · list-timers: shows timer units in order of next run time - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · -u shows CPU and -d shows disk I/O per process #### so-os-update · Performance changes after OS, kernel, driver, or firmware updates · Performance regression after OS / kernel / driver / firmware update The game code hasn’t changed, but the server has been slower since an OS, kernel, driver, or firmware update. Updates can change defaults, the scheduler, CPU vulnerability mitigations, and driver behavior. - Why → Effect → On screen: A routine security patch or a new server image changes the kernel, drivers, or firmware → Changed defaults or scheduler, or newly enabled vulnerability mitigations, make the same work take more CPU time and change the order in which threads get the CPU → A server that ran fine is a little slower all the time from the day of the update: input lag, plus stutter and slow motion when crowds gather - Symptoms: Input lag, Stutter, Slow motion / Factors: Latency, Stall, Jitter - Who: Whole server / When: Always, When crowds gather - Primary owner: Infra team (Server infrastructure) - Infra team action items: Apply updates to a few servers first and compare tick time, latency, and CPU utilization against the previous version before rolling out wider, deploy on a different day from game patches, record kernel, driver, and firmware versions and key sysctl values before and after the update, boot into the previous kernel to confirm when problems appear, weigh the security risk before turning mitigations off (mitigations=off). - Ballpark numbers: A new kernel version brings new default behavior. For example, Linux began moving its scheduler from CFS to EEVDF in 6.6, and the default connection queue cap (somaxconn) changed from 128 to 4,096 in 5.4. CPU vulnerability mitigations add work, such as flushing internal CPU buffers when returning from the kernel to a program (at the end of every system call) and on context switches and VM transitions, so network servers that make a system call per packet are hit harder. Fully blocking some vulnerabilities requires turning off SMT (the feature that runs one core as two threads), and turning off SMT can cut performance sharply depending on the workload. The kernel parameter mitigations=off turns all these mitigations off and recovers the performance, but leaves the system exposed to the vulnerabilities. - On the graph: Step change (Server tick time, CPU utilization, latency under the same load) - Where to look: Update history from the package manager and reboot times, the kernel version from uname -r, and NIC driver info from ethtool -i, lined up with when latency rose. Updated and non-updated servers compared under the same load with mpstat and pidstat, along with the mitigation status in /sys/devices/system/cpu/vulnerabilities/ - Confirmed if: Latency and CPU utilization step up from the reboot after the update and stay there, and under the same load only the updated servers run high. Booting into the previous kernel or driver brings them back - Ruled out if: Updated and non-updated servers are equally slow under the same load: not this cause. A game patch went out the same day and packets per player or packet size changed: “Patch changes the traffic pattern” - Check with: Infra tools (no game code needed) - Learn more: Check mitigation status in the files under /sys/devices/system/cpu/vulnerabilities/. The default (mitigations=auto) mitigates with SMT left on, but auto,nosmt turns SMT off on vulnerable CPUs, so the logical core count can drop by half after a kernel upgrade. Updating the OS on the same day as a game patch makes it hard to tell which one caused a problem, so deploy them separately. - Sources: - [The kernel’s command-line parameters](https://docs.kernel.org/admin-guide/kernel-parameters.html) · Linux kernel · mitigations=: off disables all CPU vulnerability mitigations for more performance but leaves the system exposed, the default auto mitigates with SMT on, auto,nosmt turns SMT off when needed - [MDS - Microarchitectural Data Sampling](https://docs.kernel.org/admin-guide/hw-vuln/mds.html) · Linux kernel · Mitigations flush CPU buffers when returning from the kernel to user space and when entering a VM; the files under /sys/devices/system/cpu/vulnerabilities/ show vulnerability and mitigation status; many CPUs need SMT off for full protection, and turning SMT off can have a large performance impact depending on the workload - [Spectre Side Channels](https://docs.kernel.org/admin-guide/hw-vuln/spectre.html) · Linux kernel · As a mitigation, branch prediction buffers are flushed on context switches and VM transitions, and stronger mitigations add overhead to every program - [EEVDF Scheduler](https://docs.kernel.org/scheduler/sched-eevdf.html) · Linux kernel · Linux began moving from CFS to the EEVDF scheduler in 6.6 - [listen(2) — Linux manual page](https://man7.org/linux/man-pages/man2/listen.2.html) · Linux man-pages · The somaxconn default changed from 128 to 4,096 in Linux 5.4 - [ethtool(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ethtool.8.html) · ethtool · ethtool -i shows driver information for a network device #### so-conntrack · Server conntrack table full · conntrack table full When the connection tracking (conntrack) table, where the Linux firewall records every connection, reaches its limit, new packets are dropped. - Why → Effect → On screen: Connection surges and repeated short-lived connections pile up connection entries → The table fills up, and new connections and some packets are dropped → Can’t connect, and teleporting from unexplained packet loss - Symptoms: Can’t connect / infinite loading, Teleporting / Factors: Packet loss - Who: Whole server / When: Right after login or maintenance, When crowds gather - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development), Game team (Client development) - Game team action items: Server: cut down short-lived connections (reuse connections for server-to-server calls). Client: when a connection fails or drops, retry with growing, randomized intervals. - Infra team action items: Raise the table size (nf_conntrack_max), exclude game ports from tracking (NOTRACK in the raw table), alert on usage. - Ballpark numbers: The default limit is about 60,000 to 260,000 entries depending on server memory. When it overflows, the kernel log shows “nf_conntrack: table full, dropping packet”. - On the graph: Hits a ceiling (conntrack entry count (nf_conntrack_count)) - Where to look: net.netfilter.nf_conntrack_count (current entries) from sysctl on the same graph as nf_conntrack_max, and “nf_conntrack: table full, dropping packet” in dmesg - Confirmed if: nf_conntrack_count flattens at max, and from that moment the kernel log shows table full - Ruled out if: Entry count well below max: not this cause. For the AWS instance’s own connection tracking limit, check conntrack_allowance_exceeded (see “Cloud PPS limit exceeded”) - Check with: Infra tools (no game code needed) - Sources: - [Netfilter Conntrack Sysfs variables](https://docs.kernel.org/networking/nf_conntrack-sysctl.html) · Linux kernel · nf_conntrack_max defaults to the number of hash buckets (nf_conntrack_buckets), which is set by memory size - [net/netfilter/nf_conntrack_core.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/netfilter/nf_conntrack_core.c?h=v6.12) · Linux kernel · Default size is 65,536 with more than 1 GB of memory and 262,144 with more than 4 GB (64-bit); when full, it logs “nf_conntrack: table full, dropping packet” and drops - [iptables-extensions(8) — Linux manual page](https://man7.org/linux/man-pages/man8/iptables-extensions.8.html) · netfilter · CT --notrack in the raw table excludes traffic from connection tracking #### so-ports · Ephemeral port exhaustion on server-to-server connections · Ephemeral port exhaustion (TIME_WAIT) When a game server opens and closes short connections to the DB or other servers very often, closed connections hold their ports for a while, and new connections can’t be opened. - Why → Effect → On screen: A new connection is opened and closed for every request → The side that closes first holds the port for about 60 seconds on Linux (TIME_WAIT), and the pool of usable ports runs dry → Internal requests fail: failed saves, broken features - Symptoms: Dropped action / rollback, Can’t connect / infinite loading / Factors: Packet loss - Who: Whole server, One feature only / When: When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Reuse connections (connection pool), stop opening and closing a new connection for every request. - Infra team action items: Widen the port range (ip_local_port_range), consider TIME_WAIT reuse for outgoing connections (tcp_tw_reuse on Linux), monitor the TIME_WAIT count. - Ballpark numbers: The default Linux port range (32768–60999) holds about 28,000 ports. More than 470 new connections per second to the same destination address use them up. Windows has about 16,000 ports by default (49152–65535) and a longer TIME_WAIT, so it runs out even faster. - On the graph: Hits a ceiling (TIME_WAIT sockets, internal connection failures) - Where to look: TIME_WAIT sockets counted per destination address with ss -tan state time-wait, and connect failures (EADDRNOTAVAIL) in the game server log - Confirmed if: TIME_WAIT sockets to the same destination (the DB, for example) flatten near the size of the ephemeral port range (about 28,000 by default), and connect fails with EADDRNOTAVAIL - Ruled out if: Few TIME_WAIT sockets, but only outbound connections to the outside fail: “Cloud NAT gateway connection and port limits” - Check with: Infra tools (no game code needed) - Learn more: On Linux, the 60-second TIME_WAIT is hard-coded in the kernel. Lowering tcp_fin_timeout, which has a similar name, doesn’t shorten TIME_WAIT. - Sources: - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · ip_local_port_range default 32768–60999; tcp_tw_reuse; tcp_fin_timeout is how long the FIN_WAIT_2 state is kept - [include/net/tcp.h (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/net/tcp.h?h=v6.12) · Linux kernel · TCP_TIMEWAIT_LEN (60*HZ): the roughly 60-second TIME_WAIT is a kernel constant - [TCP/IP port exhaustion troubleshooting](https://learn.microsoft.com/en-us/troubleshoot/windows-client/networking/tcp-ip-port-exhaustion-troubleshooting) · Microsoft · Windows dynamic ports default to 49152–65535, and a closed connection holds its port in TIME_WAIT for 4 minutes by default - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · The state filter state time-wait shows only TIME_WAIT sockets - [connect(2) — Linux manual page](https://man7.org/linux/man-pages/man2/connect.2.html) · Linux man-pages · EADDRNOTAVAIL: every port in the ephemeral port range is in use, so the connection can’t be opened ### L8 Sockets and protocols (causes: 14) #### sk-hol · TCP head-of-line blocking · Head-of-line blocking To keep data in order, TCP holds back every packet that arrived after a lost one until the lost packet is received again. - Why → Effect → On screen: One packet goes missing → The packets behind it have arrived but wait in the receive buffer → Everything stops, then releases all at once: fast-forward - Symptoms: Freeze, Fast-forward / Factors: Packet loss, Stall - Who: Just me / When: Randomly - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: send real-time positions over UDP, use reliable delivery only for what truly needs it, split traffic into multiple streams. Client: change network handling to match the server (UDP, separate channels). - Ballpark numbers: Losing one packet stalls things for at least one round-trip time plus a little more; if the retransmission is lost too, the stall lasts hundreds of ms to several seconds. - On the graph: Gap then burst (Bytes received per connection, retransmissions) - Where to look: Retransmitted packets and the gaps around them on that player’s connection in a server-side packet capture (tcpdump, Wireshark); for the whole server, the increase in TcpRetransSegs from nstat -az - Confirmed if: Each stall starts with the retransmission of a single packet, and right after the retransmitted packet arrives, the backlog is processed all at once (bytes received sit at 0, then surge) - Ruled out if: Games that talk over UDP: not applicable. Stalls with no retransmissions: look at the server tick (“Tick overrun”) - Check with: Infra tools (no game code needed) - Sources: - [RFC 9293: Transmission Control Protocol (TCP)](https://www.rfc-editor.org/rfc/rfc9293) · IETF · TCP is a reliable, in-order byte stream service - [RFC 5681: TCP Congestion Control](https://www.rfc-editor.org/rfc/rfc5681) · IETF · Loss detected by 3 duplicate ACKs triggers a fast retransmit; otherwise TCP waits for the retransmission timer - [RFC 6298: Computing TCP's Retransmission Timer](https://www.rfc-editor.org/rfc/rfc6298) · IETF · Backoff that doubles the timeout each time the retransmission timer expires - [net/ipv4/proc.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/proc.c?h=v6.12) · Linux kernel · Counter name shown by nstat: RetransSegs (segments retransmitted) in the Tcp group #### sk-rto · TCP RTO and exponential backoff · RTO and exponential backoff Each time a retransmission fails again, the wait doubles, so a brief connection drop turns into a long stall. - Why → Effect → On screen: The connection drops briefly, and retransmissions fail one after another → The wait before the next attempt doubles each time: 0.3 → 0.6 → 1.2 → 2.4 s (at 100 ms ping) → The connection was down for 1 second, but the game stalls for over 2 seconds. A longer drop eventually ends in a disconnect - Symptoms: Freeze, Disconnect / Factors: Packet loss, Stall - Who: Just me / When: Randomly, While moving or changing zones - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: answer heartbeats and close the connection yourself if nothing arrives for a set time (bring the give-up point forward with TCP_USER_TIMEOUT), resume sessions with a session token, use reliable UDP. Client: send heartbeats at short intervals and reconnect quickly when responses stop, without waiting on TCP retransmission. - Ballpark numbers: The Linux RTO (retransmission timeout) has a floor of “ping + 200 ms”, and starts at 1 second while a connection is being set up. With the default setting (tcp_retries2=15), the connection is abandoned only after about 15 minutes of failed retransmissions. - On the graph: Gap then burst (Per-connection RTO and backoff, RTO expirations) - Where to look: rto (retransmission wait in ms) and backoff (consecutive expirations) of the stalled connection from ss -ti; for the whole server, the increase in TcpExtTCPTimeouts (retransmission timer expirations) from nstat -az - Confirmed if: The stalled connection shows backoff of 1 or more with rto grown to seconds, and TCPTimeouts rises at that time - Ruled out if: If retransmissions finish as fast retransmits with no RTO expiration, stalls are short. That points to “TCP head-of-line blocking” - Check with: Infra tools (no game code needed) - Sources: - [RFC 6298: Computing TCP's Retransmission Timer](https://www.rfc-editor.org/rfc/rfc6298) · IETF · Initial RTO 1 s, doubled on every timer expiration (exponential backoff) - [net/ipv4/tcp_input.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_input.c?h=v6.12) · Linux kernel · Linux RTO = smoothed RTT + RTT variation, and the variation term has a floor of tcp_rto_min (200 ms), so RTO is at least RTT + 200 ms - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_rto_min_us defaults to 200 ms, initial RTO for connection requests is 1 s, with tcp_retries2=15 it takes at least 924.6 s (about 15 minutes) to give up - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · rto (retransmission timer, ms) and backoff (exponential backoff count) in -i - [net/ipv4/proc.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/proc.c?h=v6.12) · Linux kernel · Counter name shown by nstat: TCPTimeouts in the TcpExt group - [net/ipv4/tcp_timer.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_timer.c?h=v6.12) · Linux kernel · Each time the retransmission timer expires, TCPTimeouts goes up, backoff increases by one, and RTO doubles (up to the maximum) #### sk-nagle · Nagle’s algorithm + delayed ACK · Nagle + delayed ACK (TCP_NODELAY off) Nagle’s algorithm, which batches small packets, and delayed ACK, which sends ACKs late, interact so that each message written in pieces is delayed by 40–200 ms. - Why → Effect → On screen: Small messages are written in pieces without TCP_NODELAY turned on → The sender waits for an ACK, and the receiver sends its ACK late → Ping is low, yet every action is consistently sluggish: input lag - Symptoms: Input lag / Factors: Latency - Who: Whole server, Just me / When: Always, During specific actions - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: turn on TCP_NODELAY, collect one tick’s worth of messages and write them at once, don’t rely on turning off delayed ACK on the receiving side (Linux TCP_QUICKACK only lasts briefly, and Windows needs a registry change on every PC) because the game can’t reliably control it. Client: turn on TCP_NODELAY, collect one frame’s worth of messages and write them at once. - Ballpark numbers: Linux usually delays ACKs by 40 ms (up to 200 ms depending on the situation). Older Windows versions used 200 ms, and current versions use 40 ms (40 ms in the Windows Server 2019 default template). The receiving OS decides the delayed ACK, so if the server sends messages in pieces with Nagle on, each can be delayed by 40–200 ms depending on the receiving PC. - On the graph: Always high (Action response time (in-game RTT)) - Where to look: Gaps between requests and responses in a server-side packet capture (tcpdump, Wireshark), and whether the server and client code turn on TCP_NODELAY - Confirmed if: Ping is low, but gaps of around 40 ms (200 ms on older Windows) keep appearing between small packets, and each gap ends right after the other side’s ACK arrives. Turning on TCP_NODELAY makes them disappear - Ruled out if: Response gaps close to ping: not this cause. Game server slow to produce responses: server processing (“Message queue backlog”) - Check with: Infra tools (no game code needed) - Sources: - [RFC 9293: Transmission Control Protocol (TCP)](https://www.rfc-editor.org/rfc/rfc9293) · IETF · Nagle holds back small data while unacknowledged data is outstanding; it must be possible to turn off per connection; delayed ACK under 0.5 s; the problem of the two interacting - [include/net/tcp.h (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/net/tcp.h?h=v6.12) · Linux kernel · Linux delayed ACK minimum TCP_DELACK_MIN (HZ/25 = 40 ms), maximum TCP_DELACK_MAX (HZ/5 = 200 ms) - [TCP improvements in the Windows network stack (IETF 98 TCPM)](https://datatracker.ietf.org/meeting/98/materials/slides-98-tcpm-tcp-improvements-in-windows-01) · Microsoft · Default delayed ACK timeout on Windows changed to 40 ms (announced in 2017) - [TCP Templates for Windows Server 2019 – How to tune your Windows Server Transports (Advanced users only 😉)](https://techcommunity.microsoft.com/blog/networkingblog/tcp-templates-for-windows-server-2019-8211-how-to-tune-your-windows-server-trans/339795) · Microsoft · Server 2019 template: DelayedAckTimeout 40 ms, MaxSynRetransmissions 2, InitialRto 3000 ms - [Design issues - Sending small data segments over TCP with Winsock](https://learn.microsoft.com/en-us/previous-versions/troubleshoot/windows/win32/data-segment-tcp-winsock) · Microsoft · Older Windows TCP starts a 200 ms delayed ACK timer when data arrives, and Nagle is on by default, so small packets wait for the ACK; fixed with TCP_NODELAY #### sk-block-send · Blocking sends caused by slow clients · Blocking send on a full socket When one player on a slow connection has a full send buffer and the server sends in blocking mode (a send call that doesn’t return until the buffer has room), the server thread waits on that one player. - Why → Effect → On screen: A slow client’s send buffer is full → With blocking sends, the server thread waits until the buffer has room → Everyone that thread handles gets a freeze or slow motion - Symptoms: Freeze, Slow motion / Factors: Stall - Who: Specific zone/channel, Whole server / When: Randomly, When crowds gather - Primary owner: Game team (Server development) - Game team action items: Use non-blocking sends, cap the send queue per client, drop stale updates. - On the graph: Random spikes (Server tick time, Send-Q per connection) - Where to look: Connections whose Send-Q (bytes not yet ACKed or not yet sent) has filled up to the send buffer size, found with ss -tn; in a thread dump (stacks) of the game server taken when ticks spike, any thread stuck in a send call - Confirmed if: While a slow connection has a full Send-Q, the thread handling it is stuck in send, and only the players on that same thread stall with it - Ruled out if: Stalled threads waiting outside send (locks, DB calls): “Lock contention,” “Blocking calls on the game thread” - Check with: Infra tools (no game code needed) - Sources: - [send(2) — Linux manual page](https://man7.org/linux/man-pages/man2/send.2.html) · Linux man-pages · send() blocks when the send buffer has no room, and returns immediately with EAGAIN in non-blocking mode - [send function (winsock2.h)](https://learn.microsoft.com/en-us/windows/win32/api/winsock2/nf-winsock2-send) · Microsoft · Winsock send also blocks when buffer space runs out, unless the socket is in non-blocking mode - [net/ipv4/tcp_diag.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_diag.c?h=v6.12) · Linux kernel · Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed #### sk-slow-client · Slow client (slow consumer) handling policy · Slow-consumer policy When a client’s outgoing data keeps piling up, the server drops stale updates or disconnects it. - Why → Effect → On screen: The client’s connection can’t keep up with what the server sends → The server drops stale updates, or disconnects the client once a limit is exceeded → Just that player sees teleporting or gets a disconnect - Symptoms: Teleporting, Disconnect / Factors: Packet loss - Who: Just me / When: When crowds gather - Primary owner: Game team (Server development) - Game team action items: Send less (update rate by distance), keep sending at lower quality, keep less data queued in the kernel (TCP_NOTSENT_LOWAT on Linux). - Ballpark numbers: With a 256 KB send buffer, a 30 KB/s connection builds up more than 8 seconds of backlogged data. Linux may also grow this buffer automatically to several MB. - On the graph: Outliers only (Send-Q per connection, dropped updates per client) - Where to look: Per-client send queue length, dropped updates, and disconnect reasons logged by the game server; on the server, that connection’s Send-Q and cwnd together from ss -tni - Confirmed if: Only the connections of players who teleport or disconnect keep a full Send-Q, and the game log shows dropped updates or a disconnect for exceeding the send queue for those players - Ruled out if: Send-Q empty while the player still teleports: not a server send-side problem. Look at that player’s connection loss (“Wireless link loss”) or on-screen interpolation - Check with: Game server or client logs and metrics - Sources: - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_wmem: the maximum auto-tuned send buffer defaults to 64 KB–4 MB (depending on memory); tcp_notsent_lowat and TCP_NOTSENT_LOWAT limit the amount of unsent data - [net/ipv4/tcp_diag.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_diag.c?h=v6.12) · Linux kernel · Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed #### sk-keepalive · Keepalive default of 2 hours · TCP keepalive defaults When the other side vanishes without a close signal, TCP notices only much later. Keepalive (a TCP feature that checks whether an idle connection is still alive) is off by default, and even when it’s on, checks start only after 2 hours of idle time. - Why → Effect → On screen: The client vanishes without a close signal because its power went off or its connection dropped → The server assumes the connection is still alive (keepalive default 7,200 s; if data was being sent, about 15 minutes until retransmission gives up) → A ghost character stays behind, and reconnecting fails with an “Already logged in” error - Symptoms: Can’t connect / infinite loading, Invisible / ghost entities / Factors: Packet loss - Who: Just me / When: After sitting idle, Right after login or maintenance - Primary owner: Game team (Server development) / Also: Game team (Client development), Infra team (Server infrastructure) - Game team action items: Server: answer heartbeats and close the connection yourself if nothing arrives for a set time (tune TCP_KEEPIDLE and TCP_USER_TIMEOUT), and on reconnect use the session token to replace the old session and resume it. Client: send game-level heartbeats every few seconds to tens of seconds (half the shortest idle timeout or less), reconnect automatically when disconnected. - Infra team action items: Lower the kernel defaults (tcp_keepalive_time and so on) for sockets whose code doesn’t set its own values (applies only to sockets with SO_KEEPALIVE on). - Ballpark numbers: By default, Linux starts checking after 7,200 seconds of idle time, sends 9 probes 75 seconds apart, and drops the connection if none are answered. That adds up to about 2 hours 11 minutes. Windows also waits for 2 hours of idle time by default before it starts checking. - On the graph: Outliers only (Time since last receive, per connection) - Where to look: lastrcv (ms since the last receive) and the keepalive timer (timer:(keepalive,…)) per connection from ss -tnoi, lined up with the game server’s “Already logged in” rejections - Confirmed if: ESTABLISHED connections with lastrcv of minutes to hours are still there, and reconnects for those accounts are rejected with “Already logged in” - Ruled out if: No long-silent connections, yet “Already logged in” still appears: the game server’s session cleanup code - Check with: Infra tools (no game code needed) - Sources: - [tcp(7) — Linux manual page](https://man7.org/linux/man-pages/man7/tcp.7.html) · Linux man-pages · 9 probes 75 s apart after 7,200 s of idle time (about 11 more minutes), applies only to sockets with SO_KEEPALIVE on, TCP_KEEPIDLE and TCP_USER_TIMEOUT - [RFC 9293: Transmission Control Protocol (TCP)](https://www.rfc-editor.org/rfc/rfc9293) · IETF · Keepalive must be off by default, and the default idle interval must be 2 hours or more - [SO_KEEPALIVE socket option](https://learn.microsoft.com/en-us/windows/win32/winsock/so-keepalive) · Microsoft · Default Windows TCP keepalive timeout: 2 hours - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · lastrcv (ms since last receive) in -i, timer:(keepalive,…) in -o #### sk-fragment · IP fragmentation of UDP packets · IP fragmentation of large UDP A UDP packet larger than the MTU (the largest size that can be sent in one piece) is fragmented at the IP layer, and losing just one fragment throws away the whole packet. - Why → Effect → On screen: Snapshots in crowded areas exceed 1,500 bytes → They go out split into several fragments, and losing any one of them discards the whole packet → Large packets are lost several times as often. Teleporting only in crowded areas - Symptoms: Teleporting / Factors: Packet loss - Who: Specific zone/channel, Specific region/ISP / When: When crowds gather - Primary owner: Game team (Server development) - Game team action items: Split packets yourself to 1,200 bytes or less, send only what changed. - Ballpark numbers: On a connection with 2% loss, about 8% of packets split into 4 fragments are lost. Some firewalls and ISPs drop fragmented packets outright, so those players never receive a single large packet. - On the graph: Rises with load (IP fragments created (IpFragCreates), snapshot size) - Where to look: On the server, the increase in IpFragCreates (fragments created while sending) from nstat -az; on the receiving side, IpReasmFails (reassembly failures). UDP packet size distribution from game server logs or a packet capture - Confirmed if: IpFragCreates rises where crowds gather, UDP packets larger than 1,500 bytes show up, and teleporting reports go up at the same time - Ruled out if: IpFragCreates not rising: no fragmentation on the server’s sending side - Check with: Infra tools (no game code needed) - Sources: - [RFC 8085: UDP Usage Guidelines](https://www.rfc-editor.org/rfc/rfc8085) · IETF · Losing one fragment means the packet can’t be reassembled and is lost entirely; UDP apps should avoid IP fragmentation - [RFC 8900: IP Fragmentation Considered Fragile](https://www.rfc-editor.org/rfc/rfc8900) · IETF · Cases of firewalls and some networks dropping IP fragments - [net/ipv4/proc.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/proc.c?h=v6.12) · Linux kernel · Counter names shown by nstat: FragCreates (fragments created) and ReasmFails (reassembly failures) in the Ip group #### sk-reliable-udp · Reliable UDP retransmission settings · Reliable-UDP tuning (KCP, ENet…) When the retransmission rules you built on top of UDP are too conservative, recovery is slow; when they’re too aggressive, they clog the connection even more. - Why → Effect → On screen: Retransmission interval, retry count, and window size don’t suit the connection → Slow recovery, or duplicate sends that make congestion worse → Skills not going off, fast-forward, worse lag during congestion - Symptoms: Dropped action / rollback, Fast-forward / Factors: Packet loss, Latency - Who: Just me / When: Randomly - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: base retransmission on measured round-trip time, separate channels by importance. Client: apply the same retransmission and channel settings as the server. - On the graph: Random spikes (Reliable UDP retransmission rate, in-game RTT) - Where to look: Per-connection statistics from the library in use (retransmissions, estimated round-trip time, retransmission timeout), logged on server and client and compared with the same player’s actual connection loss rate (measured with mtr) - Confirmed if: Retransmission rate several times the actual loss rate: settings too aggressive. Retransmission timeout several times the measured round-trip time: settings too conservative - Ruled out if: Retransmission rate close to the loss rate and timeout in line with round-trip time: not a settings problem. Look at the connection loss itself - Check with: Game server or client logs and metrics - Sources: - [RFC 8085: UDP Usage Guidelines](https://www.rfc-editor.org/rfc/rfc8085) · IETF · Retransmissions can add to congestion, so they fall under congestion control; round-trip time is estimated as an average of several measurements (EWMA), initial value 1 s, lower the sending rate when the timer expires #### sk-slowstart · Slow start after idle · Slow start after idle When a connection has been idle for a while, TCP shrinks the congestion window (how much it can send at once) again, so a sudden large send goes out in several rounds. - Why → Effect → On screen: A large burst of data (on entering a town, for example) goes out over a connection that was idle → The congestion window has shrunk, so the data is spread over several round trips → Right after entering, nearby characters and NPCs appear a few round trips late (more noticeable on distant servers) - Symptoms: Input lag, Invisible / ghost entities / Factors: Latency - Who: Just me / When: While moving or changing zones, After sitting idle - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: Shrink the entry data (send what’s essential first). - Infra team action items: Turn off tcp_slow_start_after_idle (Linux, server-wide setting). - Ballpark numbers: After an idle period longer than the RTO, the congestion window starts shrinking, and after a long idle period it drops to about 14 KB (10 packets). Then 100 KB can’t go out at once and takes 3 round trips. - On the graph: Outliers only (Transfer time right after entering (players with long RTT)) - Where to look: The value of sysctl net.ipv4.tcp_slow_start_after_idle, and whether cwnd (congestion window) in ss -ti for that connection has shrunk at the moment a player enters an area after being idle - Confirmed if: The setting is 1 (the default), and on entering after idle, cwnd drops to around 10 and the transfer is split over several round trips. The longer a player’s RTT, the later things appear, and setting it to 0 makes the problem go away - Ruled out if: cwnd stays large, yet things still appear late: server-side entry handling (“Spawn burst when entering a crowded area”) - Check with: Infra tools (no game code needed) - Sources: - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_slow_start_after_idle on by default, the congestion window shrinks after an RTO of idle time (RFC 2861 method) - [RFC 5681: TCP Congestion Control](https://www.rfc-editor.org/rfc/rfc5681) · IETF · If no data has been sent for longer than the RTO, the congestion window is cut to at most the restart window min(IW, cwnd) and slow start begins again - [RFC 6928: Increasing TCP's Initial Window](https://www.rfc-editor.org/rfc/rfc6928) · IETF · Initial window 10 segments, up to 14,600 bytes - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · cwnd (congestion window) and ssthresh (slow start threshold) in -i #### sk-congestion · Sending rate plunges under congestion control · Congestion control backoff TCP treats loss as a sign of congestion and cuts its sending rate by 30–50%. It reacts the same way to Wi-Fi loss. - Why → Effect → On screen: A little loss on Wi-Fi or the connection while there’s a lot to send → TCP cuts its sending rate sharply and recovers slowly (CUBIC, the Linux and Windows default, cuts by 30%) → Updates fall behind in crowded areas: fast-forward, input lag - Symptoms: Fast-forward, Input lag / Factors: Latency, Stall - Who: Just me / When: When crowds gather - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: Send less (area of interest, changes only), spread sends out so they don’t go out in one burst. - Infra team action items: Switch to a congestion control algorithm such as BBR (tcp_congestion_control). - On the graph: Sawtooth (Per-connection congestion window (cwnd) and sending rate) - Where to look: Several ss -ti snapshots of a lagging player’s connection for changes in cwnd and ssthresh and the congestion control name (cubic, bbr), plus whether Send-Q builds up - Confirmed if: After each loss, cwnd drops sharply and climbs back slowly, and Send-Q builds up while it’s low, lining up with the times of fast-forward and input lag reports - Ruled out if: cwnd is ample, yet updates still fall behind: the receiver’s window (“Zero window (a stall that looks like retransmission)”) or the server’s sending side - Check with: Infra tools (no game code needed) - Sources: - [RFC 9438: CUBIC for Fast and Long-Distance Networks](https://www.rfc-editor.org/rfc/rfc9438) · IETF · On loss, CUBIC scales the window by 0.7 (a 30% cut) and Reno by 0.5; CUBIC is the default on Linux, Windows, and Apple platforms - [TCP BBR congestion control comes to GCP – your Internet just got faster](https://cloud.google.com/blog/products/networking/tcp-bbr-congestion-control-comes-to-gcp-your-internet-just-got-faster) · Google Cloud · Loss-based congestion control cuts the sending rate sharply even for loss that congestion didn’t cause; BBR decides based on delivery rate and RTT - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_congestion_control selects the congestion control algorithm for new connections - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · cwnd, ssthresh, and the congestion control algorithm name in -i - [net/ipv4/tcp_diag.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_diag.c?h=v6.12) · Linux kernel · Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed #### sk-linger · Last data lost to an abortive close (RST) · SO_LINGER, abrupt RST When the server cuts a connection abruptly, the final notice or save-complete signal it sent is lost. - Why → Effect → On screen: The server closes the connection with an abortive close (RST). This happens when SO_LINGER is set to 0 seconds, or when the socket is closed before all received data has been read → The kick reason and final data still in transit are thrown away → An unexplained “Connection closed due to an unknown error” - Symptoms: Disconnect / Factors: Packet loss - Who: Just me / When: Randomly - Primary owner: Game team (Server development) - Game team action items: Send the reason, then close only the sending direction (shutdown), read incoming data until the other side closes and only then close the socket, avoid SO_LINGER of 0 seconds. - On the graph: Random spikes (Connections ended by RST) - Where to look: Increase in TcpExtTCPAbortOnData (closed with RST while data was left to send, SO_LINGER 0 s) and TcpExtTCPAbortOnClose (closed with unread data left) from nstat -az, and whether a server-side packet capture at the moment of disconnect shows an RST going out where a FIN should be - Confirmed if: At the times of “Connection closed due to an unknown error” reports, the server sends RST, and AbortOnData and AbortOnClose rise - Ruled out if: Server closed normally with FIN, yet the reason still doesn’t show: the client’s close handling - Check with: Infra tools (no game code needed) - Sources: - [closesocket function (winsock.h)](https://learn.microsoft.com/en-us/windows/win32/api/winsock/nf-winsock-closesocket) · Microsoft · Turning SO_LINGER on with a time of 0 makes close an abortive close that resets the connection immediately, and unsent data is lost - [RFC 9293: Transmission Control Protocol (TCP)](https://www.rfc-editor.org/rfc/rfc9293) · IETF · Closing while received data is still unread sends an RST to signal data loss - [Graceful Shutdown, Linger Options, and Socket Closure](https://learn.microsoft.com/en-us/windows/win32/winsock/graceful-shutdown-linger-options-and-socket-closure-2) · Microsoft · The sequence: shut down only the sending side first with shutdown, then close the socket after receiving the other side’s close notification - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · TcpExtTCPAbortOnData: closed with RST while data was left to send (SO_LINGER 0 s, for example); TcpExtTCPAbortOnClose: closed with unread data left, sending RST #### sk-blocking-io · Blocking I/O design · Blocking I/O model In a design where a thread can’t do anything else while it waits on one socket, everything slows down as the player count grows. - Why → Effect → On screen: Each connection waits on its own reads and writes → A delay on one connection spreads to the other connections on the same thread → As concurrent users grow, everyone gets slow motion and input lag - Symptoms: Slow motion, Input lag / Factors: Stall - Who: Whole server / When: When crowds gather, Evening peak hours - Primary owner: Game team (Server development) - Game team action items: Move to asynchronous I/O based on epoll, IOCP, or io_uring. - On the graph: Rises with load (Response time, thread count) - Where to look: Game server thread count and per-thread voluntary context switches (cswch/s, times a thread stopped to wait for a resource) from pidstat -w -t, compared with response time as concurrent users grow - Confirmed if: Response time climbs steeply as concurrent users grow, and most of the threads (one per connection) show many voluntary switches while barely using CPU (waiting on sockets) - Ruled out if: Threads use CPU continuously without waiting: compute overload (“Tick overrun”) - Check with: Infra tools (no game code needed) - Sources: - [epoll(7) — Linux manual page](https://man7.org/linux/man-pages/man7/epoll.7.html) · Linux man-pages · I/O event notification that scales to watching many fds at once - [I/O Completion Ports](https://learn.microsoft.com/en-us/windows/win32/fileio/i-o-completion-ports) · Microsoft · The Windows approach of handling many asynchronous I/Os with a pre-created thread pool - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · cswch/s in -w: voluntary context switches where the task stopped on its own to wait for a resource; -t shows them per thread #### sk-reuseport · Uneven SO_REUSEPORT distribution · SO_REUSEPORT imbalance, stuck worker When several processes share one port, the kernel assigns each connection to a process by address hash and never reassigns it. If one of those processes stalls, only the players assigned to it wait. - Why → Effect → On screen: A gateway or login server runs several processes on one port with SO_REUSEPORT → Even when one process stalls from GC or overload, the new connections and UDP packets assigned to it don’t move to another process → Only some players can’t connect or freeze. During a restart that changes the process count, some UDP sessions drop - Symptoms: Can’t connect / infinite loading, Freeze, Disconnect / Factors: Stall, Packet loss - Who: Just me, Whole server / When: Right after login or maintenance, Randomly - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Never let the receiving thread stall, implement a procedure for handing over sessions on restart. - Infra team action items: Monitor the connection queue per process (Recv-Q in ss), make sure deploys that change the process count follow the session handover procedure. - On the graph: Outliers only (Connection queue (Recv-Q) per listening socket) - Where to look: Recv-Q (connections waiting for accept) and the owning process of each listening socket on the same port from ss -ltnp, plus throughput compared per process - Confirmed if: Of the sockets on the same port, only one keeps building up Recv-Q, and its process is stalled or its throughput is near 0 - Ruled out if: Recv-Q builds up evenly on all sockets: overall overload (“Connection queue (listen backlog) overflow”) - Check with: Infra tools (no game code needed) - Sources: - [socket(7) — Linux manual page](https://man7.org/linux/man-pages/man7/socket.7.html) · Linux man-pages · SO_REUSEPORT lets several sockets bind to the same address and share incoming TCP connections and UDP packets - [net/core/sock_reuseport.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/core/sock_reuseport.c?h=v6.12) · Linux kernel · Without a BPF program, the packet hash is mapped onto the number of sockets in the group to pick the socket - [Why does one NGINX worker take all the load?](https://blog.cloudflare.com/the-sad-state-of-linux-socket-balancing/) · Cloudflare · SO_REUSEPORT splits queues per worker with a simple hash, so if one worker is blocked, every connection queued to it stalls - [net/ipv4/tcp_diag.c (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_diag.c?h=v6.12) · Linux kernel · Recv-Q and Send-Q in ss: for a listening socket, connections waiting for accept and the backlog limit; for a connected socket, bytes the app hasn’t read yet and sent bytes not yet ACKed #### sk-udp-connreset · WSAECONNRESET errors on Windows UDP sockets · WSAECONNRESET on a Windows UDP socket When a Windows server sends UDP to a client that has already left, a “port unreachable” (ICMP) message comes back. That message makes the next receive call fail with an error, and if the server code treats the error as a failure of the socket itself, everyone using that socket is affected. - Why → Effect → On screen: UDP keeps going to the address of a client that just left, and a “port unreachable” (ICMP) message comes back → Windows fails the next receive call with WSAECONNRESET (10054), and the server code stops receiving or closes the socket → Everyone who was using that socket freezes or disconnects at once - Symptoms: Disconnect, Freeze / Factors: Packet loss, Stall - Who: Whole server / When: Randomly, Right after login or maintenance - Primary owner: Game team (Server development) - Game team action items: Turn off SIO_UDP_CONNRESET with WSAIoctl so these messages aren’t reported, just log receive errors and keep receiving. - On the graph: Mass disconnect (Connection count, receive error log) - Where to look: UDP receive error codes (WSAECONNRESET, 10054) in the game server log and records of the receive loop stopping or the socket being closed; in a server-side packet capture, whether an ICMP Port Unreachable arrived just before - Confirmed if: A WSAECONNRESET receive error is logged right before connections drop all at once, and before that an ICMP Port Unreachable arrives from the address of a client that just left - Ruled out if: Linux server, or code that turns off SIO_UDP_CONNRESET: not applicable - Check with: Game server or client logs and metrics - Sources: - [Winsock IOCTLs](https://learn.microsoft.com/en-us/windows/win32/winsock/winsock-ioctls) · Microsoft · SIO_UDP_CONNRESET turns reporting of UDP “port unreachable” (PORT_UNREACHABLE) messages on and off - [recvfrom function (winsock.h)](https://learn.microsoft.com/en-us/windows/win32/api/winsock/nf-winsock-recvfrom) · Microsoft · WSAECONNRESET on a UDP socket means a previous send got an ICMP Port Unreachable back - [Windows Sockets Error Codes](https://learn.microsoft.com/en-us/windows/win32/winsock/windows-sockets-error-codes-2) · Microsoft · WSAECONNRESET error number 10054 ### L9 Server game process (causes: 18) #### sp-tick-overrun · Tick overrun · Tick overrun When one tick has more work than its budget, the server’s tick interval stretches, and the whole area slows down or stutters. - Why → Effect → On screen: One tick (e.g., 50 ms) has more work than its budget → Game state meant to update 20 times a second updates only 8 times → Slow motion across the zone (stutter on some server designs), sluggish skill response - Symptoms: Slow motion, Input lag, Stutter / Factors: Stall - Who: Specific zone/channel, Whole server / When: When crowds gather, Evening peak hours - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Cut expensive computation, split the tick across threads, spread players out (channels), record tick processing time as a metric. - Infra team action items: Add tick time and per-core CPU utilization to monitoring and alerts, consider CPUs and instances with high single-core performance (clock speed). - Ballpark numbers: The budget is 50 ms on a 20-tick server, 33 ms at 30 ticks, and 16.7 ms at 60 ticks. To be ready for sudden crowds, it’s safer to leave headroom and normally use only about half the budget. - On the graph: Rises with load (Server tick time, player count per zone/channel, game thread CPU) - Where to look: Tick time (p99) and tick-overrun count logged by the server, on one graph with player count per zone/channel. Without tick metrics, CPU utilization of the game thread from pidstat -t 1 - Confirmed if: Tick time goes over budget (50 ms at 20 ticks) when players crowd in, while the game thread sits near 100% CPU - Ruled out if: Ticks overrun while game-thread CPU is low: points to waiting (GC pause, locks, blocking calls). Long run queue latency in bcc runqlat means threads aren’t getting CPU time: points to a CPU shortage or too many threads - Check with: Game server or client logs and metrics - Learn more: What a late tick looks like depends on the server design. A server that advances game state by a fixed amount of time each tick (e.g., 50 ms) slows down game time itself, so players see slow motion. A server that moves everything by the actual elapsed time in one step keeps the game running at normal speed, but packets become sparse and each move is large, so it shows up as stutter or teleporting. Either way, input response gets slower. If one game thread runs the whole server, the whole server slows down; if each area has its own thread, only that area does. Some games, such as EVE Online, deliberately slow game time by up to 10 times during large battles (Time Dilation) so the computation can keep up. - Real incidents: eve-hedgp-2014 - Sources: - [VALORANT's 128-Tick Servers](https://www.riotgames.com/en/news/valorants-128-tick-servers) · Riot Games · A 128-tick server must finish each frame within 7.8125 ms; server frame time is measured per subsystem and the budget is split among them - [HED-GP Technical Retrospective: What a HED-ache](https://www.eveonline.com/news/view/what-a-hed-ache) · CCP Games · Under overload, EVE Online slows game time with Time Dilation down to a floor of 10% (10 times slower); node CPU normally stays below 80% - [Handling variation in time](https://docs.unity3d.com/Manual/time-handling-variations.html) · Unity · When a fixed-step simulation falls behind, it runs the catch-up steps in a batch and discards time beyond the limit, so game time runs slower than real time - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · -t shows per-thread statistics (CPU utilization and so on) for the threads in a process - [Demonstrations of runqlat, the Linux eBPF/bcc version](https://raw.githubusercontent.com/iovisor/bcc/master/tools/runqlat_example.txt) · IO Visor · Shows scheduler run queue latency (how long a task waited to get a CPU) as a histogram #### sp-aoi · AOI calculation blowup (N²) · Area-of-interest explosion If you compare everyone against everyone to work out who can see whom, 10 times the players means 100 times the computation. - Why → Effect → On screen: Every character’s distance is checked against every other character, or, even with a grid, hundreds of players crowd around one cell → About 10,000 comparisons for 100 players, about 1 million for 1,000 → Ticks spike where crowds gather, such as world bosses and sieges: slow motion, stutter - Symptoms: Slow motion, Stutter / Factors: Stall - Who: Specific zone/channel, Whole server / When: When crowds gather - Primary owner: Game team (Server development) - Game team action items: Split the world into a grid or regions and compare only nearby entities, update distant entities less often, cap how many players one player can see. - Ballpark numbers: If a distance check plus the update to the visible and not-visible lists takes 0.1 µs (one ten-millionth of a second) per pair of players, 1,000 players (about 1 million pairs) cost 100 ms per tick. That’s twice the 20-tick budget (50 ms). - On the graph: Rises with load (Server tick time, players gathered in one place) - Where to look: Player count and tick time per zone/channel on the same graph, plus the time spent on AOI calculation within the tick, measured separately. Without a separate measurement, per-function CPU share of the game process from perf top -p - Confirmed if: When the number of players in one place doubles, tick time grows nearly 4 times, and view range and distance functions take most of the CPU time - Ruled out if: Tick time grows in proportion to player count, or send and serialization functions take a large share: points to “Broadcast fan-out overload” or serialization and compression cost - Check with: Game server or client logs and metrics - Sources: - [Comparing Interest Management Algorithms for Massively Multiplayer Games](https://www.sable.mcgill.ca/~clump/papers/boulanger-06-comparing.pdf) · ACM · NetGames 2006 paper (authors’ copy). Measuring the distance of every pair can’t keep up as player count grows; with a square grid, only the 9 surrounding cells are checked - [Replication Graph in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/replication-graph-in-unreal-engine) · Epic Games · The default approach of checking every connection for each actor becomes a server CPU bottleneck with many players and actors; MMORPGs and similar games split the world into a grid and reuse per-cell lists - [perf-top(1) — Linux manual page](https://man7.org/linux/man-pages/man1/perf-top.1.html) · perf · Shows live CPU share per function (symbol) for a running process (-p) or thread (-t) #### sp-broadcast · Broadcast fan-out overload · Broadcast fan-out (N×N) Sending one player’s movement to everyone who can see them creates updates on the order of the square of the crowd size. - Why → Effect → On screen: Each player’s changes are sent to everyone who can see them → 1,000 players who all see each other means 1 million updates per tick → The send queue and bandwidth saturate, causing delay and loss (input lag, fast-forward, teleporting) - Symptoms: Input lag, Teleporting, Fast-forward / Factors: Latency, Packet loss - Who: Specific zone/channel, Whole server / When: When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Lower the update rate by distance and importance (nearby enemies every tick, distant players a few times a second), cap how much is sent to each player and fill it with the most important updates first, pack several updates into one packet, cap the number of players displayed. - Infra team action items: Alert on each server’s outbound bandwidth and packets per second against the NIC and instance network limits, check headroom before large events. - Ballpark numbers: 1,000 players × 1,000 players × 20 ticks = 20 million updates per second. At 40 bytes each, that’s about 6.4 Gbps for the whole server and about 6.4 Mbps per receiving player. Capping the visible player count at 150 brings it down to about 1 Gbps total and about 1 Mbps per player. - On the graph: Rises with load (Server outbound packets and bytes, players gathered in one place) - Where to look: txpck/s and txkB/s (packets and KB sent per second by the server NIC) from sar -n DEV 1 alongside the player count graph. On cloud instances, the allowance-exceeded counters in ethtool -S (bw_out_allowance_exceeded and pps_allowance_exceeded on AWS ENA) - Confirmed if: As the crowd grows, outbound packets and bytes rise faster than the player count (close to its square), and from the moment they hit the limit, the allowance-exceeded counters or transmit drops increase - Ruled out if: Outbound volume unchanged while only tick time grows: points to AOI calculation or game logic - Check with: Infra tools (no game code needed) - Real incidents: eve-hedgp-2014 - Sources: - [HED-GP Technical Retrospective: What a HED-ache](https://www.eveonline.com/news/view/what-a-hed-ache) · CCP Games · O(n²) traffic, where the actions of n players must be seen by n players, is the unavoidable limiting factor in large fleet battles - [Actor Priority in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-priority-in-unreal-engine) · Epic Games · When a connection’s bandwidth is saturated, each actor gets a priority (distance, line of sight, time since last sent) and bandwidth goes to the most important first - [Detailed Actor Replication Flow in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/detailed-actor-replication-flow-in-unreal-engine) · Epic Games · NetUpdateFrequency sets the update rate per actor; actors are sent in priority order, and once the connection is saturated, the rest wait for the next tick - [sar(1) — Linux manual page](https://man7.org/linux/man-pages/man1/sar.1.html) · sysstat · rxpck/s and txpck/s (packets received and sent per second), rxkB/s and txkB/s (KB received and sent per second) in sar -n DEV - [Monitor network performance for ENA settings on your EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-network-performance-ena.html) · AWS · bw_out_allowance_exceeded (outbound bandwidth allowance exceeded) and pps_allowance_exceeded (PPS allowance exceeded): number of packets queued or dropped #### sp-hotzone · Single-threaded zone overload (hotspot) · Single-threaded hot zone When each area runs on a single thread and players crowd into one place, only that one core hits 100%. - Why → Effect → On screen: One thread runs each area (channel) → When players crowd into one place, only that core saturates while the other cores have room to spare → Only that area lags; other areas are fine - Symptoms: Slow motion, Input lag / Factors: Stall - Who: Specific zone/channel / When: When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Spread players across channels, parallelize work within an area, cap player count. - Infra team action items: Add per-core CPU utilization to monitoring and alerts (a single saturated core gets buried in the server-wide average). - Ballpark numbers: On a 16-core server, one core at 100% shows up as only about 6% total CPU utilization. You have to look at per-core utilization to find it. - On the graph: Rises with load (Per-core CPU utilization, per-thread CPU) - Where to look: Per-core utilization from mpstat -P ALL 1 and per-thread CPU of the game process from pidstat -t 1, compared with the player count in the zone the busiest thread runs - Confirmed if: Total server CPU is low, but one thread (one core) sits near 100%, and at that moment players are crowded into the zone that thread runs - Ruled out if: Several cores evenly high: overall server overload. Only one core’s %soft (receive processing) high: points to NIC interrupts concentrated on one core - Check with: Infra tools (no game code needed) - Learn more: Other areas stay fine only when each area runs its own tick independently. If the area threads wait for each other every tick and move on to the next tick together, the busiest area slows the tick for the whole server. - Real incidents: eve-hedgp-2014 - Sources: - [Time Dilation – How’s That Going?](https://www.eveonline.com/news/view/time-dilation-hows-that-going) · CCP Games · EVE Online’s Time Dilation works per node, so even distant star systems on the same node slow down; big battles run on reinforced nodes that host only 4 star systems - [mpstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/mpstat.1.html) · sysstat · Shows per-processor utilization separately from the overall average (-P ALL); %soft is the share of time spent handling software interrupts - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · -t shows per-thread statistics (CPU utilization and so on) for the threads in a process #### sp-lock · Lock contention · Lock contention When several threads wait on one lock to write the same data, they run one at a time no matter how many threads you add. - Why → Effect → On screen: Several threads use shared data at once, such as the auction house or guild storage → The others wait until the thread holding the lock finishes → Only certain features are slow; in bad cases, the whole tick is delayed - Symptoms: Input lag, Freeze / Factors: Stall - Who: One feature only, Whole server / When: When crowds gather - Primary owner: Game team (Server development) - Game team action items: Split locks into finer-grained ones, do less work inside locks, move to a message-based design (give each piece of data an owning thread, and have other threads only send it requests as messages). - Ballpark numbers: If 20% of the work happens inside the lock, throughput tops out at 5 times that of one thread no matter how many threads you add; at 40%, it stops at 2.5 times. - On the graph: Rises with load (Request processing time, per-thread CPU and context switches) - Where to look: Per-thread voluntary context switches (cswch/s, times a thread stopped to wait for a resource) from pidstat -w -t 1; where threads wait after leaving the CPU (wait time per call stack) from bcc offcputime -p. For .NET, the lock contention count in dotnet-counters (dotnet.monitor.lock_contentions on .NET 9 and later, Monitor Lock Contention Count on 8 and earlier) - Confirmed if: Processing time grows as load rises while CPU utilization stays low, most of the wait time is concentrated in call stacks trying to acquire a lock, and the lock contention count rises along with it - Ruled out if: CPU maxed out: a compute problem (tick overrun, single-threaded zone overload). Waiting on DB or file calls: points to blocking calls on the game thread - Check with: Infra tools (no game code needed) - Learn more: This happens in designs where several threads modify game data together. A design where one thread owns each area or feature and threads communicate only through messages has almost no locks, but you have to watch for work piling up on one thread (single-threaded zone overload). If the game thread waits for a lock held by a slow save operation, that whole tick stalls. - Real incidents: roblox-2021 - Sources: - [Amdahl's Law in the Multicore Era](https://research.cs.wisc.edu/multifacet/papers/ieeecomputer08_amdahl_multicore.pdf) · IEEE · IEEE Computer 2008 paper (authors’ copy). If the fraction that can’t be parallelized is 1−f, the speedup can never exceed 1/(1−f) no matter how many cores you add (Amdahl’s law) - [Request scheduling](https://learn.microsoft.com/en-us/dotnet/orleans/grains/request-scheduling) · Microsoft · Orleans grains (actors) use a single-threaded execution model that runs each request to completion one at a time, so state is never modified concurrently; grains waiting on each other’s responses can deadlock - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · cswch/s in -w is the number of voluntary context switches from stopping to wait for a resource; -t shows it per thread - [Demonstrations of offcputime, the Linux eBPF/bcc version](https://raw.githubusercontent.com/iovisor/bcc/master/tools/offcputime_example.txt) · IO Visor · Sums the time threads spent blocked off the CPU (off-CPU) per call stack; -p selects the process - [Well-known EventCounters in .NET](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/available-counters) · Microsoft · Monitor Lock Contention Count (monitor-lock-contention-count): number of times contention occurred when trying to acquire a monitor lock - [.NET runtime metrics](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/built-in-metrics-runtime) · .NET · dotnet.monitor.lock_contentions since .NET 9: number of times contention occurred when trying to acquire a monitor lock since the process started #### sp-deadlock · Deadlock · Deadlock When two threads each wait for a lock the other holds, both stop forever. - Why → Effect → On screen: Thread A holds lock 1 and waits for lock 2, while B holds lock 2 and waits for lock 1 → Both stop forever, and related threads stop one after another → The whole server stops, and everyone disconnects when the watchdog restarts it - Symptoms: Freeze, Disconnect / Factors: Stall - Who: Whole server, One feature only / When: Randomly, When crowds gather - Primary owner: Game team (Server development) - Game team action items: Set lock ordering rules, use locks with timeouts, add a watchdog that captures a thread dump at the moment of the hang. - On the graph: Mass disconnect (Connection count, server outbound traffic) - Where to look: Call stacks of every thread captured during the hang: jstack for the JVM (finds and flags deadlocks automatically), dotnet-stack for .NET, gdb’s thread apply all bt for native servers, or dump a core file with gcore, restart, and analyze it afterward - Confirmed if: Two or more threads are stuck with stacks waiting for locks held by each other, and process CPU utilization is near 0 the whole time - Ruled out if: One thread spinning at 100% CPU during the hang: infinite loop. Threads waiting on DB or external responses: points to blocking calls or thread pool exhaustion - Check with: Infra tools (no game code needed) - Sources: - [Runtime locking correctness validator](https://docs.kernel.org/locking/lockdep-design.html) · Linux kernel · Taking two locks in opposite orders causes circular waiting and a deadlock (lock inversion deadlock); the Linux kernel checks lock ordering and warns in advance - [Liveness, Readiness, and Startup Probes](https://kubernetes.io/docs/concepts/workloads/pods/probes/) · Kubernetes · A liveness probe catches a deadlock, where the app is running but can’t make progress, and restarts the container - [Diagnostic Tools (Java SE 21 Troubleshooting Guide)](https://docs.oracle.com/en/java/javase/21/troubleshoot/diagnostic-tools.html) · Oracle · jstack prints the stacks of all threads in a running JVM and also finds and flags deadlocks (Found one Java-level deadlock) - [dotnet-stack diagnostic tool - .NET CLI](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/dotnet-stack) · Microsoft · Captures and prints the managed stacks of all threads in a .NET process - [Threads (Debugging with GDB)](https://sourceware.org/gdb/current/onlinedocs/gdb.html/Threads.html) · GNU Project · thread apply all runs the same command (bt: print the call stack) on every thread - [gcore(1) — Linux manual page](https://man7.org/linux/man-pages/man1/gcore.1.html) · gdb · Creates a core file of a running program, and the program keeps running afterward #### sp-sync-call · Blocking calls on the game thread · Synchronous DB / file I/O on the game loop If the server waits for a DB response or a file write in the middle of a tick, all game progress on the server stops for that long. - Why → Effect → On screen: The tick waits on DB reads and writes, log writes, or external API calls → If the DB takes 100 ms, the tick stalls for 100 ms too → Every time the DB or disk slows down, the whole field hitches - Symptoms: Freeze, Stutter / Factors: Stall - Who: Specific zone/channel, Whole server / When: During specific actions, Randomly - Primary owner: Game team (Server development) - Game team action items: Hand all slow work (DB reads and writes, log writes, external API calls) off to run asynchronously and apply the results on the next tick (a timeout alone still leaves the tick stalled while it waits). - Ballpark numbers: Even a 0.5 ms DB round trip within the same data center adds up to 50 ms if it’s called 100 times in one tick. That alone uses up the entire 20-tick budget. - On the graph: Random spikes (Server tick time, DB query latency) - Where to look: The tick time graph on the same time axis as DB query latency (slow query log and so on) and disk latency. Without tick metrics, bcc offcputime -p to see where the game thread waits - Confirmed if: Tick spikes coincide with DB or file latency spikes, and the game thread’s wait time is concentrated in call stacks that receive DB responses or write files - Ruled out if: DB and disk latency quiet while ticks spike: points to a GC pause or lock contention - Check with: Game server or client logs and metrics - Sources: - [Designs, Lessons and Advice from Building Large Distributed Systems](https://www.cs.cornell.edu/projects/ladis2009/talks/dean-keynote-ladis2009.pdf) · Google · LADIS 2009 keynote (Jeff Dean). Round trip within the same data center about 0.5 ms (500,000 ns) - [ASP.NET Core Best Practices](https://learn.microsoft.com/en-us/aspnet/core/fundamentals/best-practices) · Microsoft · Call data access, I/O, and long-running work asynchronously; synchronous blocking calls lead to thread pool exhaustion and slow responses - [Demonstrations of offcputime, the Linux eBPF/bcc version](https://raw.githubusercontent.com/iovisor/bcc/master/tools/offcputime_example.txt) · IO Visor · Sums the time threads spent blocked off the CPU (off-CPU) per call stack; -p selects the process #### sp-queue · Message queue backlog · Mailbox / job queue backlog When requests arrive faster than they’re processed and pile up in the queue, the ones at the back get processed only seconds later or are dropped. - Why → Effect → On screen: Requests arrive faster than they can be processed → The queue grows, and messages are dropped once it passes its limit → Skills and trades respond late or are dropped - Symptoms: Input lag, Dropped action / rollback / Factors: Latency, Packet loss - Who: Specific zone/channel, One feature only / When: When crowds gather - Primary owner: Game team (Server development) - Game team action items: Monitor queue length, adopt a policy that drops the oldest requests first, parallelize processing. - On the graph: Hits a ceiling (Queue length and age of the oldest message, messages processed per second) - Where to look: Per-queue length, age of the oldest message, and messages received, processed, and dropped per second, as logged by the server. Without code metrics, the Recv-Q of the game socket from ss (or netstat): data the kernel has received but the process hasn’t read yet - Confirmed if: While arrivals exceed processing, the processing rate stays flat at one value, and queue length, message age, and drops keep growing - Ruled out if: Queue short and messages young, yet responses are slow: points to connection latency or delay in the tick itself - Check with: Game server or client logs and metrics - Real incidents: eve-hedgp-2014 - Sources: - [Avoiding insurmountable queue backlogs](https://d1.awsstatic.com/builderslibrary/pdfs/avoiding-insurmountable-queue-backlogs.pdf) · AWS · Amazon Builders’ Library. Monitor backlog by the age of waiting messages; real-time systems process the newest data first (closer to LIFO) and may drop old messages - [Site Reliability Engineering, Chapter 22: Addressing Cascading Failures](https://sre.google/sre-book/addressing-cascading-failures/) · Google · When requests arrive faster than they’re processed, the queue fills and latency grows; LIFO or CoDel in place of FIFO sheds old requests that are already useless - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · A tool that shows socket statistics (similar information to netstat); -p shows the process using each socket - [netstat(8) — Linux manual page](https://man7.org/linux/man-pages/man8/netstat.8.html) · net-tools · Recv-Q: bytes on a connected socket not yet picked up by the user program #### sp-timer-burst · Timers firing all at once · Synchronized timers When every monster respawn, every buff expiry, and the on-the-hour reward all land on the same tick, that one tick becomes tens of times heavier. - Why → Effect → On screen: Respawn, expiry, reward, and autosave timers are all set to the same moment → That one tick has tens of times its usual work → A hitch at each of those scheduled times - Symptoms: Freeze, Stutter / Factors: Stall - Who: Specific zone/channel, Whole server / When: At regular intervals - Primary owner: Game team (Server development) - Game team action items: Spread timer times out with a little randomness, split the work across several ticks. - On the graph: Periodic spikes (Server tick time) - Where to look: The times ticks spiked, collected to check the intervals (on the hour, every 5 minutes, and so on), and cross-checked against the list of respawn, buff expiry, reward, and autosave timers that run at the same moment - Confirmed if: Ticks spike at the same times or intervals every time, and some game timer work fires all at once at those times - Ruled out if: Periodic, but lining up with pause times in the GC log or the server’s cron and backup times: points to “Server GC stop-the-world pause” or “Scheduled jobs” - Check with: Game server or client logs and metrics - Sources: - [Timeouts, retries, and backoff with jitter](https://d1.awsstatic.com/builderslibrary/pdfs/timeouts-retries-and-backoff-with-jitter.pdf) · AWS · Amazon Builders’ Library. Add jitter to all timers, periodic jobs, and delayed work to spread out load that would otherwise land at the same moment; a case where one-minute requests from many servers piled into the first few seconds of every minute #### sp-pathfinding · Pathfinding storm · Pathfinding storms When hundreds of monsters chase players and compute paths at the same time, it takes a lot of CPU. - Why → Effect → On screen: Large mob pulls or mass spawns send many monsters chasing players at once → Each monster runs its own pathfinding → Slow motion in that hunting ground only - Symptoms: Slow motion / Factors: Stall - Who: Specific zone/channel / When: When crowds gather - Primary owner: Game team (Server development) - Game team action items: Cache paths, cap the number of path computations, split the work across several ticks. - On the graph: Rises with load (Server tick time, monsters per zone) - Where to look: Number of monsters chasing players and tick time, per zone. Without a separate count, per-function CPU share of the game process from perf top -p - Confirmed if: Tick time grows during large mob pulls and mass spawns, and pathfinding (path search) functions take a large share of CPU time - Ruled out if: Ticks grow when there are few monsters but many players: points to AOI calculation or broadcast fan-out - Check with: Game server or client logs and metrics - Sources: - [AI.NavMesh.pathfindingIterationsPerFrame](https://docs.unity3d.com/ScriptReference/AI.NavMesh-pathfindingIterationsPerFrame.html) · Unity · Pathfinding processes only a set number of nodes per frame, spreading the work over several frames, so the game stays smooth even with long paths or many requests at once - [perf-top(1) — Linux manual page](https://man7.org/linux/man-pages/man1/perf-top.1.html) · perf · Shows live CPU share per function (symbol) for a running process (-p) #### sp-serialize · Serialization and compression cost · Serialization / compression cost Turning outgoing data into bytes and compressing it takes CPU too, and with many players this cost explodes. - Why → Effect → On screen: Structs are converted to bytes and compressed for every update → Cost grows with the square of the player count → Sends go out late: input lag - Symptoms: Input lag / Factors: Stall, Latency - Who: Specific zone/channel / When: When crowds gather - Primary owner: Game team (Server development) - Game team action items: Reuse a packet built once for many players, use a lightweight format. - On the graph: Rises with load (Server CPU utilization, CPU of the thread that builds packets) - Where to look: Share of the game process’s CPU time spent in serialization, compression, and encryption functions (including library functions such as zlib, LZ4, and OpenSSL) from perf top -p, compared between quiet and crowded times - Confirmed if: The more players crowd in, the bigger the share of serialization, compression, and encryption functions, and the thread that builds packets saturates first - Ruled out if: These functions take a small share: points to AOI calculation or game logic - Check with: Infra tools (no game code needed) - Learn more: If the connection encrypts packets (TLS, DTLS, and so on), encryption and decryption use CPU too. Encryption is done separately for each connection, so even when one built packet is reused for many players, the encryption cost scales with the number of recipients. Symmetric ciphers such as AES-GCM are fast enough for one core to handle several GB per second, so their share is usually small, but their speed varies widely with the size of the unit encrypted at once (the record), and with many small packets, as in games, the cost per byte goes up. In the handshake done once per connection, the server signs with its certificate key and computes the key exchange (ECDHE). One core can do about 1,100 (RSA 2048) to 18,000 (ECDSA P-256) signatures and about 9,000 key exchanges per second, which becomes a burden when logins surge. - Sources: - [Introduction to Iris in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/introduction-to-iris-in-unreal-engine) · Epic Games · Holds the state to replicate as a single quantized copy to cut expensive work, and shares that work across connections - [VALORANT's 128-Tick Servers](https://www.riotgames.com/en/news/valorants-128-tick-servers) · Riot Games · Comparing replicated variables for each client every frame and bundling the changed values reads memory all over the place, which is slow and costs a lot of server CPU - [How "expensive" is crypto anyway?](https://blog.cloudflare.com/how-expensive-is-crypto-anyway/) · Cloudflare · BoringSSL measurements: AES-128-GCM about 3.7 GB per second (varies widely with record size); per core per second, 1,120 RSA 2048 signatures, 18,477 ECDSA P-256 signatures, and 9,394 P-256 ECDHE operations; on Cloudflare edge servers, the TLS library used about 1.8% of CPU - [perf-top(1) — Linux manual page](https://man7.org/linux/man-pages/man1/perf-top.1.html) · perf · Shows live CPU share per function (symbol) for a running process (-p) #### sp-crash · Server crash · Server process crash When the server process dies from an unhandled error, everyone on that server disconnects at the same time. - Why → Effect → On screen: A fatal error such as a reference to something that doesn’t exist (null reference), bad data, or running out of memory → The server (or zone) process exits → Everyone disconnects at once, and progress since the last save may be rolled back - Symptoms: Disconnect, Dropped action / rollback / Factors: Stall - Who: Specific zone/channel, Whole server / When: Randomly, During specific actions - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Analyze crash dumps and fix the root cause, save often. - Infra team action items: Restart the process automatically, set up crash dump collection and retention, alert the moment a server goes down. - On the graph: Mass disconnect (Connection count, process restarts) - Where to look: Core dump records in coredumpctl list (time, PID, terminating signal) and the service manager’s (systemd) records of abnormal exits and restarts. On Windows servers, dump files saved by WER - Confirmed if: At the moment the connection count dropped to near 0, the game server process exited abnormally and left a core dump - Ruled out if: Process stayed alive but connections dropped: points to network equipment or an idle timeout. A watchdog restart after a long hang: points to an infinite loop or deadlock - Check with: Infra tools (no game code needed) - Sources: - [Collecting User-Mode Dumps](https://learn.microsoft.com/en-us/windows/win32/wer/collecting-user-mode-dumps) · Microsoft · Configure Windows Error Reporting (WER) to collect full or mini dumps locally when a user-mode program crashes - [systemd.service(5) — Linux manual page](https://man7.org/linux/man-pages/man5/systemd.service.5.html) · systemd · Restart=on-failure automatically restarts the service after an abnormal exit, a kill by signal (including core dumps), or a watchdog timeout; recommended for long-running services - [coredumpctl(1) — Linux manual page](https://man7.org/linux/man-pages/man1/coredumpctl.1.html) · systemd · list queries core dumps saved by systemd-coredump, showing crash time, PID, and the signal that caused the crash #### sp-threadpool · Thread pool exhaustion · Thread pool starvation When every worker thread is tied up in slow work, new requests just wait with no end in sight. - Why → Effect → On screen: Worker threads are tied up waiting on external API or DB responses → No thread is free to take a new request → Infinite loading in specific features such as login or the shop - Symptoms: Can’t connect / infinite loading, Input lag, Freeze / Factors: Stall - Who: One feature only, Whole server / When: When crowds gather, Right after login or maintenance - Primary owner: Game team (Server development) - Game team action items: Put timeouts on slow calls, separate thread pools by feature, make calls asynchronous. - On the graph: Hits a ceiling (Thread pool thread count and queue length, request processing time) - Where to look: For .NET, thread pool thread count and queue length in dotnet-counters monitor (dotnet.thread_pool.thread.count and dotnet.thread_pool.queue.length on .NET 9 and later, ThreadPool Thread Count and ThreadPool Queue Length on 8 and earlier), and where the worker threads are waiting, from dotnet-stack. For JVM and native servers, the same from a thread dump - Confirmed if: CPU utilization is well below 100%, yet the thread count keeps creeping up or sits at its cap, the queue builds up, and most workers are waiting on responses from the same external call (DB, HTTP) - Ruled out if: Queue empty yet still slow: the called service itself is slow, which points to a cascading failure or an external service dependency - Check with: Infra tools (no game code needed) - Learn more: If packet receiving and game logic share the same worker thread pool, the moment a few slow jobs occupy every worker, packet processing stops for the whole server. - Real incidents: riot-euw-2021 - Sources: - [Debug ThreadPool Starvation](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/debug-threadpool-starvation) · Microsoft · When the pool has no threads left and new work has to wait, responses slow down; blocking code that holds threads is the cause. In dotnet-counters, dotnet.thread_pool.thread.count creeping up while CPU is well below 100% signals exhaustion (dotnet.thread_pool.queue.length is often large too); dotnet-stack shows where threads are waiting - [Avoiding insurmountable queue backlogs](https://d1.awsstatic.com/builderslibrary/pdfs/avoiding-insurmountable-queue-backlogs.pdf) · AWS · Concurrent requests = arrival rate × latency (Little’s law). At 100 requests per second, latency growing from 100 ms to 10 s turns 10 threads into 1,000 and the pool runs dry - [Bulkhead Pattern](https://learn.microsoft.com/en-us/azure/architecture/patterns/bulkhead) · Microsoft Azure · With separate connection and thread pools for each called service, a failure in one service blocks only its own pool - [.NET runtime metrics](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/built-in-metrics-runtime) · .NET · dotnet.thread_pool.thread.count (thread pool thread count) and dotnet.thread_pool.queue.length (queued work items) since .NET 9 - [Well-known EventCounters in .NET](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/available-counters) · Microsoft · ThreadPool Thread Count (threadpool-thread-count) and ThreadPool Queue Length (threadpool-queue-length) on .NET 8 and earlier #### sp-infinite-loop · Infinite loops and runaway logic · Infinite loop / runaway logic When a bug keeps a tick from ever finishing, the server stops, and the watchdog forces a restart. - Why → Effect → On screen: A loop that never ends because of a wrong condition, or runaway recursion → The tick never finishes, and the server stops → A freeze, then everyone disconnects - Symptoms: Freeze, Disconnect / Factors: Stall - Who: Specific zone/channel, Whole server / When: During specific actions, Randomly - Primary owner: Game team (Server development) - Game team action items: Cap loop iterations, add a watchdog, write tests that reproduce the problem input. - On the graph: Mass disconnect (Connection count, per-thread CPU) - Where to look: Per-thread CPU from pidstat -t 1 during the hang, and which function the thread spinning at 100% is in, from perf top -t (thread ID) or gdb. If it has already restarted, the watchdog timeout record (WatchdogSec in systemd, a failed Kubernetes liveness probe) - Confirmed if: While the server is stopped, one game thread sits at 100% CPU, and its stack keeps looping inside the same function or loop - Ruled out if: CPU near 0 during the hang: points to a deadlock or waiting on an external response - Check with: Infra tools (no game code needed) - Sources: - [systemd.service(5) — Linux manual page](https://man7.org/linux/man-pages/man5/systemd.service.5.html) · systemd · WatchdogSec=: if the service doesn’t send a keep-alive signal (WATCHDOG=1) within the set time, it’s treated as failed and stopped, then restarted automatically depending on the Restart= setting - [Liveness, Readiness, and Startup Probes](https://kubernetes.io/docs/concepts/workloads/pods/probes/) · Kubernetes · A liveness probe catches a state where the app is running but can’t make progress and restarts it; by default it checks every 10 seconds and restarts after 3 consecutive failures - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · -t shows per-thread statistics (CPU utilization and so on) for the threads in a process - [perf-top(1) — Linux manual page](https://man7.org/linux/man-pages/man1/perf-top.1.html) · perf · Shows live CPU share per function (symbol) for a running thread (-t) or process (-p) #### sp-hot-entity · Combat concentrated on one target (world boss) · Hot entity / combat event fan-out When hundreds of players hit one boss at the same time, the computation for that single boss piles up in one place, and hit information goes out to everyone watching. - Why → Effect → On screen: Hundreds of players use skills, buffs, and debuffs on one boss nonstop → The boss’s HP, aggro list, and debuff calculations pile up in one place, and every hit sends damage number and effect packets to everyone watching → Skills land late and damage numbers pop up in bursts; slow motion only around the boss - Symptoms: Input lag, Fast-forward, Slow motion / Factors: Stall, Latency - Who: Specific zone/channel / When: When crowds gather - Primary owner: Game team (Server development) - Game team action items: Bundle or skip other players’ damage numbers and effects, cap the number of debuffs on one target, split hit processing across several ticks. - Ballpark numbers: 800 players hitting twice a second makes 1,600 hits per second. Telling all 800 watchers about every hit means 1.28 million messages per second. - On the graph: Rises with load (Server tick time, messages sent) - Where to look: Tick time and packets sent during the boss fight, alongside the number of players around the boss, plus events per second (hits, buffs, debuffs) per target if available - Confirmed if: As players gather around the boss, tick time and outbound traffic climb steeply, and the boss alone has tens of times more events per second than any other target - Ruled out if: Things slow down just as much whenever players gather, boss or no boss: points to “AOI calculation blowup (N²)” or “Broadcast fan-out overload” - Check with: Game server or client logs and metrics - Sources: - [HED-GP Technical Retrospective: What a HED-ache](https://www.eveonline.com/news/view/what-a-hed-ache) · CCP Games · Even a single attack must be reported to every client watching, creating an O(n²) load of n players notifying n players, and message-heavy drone attacks make this load grow even faster #### sp-spawn-burst · Spawn burst when entering a crowded area · Spawn burst when entering a crowd When you teleport into a town packed with players, the server has to send the appearance, gear, and status of hundreds of newly visible players all at once. - Why → Effect → On screen: You suddenly appear somewhere crowded by teleporting, logging in, or switching channels → Full data for hundreds of players is built and sent at once, and your PC also loads it all at once → A brief pause right after arrival, characters pop in late one by one, and input lags - Symptoms: Freeze, Input lag, Fast-forward / Factors: Stall, Latency - Who: Just me, Specific zone/channel / When: While moving or changing zones, Right after login or maintenance - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: send in order of distance, spread over several ticks, cache appearance data. Client: preload during the loading screen, create received characters over several frames. - Ballpark numbers: If one player’s appearance, gear, and buff data is 300 bytes, 500 players come to about 150 KB. Tens of times the usual per-tick traffic (a few KB) arrives in a single moment. - On the graph: Surge after opening (Bytes sent per connection, client frame time) - Where to look: Bytes and packets sent over that connection in the first few seconds after arriving somewhere crowded (server log), and client frame time (net graph, client log) - Confirmed if: Right after arrival, outbound traffic on that connection shoots up to tens of times a normal tick and then settles, and client frame time spikes at the same moment - Ruled out if: The same pause when moving somewhere quiet: points to a zone transfer (handoff between servers) or client loading - Check with: Game server or client logs and metrics - Sources: - [Detailed Actor Replication Flow in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/detailed-actor-replication-flow-in-unreal-engine) · Epic Games · When an actor channel is first opened, initial data such as position and rotation is sent along with it, and once the connection is saturated, the remaining actors wait for the next tick - [Actor Priority in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-priority-in-unreal-engine) · Epic Games · Actors are prioritized by distance and view direction, and the nearest visible ones are sent first #### sp-entity-buildup · Entity buildup (items and summons never cleaned up) · Entity / timer buildup over uptime When ground items that should have disappeared, summons, and finished timers pile up without being cleaned up, every tick has more work to do the longer the server stays up. - Why → Effect → On screen: Ground items, summons, expired timers, and empty party data aren’t removed on time → The lists walked every tick grow longer day by day → Fine right after maintenance, then after a few days only that server or area gets more and more sluggish - Symptoms: Slow motion, Stutter, Input lag / Factors: Stall - Who: Whole server, Specific zone/channel / When: The longer it runs - Primary owner: Game team (Server development) - Game team action items: Record entity count per area as a metric and watch the trend, give each entity a lifetime and a count cap, clean up periodically. - Ballpark numbers: On a server that walks every entity once per tick, doubling the entity count doubles that part of the tick time. - On the graph: Sawtooth (Entity count per zone, server tick time) - Where to look: Entity count (ground items, summons, timers) and tick time per zone and server, over a period longer than the maintenance cycle (several weeks) - Confirmed if: Entity count and tick time start low after maintenance, climb every day, and drop sharply at each maintenance or restart, over and over, while memory stays ample - Ruled out if: Ticks unchanged while memory keeps climbing: points to a memory leak - Check with: Game server or client logs and metrics - Learn more: Like a memory leak, it gets worse the longer the server runs, but here memory is fine and only tick time grows. If the entity count graph forms a sawtooth with each maintenance cycle, this is the cause. - Sources: - [Actor Ticking in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-ticking-in-unreal-engine) · Epic Games · Actors and components tick once every frame unless given their own interval, and ticking can be turned off when not needed - [AActor::SetLifeSpan](https://dev.epicgames.com/documentation/en-us/unreal-engine/API/Runtime/Engine/AActor/SetLifeSpan) · Epic Games · Setting a lifespan on an actor destroys it automatically when it expires #### sp-patch-traffic · Patch changes the traffic pattern · Patch changes traffic pattern When new content, effects, or synced fields raise packet size and frequency, a server that ran fine starts hitting MTU, bandwidth, and packet-rate limits after the patch. - Why → Effect → On screen: The patch adds new skill effects, synced fields, or item data, making packets bigger or more frequent → Large packets exceed the MTU and get fragmented, and the extra volume runs into bandwidth limits, cloud PPS limits, and send buffers → Teleporting, skills not going off, and input lag in crowded places, starting right after the patch. Nothing changed in the infra, yet loss goes up - Symptoms: Teleporting, Dropped action / rollback, Input lag / Factors: Packet loss, Latency - Who: Whole server, Specific zone/channel, Specific region/ISP / When: When crowds gather, Evening peak hours, Always - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure), Infra team (Network infrastructure) - Game team action items: Split packets yourself to 1,200 bytes or less, send only changes for new synced fields and lower their rate by distance and importance, compare packets and bytes per second per player and the largest packet size against the previous build on a test server before deploying, tag traffic metrics with the build version. - Infra team action items: Servers/OS: mark deploy times on graphs and compare packets and bytes per second per player and average packet size before and after the deploy, alert on instance allowance-exceeded counters, move to a bigger instance if needed. Network: check the capacity limits of firewalls, load balancers, and DDoS protection gear, and whether they block fragments. - Ballpark numbers: UDP packets are safe at 1,200 bytes or less; the path MTU on the internet is usually 1,500 bytes, and smaller through tunnels (1,476 bytes through a GRE tunnel). Packets larger than the path MTU are fragmented or dropped, and a fragmented packet is lost entirely if even one fragment is lost. If packets per second per player rise 20%, the server total rises 20% too, and an instance running close to its limit overflows right away. - On the graph: Step change (Packets and bytes per second per player, average packet size) - Where to look: Packets and bytes per second on the server NIC (rxpck/s, txpck/s, rxkB/s, and txkB/s from sar -n DEV; NetworkPacketsOut and NetworkOut on EC2) divided by concurrent users, compared before and after the deploy. Average packet size is bytes ÷ packets; for the size distribution, run Wireshark’s Packet Lengths statistics on a packet capture - Confirmed if: From right after the deploy, packets and bytes per player or average packet size step up and stay there, and from the same moment the number of fragments the server creates (fragcrt/s in sar -n IP) or the instance allowance-exceeded counters (pps_allowance_exceeded and bw_out_allowance_exceeded on AWS ENA) rise - Ruled out if: Traffic pattern the same before and after the deploy, but only latency and loss went up: look at infra changes made at the same time (configuration, routing, equipment, OS or kernel updates) - Check with: Infra tools (no game code needed) - Learn more: When a report says “it worked fine before the patch,” this is the game-side cause to check first, along with infra changes. Even if the patch notes list no network changes, one new effect or synced field gets multiplied by hundreds of players in crowded places. Where the extra traffic actually hits a limit is covered in “IP fragmentation of UDP packets,” “Cloud PPS limit exceeded,” “NIC bandwidth saturation,” “Kernel socket buffers too small,” and “Middlebox over capacity (firewall, IPS, DDoS protection).” This entry covers the case where a game patch is what pushed traffic into those limits, so cut the traffic the patch added before raising the limits. If the OS or kernel was also updated at the same time, tell this cause apart from “Performance changes after OS, kernel, driver, or firmware updates” by whether per-player traffic changed. - Sources: - [RFC 8085: UDP Usage Guidelines](https://www.rfc-editor.org/rfc/rfc8085) · IETF · UDP apps SHOULD NOT send datagrams larger than the path MTU; losing one fragment loses the whole fragmented packet, and some NATs and firewalls drop all fragments - [RFC 8899: Packetization Layer Path MTU Discovery for Datagram Transports](https://www.rfc-editor.org/rfc/rfc8899) · IETF · Recommends 1,200 bytes as the default safe size (BASE_PLPMTU) for datagram transports such as UDP - [Maximum transmission unit and maximum segment size](https://developers.cloudflare.com/magic-transit/reference/mtu-mss/) · Cloudflare · Internet path MTU 1,500; 1,476 through a GRE tunnel - [Monitor network performance for ENA settings on your EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-network-performance-ena.html) · AWS · pps_allowance_exceeded and bw_out_allowance_exceeded: packets queued or dropped because they exceeded the instance’s PPS or outbound bandwidth allowance - [sar(1) — Linux manual page](https://man7.org/linux/man-pages/man1/sar.1.html) · sysstat · rxpck/s and txpck/s (packets per second) and rxkB/s and txkB/s (KB per second) in sar -n DEV; fragcrt/s (IP fragments created per second, ipFragCreates) in sar -n IP - [CloudWatch metrics that are available for your instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/viewing_metrics_with_cloudwatch.html) · AWS · NetworkPacketsOut (packets the instance sent on all network interfaces) and NetworkOut (bytes sent) - [8.7. Packet Lengths](https://www.wireshark.org/docs/wsug_html_chunked/ChStatPacketLengths.html) · Wireshark · Splits captured packets into length ranges and shows the count, average, minimum, and maximum for each ### L10 Memory (causes: 9) #### mem-gc · Server GC stop-the-world pause · Stop-the-world GC pause While a Java or C# server halts every thread to collect garbage (stop-the-world), the whole server stalls. - Why → Effect → On screen: The heap fills up and GC starts → Every game thread stops while GC collects (longer the more live data there is) → Everyone on the server freezes at the same moment, then the game fast-forwards - Symptoms: Freeze, Fast-forward / Factors: Stall - Who: Whole server / When: At regular intervals, The longer it runs - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Explicitly select a low-pause GC (ZGC, Shenandoah, or G1 with a lower pause target) in the launch options, reduce allocations, tune heap size. - Infra team action items: Use instances with enough memory for a generous heap, give containers at least 2 CPUs and about 1.8 GB of memory (with less, JDK 26 and earlier pick Serial GC by default), monitor GC pause times. - Ballpark numbers: A minor GC that collects only new objects (the young generation) takes a few to tens of milliseconds. A full GC over a whole heap holding several GB of live data can take over 1 second. ZGC stays under 1 ms almost regardless of heap size, and Shenandoah pauses are also short because they don’t grow with heap size. - On the graph: Periodic spikes (Server tick time, GC pause time) - Where to look: Turn on GC logging and overlay pause times and lengths on the server tick-time graph. Java: Pause lines from the startup option -Xlog:gc* (-XX:+PrintGCDetails on JDK 8 and earlier). .NET: the GC pause metric in dotnet-counters (dotnet.gc.pause.time on .NET 9 and later, % Time in GC since last GC on 8 and earlier). Go: the line GODEBUG=gctrace=1 writes for each GC - Confirmed if: Tick spikes line up with GC pauses, and pause length matches spike length. Every zone and channel on the server spikes at the same moment - Ruled out if: Ticks spike with no long pauses in the GC log: another cause such as locks, blocking calls, or disk writes. Only one zone spikes: the script engine’s GC (mem-script-gc) or that zone’s load - Check with: Infra tools (no game code needed) - Learn more: Java’s G1 has a default target of 200 ms per pause, which is 4 ticks on a 20-tick server. If a container gets fewer than 2 CPUs or less than about 1.8 GB of memory, Java on JDK 26 and earlier picks Serial GC, which collects on a single thread, as the default GC, and pauses get much longer. C# (.NET) servers usually turn on server GC and background GC, but collections of generations 0 and 1 (Gen0/1), where new objects live, and full GCs with compaction still stop every thread. Go pauses are usually under 1 ms, but under heavy allocation the code requesting memory has to take on part of the GC work, so ticks slow down. With any GC, if allocation outpaces collection, the game thread eventually stops. G1 falls back to a full GC, and ZGC stalls the thread requesting memory until the collection finishes. - Real incidents: riot-euw-2021 - Sources: - [Garbage-First (G1) Garbage Collector](https://docs.oracle.com/en/java/javase/25/gctuning/garbage-first-g1-garbage-collector1.html) · Oracle · G1 default pause target 200 ms (MaxGCPauseMillis); if memory runs out during collection, it falls back to a full GC that stops and compacts the whole heap - [JEP 523: Make G1 the Default Garbage Collector in All Environments](https://openjdk.org/jeps/523) · OpenJDK · Previously, Serial GC was the default with 1 CPU or less than 1792 MB of memory; from JDK 27, G1 is the default everywhere - [JEP 439: Generational ZGC](https://openjdk.org/jeps/439) · OpenJDK · ZGC pauses are 1 ms or less regardless of heap size; G1 pauses range from a few ms to a few seconds. Risk of allocation stalls when allocation outpaces reclamation - [JEP 189: Shenandoah: A Low-Pause-Time Garbage Collector (Experimental)](https://openjdk.org/jeps/189) · OpenJDK · Shenandoah pause times are about the same whether the heap is 200 MB or 200 GB - [Background garbage collection](https://learn.microsoft.com/en-us/dotnet/standard/garbage-collection/background-gc) · .NET · Background GC applies only to gen2 collections; gen0 and gen1 collections (foreground GC) stop all managed threads - [A Guide to the Go Garbage Collector](https://go.dev/doc/gc-guide) · Go · Go’s GC runs mostly concurrently with only short stop-the-world pauses; under heavy allocation, goroutines take on GC work (assist), which adds latency - [JEP 271: Unified GC Logging](https://openjdk.org/jeps/271) · OpenJDK · JDK 9 reimplemented GC logging on unified logging (-Xlog); -Xlog:gc writes one line per GC, like the old -XX:+PrintGC - [The java Command](https://docs.oracle.com/en/java/javase/25/docs/specs/man/java.html) · Oracle · Table mapping old GC log options to -Xlog: -XX:+PrintGCDetails becomes -Xlog:gc* - [dotnet-counters diagnostic tool](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/dotnet-counters) · .NET · .NET 9 and later show System.Runtime meters (dotnet.gc.pause.time and others); .NET 8 and earlier show the older EventCounters (% Time in GC since last GC and others) - [runtime package](https://pkg.go.dev/runtime) · Go · GODEBUG=gctrace=1: one line per GC with wall-clock time per phase, heap size at GC start and end, and the heap goal #### mem-script-gc · Script engine GC pause · Scripting VM GC (Lua, etc.) Even on a C++ server, if quests, AI, and skills run in a scripting language such as Lua, the zone stops while the script engine’s GC runs. - Why → Effect → On screen: Each zone’s script engine creates large numbers of temporary objects while running quests, AI, and events → When the script engine’s GC collects a lot at once, that zone’s tick stops → Periodic hitches only in certain zones or during certain events - Symptoms: Stutter, Freeze / Factors: Stall - Who: Specific zone/channel / When: When crowds gather, At regular intervals - Primary owner: Game team (Server development) - Game team action items: Configure incremental or generational GC, advance GC a little every tick, reduce temporary objects in scripts. - Ballpark numbers: Once a script heap grows to hundreds of MB, an all-at-once collection (with incremental collection turned off, or a full collection in generational mode) can take tens to hundreds of milliseconds. - On the graph: Periodic spikes (Tick time per zone, script engine memory) - Where to look: Record each zone’s tick time and the memory use of that zone’s script engine (collectgarbage("count") in Lua) every tick and overlay them on one graph - Confirmed if: The moments script memory drops sharply (an all-at-once collection) line up with that zone’s tick spikes, while other zones are fine - Ruled out if: Spikes with no change in script memory: that zone’s load or locks. Every zone on the server spikes at once: server GC (mem-gc) or swap (mem-swap) - Check with: Game server or client logs and metrics - Sources: - [Lua 5.4 Reference Manual](https://www.lua.org/manual/5.4/manual.html) · Lua.org · Incremental mode splits collection into small steps interleaved with execution (large steps make it stop-the-world); a major collection in generational mode is a stop-the-world pass over every object; collectgarbage("count") returns the total memory Lua is using (KB) #### mem-alloc · Allocation surge · Allocation storms Creating large numbers of temporary objects during an event makes GC run far more often than usual. - Why → Effect → On screen: Item drops, combat logs, and event rewards create a flood of temporary objects → GC runs several times as often, and objects not yet discarded get promoted to the old generation, so full GCs come sooner too → Periodic hitches only during events - Symptoms: Stutter, Freeze / Factors: Stall - Who: Whole server, Specific zone/channel / When: When crowds gather - Primary owner: Game team (Server development) - Game team action items: Use object pools and reusable buffers, profile allocations. - On the graph: Rises with load (GC count, allocation rate) - Where to look: Count GCs per minute from GC logs (Java -Xlog:gc, Go GODEBUG=gctrace=1); on .NET, check allocation volume and GC count in dotnet-counters (dotnet.gc.heap.total_allocated and dotnet.gc.collections on .NET 9 and later, Allocation Rate and Gen 0 GC Count on 8 and earlier). Overlay with concurrent users and event times - Confirmed if: When the event starts, allocation rate and GC count climb faster than player count, and short pauses become frequent. Back to normal once the event ends - Ruled out if: GC count unchanged but each pause gets longer: live data has grown (mem-gc-thrash, mem-leak) - Check with: Infra tools (no game code needed) - Sources: - [Garbage Collector Implementation](https://docs.oracle.com/en/java/javase/25/gctuning/garbage-collector-implementation.html) · Oracle · Minor GC when the young generation fills; some surviving objects move to the old generation, and when the old generation fills, the whole heap is collected (much slower than a minor GC); -Xlog:gc writes one line per GC - [A Guide to the Go Garbage Collector](https://go.dev/doc/gc-guide) · Go · The higher the allocation rate, the more often GC cycles run; GODEBUG=gctrace=1 prints GC trace output - [dotnet-counters diagnostic tool](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/dotnet-counters) · .NET · .NET 9 and later show dotnet.gc.heap.total_allocated and dotnet.gc.collections; .NET 8 and earlier show Allocation Rate and Gen 0 GC Count #### mem-leak · Memory leak · Memory leak Memory that is never freed piles up little by little and, days later, leads to GC storms, swapping, or the process getting killed. - Why → Effect → On screen: Data for logged-out characters and event handlers is never freed → Free memory shrinks over several days → Fine right after maintenance, laggier every day, and eventually the server goes down - Symptoms: Slow motion, Freeze, Disconnect / Factors: Stall - Who: Whole server / When: The longer it runs, Evening peak hours - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Analyze heap dumps, run long-duration load tests. - Infra team action items: Monitor and alert on per-process memory usage trends. - On the graph: Slow climb (Process memory (RSS), heap after GC) - Where to look: Watch the game server process’s memory (RSS from pidstat -r) over several days; on servers with GC, watch the heap left right after GC. Java: the after-GC value in the before/after usage of -Xlog:gc lines. .NET: the post-GC heap size in dotnet-counters (dotnet.gc.last_collection.heap.size on .NET 9 and later, GC Heap Size on 8 and earlier) - Confirmed if: The heap left right after GC (the baseline) climbs every day after a restart and doesn’t come down even in the quiet early-morning hours - Ruled out if: Heap baseline flat while only RSS climbs: fragmentation (mem-fragment) or native memory. Rises and falls with player count: normal usage - Check with: Infra tools (no game code needed) - Learn more: Weekly restarts during scheduled maintenance hide a leak, so it can go unnoticed for a long time. It often shows up suddenly when maintenance is postponed once or an event brings in more players. - Sources: - [Troubleshoot Memory Leaks](https://docs.oracle.com/en/java/javase/25/troubleshoot/troubleshooting-memory-leaks.html) · Oracle · Suspect a leak when execution gets gradually slower; memory eventually runs out and the program terminates abnormally. The key data for leak analysis is a heap dump - [Debug a memory leak in .NET](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/debug-memory-leak) · .NET · Even with GC, holding references to objects no longer needed is a leak, leading to performance degradation and OutOfMemoryException. Check memory trends and analyze dumps - [Garbage Collector Implementation](https://docs.oracle.com/en/java/javase/25/gctuning/garbage-collector-implementation.html) · Oracle · -Xlog:gc lines use the format “before-GC usage->after-GC usage(heap size)” - [dotnet-counters diagnostic tool](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/dotnet-counters) · .NET · .NET 9 and later show dotnet.gc.last_collection.heap.size; .NET 8 and earlier show GC Heap Size - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · -r: per-process RSS (memory actually resident in RAM) and page faults #### mem-gc-thrash · GC thrashing (too little heap headroom) · GC thrashing (heap nearly full) When live data gets close to the heap limit, each GC reclaims almost nothing, so GC runs over and over without a break. - Why → Effect → On screen: An event crowd or a leak pushes live data close to the heap limit → GC reclaims only a little, so another full GC follows right away and GC uses most of the CPU → The whole server alternates between slow motion and freezes for several minutes, then dies from running out of memory - Symptoms: Slow motion, Freeze, Disconnect / Factors: Stall - Who: Whole server / When: Evening peak hours, When crowds gather, The longer it runs - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Size the heap well above peak live data (usually 2× or more), reduce long-held references and leaks. - Infra team action items: Alert on GC time ratio, restart promptly without waiting for it to recover, use instances with enough memory to grow the heap. - Ballpark numbers: GC taking more than 10% of run time is usually treated as a warning sign. Some Java GCs throw an out-of-memory error when they spend 98% of the time in GC and still reclaim almost nothing. - On the graph: Hits a ceiling (Heap after GC, GC time ratio) - Where to look: How close the heap left after GC is to the max heap, and the share of time spent in GC. Java: “after GC (heap size)” in -Xlog:gc lines and how often Pause Full lines appear. .NET: dotnet-counters (% Time in GC since last GC on .NET 8 and earlier, growth of dotnet.gc.pause.time on .NET 9 and later). Go: the interval between GODEBUG=gctrace=1 lines - Confirmed if: The heap stays near its max even right after GC, full GCs run back to back, and the GC time ratio climbs well above normal (usually past 10%). Ticks slow down across the whole server meanwhile - Ruled out if: Plenty of heap left after GC but pauses are long: GC type or settings (mem-gc) - Check with: Infra tools (no game code needed) - Sources: - [The Parallel Collector](https://docs.oracle.com/en/java/javase/25/gctuning/parallel-collector1.html) · Oracle · Parallel GC throws OutOfMemoryError if it spends more than 98% of total time in GC and recovers less than 2% of the heap - [Garbage-First Garbage Collector Tuning](https://docs.oracle.com/en/java/javase/25/gctuning/garbage-first-garbage-collector-tuning.html) · Oracle · By default (GCTimeRatio=12), G1 sizes the heap to keep GC time at about 8% of total time or less; full GCs caused by excessive heap occupancy show up in the log as Pause Full (G1 Compaction Pause) - [A Guide to the Go Garbage Collector](https://go.dev/doc/gc-guide) · Go · With the default GOGC=100, the heap goal is about 2× the live heap; near the memory limit, GC runs nonstop (thrashing); GODEBUG=gctrace=1 prints GC trace output - [Garbage Collector Implementation](https://docs.oracle.com/en/java/javase/25/gctuning/garbage-collector-implementation.html) · Oracle · -Xlog:gc lines show the GC type (Pause Young, Pause Full), “before-GC usage->after-GC usage(heap size)”, and the pause time - [dotnet-counters diagnostic tool](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/dotnet-counters) · .NET · .NET 9 and later show dotnet.gc.pause.time; .NET 8 and earlier show % Time in GC since last GC #### mem-swap · Swap · Swapping When memory runs short and the OS moves part of it out to disk, every access to that memory waits on a disk more than 1,000 times slower. - Why → Effect → On screen: Memory in use exceeds physical RAM → The OS moves part of it to disk and reads it back when needed → Ticks balloon to hundreds of ms, and every player on the server sees slow motion and freezes - Symptoms: Slow motion, Freeze / Factors: Stall - Who: Whole server / When: The longer it runs, Evening peak hours - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: Put a cap on process memory use (heap size and so on), check for leaks. - Infra team action items: Configure game servers not to use swap, act on memory alerts, provision RAM well above peak usage, since with swap off the process is killed (OOM) the moment memory runs out. - Ballpark numbers: A RAM read takes about 100 ns; reading back from an SSD takes about 100 µs (1,000×), a cloud disk over the network about 1 ms (10,000×), and an HDD 10 ms (100,000×). - On the graph: Slow climb (Swap usage, swap-in/out) - Where to look: Overlay on tick time: the si and so columns of vmstat 1 (amount swapped in and out per second), some and full in /proc/pressure/memory (share of time stalled waiting for memory), and pidstat -r majflt/s for the game server process (page faults that had to read from disk) - Confirmed if: si above 0 at the time of the lag, with the game server’s majflt/s and the memory full value rising together - Ruled out if: si and so at 0 and memory pressure (PSI) near 0: swap isn’t the cause. No swap but majflt/s and PSI rising: memory is running out and code pages are being reread, so free up memory first - Check with: Infra tools (no game code needed) - Learn more: A server with GC reads all over the heap when it collects, so if even part of the heap is swapped out, a single GC can stretch to seconds or tens of seconds. With swap off, there is no slow swapping stage and the process goes straight to being killed (OOM), so secure spare memory first. Even without swap, when memory is nearly exhausted the OS may drop the executable’s code pages from memory and read them back again, so the whole server can slow down badly for a while before the OOM kill. - Sources: - [Designs, Lessons and Advice from Building Large Distributed Systems (LADIS 2009 keynote)](https://www.cs.cornell.edu/projects/ladis2009/talks/dean-keynote-ladis2009.pdf) · Google · “Numbers Everyone Should Know”: main memory reference 100 ns, disk seek 10 ms (as of 2009) - [Documentation for /proc/sys/vm/](https://docs.kernel.org/admin-guide/sysctl/vm.html) · Linux kernel · swappiness: relative cost of swapping vs. reclaiming file pages; swap is random I/O and therefore expensive - [Concepts overview](https://docs.kernel.org/admin-guide/mm/concepts.html) · Linux kernel · Reclaims page cache backed by files on disk and swappable pages; if that’s still not enough, the OOM killer kills a process - [Solidigm™ D7-P5520 and D7-P5620 Product Brief](https://www.solidigm.com/products/data-center/product-briefs/d7-p5520-p5620-product-brief.html) · Solidigm · 99.99th percentile latency (four-nines latency) of 130 µs for server NVMe SSDs: basis for a single SSD read taking around 100 µs - [Amazon EBS General Purpose SSD volumes](https://docs.aws.amazon.com/ebs/latest/userguide/general-purpose.html) · AWS · Latency of the default cloud disk (gp3) is in single-digit milliseconds - [vmstat(8) — Linux manual page](https://man7.org/linux/man-pages/man8/vmstat.8.html) · procps-ng · si: memory swapped in per second; so: memory swapped out per second - [PSI - Pressure Stall Information](https://docs.kernel.org/accounting/psi.html) · Linux kernel · some (share of time some tasks were stalled) and full (share of time all tasks were stalled at once) in /proc/pressure/memory - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · -r: majflt/s (faults that had to read a page from disk) #### mem-cache-miss · Cache miss · CPU cache misses When data is scattered all over memory, the CPU has to go all the way out to slow RAM and wait every time. - Why → Effect → On screen: Objects scattered behind pointers and accessed in no particular order → Data isn’t in the CPU cache, so every read goes to RAM (roughly 100 times slower) → The same work costs several times more tick time; in bad cases, slow motion - Symptoms: Slow motion / Factors: Stall - Who: Whole server / When: Always, When crowds gather - Primary owner: Game team (Server development) - Game team action items: Lay out data that is used together contiguously (data-oriented design). - On the graph: Always high (Tick time, CPU utilization) - Where to look: Attach perf stat -d -p PID to the game server process to measure instructions per cycle (insn per cycle) and L1/LLC cache misses, and view them alongside tick time and CPU utilization - Confirmed if: CPU stays busy, but insn per cycle is low and LLC misses are high. Confirmed if a build with a changed data layout cuts tick time sharply at the same player count - Ruled out if: Low CPU utilization but slow ticks: a cause that waits outside the CPU, such as locks or I/O waits - Check with: Infra tools (no game code needed) - Sources: - [Designs, Lessons and Advice from Building Large Distributed Systems (LADIS 2009 keynote)](https://www.cs.cornell.edu/projects/ladis2009/talks/dean-keynote-ladis2009.pdf) · Google · L1 cache 0.5 ns, L2 cache 7 ns, main memory 100 ns (as of 2009): going out to RAM is one to two orders of magnitude slower than cache - [perf-stat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/perf-stat.1.html) · perf · -p counts hardware events for a running process and shows insn per cycle; -d adds L1 and LLC data cache events #### mem-fragment · Memory fragmentation · Heap fragmentation When repeated allocation and freeing chops free space into small pieces, the process holds far more memory than it actually uses. - Why → Effect → On screen: Many threads allocate and free blocks of varying sizes over a long time → Free space ends up scattered in small pieces that can’t be returned to the OS, so usage keeps growing like a leak → The longer it runs, the slower it gets from swapping and memory shortage, until it gets killed - Symptoms: Slow motion, Disconnect / Factors: Stall - Who: Whole server / When: The longer it runs - Primary owner: Game team (Server development) - Game team action items: Use per-size memory pools and fragmentation-resistant allocators (jemalloc, mimalloc, and so on). - On the graph: Slow climb (Process memory (RSS)) - Where to look: Run two servers on the same build, and on one of them either reduce the number of glibc arenas with the MALLOC_ARENA_MAX environment variable or switch to another allocator such as jemalloc. Compare RSS from pidstat -r over several days - Confirmed if: With similar player and entity counts, only the changed server’s RSS stops growing or grows much more slowly - Ruled out if: Still climbs the same way after switching allocators: memory that is never freed (mem-leak) - Check with: Infra tools (no game code needed) - Learn more: It looks just like a leak on the graph, yet heap analysis finds no leak site. The default Linux allocator (glibc) is especially bad on servers with many threads, and just switching allocators can cut memory use significantly. - Sources: - [mallopt(3) — Linux manual page](https://man7.org/linux/man-pages/man3/mallopt.3.html) · Linux man-pages · To reduce thread contention, glibc malloc creates arenas up to a multiple of the CPU count, and more arenas mean more memory use (limit with M_ARENA_MAX, or set the MALLOC_ARENA_MAX environment variable) - [jemalloc memory allocator](https://jemalloc.net/) · jemalloc · A general-purpose malloc implementation that emphasizes fragmentation avoidance and scalable concurrency - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · -r: per-process RSS (memory actually resident in RAM) #### mem-numa · Remote NUMA memory · Remote NUMA access On a server with two CPUs, using memory attached to the other CPU slows access down. - Why → Effect → On screen: Threads and their memory sit on different CPU sockets → Memory access slows down (1.5–2× depending on hardware) → Same specs, but performance differs from process to process - Symptoms: Slow motion / Factors: Stall - Who: Whole server / When: Always - Primary owner: Infra team (Server infrastructure) - Infra team action items: Pin processes and their memory to one socket (numactl), on a two-socket machine run separate game server processes per socket. - On the graph: Outliers only (Tick time per process, memory per node) - Where to look: Use numastat -p PID to see which NUMA node holds the game server process’s memory, check whether numa_miss and other_node in numastat are rising, and compare with the node of the CPU the process runs on - Confirmed if: Only the slow processes have most of their memory on a different node from the CPU they run on, and the gap disappears after restarting them with CPU and memory pinned to one node via numactl - Ruled out if: Slow even with the same node placement as the fast processes: another cause such as a noisy neighbor, CPU throttling, or that process’s load - Check with: Infra tools (no game code needed) - Sources: - [What is NUMA?](https://docs.kernel.org/mm/numa.html) · Linux kernel · Memory in the same cell is faster and has higher bandwidth; memory in another (remote) cell is slower to access - [numactl(8) — Linux manual page](https://man7.org/linux/man-pages/man8/numactl.8.html) · numactl · --cpunodebind and --membind pin a process’s CPU and memory to specific NUMA nodes - [numastat(8) — Linux manual page](https://man7.org/linux/man-pages/man8/numastat.8.html) · numactl · numa_miss (allocated on a node other than the intended one) and other_node (allocated on this node by a process running on another node) counters; -p shows a process’s memory per node ### L11 Disk (causes: 9) #### dk-sync-log · Synchronous log writes · Synchronous logging If the game thread waits for the disk to finish every log line, the game stalls too whenever the disk is busy. - Why → Effect → On screen: Combat and trade logs are written straight to a file from the game thread → When a durable write (fsync) is required or the OS write buffer (page cache) hits its limit, a single write takes tens of ms while the disk is busy → Hitches in log-heavy fights - Symptoms: Stutter, Freeze / Factors: Stall - Who: Specific zone/channel, Whole server / When: When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Log asynchronously (memory buffer + separate thread), reduce log volume, don’t fsync on the game thread. - Infra team action items: Run log rotation and compression at low I/O priority, keep logs on a different disk from data, monitor disk write latency. - On the graph: Random spikes (Server tick time, disk write latency) - Where to look: Overlay w_await and aqu-sz from iostat -x 1 on tick time, and use perf trace -p PID --duration 10 to find write and fsync calls in the game server that took over 10 ms, along with their threads - Confirmed if: At the tick spikes, the game thread’s write and fsync calls take tens of ms, and disk write latency spikes at the same moment. Often lines up with log rotation or compression - Ruled out if: Ticks spike with no slow system calls on the game thread: another cause such as GC, locks, or tick overrun. Only the dedicated logging thread is slow: no effect on gameplay - Check with: Infra tools (no game code needed) - Learn more: Normally the OS accepts writes into memory (the page cache) first and flushes them to disk later, so a log line usually completes right away. Stalls happen when fsync demands a durable write, when backed-up writes exceed the limit and the OS blocks the write call, and when log files are rotated or compressed. That’s why it’s fine most of the time and spikes only at moments when the disk is busy. - Sources: - [fsync(2) — Linux manual page](https://man7.org/linux/man-pages/man2/fsync.2.html) · Linux man-pages · fsync flushes modified data all the way to the disk (including the disk cache) and blocks until the device reports completion - [Documentation for /proc/sys/vm/](https://docs.kernel.org/admin-guide/sysctl/vm.html) · Linux kernel · When backed-up (dirty) writes reach dirty_ratio, the writing process has to do the disk writeback itself - [ionice(1) — Linux manual page](https://man7.org/linux/man-pages/man1/ionice.1.html) · util-linux · A job at idle I/O priority gets disk time only when no other program is using the disk - [iostat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/iostat.1.html) · sysstat · -x: w_await (average time per write request, including time waiting in the queue), aqu-sz (average queue length, formerly avgqu-sz) - [perf-trace(1) — Linux manual page](https://man7.org/linux/man-pages/man1/perf-trace.1.html) · perf · -p traces the system calls of a running process; --duration shows only calls that took longer than the given ms #### dk-fsync · fsync surge · fsync storms Asking for data to be written to disk “for sure” takes 0.1 ms to tens of ms per request depending on the disk, and when requests pile up, the queue grows. - Why → Effect → On screen: Scheduled saves and logout rushes send a flood of durable write requests → The disk queue grows → Lag at every save time, slow logouts and channel changes - Symptoms: Stutter, Input lag / Factors: Stall, Latency - Who: Whole server / When: At regular intervals, When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure), Infra team (DB infrastructure) - Game team action items: Batch saves (many saves in one fsync), spread out save times. - Infra team action items: Servers/OS: use server SSDs with power-loss protection, monitor disk queue length and fsync latency. DB hosts: if saves go to the DB, put the DB log disk on the same kind of SSD, monitor commit latency. - Ballpark numbers: The time per call varies by hardware, but roughly: server SSD (with power-loss protection) 0.1 ms, regular SSD 1 to a few ms, cloud disk 1–2 ms, HDD 10 ms or more. With one thread waiting on each call in turn, an HDD can’t manage even 100 per second. - On the graph: Periodic spikes (Disk queue length, flush and write latency) - Where to look: Overlay f/s and f_await (flushes the disk handled and how long they took), plus w/s, aqu-sz, and w_await, from iostat -x 1 on scheduled save and logout times. Older sysstat shows aqu-sz as avgqu-sz. On cloud disks, check EBS VolumeQueueLength and VolumeAvgWriteLatency - Confirmed if: At every save time and logout rush, flush count and queue length spike together, and w_await and f_await reach several times normal. Saves and channel changes slow down at those moments - Ruled out if: Queue spikes at times unrelated to saves or logouts: backup or compression (dk-backup) or the IOPS limit (dk-iops). Slower at the same flush count: points to burst credit depletion (dk-burst) - Check with: Infra tools (no game code needed) - Sources: - [fsync(2) — Linux manual page](https://man7.org/linux/man-pages/man2/fsync.2.html) · Linux man-pages · fsync empties even the disk cache and blocks until the device reports completion - [Reliability (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/wal-reliability.html) · PostgreSQL · Ordinary SATA disks and many SSDs have write caches that are lost on power failure; durable writes need a cache with battery or power-loss protection - [Amazon EBS General Purpose SSD volumes](https://docs.aws.amazon.com/ebs/latest/userguide/general-purpose.html) · AWS · Default cloud disk (gp3) latency is single-digit ms; io2 Block Express averages under 500 µs for 16 KiB I/O - [Designs, Lessons and Advice from Building Large Distributed Systems (LADIS 2009 keynote)](https://www.cs.cornell.edu/projects/ladis2009/talks/dean-keynote-ladis2009.pdf) · Google · One HDD seek 10 ms - [iostat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/iostat.1.html) · sysstat · -x: f/s and f_await (flush requests the disk handled and their average time), w/s, w_await, aqu-sz (formerly avgqu-sz) - [Amazon CloudWatch metrics for Amazon EBS](https://docs.aws.amazon.com/ebs/latest/userguide/using_cloudwatch_ebs.html) · AWS · VolumeQueueLength (requests waiting to complete), VolumeAvgWriteLatency (1-minute average write latency, Nitro instances) #### dk-burst · Cloud disk out of burst credits · Burst credit depletion Some cloud disks and small server sizes have burst credits that let them run faster than baseline for a while, so when a busy period drags on and the credits run out, speed drops suddenly. - Why → Effect → On screen: Sustained use above baseline performance → Burst credits run out and performance drops sharply to baseline → Lag starts a few hours into every evening - Symptoms: Stutter, Slow motion, Input lag / Factors: Stall, Latency - Who: Whole server / When: Evening peak hours, The longer it runs - Primary owner: Infra team (Server infrastructure) / Also: Infra team (DB infrastructure) - Infra team action items: Servers/OS: use disks with provisioned performance (gp3, provisioned IOPS), alert on credit balance, also check the instance’s disk bandwidth burst limit and CPU credits. DB hosts: move DB disks, including managed databases, to provisioned performance too, alert on credit balance. - Ballpark numbers: An AWS gp2 100 GB disk normally gets 300 IOPS, bursts to 3,000, and lasts about 30 minutes on a full credit balance. gp3 has no credits and always gets 3,000. Small Azure Premium SSDs also burst on credits for up to 30 minutes. - On the graph: Hits a ceiling (IOPS, burst credit balance) - Where to look: In CloudWatch, EBS BurstBalance (gp2, st1, sc1), the instance’s EBSIOBalance% and EBSByteBalance% (some instances that burst), and CPUCreditBalance on burstable instances. On Azure, burst credit usage metrics such as Data Disk Used Burst IO Credits Percentage - Confirmed if: From the moment the balance drops near 0, IOPS (VolumeReadOps, VolumeWriteOps) flattens at baseline, and VolumeQueueLength and lag rise together. Starts after the peak has lasted a few hours - Ruled out if: All balances are healthy but IOPS is flat: a fixed volume or instance limit (dk-iops) - Check with: Infra tools (no game code needed) - Learn more: Even with a healthy disk, small virtual servers have a burst limit on the instance’s own disk bandwidth (for example, at least 30 minutes a day), which produces the same pattern. Low-cost servers that run on CPU credits also slow down to baseline performance once the credits run out. - Sources: - [Amazon EBS General Purpose SSD volumes](https://docs.aws.amazon.com/ebs/latest/userguide/general-purpose.html) · AWS · gp2 baseline is 3 IOPS per GiB (minimum 100), bursting to 3,000 IOPS on I/O credits; 5.4 million credits last at least 30 minutes. gp3 always gets 3,000 IOPS with no burst - [Managed disk bursting](https://learn.microsoft.com/en-us/azure/virtual-machines/disk-bursting) · Microsoft Azure · Premium SSD P20 and smaller use credit-based bursting; a full credit balance gives 30 minutes at max burst speed - [Amazon EBS-optimized instance types](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ebs-optimized.html) · AWS · Some instances sustain maximum EBS performance for only 30 minutes once every 24 hours, then return to baseline - [Standard mode for burstable performance instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/burstable-performance-instances-standard-mode.html) · AWS · In standard mode, a burstable instance that runs out of CPU credits lowers CPU utilization to the baseline level (gradually, without a sudden drop) - [Amazon CloudWatch metrics for Amazon EBS](https://docs.aws.amazon.com/ebs/latest/userguide/using_cloudwatch_ebs.html) · AWS · BurstBalance: remaining I/O credits for gp2 and throughput credits for st1 and sc1 (%); VolumeReadOps, VolumeWriteOps, VolumeQueueLength - [CloudWatch metrics that are available for your instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/viewing_metrics_with_cloudwatch.html) · AWS · EBSIOBalance% and EBSByteBalance%: remaining EBS credits on some instances that burst for 30 minutes once every 24 hours; CPUCreditBalance: remaining CPU credits on burstable instances - [Disk metrics](https://learn.microsoft.com/en-us/azure/virtual-machines/disks-metrics) · Microsoft Azure · Disk and VM burst credit usage (5-minute intervals), such as Data Disk Used Burst IO Credits Percentage #### dk-iops · IOPS limit / queue saturation · IOPS limit / queue saturation When requests exceed what the disk can handle per second, the queue grows and latency explodes. - Why → Effect → On screen: Read and write requests approach the disk’s capacity → The queue grows (usually exploding above 90% utilization) → Slow saves and loading; freezes if the calls are blocking - Symptoms: Input lag, Freeze / Factors: Latency, Stall - Who: Whole server / When: When crowds gather, Evening peak hours - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development), Infra team (DB infrastructure) - Game team action items: Merge requests, cache frequently read data, make disk access asynchronous so the game thread never waits on the disk. - Infra team action items: Servers/OS: use faster disks, check disk bandwidth and IOPS limits per instance type, alert on disk utilization and queue, run large file copies during quiet hours. DB hosts: alert on IOPS and throughput utilization for DB disks too, check the disk limits of the DB instance size. - Ballpark numbers: HDD about 150 IOPS, SATA SSD tens of thousands, NVMe hundreds of thousands. The default cloud disk (AWS gp3) gets 3,000 IOPS and 125 MiB per second. The per-second throughput limit is separate from IOPS, and when a large file copy fills it, even small writes get stuck behind it. - On the graph: Hits a ceiling (IOPS, disk queue length) - Where to look: r/s and w/s, rkB/s and wkB/s, aqu-sz, and r_await and w_await from iostat -x 1. In the cloud, EBS VolumeReadOps, VolumeWriteOps, and VolumeQueueLength, the limit-exceeded checks VolumeIOPSExceededCheck and VolumeThroughputExceededCheck, and on the instance side InstanceEBSIOPSExceededCheck and InstanceEBSThroughputExceededCheck - Confirmed if: Requests per second or throughput flatten at the limit while aqu-sz and await shoot up together. In the cloud, the exceeded-check metric reads 1 - Ruled out if: Even at 100% %util, low await means there may still be headroom. On SSDs and RAID that process requests in parallel, %util doesn’t mean the limit has been reached. High await without hitting the limit: latency of the disk itself (dk-hdd) or fsync (dk-fsync) - Check with: Infra tools (no game code needed) - Learn more: In the cloud, each server size (instance type) has its own disk bandwidth and IOPS limits, separate from the disk’s limits. Even with an expensive disk attached, a small server gets capped at the instance limit. - Sources: - [Exos X18 Data Sheet](https://www.seagate.com/www-content/datasheets/pdfs/exos-x18-mango-DS2045-1N-2007US-en_US.pdf) · Seagate · 4K random reads on a 7,200 rpm server HDD: 170 IOPS (QD16) - [D3-S4520 SSD](https://www.solidigm.com/products/data-center/d3/s4520.html) · Solidigm · Server SATA SSD 4 KB random read/write up to 92K/48K IOPS - [Solidigm™ D7-P5520 and D7-P5620 Product Brief](https://www.solidigm.com/products/data-center/product-briefs/d7-p5520-p5620-product-brief.html) · Solidigm · Server NVMe SSD random read/write 1,000K/200K IOPS - [Amazon EBS General Purpose SSD volumes](https://docs.aws.amazon.com/ebs/latest/userguide/general-purpose.html) · AWS · gp3 baseline 3,000 IOPS and 125 MiB/s, two separate limits that can be raised independently - [Amazon EBS-optimized instance types](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ebs-optimized.html) · AWS · Each instance type has its own baseline and maximum limits for EBS bandwidth, throughput, and IOPS - [iostat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/iostat.1.html) · sysstat · -x: r/s and w/s, rkB/s and wkB/s, aqu-sz (formerly avgqu-sz), r_await and w_await, %util. On RAID and modern SSDs that process requests in parallel, %util doesn’t indicate the performance limit - [Amazon CloudWatch metrics for Amazon EBS](https://docs.aws.amazon.com/ebs/latest/userguide/using_cloudwatch_ebs.html) · AWS · VolumeIOPSExceededCheck and VolumeThroughputExceededCheck: 1 if the volume tried to exceed its IOPS or throughput limit (Nitro instances); VolumeQueueLength - [CloudWatch metrics that are available for your instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/viewing_metrics_with_cloudwatch.html) · AWS · InstanceEBSIOPSExceededCheck and InstanceEBSThroughputExceededCheck: 1 if the instance tried to exceed its EBS IOPS or throughput limit #### dk-full · Disk full · Disk full When logs and dumps pile up and fill the disk, writes fail, and without safeguards the server crashes. - Why → Effect → On screen: Logs, dumps, and temp files pile up to 100% → Writes fail. Crash if there’s no error handling, failed saves if there is → Disconnects, rolled-back progress - Symptoms: Disconnect, Dropped action / rollback / Factors: Stall - Who: Whole server / When: The longer it runs - Primary owner: Infra team (Server infrastructure) / Also: Infra team (DB infrastructure), Game team (Server development) - Game team action items: Handle write failures (retry the save and raise an alert so the server doesn’t crash), cut unnecessary logs and dumps. - Infra team action items: Servers/OS: log rotation, capacity alerts, separate log and data disks. DB hosts: watch that DB transaction logs (WAL, binlog) don’t pile up because replication stopped or a log backup was missed. - On the graph: Slow climb (Disk usage) - Where to look: Usage from df -h and inode usage from df -i, plus ENOSPC errors in server and DB logs. For the DB: slots whose active is false in PostgreSQL pg_replication_slots, file count and size from MySQL SHOW BINARY LOGS, log_reuse_wait_desc in SQL Server sys.databases, and FreeStorageSpace on RDS - Confirmed if: Usage climbs steadily over several days, the moment it hits 100% lines up with crashes or failed saves, and ENOSPC shows up in the logs - Ruled out if: Writes fail with plenty of space left: another cause such as permissions or a file size limit - Check with: Infra tools (no game code needed) - Learn more: A DB’s transaction logs (WAL, binlog, and so on) are never deleted and keep piling up if a replica stops or a log backup is missed. When that disk fills, every write on the DB stops, and saves and trades fail all at once. - Sources: - [write(2) — Linux manual page](https://man7.org/linux/man-pages/man2/write.2.html) · Linux man-pages · Writes fail with ENOSPC when the device has no space left - [Monitoring Disk Usage (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/diskusage.html) · PostgreSQL · If the WAL disk fills up, the DB server can panic and shut down - [Log-Shipping Standby Servers (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/warm-standby.html) · PostgreSQL · Replication slots keep WAL until the replica receives it, so they can fill pg_wal (limit with max_slot_wal_keep_size) - [Troubleshoot a full transaction log (SQL Server Error 9002)](https://learn.microsoft.com/en-us/sql/relational-databases/logs/troubleshoot-a-full-transaction-log-sql-server-error-9002) · Microsoft SQL Server · When the log is full, the DB is read-only and can’t be modified; missed log backups, replication lag, and long transactions are common causes that block log truncation; see what is blocking it in log_reuse_wait_desc of sys.databases - [df(1) — Linux manual page](https://man7.org/linux/man-pages/man1/df.1.html) · coreutils · Usage per file system; -i shows inode usage in place of blocks - [pg_replication_slots (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/view-pg-replication-slots.html) · PostgreSQL · active: whether the slot is currently streaming; wal_status: whether the WAL the slot retains has exceeded max_wal_size - [SHOW BINARY LOGS Statement](https://dev.mysql.com/doc/refman/8.4/en/show-binary-logs.html) · MySQL · List of the server’s binary log files and their sizes (File_size) - [Amazon CloudWatch metrics for Amazon RDS](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-metrics.html) · AWS · FreeStorageSpace: storage space left on the DB instance #### dk-backup · Backup / compression / scan jobs · Backup / compression / scans When early-morning backups, log compression, or security scans monopolize the disk, the game server’s reads and writes get held up. - Why → Effect → On screen: A scheduled backup or compression job starts → It takes most of the disk bandwidth and IOPS → Lag at the same time every day - Symptoms: Stutter, Input lag / Factors: Stall, Latency - Who: Whole server / When: At regular intervals - Primary owner: Infra team (Server infrastructure) / Also: Infra team (DB infrastructure) - Infra team action items: Servers/OS: lower the I/O priority of backup, compression, and scan jobs, stagger their times. DB hosts: take backups from a replica. - On the graph: Periodic spikes (Disk utilization, disk wait time) - Where to look: Overlay %util, await, and aqu-sz by day from the past few days of sar -d history (daily files in /var/log/sa; sadc must collect disk data with -S DISK), find the processes with the highest kB_rd/s and kB_wr/s at that time with pidstat -d 1, and match them against cron and systemd timer schedules - Confirmed if: await and %util spike at the same time every day, and backup, compression, or scan processes account for most disk reads and writes at that time - Ruled out if: Spikes at a different time each day: a scheduled job is unlikely. The game server itself does most of the I/O at that time: points to saves or logging (dk-fsync, dk-sync-log) - Check with: Infra tools (no game code needed) - Sources: - [ionice(1) — Linux manual page](https://man7.org/linux/man-pages/man1/ionice.1.html) · util-linux · A job run in the idle class gets disk time only when no other program is using the disk - [Using Replication for Backups](https://dev.mysql.com/doc/refman/8.4/en/replication-solutions-backups.html) · MySQL · Stopping a replica to take a backup doesn’t affect the primary - [sar(1) — Linux manual page](https://man7.org/linux/man-pages/man1/sar.1.html) · sysstat · -d: per-device await, aqu-sz, and %util from daily history files (default /var/log/sa); disk data must be collected with sadc’s -S DISK option - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · -d: per-process kB_rd/s and kB_wr/s (disk read and write volume per second) #### dk-lazy-load · Server-side lazy loading · Lazy loading on the server If the server reads dungeon or map data from disk the first time it’s requested, everyone freezes for that tick. - Why → Effect → On screen: Someone enters a dungeon or area for the first time → The server reads the data from disk on the game thread → Everyone on that server freezes briefly - Symptoms: Freeze / Factors: Stall - Who: Specific zone/channel, Whole server / When: While moving or changing zones - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Preload at server startup, load asynchronously. - Infra team action items: For servers just created from a snapshot, warm up the disk before putting them into service (read every block once) or use fast snapshot restore. - On the graph: Random spikes (Server tick time, disk reads) - Where to look: Match freeze times to first-entry records for dungeons and areas in the game server log, and at those moments check the game server’s disk reads (kB_rd/s from pidstat -d) and slow read and open calls with perf trace --duration. On a newly launched cloud server, compare EBS VolumeAvgReadLatency with older servers - Confirmed if: Freezes only on the first entry, with no freeze on the second entry to the same place. During the freeze, the game thread is waiting on a file read - Ruled out if: Freezes just the same in areas already loaded: another cause such as tick overrun or GC - Check with: Game server or client logs and metrics - Learn more: In the cloud, a server just created from a snapshot (a copy of a disk) fetches every block from remote storage the first time it reads it, so reads are far slower than usual. Suspect this if first entry takes unusually long only on servers newly launched by autoscaling. - Sources: - [Initialize Amazon EBS volumes](https://docs.aws.amazon.com/ebs/latest/userguide/ebs-initialize.html) · AWS · Volumes created from snapshots have higher latency and lower performance while blocks are fetched from S3; initialize them in advance by reading every block with dd or fio - [Amazon EBS fast snapshot restore](https://docs.aws.amazon.com/ebs/latest/userguide/ebs-fast-snapshot-restore.html) · AWS · Fast snapshot restore provides volumes that are fully initialized at creation, removing first-access latency - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · -d: per-process kB_rd/s (disk read volume per second) - [perf-trace(1) — Linux manual page](https://man7.org/linux/man-pages/man1/perf-trace.1.html) · perf · --duration shows only system calls that took longer than the given ms - [Amazon CloudWatch metrics for Amazon EBS](https://docs.aws.amazon.com/ebs/latest/userguide/using_cloudwatch_ebs.html) · AWS · VolumeAvgReadLatency: 1-minute average read latency (Nitro instances) #### dk-coredump · Writing a core dump · Core dump writing When the server crashes, writing several GB of memory to disk can delay the restart by several minutes. - Why → Effect → On screen: A server crash writes all of memory to a file → No restart until several GB have been written → After the server dies and players disconnect, they can’t connect again for a long time - Symptoms: Can’t connect / infinite loading / Factors: Stall - Who: Whole server / When: Randomly - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: Consider small dumps holding only the needed memory (minidumps), fix the cause of the crash. - Infra team action items: Limit dump size (OS core dump settings), use fast disks, decouple the restart from the dump (compress and upload dumps separately after the restart). - On the graph: Mass disconnect (Connection count, server restart times) - Where to look: Line up the crash time, core file size (coredumpctl list and info, or the file where core_pattern points), the time the file finished writing, and the time the service came back up, and check wkB/s from iostat -x during that window - Confirmed if: After the crash, disk writes stay near the limit while a multi-GB core file is written, and the restart begins only after the write finishes - Ruled out if: Restart is still slow with core dumps off or finished small: points to the server startup process, such as map loading or a DB cold cache (db-cold-cache) - Check with: Infra tools (no game code needed) - Sources: - [core(5) — Linux manual page](https://man7.org/linux/man-pages/man5/core.5.html) · Linux man-pages · RLIMIT_CORE caps core file size, coredump_filter selects which memory regions to include, and core dumps can be piped to a program for separate handling - [Minidump Files](https://learn.microsoft.com/en-us/windows/win32/debug/minidump-files) · Microsoft · A minidump holds only a useful subset of crash dump information, so it is fast and small - [coredumpctl(1) — Linux manual page](https://man7.org/linux/man-pages/man1/coredumpctl.1.html) · systemd · list: core dumps recorded in the journal (TIME is the crash time reported by the kernel); info: details for each dump and the size written to disk - [iostat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/iostat.1.html) · sysstat · -x: wkB/s (disk write volume per second) #### dk-hdd · HDD seek latency · HDD seek latency An HDD has to move its head across the platter (a seek), so reading or writing scattered data takes close to 10 ms each time. - Why → Effect → On screen: HDDs in old servers or low-cost storage → About 10 ms for every scattered read or write → Slow saves and loading across the board - Symptoms: Input lag / Factors: Latency - Who: Whole server / When: Always - Primary owner: Infra team (Server infrastructure) / Also: Infra team (DB infrastructure), Game team (Server development) - Game team action items: Design around sequential writes. - Infra team action items: Servers/OS: replace with SSDs (starting with storage disks that see the most scattered reads and writes). DB hosts: replace the DB disks with the most scattered reads and writes with SSDs first. - On the graph: Always high (Disk read/write latency (r_await, w_await)) - Where to look: Check whether it’s a rotating disk (HDD) with lsblk -d -o NAME,ROTA, and look at r/s, w/s, r_await, and w_await from iostat -x 1. For virtual servers, check the disk type in the cloud or storage specs - Confirmed if: A rotating disk, with r_await and w_await always at a few ms to tens of ms even at only tens to about a hundred requests per second - Ruled out if: High latency on an SSD: points to queue saturation (dk-iops) or burst credit depletion (dk-burst) - Check with: Infra tools (no game code needed) - Sources: - [Designs, Lessons and Advice from Building Large Distributed Systems (LADIS 2009 keynote)](https://www.cs.cornell.edu/projects/ladis2009/talks/dean-keynote-ladis2009.pdf) · Google · One disk seek 10 ms, 1 MB sequential read from disk 20 ms - [Exos X18 Data Sheet](https://www.seagate.com/www-content/datasheets/pdfs/exos-x18-mango-DS2045-1N-2007US-en_US.pdf) · Seagate · 7,200 rpm HDD average rotational latency 4.16 ms, 4K random reads 170 IOPS - [lsblk(8) — Linux manual page](https://man7.org/linux/man-pages/man8/lsblk.8.html) · util-linux · -o selects output columns; device topology columns include ROTA (whether the device is rotational) - [ABI stable symbols](https://docs.kernel.org/admin-guide/abi-stable.html) · Linux kernel · /sys/block/(disk)/queue/rotational: whether the device is rotational or non-rotational - [iostat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/iostat.1.html) · sysstat · -x: r/s and w/s, r_await and w_await (average time per request, including time spent waiting in the queue) ### L12 Database (causes: 16) #### db-no-index · Queries with no index · Missing index / full table scan Without an index, finding the rows that match a condition means reading the entire table (a full table scan). - Why → Effect → On screen: A new feature ships with a search on a condition that has no index → Scanning millions of rows makes a single query take hundreds of ms to several seconds → Mailbox and trade history load slowly, and tied-up connections make other requests wait too - Symptoms: Input lag, Can’t connect / infinite loading / Factors: Latency, Stall - Who: One feature only, Whole server / When: During specific actions - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Review the query plan of every new query before deploying, add indexes, check that modifying queries (UPDATE, DELETE) use an index too. - Infra team action items: Watch the slow query log, find full-table-scan queries and share them with the game team, add indexes in production with online methods that hold locks only briefly. - Ballpark numbers: With an index, a few ms. Without one, the query slows down in proportion to data size, which on a large table means hundreds to tens of thousands of times slower. - On the graph: Step change (DB query latency, rows read) - Where to look: MySQL: Rows_examined and Rows_sent in the slow query log (with log_queries_not_using_indexes on, queries that don’t use an index are logged too), SUM_NO_INDEX_USED and SUM_ROWS_EXAMINED in performance_schema events_statements_summary_by_digest, then run EXPLAIN. PostgreSQL: seq_scan and seq_tup_read in pg_stat_user_tables, then run EXPLAIN - Confirmed if: A query that appeared after the deploy reads thousands of times more rows (Rows_examined) than it returns (Rows_sent), and EXPLAIN shows a full table scan (MySQL type ALL, PostgreSQL Seq Scan). seq_tup_read on a large table climbs steeply from the deploy time - Ruled out if: Slow even while using an index: lock waits (db-hot-row, db-ddl-lock) or a query plan change (db-plan-flip). A full scan of a small table can be normal - Check with: Infra tools (no game code needed) - Learn more: Reads aren’t the only thing that slows down. A modifying query (UPDATE, DELETE) without an index can, depending on the DB, lock every row it scans and block saves for unrelated players. - Sources: - [How MySQL Uses Indexes](https://dev.mysql.com/doc/refman/8.4/en/mysql-indexes.html) · MySQL · Without an index, the whole table is read starting from the first row, and the bigger the table, the higher the cost - [Locks Set by Different SQL Statements in InnoDB](https://dev.mysql.com/doc/refman/8.4/en/innodb-locks-set.html) · MySQL · When no suitable index exists, a full table scan locks every row and blocks even other users’ inserts - [The Slow Query Log](https://dev.mysql.com/doc/refman/8.4/en/slow-query-log.html) · MySQL · Logs queries that exceed long_query_time (default 10 seconds); queries that don’t use indexes can also be logged separately - [CREATE INDEX (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/sql-createindex.html) · PostgreSQL · Building with CONCURRENTLY creates the index without blocking writes; a regular build blocks writes until it finishes - [Statement Summary Tables](https://dev.mysql.com/doc/refman/8.4/en/performance-schema-statement-summary-tables.html) · MySQL · events_statements_summary_by_digest: SUM_NO_INDEX_USED (times executed without an index) and SUM_ROWS_EXAMINED per normalized query - [EXPLAIN Output Format](https://dev.mysql.com/doc/refman/8.4/en/explain-output.html) · MySQL · type ALL means a full table scan, usually avoided by adding an index - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · seq_scan (number of sequential scans) and seq_tup_read (rows read by sequential scans) in pg_stat_user_tables - [Using EXPLAIN (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/using-explain.html) · PostgreSQL · Seq Scan: a query plan that reads every row of the table in order #### db-hot-row · Hot row lock contention · Hot row lock contention When everyone tries to modify the same row (a guild vault, a popular auction house item, a server-wide counter), only one request at a time gets the lock. - Why → Effect → On screen: An event or a popular item concentrates updates on the same row → Requests wait until they get the lock → Failed trades, “Please try again later” messages, timeouts - Symptoms: Dropped action / rollback, Input lag / Factors: Stall, Latency - Who: One feature only / When: When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Split the row (sharded counters), keep transactions short, aggregate in memory and apply in one write. - Infra team action items: Monitor row lock wait time and count, find the rows where contention concentrates, and share them. - Ballpark numbers: If one request holds the lock for 10 ms, that row can be modified at most 100 times a second. A round trip to another server inside the transaction cuts that number further. - On the graph: Rises with load (Row lock wait count and time) - Where to look: MySQL: growth of Innodb_row_lock_waits and Innodb_row_lock_time plus Innodb_row_lock_current_waits, and sys.innodb_lock_waits to find who is waiting on whom. PostgreSQL: sessions whose wait_event_type is Lock in pg_stat_activity and requests whose granted is false in pg_locks; turning on log_lock_waits (off by default) logs long lock waits - Confirmed if: Lock waits climb steeply with events and player count, and most waiting requests point to the same row (same key) in the same table - Ruled out if: Waits spread evenly across many tables and rows: points to disk or CPU saturation. One session holds a lock for a long time without releasing it: a long-open transaction (db-long-tx) - Check with: Infra tools (no game code needed) - Sources: - [InnoDB Locking](https://dev.mysql.com/doc/refman/8.4/en/innodb-locking.html) · MySQL · When one transaction locks a row (index record), other transactions can’t modify that row and wait - [How to Minimize and Handle Deadlocks](https://dev.mysql.com/doc/refman/8.4/en/innodb-deadlocks-handling.html) · MySQL · Recommends keeping transactions small and short and committing right after related changes to reduce conflicts - [Server Status Variables](https://dev.mysql.com/doc/refman/8.4/en/server-status-variables.html) · MySQL · Innodb_row_lock_waits and Innodb_row_lock_time give the count and total time of row lock waits; Innodb_row_lock_current_waits gives how many are waiting right now - [The innodb_lock_waits and x$innodb_lock_waits Views](https://dev.mysql.com/doc/refman/8.4/en/sys-innodb-lock-waits.html) · MySQL · The waiting query (waiting_query), the blocking session (blocking_pid), and the wait time (wait_age) - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · wait_event_type in pg_stat_activity: Lock means waiting for a heavyweight lock - [pg_locks (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/view-pg-locks.html) · PostgreSQL · granted false means that process is waiting to acquire the lock - [Error Reporting and Logging (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/runtime-config-logging.html) · PostgreSQL · log_lock_waits: logs a message when a lock wait lasts longer than deadlock_timeout; off by default #### db-deadlock · DB deadlock · Database deadlock When two transactions (groups of DB operations processed as one unit) each wait for a row the other has locked, the DB forcibly cancels one of them. - Why → Effect → On screen: Trade A locks in item→currency order, trade B in currency→item order → The DB detects the deadlock and rolls back one side → Trades and crafting fail now and then, items revert - Symptoms: Dropped action / rollback, Input lag / Factors: Packet loss, Stall - Who: One feature only / When: When crowds gather, During specific actions - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Use a consistent lock order, keep transactions short, retry automatically on failure. - Infra team action items: Keep deadlock detection on, collect and share deadlock records, lower the lock wait timeout (default 50 seconds) on MySQL servers with detection turned off. - Ballpark numbers: Detection is almost instant in MySQL (InnoDB), takes 1 second by default in PostgreSQL, and up to about 5 seconds in SQL Server. Both requests stay stuck in the meantime. On a MySQL server with detection turned off because of very high concurrency, requests wait until the lock wait timeout (default 50 seconds). - On the graph: Random spikes (Deadlock count, failed trade count) - Where to look: MySQL: LATEST DETECTED DEADLOCK in SHOW ENGINE INNODB STATUS (the most recent one only), every deadlock in the error log with innodb_print_all_deadlocks on, and lock_deadlocks in INFORMATION_SCHEMA.INNODB_METRICS. PostgreSQL: deadlocks in pg_stat_database. SQL Server: xml_deadlock_report from the system_health session, which is on by default. Error codes on the game server side: MySQL 1213, PostgreSQL 40P01, SQL Server 1205 - Confirmed if: Deadlock count rises at the times trades and crafting fail, and the two recorded transactions lock the same tables in opposite order - Ruled out if: Failures with an unchanged deadlock count: lock wait timeout exceeded (MySQL error 1205) or a hot row (db-hot-row) - Check with: Infra tools (no game code needed) - Sources: - [InnoDB Startup Options and System Variables](https://dev.mysql.com/doc/refman/8.4/en/innodb-parameters.html) · MySQL · With detection on (the default), InnoDB detects deadlocks immediately and rolls one back; innodb_lock_wait_timeout defaults to 50 seconds - [Deadlock Detection](https://dev.mysql.com/doc/refman/8.4/en/innodb-deadlock-detection.html) · MySQL · At very high concurrency, detection itself can slow things down, so it is sometimes turned off in favor of the lock wait timeout - [Lock Management (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/runtime-config-locks.html) · PostgreSQL · deadlock_timeout defaults to 1 second: the deadlock check runs only after waiting this long for a lock - [Deadlocks guide](https://learn.microsoft.com/en-us/sql/relational-databases/sql-server-deadlocks-guide) · Microsoft SQL Server · Deadlock checks run every 5 seconds by default, dropping to as little as 100 ms when deadlocks are frequent; the system_health session, on by default, collects xml_deadlock_report; the victim gets error 1205 - [How to Minimize and Handle Deadlocks](https://dev.mysql.com/doc/refman/8.4/en/innodb-deadlocks-handling.html) · MySQL · Always modify multiple rows and tables in the same order, retry on failure, log every deadlock with innodb_print_all_deadlocks - [InnoDB Standard Monitor and Lock Monitor Output](https://dev.mysql.com/doc/refman/8.4/en/innodb-standard-monitor.html) · MySQL · LATEST DETECTED DEADLOCK: the two transactions in the most recent deadlock, the locks they held and waited for, and which one was rolled back - [InnoDB INFORMATION_SCHEMA Metrics Table](https://dev.mysql.com/doc/refman/8.4/en/innodb-information-schema-metrics-table.html) · MySQL · lock_deadlocks counter in INNODB_METRICS (enabled by default) - [Server Error Message Reference](https://dev.mysql.com/doc/mysql-errors/8.4/en/server-error-reference.html) · MySQL · 1213 ER_LOCK_DEADLOCK (deadlock), 1205 ER_LOCK_WAIT_TIMEOUT (lock wait timeout exceeded) - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · deadlocks in pg_stat_database: number of deadlocks detected in this database - [PostgreSQL Error Codes (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/errcodes-appendix.html) · PostgreSQL · 40P01 deadlock_detected #### db-pool · Connection pool exhaustion · Connection pool exhaustion The number of connections open to the DB is fixed, so when slow queries hold connections, every other request waits. - Why → Effect → On screen: Slow queries or a flood of requests put every connection in use → New requests wait until a connection frees up → Infinite loading at login, slow saves, timeouts - Symptoms: Can’t connect / infinite loading, Input lag / Factors: Stall, Latency - Who: Whole server, One feature only / When: Right after login or maintenance, When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Remove slow queries, tune pool size and wait timeout (don’t just keep growing the pool), separate pools per feature. - Infra team action items: Check the DB’s max connections and CPU/IOPS headroom, confirm that server count × pool size stays within max connections before adding servers or autoscaling, add connection wait and lock wait metrics to monitoring. - Ballpark numbers: Estimate the connections you need as “requests per second × time each request holds a connection”. At 2,000 requests a second and 5 ms each, 10 connections are busy on average at all times. To handle bursts, pools are usually sized at two to three times that. If queries slow to 150 ms, the same load needs 300. - On the graph: Hits a ceiling (DB connections in use, connection wait time) - Where to look: Count connection states per game server on the DB side. MySQL: Host, Command (Sleep for idle connections), and Time in SHOW PROCESSLIST, plus Threads_connected, Threads_running, and refused connections in Connection_errors_max_connections. PostgreSQL: pg_stat_activity grouped by client_addr and state. If the game server’s connection pool library exports wait count and wait time, view those too - Confirmed if: All of one game server’s connections, up to the pool size, are running queries with 0 idle, and logins and saves wait meanwhile. Or the DB’s total connection count hits max_connections and new connections are refused - Ruled out if: Slow even with plenty of idle connections: latency of the queries themselves (db-no-index, db-hot-row) or saturated DB resources - Check with: Infra tools (no game code needed) - Learn more: Blindly growing the pool only adds DB CPU load and lock contention, and everyone slows down together. Also, if server count × pool size exceeds the DB’s max connections, newly added or restarted servers can’t even open a connection. This is common right after autoscaling or maintenance. - Sources: - [Number Of Database Connections](https://wiki.postgresql.org/wiki/Number_Of_Database_Connections) · PostgreSQL · Once DB resources are fully used, adding connections actually lowers throughput; matching active connections to resources and queueing the rest gives better latency and throughput - [Too many connections](https://dev.mysql.com/doc/refman/8.4/en/too-many-connections.html) · MySQL · When all max_connections are in use, new connections are refused with a Too many connections error - [Connections and Authentication (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/runtime-config-connection.html) · PostgreSQL · max_connections: cap on concurrent connections, usually 100 by default - [SHOW PROCESSLIST Statement](https://dev.mysql.com/doc/refman/8.4/en/show-processlist.html) · MySQL · Host (client address), Command (Sleep for idle sessions), Time, State - [Server Status Variables](https://dev.mysql.com/doc/refman/8.4/en/server-status-variables.html) · MySQL · Threads_connected, Threads_running, Connection_errors_max_connections (connections refused because max_connections was reached) - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · pg_stat_activity: client_addr and state (active, idle, idle in transaction, and so on) for each connection #### db-replica-lag · Replication lag · Replication lag Writes go to the primary and reads come from replicas, so when a replica falls behind, data that was just written isn’t visible yet. - Why → Effect → On screen: A burst of writes on the primary puts replicas several seconds behind → Reading just-saved data from a replica finds it missing → An item you just bought doesn’t show up, marketplace prices are stale, duplicate-reward bugs - Symptoms: Dropped action / rollback / Factors: Latency - Who: One feature only / When: When crowds gather, Evening peak hours - Primary owner: Infra team (DB infrastructure) / Also: Game team (Server development) - Game team action items: Read just-written data from the primary, check and grant rewards in one transaction on the primary (block duplicates with a unique key or a conditional UPDATE). - Infra team action items: Alert on replication lag, give replicas specs equal to or better than the primary and enable parallel replication, run bulk deletes in small chunks, manage long-running aggregate queries on replicas. - On the graph: Rises with load (Replication lag (seconds)) - Where to look: MySQL: Seconds_Behind_Source from SHOW REPLICA STATUS on the replica (SHOW SLAVE STATUS on versions before 8.0.22). PostgreSQL: write_lag, flush_lag, and replay_lag in pg_stat_replication on the primary. RDS: ReplicaLag - Confirmed if: Lag is several seconds or more at the times of “it’s not showing up” reports, and things look normal once the lag clears. Lag grows during write bursts, bulk deletes, or long aggregate queries on the replica - Ruled out if: Lag near 0 but data still not showing up: points to the game server’s cache or sync - Check with: Infra tools (no game code needed) - Learn more: Replicas can fall behind even without heavy writes. A single bulk delete that took 10 minutes on the primary puts a replica that far behind while it replays, and long-running aggregate queries on a replica also slow its catch-up. - Sources: - [SHOW REPLICA STATUS Statement](https://dev.mysql.com/doc/refman/8.4/en/show-replica-status.html) · MySQL · Seconds_Behind_Source: time elapsed since the event the replica is currently applying was written on the primary (replication lag) - [Replica Server Options and Variables](https://dev.mysql.com/doc/refman/8.4/en/replication-options-replica.html) · MySQL · replica_parallel_workers lets multiple threads apply transactions in parallel (default 4; 0 means a single thread applies them in order) - [Log-Shipping Standby Servers (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/warm-standby.html) · PostgreSQL · Streaming replication is asynchronous by default, so there is a delay between commit and the replica applying it (usually under 1 second if the replica can keep up) - [MySQL 8.0 Reference Manual: SHOW REPLICA STATUS Statement](https://dev.mysql.com/doc/refman/8.0/en/show-replica-status.html) · MySQL · From 8.0.22, SHOW REPLICA STATUS replaces SHOW SLAVE STATUS; earlier versions use SHOW SLAVE STATUS - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · write_lag, flush_lag, and replay_lag in pg_stat_replication: time from the primary writing WAL until the replica reports it has written, flushed to disk, and applied it - [Amazon CloudWatch metrics for Amazon RDS](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-metrics.html) · AWS · ReplicaLag: how far a read replica lags behind its source (seconds) #### db-checkpoint · Checkpoint / log flush · Checkpoint / log flush stalls Queries slow down at the moments the DB periodically writes its accumulated in-memory changes to disk in bulk. - Why → Effect → On screen: Changes pile up and are periodically written to disk → The disk gets busy at that moment and queries slow down → Saves and loading slow down periodically - Symptoms: Input lag, Stutter / Factors: Latency - Who: Whole server, One feature only / When: At regular intervals - Primary owner: Infra team (DB infrastructure) - Infra team action items: Spread checkpoints out in small, even steps, size the transaction log (redo log, WAL) generously, use fast disks. - On the graph: Periodic spikes (DB query latency, disk write volume) - Where to look: PostgreSQL: checkpoint times and buffers written from the log_checkpoints log (on by default in recent versions), checkpoint counts (num_timed and num_requested in pg_stat_checkpointer on 17 and later, checkpoints_timed and checkpoints_req in pg_stat_bgwriter on 16 and earlier), and checkpoint_warning messages. MySQL: the gap between Log sequence number and Last checkpoint at in the LOG section of SHOW ENGINE INNODB STATUS. Overlay the server’s disk write volume and write latency - Confirmed if: Query latency spikes line up with checkpoint times, with disk write volume and write latency spiking at the same moments. In PostgreSQL, far more requested checkpoints (num_requested) than timed ones (num_timed) means WAL keeps hitting max_wal_size and checkpoints come early - Ruled out if: Spikes on a cycle unrelated to checkpoint times: backups or batch jobs (dk-backup, db-batch) - Check with: Infra tools (no game code needed) - Learn more: If the transaction log that records changes (the redo log in MySQL, WAL in PostgreSQL) is too small, the DB has to rush a checkpoint every time the log fills, and write throughput drops sharply for short periods. - Sources: - [WAL Configuration (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/wal-configuration.html) · PostgreSQL · By default, a checkpoint runs every 5 minutes or every 1 GB of WAL (max_wal_size) and is expensive because it writes all dirty pages. checkpoint_completion_target spreads the writes out to avoid I/O bursts. If checkpoints come closer together than checkpoint_warning, the log suggests raising max_wal_size - [Configuring Buffer Pool Flushing](https://dev.mysql.com/doc/refman/8.4/en/innodb-buffer-pool-flushing.html) · MySQL · When the redo log fills, a sharp checkpoint briefly drops throughput; adaptive flushing spreads the writes out evenly - [Error Reporting and Logging (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/runtime-config-logging.html) · PostgreSQL · log_checkpoints: logs the buffers written and time taken for each checkpoint; on by default - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · num_timed (checkpoints run because their scheduled time came) and num_requested (requested checkpoints) in pg_stat_checkpointer - [PostgreSQL 17 Release Notes](https://www.postgresql.org/docs/release/17.0/) · PostgreSQL · pg_stat_checkpointer added; checkpoint-related columns moved out of pg_stat_bgwriter - [The Cumulative Statistics System (PostgreSQL 16 Documentation)](https://www.postgresql.org/docs/16/monitoring-stats.html) · PostgreSQL · Up to 16: checkpoints_timed and checkpoints_req in pg_stat_bgwriter - [InnoDB Standard Monitor and Lock Monitor Output](https://dev.mysql.com/doc/refman/8.4/en/innodb-standard-monitor.html) · MySQL · LOG section: current log sequence number and last checkpoint position #### db-cold-cache · Cold cache (right after a restart) · Cold buffer pool after restart After a DB restart, the memory cache is empty, so for a while every lookup reads from disk. - Why → Effect → On screen: The DB restarts for maintenance → Frequently used data isn’t in memory, so it’s read from disk → Logins and loading are slow for a while right after maintenance - Symptoms: Can’t connect / infinite loading, Input lag / Factors: Latency - Who: Whole server / When: Right after login or maintenance - Primary owner: Infra team (DB infrastructure) / Also: Game team (Server development) - Game team action items: Open gradually (use a login queue to ramp up the number of players logging in step by step). - Infra team action items: Warm the cache after a restart (check the buffer pool save/restore settings), pre-read the disk too on a DB restored from a snapshot. - On the graph: Surge after opening (Disk reads, buffer cache hit ratio) - Where to look: MySQL: the ratio of Innodb_buffer_pool_reads (reads that missed the buffer pool and went to disk) to Innodb_buffer_pool_read_requests, and warm-up progress in Innodb_buffer_pool_load_status. PostgreSQL: blks_read and blks_hit in pg_stat_database. Also check disk reads on the DB server - Confirmed if: Right after the restart, disk reads spike and the hit ratio is low, recovering over time, and logins and loading are slow during that window - Ruled out if: Hit ratio normal but slow right after maintenance: a login storm and N+1 queries (db-login-storm) or the connection pool (db-pool) - Check with: Infra tools (no game code needed) - Learn more: MySQL saves the list of buffer pool pages at shutdown and reloads them in the background at startup, but filling the pool takes time. If a cloud DB was restored from a snapshot (a copy of a disk), the disk itself is also slow for every block read for the first time, so the slowdown lasts even longer. - Real incidents: roblox-2021 - Sources: - [Saving and Restoring the Buffer Pool State](https://dev.mysql.com/doc/refman/8.4/en/innodb-preload-buffer-pool.html) · MySQL · Saves the list of recently used pages (25% by default) at shutdown and reads them back at startup to shorten warm-up after a restart; both are on by default - [pg_prewarm — preload relation data into buffer caches](https://www.postgresql.org/docs/current/pgprewarm.html) · PostgreSQL · Periodically records shared buffer contents and loads them back after a restart (autoprewarm) - [Initialize Amazon EBS volumes](https://docs.aws.amazon.com/ebs/latest/userguide/ebs-initialize.html) · AWS · Volumes created from snapshots have higher latency and lower performance until all blocks have been fetched - [Server Status Variables](https://dev.mysql.com/doc/refman/8.4/en/server-status-variables.html) · MySQL · Innodb_buffer_pool_reads (logical reads that missed the buffer pool and read straight from disk), Innodb_buffer_pool_read_requests, Innodb_buffer_pool_load_status (warm-up progress) - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · blks_read (blocks read from disk) and blks_hit (blocks found in the buffer cache) in pg_stat_database #### db-login-storm · Login storm and N+1 queries · Login storm, N+1 queries If loading one character takes dozens of separate queries, tens of thousands of simultaneous logins turn into millions of queries. - Why → Effect → On screen: Character loading queries items, skills, and quests one by one → Simultaneous logins right after maintenance make the query count explode → Infinite loading at login, and even saves for players already in the game get held up - Symptoms: Can’t connect / infinite loading, Input lag / Factors: Stall, Latency - Who: Whole server / When: Right after login or maintenance - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Fetch in batched queries, use a login queue and caching, check how many queries ORM lazy loading generates. - Infra team action items: Pull a ranking of the most frequently called queries and share it, monitor query and connection counts during the login window right after maintenance. - On the graph: Surge after opening (DB queries per second, login count) - Where to look: Overlay the login count right after maintenance with the DB’s queries per second (growth of Questions in MySQL) and compute queries per login. Pull the most frequently called queries from COUNT_STAR in MySQL events_statements_summary_by_digest or calls in PostgreSQL pg_stat_statements - Confirmed if: Dozens of queries per login, and the top queries are short queries of the same shape that look up by a single character ID. If queries per login went up after a patch, that patch is the starting point - Ruled out if: Few queries per login but each one is slow: cold cache (db-cold-cache) or indexes (db-no-index) - Check with: Infra tools (no game code needed) - Learn more: Lazy loading in an ORM (a library that builds DB queries for you) generates queries like these without developers even noticing. On a dev server with only a few characters it goes unnoticed, and it first shows up with simultaneous logins on live servers. - Sources: - [Efficient Querying](https://learn.microsoft.com/en-us/ef/core/performance/efficient-querying) · .NET · ORM lazy loading creates the N+1 problem, sending one more query per item and badly hurting performance; recommends loading in one batch (eager loading) - [pg_stat_statements — track statistics of SQL planning and execution](https://www.postgresql.org/docs/current/pgstatstatements.html) · PostgreSQL · Collects execution count (calls) and total execution time per statement to rank the most frequently called queries - [Performance Schema Statement Digests and Sampling](https://dev.mysql.com/doc/refman/8.4/en/performance-schema-statement-digests.html) · MySQL · events_statements_summary_by_digest groups queries of the same shape and aggregates counts and times - [Statement Summary Tables](https://dev.mysql.com/doc/refman/8.4/en/performance-schema-statement-summary-tables.html) · MySQL · COUNT_STAR (execution count) and SUM_TIMER_WAIT (total time) in summary tables - [Server Status Variables](https://dev.mysql.com/doc/refman/8.4/en/server-status-variables.html) · MySQL · Questions: number of statements sent by clients #### db-batch · Bulk batch jobs · Batch jobs during service Running ranking aggregation, mass mail sends, or old-data cleanup during live service ties up locks and the disk. - Why → Effect → On screen: Bulk jobs run during service hours → Wide-range locks, disk and CPU tied up → Failed trades and saves, slow loading at certain times of day - Symptoms: Input lag, Dropped action / rollback / Factors: Latency, Stall - Who: One feature only, Whole server / When: At regular intervals - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Split jobs into small chunks and run them a little at a time, run aggregation on a replica. - Infra team action items: Provide a replica for aggregation, schedule batches for quiet hours, watch for lock escalation and gap lock waits. - On the graph: Periodic spikes (DB query latency, lock waits) - Where to look: Find long queries running at the time of the lag. MySQL: slow query log. PostgreSQL: query_start and query in pg_stat_activity. Match them against lock wait metrics at the same time and the batch schedule (cron, DB event scheduler). SQL Server: record lock escalation with the lock_escalation extended event - Confirmed if: Large UPDATE, DELETE, or aggregation queries run at the same time every time, and lock waits and disk utilization rise together meanwhile - Ruled out if: No long queries at that time: checkpoints (db-checkpoint) or a server backup (dk-backup) - Check with: Infra tools (no game code needed) - Learn more: SQL Server converts to a table lock when one statement holds more than about 5,000 row locks (lock escalation). At that moment, every request using the same table stops. MySQL, with default settings, also locks the gaps between rows when modifying by a range condition (gap locks), blocking inserts of new rows. - Sources: - [Transaction Locking and Row Versioning Guide](https://learn.microsoft.com/en-us/sql/relational-databases/sql-server-transaction-locking-and-row-versioning-guide) · Microsoft SQL Server · Lock escalation when one statement holds 5,000 or more locks on one table (or index); recorded with the lock_escalation extended event - [InnoDB Locking](https://dev.mysql.com/doc/refman/8.4/en/innodb-locking.html) · MySQL · At InnoDB’s default isolation level, REPEATABLE READ, searches and scans use next-key locks, so gap locks block inserts of new rows into those gaps - [The Slow Query Log](https://dev.mysql.com/doc/refman/8.4/en/slow-query-log.html) · MySQL · Logs queries exceeding long_query_time with execution time (Query_time), lock time (Lock_time), and rows read - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · pg_stat_activity: the query each session is running now (query) and when it started (query_start) #### db-failover · DB failover · Database failover When the primary DB dies, writes stop while it fails over to a standby, and the last data that hadn’t been replicated yet can be lost. - Why → Effect → On screen: The primary DB fails and a standby is promoted → No writes for seconds to minutes during the switch; with asynchronous replication, unreplicated data may be lost → Every save fails for a moment, items and XP roll back - Symptoms: Dropped action / rollback, Freeze, Disconnect, Can’t connect / infinite loading / Factors: Stall, Packet loss - Who: Whole server / When: Randomly - Primary owner: Infra team (DB infrastructure) / Also: Game team (Server development) - Game team action items: Make saves retryable, configure the connection pool and DNS cache to drop broken connections quickly and reconnect to the new address, verify reconnection during failover drills. - Infra team action items: Use synchronous or semi-synchronous replication (at the cost of write latency), run failover drills, monitor failover time and replication lag. - Ballpark numbers: Automatic failover on a managed DB usually takes tens of seconds to 2 minutes. With asynchronous replication, you can lose recent saves equal to the replication lag (under 1 second to a few seconds). - On the graph: Mass disconnect (DB connection count, write error count) - Where to look: Put the DB’s failover records (on RDS, events RDS-EVENT-0013 failover started and RDS-EVENT-0049 failover completed; on self-managed DBs, promotion logs) and the game server’s DB connection count and connection error count on one graph. With asynchronous replication, also check replication lag just before the failure (RDS ReplicaLag, replay_lag in PostgreSQL pg_stat_replication) - Confirmed if: Save failures cluster in one window that lines up with the span between failover start and completion. The amount rolled back is close to the replication lag just before the failure. A game server whose errors continue after the failover finishes is still using connections opened to the old address - Ruled out if: Disconnects at times with no failover record: points to the network or DB overload - Check with: Infra tools (no game code needed) - Real incidents: riot-euw-2021 - Sources: - [Failing over a Multi-AZ DB instance for Amazon RDS](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Concepts.MultiAZ.Failover.html) · AWS · Multi-AZ failover usually takes 60–120 seconds; connections must be reestablished afterward, and a JVM DNS cache TTL of 60 seconds or less is recommended - [High availability for Amazon Aurora](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/Concepts.AuroraHighAvailability.html) · AWS · Reads and writes fail during the outage, and recovery usually takes under 60 seconds (often under 30) - [Semisynchronous Replication](https://dev.mysql.com/doc/refman/8.4/en/replication-semisync.html) · MySQL · With asynchronous replication, committed transactions may be missing from the replica when the primary dies; semi-synchronous replication narrows this by waiting for one replica to acknowledge receipt, at the cost of higher latency - [Log-Shipping Standby Servers (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/warm-standby.html) · PostgreSQL · Log shipping is asynchronous, so transactions not yet sent are lost if the primary dies; streaming replication lag is usually under 1 second - [Amazon RDS event categories and event messages](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_Events.Messages.html) · AWS · RDS-EVENT-0013: Multi-AZ failover started, RDS-EVENT-0049: Multi-AZ failover completed - [Amazon CloudWatch metrics for Amazon RDS](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/rds-metrics.html) · AWS · ReplicaLag: how far a read replica lags behind its source (seconds) - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · replay_lag in pg_stat_replication: time from the primary writing WAL until the replica reports having applied it #### db-save-interval · Lost progress from a long save interval · Periodic save window If the server saves only once every few minutes to reduce load, progress is lost when the server dies in between. - Why → Effect → On screen: Character state is saved once every few minutes → A server crash or outage hits in between → After reconnecting, the character is back to where it was minutes ago (rollback) - Symptoms: Dropped action / rollback / Factors: Packet loss - Who: Whole server, Specific zone/channel / When: Randomly - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Save important events (trades, rare drops) immediately, keep a change log. - Infra team action items: Confirm the DB has enough IOPS and CPU headroom for the extra writes a shorter save interval brings. - On the graph: Mass disconnect (Connection count, rollback reports) - Where to look: Line up crash and outage times with the last save time of the characters that reported rollbacks (the game server’s save log or the DB’s last-modified column) - Confirmed if: The point the character reverted to matches the last save before the crash, and the lost time is shorter than the save interval - Ruled out if: Reverted even though the game server log says the save completed: data loss from a DB failover (db-failover) or a stale value read from a replica (db-replica-lag) - Check with: Game server or client logs and metrics - Sources: - [Asynchronous Commit (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/wal-async-commit.html) · PostgreSQL · Batching writes and flushing them late raises throughput, but the most recent transactions can be lost in a failure (the same tradeoff) - [Redis persistence](https://redis.io/docs/latest/operate/oss_and_stack/management/persistence/) · Redis · Taking RDB snapshots every few minutes means accepting the loss of the last few minutes of data on an abnormal shutdown #### db-cache-stampede · Cache stampede · Cache stampede / thundering herd When cache entries for popular data expire at the same time, thousands of requests hit the DB all at once. - Why → Effect → On screen: Popular data stored in Redis or a similar cache expires at the same time → Requests trying to rebuild the same data rush to the DB all at once → DB overload makes one feature after another slow down or freeze - Symptoms: Input lag, Freeze, Can’t connect / infinite loading / Factors: Stall, Latency - Who: Whole server / When: At regular intervals, When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Randomize expiry times, let a single request refresh the value while the rest keep using the old value. - Infra team action items: Set up Redis replicas and automatic failover so the cache doesn’t empty entirely on a restart or failure, confirm the DB has enough headroom to survive an empty cache. - On the graph: Periodic spikes (Cache hit ratio, DB queries per second) - Where to look: Overlay keyspace_hits and keyspace_misses (hit ratio), expired_keys, and restarts (uptime_in_seconds) from Redis INFO on the DB’s queries per second, and count how many copies of the same query run at once on the DB at that moment (MySQL SHOW PROCESSLIST, PostgreSQL pg_stat_activity) - Confirmed if: At the moment cache misses shoot up, the DB query count spikes with them, and most concurrent queries are the same query reading the same data. Lines up with the expiry cycle of popular keys or a Redis restart - Ruled out if: Cache misses normal but only DB queries rise: a login storm (db-login-storm) or a batch job (db-batch) - Check with: Infra tools (no game code needed) - Learn more: The same thing happens when a Redis restart or failure empties the whole cache. The more an architecture relies on the cache and keeps the DB small, the greater the risk. - Sources: - [Scaling Memcache at Facebook (NSDI '13)](https://www.usenix.org/conference/nsdi13/technical-sessions/presentation/nishtala) · USENIX · When a hot key is invalidated, many reads rush to the DB (thundering herd); prevented with leases (only one client refreshes) and by returning stale values; clusters with an empty cache are warmed up separately - [Optimal Probabilistic Cache Stampede Prevention](https://www.vldb.org/pvldb/vol8/p886-vattani.pdf) · VLDB Endowment · When a popular item expires, many requests regenerate it at once (cache stampede); prevented by probabilistically refreshing early, before expiry - [High availability with Redis Sentinel](https://redis.io/docs/latest/operate/oss_and_stack/management/sentinel/) · Redis · Automatic failover that promotes a replica when the primary dies - [INFO](https://redis.io/docs/latest/commands/info/) · Redis · keyspace_hits and keyspace_misses (successful and failed key lookups), expired_keys (number of expired keys), uptime_in_seconds (time since startup) - [SHOW PROCESSLIST Statement](https://dev.mysql.com/doc/refman/8.4/en/show-processlist.html) · MySQL · The statement each session is running (Info) and the time spent in its current state (Time, seconds) - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · pg_stat_activity: the query each session is running now (query) #### db-long-tx · Transaction left open too long · Long-running transaction / MVCC purge lag When a transaction stays open for a long time, it keeps holding its locks and the DB can’t clean up (purge) old versions of data, so everything gradually slows down. - Why → Effect → On screen: A transaction stays open while waiting for another server’s response, or a long aggregate query runs on the primary during service → Its locks are never released, and old row versions awaiting cleanup keep piling up → Features that use those rows time out, and saves and lookups slow down across the board over several hours - Symptoms: Input lag, Dropped action / rollback / Factors: Latency, Stall - Who: One feature only, Whole server / When: The longer it runs, Randomly - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Don’t wait on network calls or user input inside a transaction, run aggregate queries on a replica. - Infra team action items: Alert on and kill long-open transactions, provide a replica for aggregation, watch undo log and dead row growth. - On the graph: Slow climb (Undo log length (History list length), dead row count) - Where to look: MySQL: find the oldest transaction by trx_started in INFORMATION_SCHEMA.INNODB_TRX, and check History list length (undo log not yet purged) in the TRANSACTIONS section of SHOW ENGINE INNODB STATUS. PostgreSQL: xact_start in pg_stat_activity, sessions whose state is idle in transaction, and n_dead_tup in pg_stat_user_tables - Confirmed if: A transaction minutes to hours old exists, History list length or n_dead_tup keeps climbing while it’s open, then falls as cleanup (purge, VACUUM) runs after that transaction ends - Ruled out if: Slow across the board with no old transactions: checkpoints (db-checkpoint) or the disk - Check with: Infra tools (no game code needed) - Learn more: The DB keeps old versions so readers can see data as it was before a change (MVCC). These records can only be removed once the oldest transaction ends, so if one transaction stays open for hours, MySQL piles up undo logs and PostgreSQL piles up dead rows (dead tuples) that VACUUM can’t clean up. In SQL Server, the transaction log can’t shrink and may fill the disk. - Sources: - [InnoDB Multi-Versioning](https://dev.mysql.com/doc/refman/8.4/en/innodb-multi-versioning.html) · MySQL · While a transaction that can see old versions remains, update undo logs can’t be discarded and the rollback segment grows; recommends committing often, even for read-only transactions - [Routine Vacuuming (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/routine-vacuuming.html) · PostgreSQL · Old row versions can’t be removed while other transactions can still see them; long-open transactions must be ended or their sessions terminated - [Client Connection Defaults (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/runtime-config-client.html) · PostgreSQL · idle_in_transaction_session_timeout: terminates sessions sitting idle with a transaction open so they can’t hold locks for long - [Troubleshoot a full transaction log (SQL Server Error 9002)](https://learn.microsoft.com/en-us/sql/relational-databases/logs/troubleshoot-a-full-transaction-log-sql-server-error-9002) · Microsoft SQL Server · A long-running active transaction blocks transaction log truncation - [The INFORMATION_SCHEMA INNODB_TRX Table](https://dev.mysql.com/doc/refman/8.4/en/information-schema-innodb-trx-table.html) · MySQL · TRX_STARTED: transaction start time - [Purge Configuration](https://dev.mysql.com/doc/refman/8.4/en/innodb-purge-configuration.html) · MySQL · Purge cleans up the list of undo logs from committed transactions (history list); the backlog shows as History list length in the TRANSACTIONS section of SHOW ENGINE INNODB STATUS - [The Cumulative Statistics System (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/monitoring-stats.html) · PostgreSQL · xact_start (transaction start time) and state (idle in transaction) in pg_stat_activity, n_dead_tup (estimated dead rows) in pg_stat_user_tables #### db-redis-block · Slow Redis commands · Redis blocking commands (single-threaded) Redis processes commands one at a time, so a single slow command blocks every request behind it. - Why → Effect → On screen: A full key search with KEYS in production, or reading or deleting a ranking or list with millions of elements in one go → Every other request waits until that command finishes (tens of ms to several seconds) → Features that use sessions, rankings, or the cache all hitch at once, slow logins - Symptoms: Freeze, Input lag, Can’t connect / infinite loading / Factors: Stall, Latency - Who: Whole server, One feature only / When: Randomly, At regular intervals - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Replace KEYS with SCAN, split large keys, delete with UNLINK (background deletion), spread out expiry times that cluster in the same second. - Infra team action items: Watch the slow command log (SLOWLOG), block dangerous commands such as KEYS on production servers, check for large keys regularly, turn off THP and keep enough spare memory for fork, run RDB and AOF persistence on replicas. - Ballpark numbers: A typical command takes under 1 ms. Handling millions of elements at once can take anywhere from hundreds of ms to several seconds. - On the graph: Random spikes (Redis response latency, slow command count) - Where to look: SLOWLOG GET for commands over slowlog-log-slower-than; turn on the latency monitor (off by default) with CONFIG SET latency-monitor-threshold, then check per-event latency such as fork and expire-cycle with LATENCY LATEST and LATENCY DOCTOR. Also check fork time and large keys with latest_fork_usec in INFO and redis-cli --bigkeys - Confirmed if: At the time of the stall, SLOWLOG shows KEYS or commands handling a whole large key, or LATENCY records fork or expire-cycle events of tens of ms or more at the same time - Ruled out if: SLOWLOG and LATENCY are empty but it’s slow only as seen from the game server: the network or waiting inside the game server (SLOWLOG measures only command execution time, excluding time spent talking to the client) - Check with: Infra tools (no game code needed) - Learn more: Redis also stalls the moment it forks the process to create a save file (RDB snapshot) or rewrite the AOF. On modern servers this takes about 10 ms per GB of memory, so about 300 ms for 30 GB. With transparent huge pages (THP) turned on, every write after a fork copies an entire huge page (copy-on-write), sharply increasing stalls and memory use, so THP is usually turned off and plenty of spare memory is kept. Redis also stalls briefly to delete keys when a very large number of them expire in the same second. - Sources: - [Diagnosing latency issues](https://redis.io/docs/latest/operate/oss_and_stack/management/optimization/latency/) · Redis · One thread processes requests in turn, so a slow command blocks everything behind it; use SCAN in place of KEYS; fork measured at about 9–13 ms per GB on physical servers and modern VMs; THP causes latency and memory spikes from copying after fork; mass expiry in the same second causes stalls - [KEYS](https://redis.io/docs/latest/commands/keys/) · Redis · Use with extreme care in production; can ruin performance on large databases (40 ms for 1 million keys on an entry-level laptop) - [UNLINK](https://redis.io/docs/latest/commands/unlink/) · Redis · Asynchronous deletion: unlinks the key immediately and reclaims its memory in another thread - [SLOWLOG](https://redis.io/docs/latest/commands/slowlog/) · Redis · Slow command log that records commands exceeding slowlog-log-slower-than; execution time excludes I/O with the client - [Redis latency monitoring](https://redis.io/docs/latest/operate/oss_and_stack/management/optimization/latency-monitor/) · Redis · latency-monitor-threshold defaults to 0 (off); LATENCY LATEST and LATENCY DOCTOR; records latency per event such as fork and expire-cycle - [INFO](https://redis.io/docs/latest/commands/info/) · Redis · latest_fork_usec: time taken by the last fork (microseconds) - [Redis CLI](https://redis.io/docs/latest/develop/tools/cli/) · Redis · --bigkeys: scans the keyspace to find large keys #### db-plan-flip · Query slowdown from a query plan change · Query plan regression (stats, parameter sniffing) Even with the code unchanged, if the DB changes how it executes a query (its query plan), a query that took 2 ms yesterday takes hundreds of ms today. - Why → Effect → On screen: Automatic statistics updates, a DB restart, or shifts in data distribution make the DB build a new query plan → A plan that skips the index gets picked, the same query becomes tens to hundreds of times slower, and connections get tied up → With no deploy at all, loading for a specific feature suddenly slows down and other requests wait too - Symptoms: Input lag, Can’t connect / infinite loading / Factors: Latency, Stall - Who: One feature only, Whole server / When: Randomly - Primary owner: Infra team (DB infrastructure) / Also: Game team (Server development) - Game team action items: For queries whose result count varies widely by parameter value, split them or consider plan hints, design queries that reliably use an index. - Infra team action items: Watch slow queries and query plan history, pin good plans (Query Store in SQL Server, and so on), manage when statistics get updated. - On the graph: Step change (Average execution time per query) - Where to look: Collect the average time per normalized query periodically and watch the trend. MySQL: AVG_TIMER_WAIT in events_statements_summary_by_digest. PostgreSQL: mean_exec_time in pg_stat_statements (mean_time on 12 and earlier). Compare query plans from before and after the slowdown with EXPLAIN or PostgreSQL auto_explain, and in SQL Server with the Regressed Queries view in Query Store - Confirmed if: With no deploy at the time, one query’s average time steps up tens of times, the moment lines up with a statistics update or a DB restart, and the query plan has changed - Ruled out if: Slower with the query plan unchanged: data growth, lock waits (db-hot-row), or the disk - Check with: Infra tools (no game code needed) - Learn more: SQL Server reuses a plan built for the first value it sees (parameter sniffing). A plan built for a new character with only a few items gets very slow when used for an old character with tens of thousands of items, and the reverse is common too. It may recover when a restart clears the plan, then go bad again. - Sources: - [Query Processing Architecture Guide](https://learn.microsoft.com/en-us/sql/relational-databases/query-processing-architecture-guide) · Microsoft SQL Server · Parameter sniffing: the query plan is built for the parameter values passed at compile or recompile time - [Parameter Sensitive Plan Optimization](https://learn.microsoft.com/en-us/sql/relational-databases/performance/parameter-sensitive-plan-optimization) · Microsoft SQL Server · When data distribution is uneven, one cached plan doesn’t fit every parameter value - [Monitor performance by using the Query Store](https://learn.microsoft.com/en-us/sql/relational-databases/performance/monitoring-performance-by-using-the-query-store) · Microsoft SQL Server · Plans change with statistics, schema, and index changes, and the plan cache keeps only the latest plan; Query Store’s plan forcing pins a good plan; the Regressed Queries view compares slowed-down queries and their plans - [Statement Summary Tables](https://dev.mysql.com/doc/refman/8.4/en/performance-schema-statement-summary-tables.html) · MySQL · events_statements_summary_by_digest: COUNT_STAR and AVG_TIMER_WAIT (average time) per normalized query - [pg_stat_statements — track statistics of SQL planning and execution](https://www.postgresql.org/docs/current/pgstatstatements.html) · PostgreSQL · calls, total_exec_time, and mean_exec_time (average execution time) per statement - [pg_stat_statements (PostgreSQL 12 Documentation)](https://www.postgresql.org/docs/12/pgstatstatements.html) · PostgreSQL · Up to 12, the columns are named total_time and mean_time - [auto_explain — log execution plans of slow queries](https://www.postgresql.org/docs/current/auto-explain.html) · PostgreSQL · Logs the query plans of queries that took longer than auto_explain.log_min_duration #### db-ddl-lock · Schema change (DDL) lock during live service · Schema change lock (DDL / metadata lock) Adding a column or index to a table during live service can make every request that uses that table wait, all because of one lock that’s needed only briefly. - Why → Effect → On screen: A hotfix adds a column or index to a table in live use → The schema change waits for a long transaction opened earlier, and every request that comes after waits for the schema change → Features that use that table (inventory, mail, and so on) stop entirely and time out - Symptoms: Input lag, Dropped action / rollback, Can’t connect / infinite loading / Factors: Stall - Who: One feature only, Whole server / When: Randomly, Right after login or maintenance - Primary owner: Infra team (DB infrastructure) / Also: Game team (Server development) - Game team action items: Coordinate the timing of hotfixes that include schema changes with DB infrastructure, deploy code that works without the new column first. - Infra team action items: Set a short lock wait timeout and retry on failure, run it when there are no long transactions, use online schema change tools, change large tables during maintenance. - On the graph: Step change (Sessions waiting on locks, query latency on that table) - Where to look: MySQL: count sessions in SHOW PROCESSLIST whose State is Waiting for table metadata lock, and find the blocking session (blocking_pid) with sys.schema_table_lock_waits. PostgreSQL: requests whose granted is false and AccessExclusiveLock in pg_locks, and find the blocking session with pg_blocking_pids() - Confirmed if: From the moment the schema change starts, every query on that table piles up waiting on the lock, with an unfinished transaction or the schema change statement at the front - Ruled out if: Waits concentrate on specific rows while other rows in the same table go through fine: a hot row (db-hot-row) - Check with: Infra tools (no game code needed) - Learn more: MySQL briefly takes a metadata lock when changing a schema, and PostgreSQL briefly takes its strongest table lock. Even if the change itself is instant, a single unfinished transaction ahead of it makes every request behind it wait. - Sources: - [Online DDL Performance and Concurrency](https://dev.mysql.com/doc/refman/8.4/en/innodb-online-ddl-performance.html) · MySQL · Even online DDL briefly needs an exclusive metadata lock to finish; it waits if there is a long transaction, and the waiting lock request blocks every transaction behind it - [Server System Variables](https://dev.mysql.com/doc/refman/8.4/en/server-system-variables.html) · MySQL · lock_wait_timeout: metadata lock wait limit, default 31,536,000 seconds (1 year) - [ALTER TABLE (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/sql-altertable.html) · PostgreSQL · ALTER TABLE takes the strongest ACCESS EXCLUSIVE lock unless otherwise noted - [Client Connection Defaults (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/runtime-config-client.html) · PostgreSQL · lock_timeout: aborts the statement if it waits for a lock longer than this - [General Thread States](https://dev.mysql.com/doc/refman/8.4/en/general-thread-states.html) · MySQL · Waiting for table metadata lock: thread state while waiting for a metadata lock - [The schema_table_lock_waits and x$schema_table_lock_waits Views](https://dev.mysql.com/doc/refman/8.4/en/sys-schema-table-lock-waits.html) · MySQL · The session waiting on a metadata lock (waiting_query) and the blocking session (blocking_pid) - [pg_locks (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/view-pg-locks.html) · PostgreSQL · granted false means waiting for the lock; mode shows the lock type, such as AccessExclusiveLock - [System Information Functions and Operators (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/functions-info.html) · PostgreSQL · pg_blocking_pids(): list of sessions blocking the given session from acquiring a lock ### L13 Server architecture and operations (causes: 13) #### in-gateway · Routing through a gateway or proxy · Gateway / proxy hop Putting an intermediate server between the client and the game server adds processing time at every hop, and that server becomes a single point of failure. - Why → Effect → On screen: Client ↔ gateway ↔ game server architecture → The intermediate server adds processing and queueing time, and when it’s overloaded everyone is affected → Higher ping for everyone; if a gateway fails, every player routed through it disconnects - Symptoms: Input lag, Disconnect / Factors: Latency, Stall - Who: Whole server / When: When crowds gather, Always - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure), Game team (Client development) - Game team action items: Server: make the gateway tier scale out to more machines, let a character carry on unchanged when it reconnects through another gateway after its gateway dies (session reconnection). Client: reconnect automatically when the gateway connection drops. - Infra team action items: Scale gateways horizontally (add machines), monitor CPU, connection count, and processing latency per gateway. - Ballpark numbers: Inside the same data center, each hop normally adds less than 1 ms. When the gateway is overloaded, that grows to tens to hundreds of ms. - On the graph: Rises with load (Gateway processing latency, gateway CPU and connection count) - Where to look: Gateway CPU and connection count, Recv-Q on the gateway’s sockets (ss, netstat), and the latency difference before and after the gateway. For HTTP/gRPC calls through a service mesh, compare the Istio standard metric istio_request_duration_milliseconds split by sender (reporter=source) and receiver (reporter=destination) - Confirmed if: Game server processing time is unchanged but latency grows only across the gateway hop, and at the same time gateway CPU is saturated or Recv-Q builds up - Ruled out if: Paths that skip the gateway (direct connection, another gateway) are just as slow: points to the connection or the game server - Check with: Infra tools (no game code needed) - Learn more: With a service mesh such as Istio, the sidecar proxy (Envoy) running next to each server adds one more hop. A request between services passes through the sender’s sidecar and then the receiver’s sidecar, and every feature you add to the proxy, such as log and metric collection, adds processing and queueing time. - Real incidents: riot-edge-2020 - Sources: - [The Unique Architecture behind Amazon Games’ Seamless MMO New World](https://aws.amazon.com/blogs/gametech/the-unique-architecture-behind-amazon-games-seamless-mmo-new-world/) · AWS · In New World, the client connects to one of four entry servers (REP) with public addresses, which talk to the simulation servers (hubs) behind them - [Designs, Lessons and Advice from Building Large Distributed Systems](https://www.cs.cornell.edu/projects/ladis2009/talks/dean-keynote-ladis2009.pdf) · Google · LADIS 2009 keynote (Jeff Dean). Round trip within the same data center is about 0.5 ms - [Site Reliability Engineering, Chapter 22: Addressing Cascading Failures](https://sre.google/sre-book/addressing-cascading-failures/) · Google · When overload lengthens the queue, wait time grows to several times the processing time (100 ms of processing with a queue 10 times the thread count means 1.1 s) - [Performance and Scalability](https://istio.io/latest/docs/ops/deployment/performance-and-scalability/) · Istio · In sidecar mode, a request passes through the sender’s sidecar proxy and then the receiver’s; every added feature lengthens the processing path inside the proxy, and telemetry collection adds queueing time to the next request - [What is Envoy](https://www.envoyproxy.io/docs/envoy/latest/intro/what_is_envoy) · Envoy · Envoy is a separate process running alongside every application server, and the app sends and receives through Envoy on localhost - [Istio Standard Metrics](https://istio.io/latest/docs/reference/config/metrics/) · Istio · istio_request_duration_milliseconds (distribution of HTTP/gRPC request duration); the reporter label separates the sending (source) and receiving (destination) proxy - [netstat(8) — Linux manual page](https://man7.org/linux/man-pages/man8/netstat.8.html) · net-tools · Recv-Q: bytes on a connected socket not yet picked up by the user program #### in-zone-transfer · Zone transfer (handoff between servers) · Zone / server handoff Entering another area or dungeon means handing the character’s data to another server, and that handoff can be slow or fail. - Why → Effect → On screen: Entering a dungeon or traveling to another continent changes which server is responsible → Save → transfer → load, with a wait if the target server is busy or has no free dungeon instance → Long loading screens, failed entry, disconnects mid-transfer - Symptoms: Can’t connect / infinite loading, Freeze, Disconnect, Rubber-banding / Factors: Latency, Stall - Who: Just me, Specific zone/channel / When: While moving or changing zones - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Shrink the transferred data, reserve the target server in advance, send the player back to where they were if the transfer fails. - Infra team action items: Monitor free instance headroom on dungeon and zone servers, add servers before peak hours. - On the graph: Rises with load (Zone transfer duration and failure count) - Where to look: Server-side timing for each transfer stage (save, transfer, load) and failure reasons, plus player count and free instance count on the target server - Confirmed if: At the times long loading or failed entry is reported, transfer time rises or failures cluster, and the target server is crowded or out of free instances - Ruled out if: The transfer finishes quickly but the game freezes after arrival: points to a spawn burst when entering a crowded area, or client-side loading - Check with: Game server or client logs and metrics - Learn more: Seamless worlds without loading screens also switch the responsible server when you cross a server boundary. Near the boundary you may see a brief hitch or rubber-banding. - Sources: - [The Unique Architecture behind Amazon Games’ Seamless MMO New World](https://aws.amazon.com/blogs/gametech/the-unique-architecture-behind-amazon-games-seamless-mmo-new-world/) · AWS · The seamless world is split into a grid handled by several servers (hubs), and player state is handed from hub to hub as players move; session-based modes take spare servers from a shared pool #### in-cascade · Cascading failure · Cascading failure When one service slows down, the servers that call it get tied up waiting for responses, and even unrelated features stop. - Why → Effect → On screen: One service, such as the DB or authentication, slows down → Threads and connections on the calling servers are tied up waiting for responses, and retries of failed requests add more load → Everything slows down or stops, even features that look unrelated - Symptoms: Freeze, Input lag, Can’t connect / infinite loading / Factors: Stall - Who: Whole server / When: When crowds gather, Randomly - Primary owner: Game team (Server development) / Also: Infra team (Network infrastructure) - Game team action items: Put a timeout on every call, add circuit breakers and per-feature isolation (bulkheads), retry with growing intervals and a capped count, keep health check responses separate from busy work. - Infra team action items: Give load balancer health checks slack in failure count and interval so a briefly slow server isn’t pulled right away, limit how many servers can be pulled at once. - On the graph: Hits a ceiling (Per-service response time and error rate, thread and connection usage) - Where to look: Per-service response time, error rate, and retry count on one screen with aligned time axes, to find what slowed down first. Behind a load balancer: target response time (TargetResponseTime on AWS ALB), target 5xx count (HTTPCode_Target_5XX_Count), and number of targets pulled as unhealthy (UnHealthyHostCount) - Confirmed if: One service’s latency rises first, then thread and connection usage on its callers hits the limit, errors spread to other services, and retry count and pulled-target count rise together - Ruled out if: Several services slowed down at the same instant: check shared resources (DB, network, hosts) first - Check with: Infra tools (no game code needed) - Learn more: Health checks (probes that confirm a server is alive) also make cascades worse. When a busy server answers a check late, the load balancer pulls a server that is actually working, its traffic piles onto the remaining servers, and the next server falls behind too. - Real incidents: riot-edge-2020, riot-euw-2021, roblox-2021, aws-2021, aws-2025 - Sources: - [Site Reliability Engineering, Chapter 22: Addressing Cascading Failures](https://sre.google/sre-book/addressing-cascading-failures/) · Google · An overloaded server that fails health checks gets pulled, load piles onto the rest, and retries amplify it; recommends capped retries, randomized exponential backoff, and deadlines - [Circuit Breaker Pattern](https://learn.microsoft.com/en-us/azure/architecture/patterns/circuit-breaker) · Microsoft Azure · Requests blocked until their timeout hold threads and DB connections and make unrelated features fail; once failures pile up within a set time, calls are rejected immediately - [Timeouts, retries, and backoff with jitter](https://d1.awsstatic.com/builderslibrary/pdfs/timeouts-retries-and-backoff-with-jitter.pdf) · AWS · Amazon Builders’ Library. With 3 retries at each layer of a 5-deep call chain, DB load grows 243 times; retry at only one layer and cap retries with a token bucket - [CloudWatch metrics for your Application Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/application/load-balancer-cloudwatch-metrics.html) · AWS · TargetResponseTime (time from the request leaving the load balancer until the target starts responding), HTTPCode_Target_5XX_Count (5xx responses generated by targets), UnHealthyHostCount (number of unhealthy targets) #### in-subservice · Auxiliary server outage · Auxiliary service outage When a server that runs separately from the game server, such as chat, party, or auction house, fails, only that feature stops working. - Why → Effect → On screen: A server dedicated to one feature slows down or dies → Only requests for that feature get no response → Chat doesn’t work, party invites do nothing, the marketplace loads forever (combat is fine) - Symptoms: Dropped action / rollback, Can’t connect / infinite loading / Factors: Stall, Packet loss - Who: One feature only / When: Randomly, When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Design the game to keep running when a feature fails, show status per feature, don’t pile many features onto one central server. - Infra team action items: Set up health checks and alerts for each auxiliary server, add redundancy and automatic restarts. - On the graph: Mass disconnect (Request success rate per feature, auxiliary server connection count and health checks) - Where to look: Health checks, process state, and connection count for each auxiliary server (chat, party, auction house), plus request success rate and response time per feature. Behind a load balancer, UnHealthyHostCount for the target group - Confirmed if: Only the server behind the reported feature fails health checks or shows a sharp drop in connections, while game server ticks and combat are normal - Ruled out if: Several features stopped at once: points to a central server that relays them all, or a cascading failure - Check with: Infra tools (no game code needed) - Learn more: If one central server (a world or manager server) relays parties, guilds, whispers, and cross-server moves, several features stop at once when that one server slows down. - Sources: - [Bulkhead Pattern](https://learn.microsoft.com/en-us/azure/architecture/patterns/bulkhead) · Microsoft Azure · Isolating components into pools lets the rest keep working when one fails and keeps the failure from spreading - [REL05-BP01 Implement graceful degradation to transform applicable hard dependencies into soft dependencies](https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/rel_mitigate_interaction_failure_graceful_degradation.html) · AWS · AWS Well-Architected. Design core functions to keep working on slightly stale or substitute data when a dependency fails - [CloudWatch metrics for your Network Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/load-balancer-cloudwatch-metrics.html) · AWS · UnHealthyHostCount: number of targets that health checks judged unhealthy #### in-deploy · Deploys and restarts · Deploy / rolling restart If you restart a server for an update without moving its connections, everyone on it disconnects, and the final saves before shutdown and the reconnects all hit at once. - Why → Effect → On screen: Servers restart one after another to roll out a hotfix → Each server shuts down without moving its connections, and saves for every player on it hit the DB at once → Disconnects without notice, a surge of reconnects - Symptoms: Disconnect, Can’t connect / infinite loading, Input lag / Factors: Stall - Who: Whole server, Specific zone/channel / When: Randomly, Right after login or maintenance - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Add draining (block new connections only and wait for current players to leave), move characters to another server, spread out saves before shutdown, report ready after a restart only once cache loading and JIT warm-up are done, for hot reload load the new data ahead of time on a separate thread and swap it in all at once between ticks. - Infra team action items: Have the deploy tool wait for each server to drain before restarting it, route traffic to a restarted server only after confirming it’s ready (warm-up done), announce deploy times. - Ballpark numbers: With 5,000 players on one server, 5,000 saves hit the DB within a few seconds before shutdown. - On the graph: Mass disconnect (Connections per server, DB write count) - Where to look: Deploy tool job history (restart time per server) overlaid as vertical lines (annotations) on graphs of connections, disconnects, DB writes, and login requests - Confirmed if: Connections per server drop sharply one server at a time at each restart, with DB writes spiking just before and login requests right after - Ruled out if: Disconnect times don’t line up with the deploy or restart history: points to a server crash or network equipment - Check with: Infra tools (no game code needed) - Learn more: Servers are also slow for a few minutes right after they come back. The cache is empty so DB queries pile up, and Java and C# servers haven’t yet finished optimizing code as it runs (JIT warm-up), so the same work takes longer. Reloading scripts or data tables without shutting down (hot reload) also stops the tick while it reads, causing a brief pause. - Sources: - [Site Reliability Engineering, Chapter 20: Load Balancing in the Datacenter](https://sre.google/sre-book/load-balancing-datacenter/) · Google · A server that receives SIGTERM goes into lame duck state, sending new requests to other servers and finishing only the ones in progress; for the first few minutes after a restart, before JIT optimization, it uses more resources, so it takes traffic only after warming up - [Liveness, Readiness, and Startup Probes](https://kubernetes.io/docs/concepts/workloads/pods/probes/) · Kubernetes · A readiness check holds back traffic until connections are established, files are loaded, and caches are warmed - [Edit target group attributes for your Network Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/edit-target-group-attributes.html) · AWS · Deregistering a target stops new connections to it and drains existing ones (300 s by default) #### in-autoscale · Autoscaling delay · Autoscaling lag When players flood in, servers are added automatically, but getting them ready takes several minutes, and the existing servers are overloaded in the meantime. - Why → Effect → On screen: Connections spike when an event starts → Several minutes pass before new servers boot and are ready → Slow motion, and players who can’t connect, for the first few minutes after the event starts - Symptoms: Slow motion, Can’t connect / infinite loading / Factors: Stall - Who: Whole server / When: When crowds gather, Right after login or maintenance - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: Spread players across channels (players already in a crowded channel can’t be moved to a new server), cut startup and data loading time for new servers. - Infra team action items: Scale out ahead of events, keep warmed-up spare servers, when scaling in shut a server down only after its remaining players leave. - Ballpark numbers: 1 to a few minutes to detect the load (because metrics are averaged over several minutes), then several more minutes to boot a new server, read game data, and fill caches. - On the graph: Surge after opening (Instance count, CPU utilization, queued logins) - Where to look: Autoscaling activity history (when scale-out was decided, when new instances went into service) overlaid on CPU utilization and connection graphs. On AWS, the Auto Scaling group metrics (must be enabled to appear) GroupDesiredCapacity (target count), GroupPendingInstances (starting up), and GroupInServiceInstances (in service) - Confirmed if: For several minutes after a connection spike, only the desired count and pending instances rise while existing servers sit at their CPU limit, and things ease from the moment in-service instances increase - Ruled out if: Still slow after new instances come in: points to a cause other than server count (a shared resource such as the DB, a cascading failure) - Check with: Infra tools (no game code needed) - Learn more: Autoscaling is mostly used where a new server can simply take new players, such as login, gateway, and dungeon servers. Scaling in causes trouble too. If you scale in during the early-morning lull and shut servers down without waiting for the remaining players to leave, those players disconnect. - Real incidents: aws-2021, aws-2025 - Sources: - [Target tracking scaling policies for Amazon EC2 Auto Scaling](https://docs.aws.amazon.com/autoscaling/ec2/userguide/as-scaling-target-tracking.html) · AWS · EC2 basic metrics come at 5-minute intervals (1 minute with detailed monitoring), so metrics at 1-minute or finer intervals are recommended for a fast reaction - [Amazon CloudWatch metrics for Amazon EC2 Auto Scaling](https://docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-auto-scaling-metrics.html) · AWS · Group metrics are published every minute only when enabled; GroupDesiredCapacity (the count the group tries to maintain), GroupPendingInstances (instances not yet in service), GroupInServiceInstances (instances in service) - [Scheduled scaling for Amazon EC2 Auto Scaling](https://docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-auto-scaling-scheduled-scaling.html) · AWS · Adds and removes capacity ahead of time, at set times, to match predictable load changes - [Decrease latency for applications with long boot times using warm pools](https://docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-auto-scaling-warm-pools.html) · AWS · For apps that take a long time to boot, a pool of pre-initialized instances (warm pool) cuts scale-out delay #### in-monitoring · Logging and monitoring overload · Logging / monitoring overhead During an outage, log volume explodes, and servers that ship logs synchronously get even slower because of the logging. - Why → Effect → On screen: Errors make log and metric volume explode → The log collector falls behind, and servers that send synchronously wait on it → Stutter and freezes during an outage get worse because of logging - Symptoms: Stutter, Freeze / Factors: Stall - Who: Whole server / When: When crowds gather, Randomly - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure) - Game team action items: Send asynchronously, sample, drop when the buffer overflows, batch repeated error logs into one. - Infra team action items: Size log collector capacity for the burst volume seen during outages, alert on collector backlog. - On the graph: Random spikes (Log volume, log collector queue) - Where to look: Log lines and bytes per second on the server, plus the log collection agent’s queue and dropped count, alongside tick time. If a thread is stalled, use bcc offcputime -p to see whether it’s waiting on log writes or shipping - Confirmed if: When ticks spike, log volume jumps to tens of times normal, and the game thread’s wait time concentrates in log write/send call stacks - Ruled out if: Log volume is normal or the game thread isn’t waiting on logging: the log burst is only a result of the outage, so look separately for whatever caused the first errors - Check with: Infra tools (no game code needed) - Sources: - [Logging in C#](https://learn.microsoft.com/en-us/dotnet/core/extensions/logging) · Microsoft · .NET logging methods are synchronous, so with slow storage it recommends writing to fast storage first and moving the logs later - [Asynchronous loggers](https://logging.apache.org/log4j/2.x/manual/async.html) · Apache Software Foundation · Asynchronous logging absorbs short bursts in a queue, but if output stays slow, the queue fills and logging drops to the speed of the slowest output, or logs are dropped (Discard) depending on policy - [Demonstrations of offcputime, the Linux eBPF/bcc version](https://raw.githubusercontent.com/iovisor/bcc/master/tools/offcputime_example.txt) · IO Visor · Sums the time threads spent blocked off the CPU (off-CPU) per call stack; -p selects the process #### in-clock-skew · Clock skew between servers · Clock skew between servers When each server’s clock is slightly off from the others, cooldown, buff, and event start checks disagree from server to server. - Why → Effect → On screen: A server whose time sync stopped drifts hundreds of ms to several seconds away from the other servers → Passing absolute times, such as when a buff ends, between servers makes their checks disagree → A buff disappears or a cooldown starts over after moving to another server - Symptoms: Dropped action / rollback / Factors: Latency - Who: Just me / When: While moving or changing zones - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: Pass remaining time between servers in place of absolute timestamps. - Infra team action items: Monitor time sync (NTP, chrony), alert on clock differences between servers. - Ballpark numbers: With healthy time sync (NTP, chrony), servers in the same data center usually stay within a few ms of each other. If sync stops, or a virtual machine is paused for a long time and then resumes, the gap grows to hundreds of ms to several seconds. - On the graph: Slow climb (Clock offset per server) - Where to look: chronyc tracking on every server, comparing System time (difference between the system clock and NTP time), Last offset, and Ref time (when a measurement from the time source was last applied) - Confirmed if: The problem server’s offset is hundreds of ms or more away from the other servers, or its Ref time stopped long ago, and the mismatched checks happen only on moves to or from that server - Ruled out if: All servers’ offsets are within a few ms: points to the game’s own time calculation or client clock sync error - Check with: Infra tools (no game code needed) - Learn more: A single server’s clock jumping forward or backward all at once is covered in the server OS layer under “System clock jump (NTP step).” - Sources: - [RFC 5905: Network Time Protocol Version 4: Protocol and Algorithms Specification](https://www.rfc-editor.org/rfc/rfc5905) · IETF · NTP clients on a fast LAN usually stay within a few hundred µs - [chrony – Frequently Asked Questions](https://chrony-project.org/faq.html) · chrony · Ordinary computer clocks drift less than 100 ppm, but virtual machines can drift more; a VM that was paused and resumed can be far enough off to need a step correction - [clock_gettime(2) — Linux manual page](https://man7.org/linux/man-pages/man2/clock_gettime.2.html) · Linux man-pages · CLOCK_REALTIME can jump discontinuously from manual changes or NTP adjustments; CLOCK_MONOTONIC is not affected by such jumps - [chronyc(1)](https://chrony-project.org/doc/4.6/chronyc.html) · chrony · chronyc tracking fields: System time (difference between NTP time and the system clock), Last offset (offset estimated at the last update), Ref time (when the last measurement from the time source was applied) #### in-bots · Too many macros and bots · Bots and macros Bots send requests far more often than people do and eat into server capacity. - Why → Effect → On screen: Large numbers of bots connected, farming, moving, and trading nonstop → Server processing load and DB load go up → A specific farming spot or the whole server slows down (slow motion, input lag) - Symptoms: Slow motion, Input lag / Factors: Stall - Who: Whole server, Specific zone/channel / When: Always, Evening peak hours - Primary owner: Game team (Server development) / Also: Infra team (Network infrastructure) - Game team action items: Detect bots, rate-limit requests per account and character. - Infra team action items: Rate-limit connections and requests per IP (loosely, since many people share one IP at PC bangs and on mobile networks), block bot IP ranges at the firewall/WAF. - On the graph: Outliers only (Requests per second per account and IP) - Where to look: Distribution and top list of requests per second per account and character from game server logs. Without code metrics, requests per IP at the firewall/WAF - Confirmed if: A handful of accounts or IPs send requests nonstop at rates no human could produce, and limiting them visibly cuts server load - Ruled out if: Requests are spread evenly across accounts: points to normal player growth (tick overrun, autoscaling delay) - Check with: Game server or client logs and metrics - Sources: - [Using rate-based rule statements in AWS WAF](https://docs.aws.amazon.com/waf/latest/developerguide/waf-rule-statement-type-rate-based.html) · AWS · Counts requests per key such as IP, and rate-limits when there are too many within a set time window - [RFC 6269: Issues with IP Address Sharing](https://www.rfc-editor.org/rfc/rfc6269) · IETF · When many subscribers share one IP, blocking or limiting by IP also blocks other users #### in-external · External service dependency · External dependencies (auth, billing, platform) When an external service such as platform login, payments, or identity verification is slow or down, players get stuck at that step. - Why → Effect → On screen: An external authentication or payment service is down or slow → That step waits for a response → Can’t log in, payments fail. Players already in the game are fine - Symptoms: Can’t connect / infinite loading, Dropped action / rollback / Factors: Stall - Who: Whole server, One feature only / When: Right after login or maintenance, During specific actions - Primary owner: External (External) / Also: Game team (Server development) - Game team action items: Put timeouts and a friendly message on external calls, cache authentication results, set up a retry and compensation process for payments. - External action items: Ask the authentication, payment, or platform provider to confirm the outage and restore service, tell players the problem is an external service outage. - On the graph: Step change (External call response time and error rate, successful logins) - Where to look: Response time, error rate, and timeout count per external call (platform login, payments, identity verification), plus the provider’s status page - Confirmed if: From the time login and payment failures pile up, errors and timeouts for one specific external call step up and stay there, and the provider’s status page shows an outage at the same time - Ruled out if: External calls are healthy but logins are blocked: points to the login server itself (thread pool exhaustion, DB) or the OS connection queue (backlog) - Check with: Game server or client logs and metrics - Real incidents: fastly-2021, aws-2021, aws-2025 - Sources: - [Timeouts, retries, and backoff with jitter](https://d1.awsstatic.com/builderslibrary/pdfs/timeouts-retries-and-backoff-with-jitter.pdf) · AWS · Amazon Builders’ Library. Waiting for a response holds resources such as threads and connections, so set timeouts, and retry APIs with side effects only when they’re idempotent - [Circuit Breaker Pattern](https://learn.microsoft.com/en-us/azure/architecture/patterns/circuit-breaker) · Microsoft Azure · Calls likely to fail are rejected immediately without waiting for the timeout, which protects response time - [REL05-BP01 Implement graceful degradation to transform applicable hard dependencies into soft dependencies](https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/rel_mitigate_interaction_failure_graceful_degradation.html) · AWS · AWS Well-Architected. Keep core functions running when a dependency fails, even on slightly stale data (the basis for caching authentication results) #### in-region-match · Matchmaking and region assignment errors · Wrong region assignment (matchmaking / GeoDNS) When a player lands on a server in a distant region while a closer region exists, that player’s ping stays high even though their connection is fine. - Why → Effect → On screen: Bad GeoIP data, a VPN, assigning a whole party by the members’ average ping, rules that widen the search to distant regions when there aren’t enough players, assignment by DNS resolver location → The player connects to a server across the ocean even though a nearby region exists → In a game with servers in several regions, only you (or only your party) always have high ping, with input lag, rubber-banding, and skills that don’t go off - Symptoms: Input lag, Rubber-banding, Dropped action / rollback / Factors: Latency - Who: Just me, Specific region/ISP / When: Always, Right after login or maintenance - Primary owner: Game team (Server development) / Also: Game team (Client development), Infra team (Network infrastructure), External (External) - Game team action items: Server: assign by the per-region ping the client measures in place of GeoIP, cap ping in the rules that widen to distant regions, for parties check the highest member ping as well as the average, log the assigned region and the ping at that moment. Client: measure per-region ping over UDP and send it with the matchmaking request, show the connected region and ping on screen, offer a manual region choice. - Infra team action items: If DNS picks the region, confirm the authoritative DNS supports EDNS Client Subnet (if the player’s resolver doesn’t send it, assignment goes by resolver location), update the GeoIP database regularly, tag connection logs on each region’s servers with GeoIP country and ASN to find the countries and ISPs that end up in distant regions. - External action items: Tell players to turn off VPNs and game boosters and reconnect, tell players on a corporate or overseas DNS to switch to their ISP’s DNS, ask the GeoIP provider to correct wrong locations. - Ballpark numbers: A Seoul player sent to the US West region when Tokyo is available sees ping rise from about 30 ms to about 130 ms. GeoIP is about 99.8% accurate at the country level, but at the city level only about 66% of lookups fall within 50 km, even in the US, and with a VPN it returns the VPN server’s location in place of the player’s. - On the graph: Outliers only (RTT (ping) per player, distribution of assigned regions) - Where to look: Tag client IPs from each region’s connection records (load balancer access logs, VPC flow logs) with GeoIP country and ASN, and count which region each country and ISP connects to. For a single player, compare the region they actually connected to with their ping to the nearest region (measured by the player, or mtr from that region’s server to the player’s IP) - Confirmed if: High-RTT players or countries are connected to a distant region even though a nearby one exists, and ping measured to the nearby region is low - Ruled out if: Correctly assigned to the nearby region but ping is still high: points to detour routing or the player’s own connection or Wi-Fi - Check with: Infra tools (no game code needed) - Learn more: Picking the region by DNS (geolocation or latency-based DNS) guesses location from the address of the player’s DNS resolver in place of the player’s own address. If the resolver doesn’t support EDNS Client Subnet, which passes along part of the player’s address, players on a corporate DNS or a distant DNS are assigned by where the resolver is. Matchmaking systems may also judge a party by the average of its members’ ping, or widen the ping limit after a long wait and assign a distant region. AWS GameLift Servers also uses the average as the default for party ping, and its example setting widens the ping cap from 50 ms to 100 ms and then 200 ms. For a player on a VPN, the extra latency of going through the relay server (“Routing through a VPN or game booster”) can overlap with a distant-region assignment; tell them apart by whether the assigned region changes when the player turns off the VPN and reconnects. Connecting to a distant region because there is no nearby region at all is covered in “Propagation delay (physical distance).” - Sources: - [RFC 7871: Client Subnet in DNS Queries](https://www.rfc-editor.org/rfc/rfc7871) · IETF · DNS that answers differently by location guesses location from the address of the resolver sending the query, and gives inappropriate answers when the player uses a central resolver far away. EDNS Client Subnet (an optional feature) passes along part of the player’s address - [How Amazon Route 53 uses EDNS0 to estimate the location of a user](https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/routing-policy-edns0.html) · AWS · If the resolver doesn’t support edns-client-subnet, the player’s location is guessed from the resolver’s address and the answer is based on the resolver’s location (same for geolocation and latency-based routing) - [Geolocation accuracy](https://support.maxmind.com/knowledge-base/articles/maxmind-geolocation-accuracy) · MaxMind · About 99.8% at the country level, about 66% at the US city level (within 50 km); with a VPN it gives the VPN server’s location in place of the end user’s; mobile network IPs are used across wide areas, so fine-grained location is unknown; the database needs continuous updates; corrections can be requested - [FlexMatch rule types](https://docs.aws.amazon.com/gameliftservers/latest/flexmatchguide/match-rules-reference-ruletype.html) · AWS · The latency rule (maxLatency) looks at player latency per location; for parties it uses the members’ average by default (partyAggregation avg); queues can place games in regions that don’t meet the latency rule - [Create a player latency policy](https://docs.aws.amazon.com/gameliftservers/latest/developerguide/queues-design-latency.html) · AWS · Places the game at the location with the lowest average latency across all players, but players with extreme latency get placed too; example policy widens the ping cap from 50 ms to 100 ms and then 200 ms - [Amazon GameLift Servers UDP ping beacons](https://docs.aws.amazon.com/gameliftservers/latest/developerguide/reference-udp-ping-beacons.html) · AWS · The game client measures latency to a UDP endpoint at each hosting location and uses it for placement and matchmaking; closer to real game traffic than ICMP ping - [Azure network round-trip latency statistics](https://learn.microsoft.com/en-us/azure/networking/azure-network-latency) · Microsoft Azure · Measured median round trip from Seoul (Korea Central): Tokyo (Japan East) 29 ms, US West 124–136 ms - [Flow log records](https://docs.aws.amazon.com/vpc/latest/userguide/flow-log-records.html) · AWS · srcaddr in a VPC flow log record: for inbound traffic, the sender’s IP address #### in-cert · Expired or misconfigured TLS certificate · TLS certificate expiry / misconfiguration When the certificate on a login, API, or patch server expires or is missing its intermediate certificate, every client that connects from that moment on fails the TLS connection. - Why → Effect → On screen: The certificate is past its validity period, the server sends it without the intermediate certificate, or the date and time on the player’s device are wrong → The client fails certificate validation and drops the TLS connection → Can’t connect / infinite loading at the login or patch step, or only HTTPS features such as the store fail. Players already connected are usually fine - Symptoms: Can’t connect / infinite loading, Dropped action / rollback / Factors: Stall - Who: Whole server, One feature only, Just me / When: Right after login or maintenance, During specific actions - Primary owner: Infra team (Network infrastructure) / Also: Infra team (Server infrastructure), Game team (Client development) - Game team action items: Log certificate errors under an error code distinct from other connection failures and show a message, if the date is wrong tell players to set their device’s date and time automatically, if you use certificate pinning include a backup key and coordinate the certificate rotation schedule with the infra team. - Infra team action items: Network: when TLS terminates at the load balancer or CDN, alert on managed certificate auto-renewal status and days remaining (DaysToExpiry on ACM), keep the validation DNS records in place. Servers/OS: when TLS terminates on the servers, automate renewal and reload the config after renewal, configure the full chain including intermediate certificates, check the remaining validity of every login, API, and patch address from outside on a schedule and alert on it. - Ballpark numbers: Let’s Encrypt certificates last 90 days and renewal every 60 days is recommended; AWS Certificate Manager checks DNS-validated certificates 45 days before expiry and renews them automatically. If automatic renewal fails silently, new connections are all blocked at exactly the expiry time. - On the graph: Mass disconnect (Successful logins, TLS handshake errors) - Where to look: openssl s_client -connect HOST:443 -showcerts to see the certificate list the server actually sends, then each certificate’s expiry date (notAfter) with openssl x509 -noout -enddate. If TLS terminates at the load balancer, TLS negotiation error count (ClientTLSNegotiationErrorCount on AWS ALB and NLB) and successful logins - Confirmed if: The expiry date has passed or the intermediate certificate is missing from the list the server sends, and errors started rising at the expiry time or when the certificate was changed - Ruled out if: Certificate list and expiry date are fine but only some players fail: check the date and time on those players’ devices or the root certificate list of an old OS - Check with: Infra tools (no game code needed) - Learn more: A config missing the intermediate certificate can look fine when you open it in a desktop browser. Browsers remember intermediate certificates picked up from other sites and fill the gap, but clients without that memory, such as Android apps, fail. Validity periods are also getting shorter. Let’s Encrypt plans to cut the default validity to 64 days in 2027 and 45 days in 2028, so a setup hard-coded to renew every 60 days leaves only four days of margin with a 64-day certificate and runs past expiry with a 45-day one. AWS Certificate Manager also doesn’t auto-renew imported certificates, and renewal fails if you delete the validation DNS record. Blocked logins look similar to “DNS failures and delays,” but a certificate problem fails at the TLS handshake after the server address has been resolved, and it starts at the expiry time or when the certificate was changed. - Sources: - [RFC 5280: Internet X.509 Public Key Infrastructure Certificate and Certificate Revocation List (CRL) Profile](https://www.rfc-editor.org/rfc/rfc5280) · IETF · A certificate is valid from notBefore to notAfter; path validation checks that the current time falls within the validity period of every certificate in the chain (it fails if the validating side’s clock is wrong) - [FAQ](https://letsencrypt.org/docs/faq/) · Let's Encrypt · Default certificate lifetime of 90 days, renewal every 60 days recommended - [Decreasing Certificate Lifetimes to 45 Days](https://letsencrypt.org/2025/12/02/from-90-to-45) · Let's Encrypt · Default lifetime cut to 64 days in February 2027 and 45 days in February 2028; a fixed 60-day renewal interval will no longer be enough, so renewal at about two-thirds of the lifetime is recommended - [Renewal for domains validated by DNS](https://docs.aws.amazon.com/acm/latest/userguide/dns-renewal-validation.html) · AWS · 45 days before expiry, checks whether the certificate is in use by an AWS service and whether the validation CNAME record exists, then renews automatically; if it can’t validate, sends notices 30, 15, 7, 3, and 1 days before expiry - [Managed certificate renewal in AWS Certificate Manager](https://docs.aws.amazon.com/acm/latest/userguide/managed-renewal.html) · AWS · Imported certificates and certificates that have already expired are not eligible for automatic renewal - [Supported CloudWatch metrics](https://docs.aws.amazon.com/acm/latest/userguide/cloudwatch-metrics.html) · AWS · DaysToExpiry: days left until the certificate expires, published twice a day until expiry - [Security with network protocols](https://developer.android.com/privacy-and-security/security-ssl) · Android (Google) · If the server omits the intermediate certificate, Android apps fail with SSLHandshakeException, while desktop browsers may fill it in from cached intermediates and show no error; check the chain the server sends with openssl s_client - [Network security configuration](https://developer.android.com/privacy-and-security/security-config) · Android (Google) · With certificate pinning, you must include backup keys to prepare for key rotation or CA changes; otherwise connections break until the app is updated - [openssl-s_client](https://docs.openssl.org/3.0/man1/openssl-s_client/) · OpenSSL · -showcerts: shows the certificates the server sent, in the order it sent them (not a validated chain) - [openssl-x509](https://docs.openssl.org/3.0/man1/openssl-x509/) · OpenSSL · -enddate: prints the certificate’s expiry date (notAfter); -checkend: checks whether it expires within the given number of seconds - [CloudWatch metrics for your Application Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/application/load-balancer-cloudwatch-metrics.html) · AWS · ClientTLSNegotiationErrorCount: number of connections that failed to establish a TLS session, for example because the client dropped the connection after failing to validate the server certificate - [CloudWatch metrics for your Network Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/load-balancer-cloudwatch-metrics.html) · AWS · ClientTLSNegotiationErrorCount: number of TLS handshakes that failed during negotiation between a client and a TLS listener #### in-login-queue · Login queue cap and insufficient reconnect grace · Login queue cap / no reconnect grace When players flood in right after launch or maintenance, the login queue hits its cap and turns away new arrivals, and players who were already waiting lose their place during a brief disconnect and go back to the end of the line. - Why → Effect → On screen: More people try to connect than the login server can take at once, so it keeps a queue, and when the queue gets too long it refuses new entries to protect the server → The longer the queue, the longer the wait, and a brief Wi-Fi or mobile network drop during that time costs the player their place → Can’t connect / infinite loading, the game quits with an error while waiting, the player starts over at the back of the line - Symptoms: Can’t connect / infinite loading, Disconnect / Factors: Stall - Who: Whole server, Just me / When: Right after login or maintenance, Evening peak hours - Primary owner: Game team (Server development) / Also: Game team (Client development), Infra team (Server infrastructure) - Game team action items: Server: set the queue cap to what the login server can actually handle, hold a disconnected player’s place for a set time (reconnect grace period), show queue position and estimated wait, record queue length, refusals, and disconnects while waiting as metrics. Client: when disconnected while waiting, reconnect automatically to the same place without quitting the game, spread out retries with exponential backoff and jitter. - Infra team action items: Servers/OS: measure login and lobby server capacity with load tests before launch, prepare spare machines that can be added quickly at launch, graph queue metrics alongside connection attempts. - Ballpark numbers: At the 2021 FINAL FANTASY XIV expansion launch, new entries were refused once the queue passed 17,000 players per logical data center (Error 2002). If a player disconnected while waiting, the lobby server waited tens of seconds to 1 minute, and a player who reconnected within that time kept their place in the queue. - On the graph: Hits a ceiling (Login queue length, refusals at the cap, disconnects while waiting) - Where to look: Queue length, average wait time, refusals at the cap, and disconnects while waiting, as recorded by the login and lobby servers, on the same graph as connection attempts - Confirmed if: Right after launch or maintenance, refusals rise while queue length flattens at the cap, and disconnects while waiting concentrate on Wi-Fi and mobile network players - Ruled out if: The queue is short but login is slow: points to the DB (db-login-storm) or the OS connection queue (so-backlog) - Check with: Game server or client logs and metrics - Learn more: A login rush that slows down the DB is covered in “Login storm and N+1 queries,” and an OS connection queue overflow in “Connection queue (listen backlog) overflow.” This entry is about the design of the login queue the game keeps on purpose. The queue cap is a safeguard that protects the login server, so you can’t remove it: refusing excess requests early is what keeps the server processing the requests it can handle. What matters is reducing what refusals and disconnects cost players, and the longer the queue gets, the more the errors fall on players with unstable connections such as Wi-Fi and mobile networks. - Real incidents: ffxiv-2021 - Sources: - [Response to Congestion (as of Dec. 11)](https://na.finalfantasyxiv.com/lodestone/news/detail/6a94b30182b6d963994fdc0b789264ac9f24986f) · Square Enix · Once the queue passed 17,000 players per logical data center, new entries were refused so the login servers wouldn’t go down (Error 2002); a player disconnected while waiting got tens of seconds to 1 minute from the lobby server to reconnect and resume mid-queue, and went to the back of the line after that - [Using load shedding to avoid overload](https://aws.amazon.com/builders-library/using-load-shedding-to-avoid-overload/) · Amazon Builders' Library · Load shedding: refusing excess requests early so the server keeps processing the requests it can handle ### Netcode design (causes: 16) #### sy-request-response · Feedback only after the server responds (request-response) · Request-response (no client-side feedback) Press a button and there’s no animation or sound until the server answers. Your ping becomes your response time. - Why → Effect → On screen: Skills, movement, and item pickups play only after the server confirms them → From the moment you press, nothing happens for a full round trip plus the tick wait → At 150 ms ping, every action feels 0.2 s sluggish - Symptoms: Input lag / Factors: Latency - Who: Just me / When: Always, During specific actions - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: start animations, sounds, and effects the moment the player presses (client-side feedback), show only results (damage, rewards) after the server confirms, predict movement and basic attacks and apply them immediately, when a position correction arrives from the server re-apply the unconfirmed inputs from that position. Server: compute movement from the received inputs and send a correction only when the difference from the client’s predicted position exceeds a threshold. - Ballpark numbers: Response time ≈ ping + half the tick interval + one frame. At 20 ticks and 150 ms ping, about 190 ms. - On the graph: Always high (Time from input to start of feedback, RTT (ping)) - Where to look: Client log in a development build with the button press time, the time the first animation or sound starts, and the time the server response arrives, next to in-game RTT. Add latency with the engine’s network emulation (Unreal NetEmulation.PktLag) or Linux tc netem on a test server and measure at different pings - Confirmed if: Feedback always starts at the same instant the server response arrives, the time from input to feedback equals RTT plus the tick wait, and it grows by exactly the latency you add - Ruled out if: Feedback starts on the press and only results such as damage numbers come late: normal design. Delay longer than a tick interval even at low ping: points to a double tick wait or a client frame problem - Check with: Game server or client logs and metrics - Learn more: For games that don’t need fast reactions, such as turn-based, card, and idle games, this approach is the simplest and safest. The problem is real-time action games that build even movement and basic attacks this way. - Sources: - [Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization (Yahn W. Bernier, GDC 2001)](https://web.cs.wpi.edu/~claypool/courses/4513-B03/papers/games/bernier.pdf) · Valve · A client that waits only for server results shows every action 500 ms late when latency is 500 ms. Solved with client-side prediction and server reconciliation - [Using Gameplay Abilities in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/using-gameplay-abilities-in-unreal-engine) · Epic Games · Local Predicted runs immediately on press and the server makes the final decision; Server Initiated has no prediction, so the player sees the delay - [Understanding Networked Movement in the Character Movement Component for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/understanding-networked-movement-in-the-character-movement-component-for-unreal-engine) · Epic Games · The client predicts and saves moves, the server corrects only when the error exceeds the tolerance (MAXPOSITIONERRORSQUARED), and the client re-applies saved moves after a correction - [Using Network Emulation in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/using-network-emulation-in-unreal-engine) · Epic Games · Test with minimum and maximum latency and a packet loss percentage on server and client; set from the console, e.g., NetEmulation.PktLag - [tc-netem(8) — Linux manual page](https://man7.org/linux/man-pages/man8/tc-netem.8.html) · iproute2 · A test tool that adds delay and jitter (delay TIME JITTER) and loss (loss random PERCENT) to outgoing packets to mimic a real network #### sy-chatty · Chatty protocol (many sequential round trips) · Chatty protocol / sequential round trips If one action needs several server round trips one after another, your ping is multiplied by that many. - Why → Effect → On screen: Open shop → request list → check price → buy → refresh inventory, each as a separate request → Each request is sent only after the answer to the previous one arrives → At 150 ms ping, a single purchase takes close to 1 s. Loading takes unusually long - Symptoms: Input lag, Can’t connect / infinite loading / Factors: Latency - Who: One feature only, Just me / When: During specific actions, Right after login or maintenance - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: change the protocol so several steps go in one request and response (e.g., include the updated inventory in the purchase response). Client: fetch the data you’ll need ahead of time, use UI that doesn’t wait for results. - Ballpark numbers: Time taken ≈ number of round trips × (ping + server processing + tick wait). With 5 round trips at 150 ms ping, about 0.85–1 s. - On the graph: Always high (Completion time per feature, round trips per action) - Where to look: Server-side packet capture (Wireshark) while a test account performs one action such as a shop purchase or login, counting how many times requests and responses alternate and the gaps between them. With server request logs, group by session ID and look at the request count and each request’s arrival and response times - Confirmed if: One action sends several requests in turn, each waiting for the previous response, completion time is roughly round trips × RTT, and the same feature is proportionally slower for players in high-ping regions - Ruled out if: Only one or two round trips but one response takes a long time: points to server processing or the DB. All players equally slow regardless of ping: check server load - Check with: Infra tools (no game code needed) - Sources: - [Chatty I/O antipattern](https://learn.microsoft.com/en-us/azure/architecture/antipatterns/chatty-io/) · Microsoft Azure · Many small I/O requests add up to latency that badly hurts responsiveness; recommends fewer, larger requests #### sy-no-queue · No skill input buffering · No input/spell queue If you can’t press the next skill until the server confirms the previous one has finished, a round trip gets inserted between every skill in a rotation. - Why → Effect → On screen: The next skill input is accepted only “after the previous skill is confirmed” → A gap as long as your ping opens between every skill → Gaps between skills in a rotation, and the higher the ping, the lower the DPS - Symptoms: Input lag, Dropped action / rollback / Factors: Latency - Who: Just me, One feature only / When: During specific actions - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: an input buffer window that accepts input within a set time before the cooldown ends (e.g., 0.3–0.4 s) and sends it to the server right away. Server: don’t reject input that arrives slightly early, run it the moment the cooldown ends. - Ballpark numbers: In a rotation with a 1-second cooldown at 150 ms ping, each gap between skills is 0.15 s or more, so you cast over 13% fewer skills in the same time. - On the graph: Always high (Gap between skills, RTT (ping)) - Where to look: Server log per character of when a skill cooldown ended, when the next skill request arrived, and when it ran, with the gap compared against the player’s RTT - Confirmed if: There is always a gap of about one RTT between the cooldown ending and the next skill running, and higher-ping players have longer gaps and cast fewer skills in the same time - Ruled out if: The gap is constant regardless of ping: global cooldown or animation length by design. The gap spikes only occasionally: check jitter/loss or tick overrun - Check with: Game server or client logs and metrics - Learn more: World of Warcraft, for example, has an input buffer window that players can adjust in the settings. When the window is longer than the round trip, ping barely shows up between skills in a rotation. - Sources: - [Latency and Player Actions in Online Games (Communications of the ACM, 2006)](https://web.cs.wpi.edu/~claypool/papers/precision-deadline/final.pdf) · ACM · In EverQuest II, damage dealt by a character dropped as latency rose, making fights longer (going from 0 to 500 ms added 5 seconds to a fight of about 2 minutes) #### sy-short-window · Short timing windows eaten up by ping · Timing window too short for latency + reaction When the time you have to react is short, as with dodges, parries, and guards, ping eats up that time and some attacks become impossible to avoid. - Why → Effect → On screen: Short timing windows, such as a 0.5 s boss attack telegraph or a 0.2 s parry window → You see the telegraph late (downstream latency + interpolation), and your input also arrives late (upstream latency + tick wait) → You get hit even though you clearly dodged, parries don’t go off - Symptoms: Dropped action / rollback, Input lag / Factors: Latency - Who: Just me, One feature only / When: During specific actions - Primary owner: Game team (Server development) / Also: Game team (Client development), Infra team (Server infrastructure) - Game team action items: Server: schedule attack telegraphs at a server time and send them in advance, widen the timing window by the player’s ping (lag compensation). Client: play received telegraphs at their scheduled server time. - Infra team action items: Put servers close to regions with many players (regional servers) to cut ping itself. - Ballpark numbers: At 150 ms ping and 100 ms interpolation, the telegraph takes about 0.18 s to appear on your screen and your input takes about 0.1 s to reach the server. Add 0.25 s of human reaction time, and a 0.5 s telegraph is nearly impossible. - On the graph: Outliers only (Dodge/parry failure rate (by ping bracket)) - Where to look: Server log of the timing window’s start and end times, when the player’s input reached the server, and that player’s RTT, with failure rate broken down by ping bracket (e.g., 50 ms steps) - Confirmed if: Failure rate is clearly higher in higher ping brackets, and failed inputs arrive shortly after the window closes (within RTT plus interpolation time) - Ruled out if: Failure rate is similar across ping brackets: the pattern is just hard. Inputs arrive inside the window but still count as failures: check the window-checking code or server validation - Check with: Game server or client logs and metrics - Sources: - [Factors influencing the latency of simple reaction time (Frontiers in Human Neuroscience, 2015)](https://doi.org/10.3389/fnhum.2015.00131) · Frontiers · Average simple reaction time about 231 ms (213 ms after correcting for equipment delay); recent large studies report 233–400 ms - [Latency and Player Actions in Online Games (Communications of the ACM, 2006)](https://web.cs.wpi.edu/~claypool/papers/precision-deadline/final.pdf) · ACM · The more precise and deadline-bound an action, the more sensitive it is to latency (limits around 100 ms for first-person, 500 ms for third-person, and 1,000 ms for omnipresent views) - [NetworkTime and ticks (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/advanced-topics/networktime-ticks.html) · Unity · Scheduling events at a server time (ServerTime) so every client plays them at the same moment #### sy-no-lagcomp · Hit registration without lag compensation · Server-now hit validation If the server checks hits only against where targets are on the server right now, what you saw on your screen and the server’s call disagree. - Why → Effect → On screen: The opponent on your screen is at a position about 0.2 s in the past (at 150 ms ping and 100 ms interpolation) → The server checks against the current position, so the target has already left the spot you aimed at → A clear hit counts as a miss. You have to lead moving targets - Symptoms: Dropped action / rollback / Factors: Latency - Who: Just me / When: During specific actions - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: rewind to what the attacker was seeing and check the hit there (lag compensation), or switch to a target-lock system. Client: send the moment you were seeing (the server time being interpolated) with each attack. - On the graph: Outliers only (Hit rate on moving targets (by ping bracket)) - Where to look: Server hit log with the attack time, the target position on the attacker’s screen (sent by the client), the target position the server used, and the attacker’s RTT. In a development build, drawing the position the server used over the client’s screen shows it right away - Confirmed if: On misses, the gap between the two positions is roughly target speed × (attacker RTT + interpolation time), and higher ping lowers the hit rate on moving targets only - Ruled out if: Stationary targets get missed too: hitbox or collision check problem. Still mismatched with rewind in place: check whether the client reports its interpolation time to the server incorrectly - Check with: Game server or client logs and metrics - Sources: - [Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization (Yahn W. Bernier, GDC 2001)](https://web.cs.wpi.edu/~claypool/courses/4513-B03/papers/games/bernier.pdf) · Valve · Without lag compensation, you have to lead targets by your latency. Lag compensation has the server rewind by latency plus interpolation time to check hits - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · The server rewinds to the game state the player saw at the moment of the shot to check the hit; the client sends the simulation time it was seeing - [Source SDK 2013: player_lagcompensation.cpp](https://raw.githubusercontent.com/ValveSoftware/source-sdk-2013/master/src/game/server/player_lagcompensation.cpp) · Valve · In the Source engine, rewind time = network latency + interpolation time #### sy-lagcomp-overreach · Too much lag compensation · Excessive lag compensation If the server rewinds too far in the attacker’s favor, the target gets hit even after they’ve already taken cover. - Why → Effect → On screen: The server rewinds a long way to check hits for a high-ping attacker → On the target’s screen, they were already behind cover → “I got shot behind a wall,” high-ping players have the advantage - Symptoms: Dropped action / rollback / Factors: Latency - Who: Just me, Specific region/ISP / When: During specific actions - Primary owner: Game team (Server development) - Game team action items: Cap the rewind (e.g., 200–250 ms), for attackers with higher ping rewind only up to the cap and let them lead their shots for the rest. - On the graph: Outliers only (Rewind time per hit (by attacker ping)) - Where to look: Server hit log with the rewind time for each hit, attacker RTT, and the server time the target entered cover. In a development build, draw the rewound hitboxes on screen (sv_showlagcompensation in the Source engine) - Confirmed if: Hits in “I got shot behind a wall” reports come mostly from attackers with long rewind times, and rewind time grows with attacker ping with no cap - Ruled out if: Behind-the-wall hits also show up on hits with short rewind times: hitbox or collision check problem. If the target has high ping, their own movement reached the server late - Check with: Game server or client logs and metrics - Learn more: Rewound hit checks follow “favor the shooter.” A “favor the target” exception has also been proposed: skip the rewind if the target had already reached safety on their own screen. - Sources: - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · Without a rewind cap, a player with 500 ms latency could land hits 0.5 s after the target took cover, so a cap is set - [A Survey and Taxonomy of Latency Compensation Techniques for Network Computer Games (ACM Computing Surveys, 2022)](https://web.cs.wpi.edu/~claypool/papers/lag-taxonomy/paper.pdf) · ACM · The “shot around the corner” effect, rewind caps in commercial FPS games, and a proposed technique that skips the rewind when the target is safe - [Source SDK 2013: player_lagcompensation.cpp](https://raw.githubusercontent.com/ValveSoftware/source-sdk-2013/master/src/game/server/player_lagcompensation.cpp) · Valve · Source engine rewind cap sv_maxunlag defaults to 1 second (maximum 1 second); sv_showlagcompensation draws the rewound hitboxes on screen #### sy-client-auth · Client authority · Client-authoritative results When each client decides its own results, your own screen feels responsive, but results disagree with other players’ screens and the game is easy to hack. - Why → Effect → On screen: The client decides position and hits, and the server only relays them → Two players each claim they hit first, and the server can’t verify either claim → Opponents teleport or pass through walls, “I hit them but it didn’t count” - Symptoms: Teleporting, Dropped action / rollback / Factors: Latency - Who: Whole server / When: Always - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: validate important results (such as hits) on the server, check movement speed and distance. Client: when the server rejects or corrects a result, roll back to the server’s value. - On the graph: Always high (Impossible movement speeds and conflicting hit reports) - Where to look: On the server, log the positions and hits clients report as-is, compute movement speed from consecutive position reports, and count reports over maximum speed and cases where two players both claim they hit first - Confirmed if: The server passes reports on to other clients without validation, and impossible speeds or conflicting hit reports show up steadily regardless of patch or region - Ruled out if: The server computes or validates results itself: not this cause. Teleporting in that case points to packet loss or the interpolation buffer - Check with: Game server or client logs and metrics - Sources: - [Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization (Yahn W. Bernier, GDC 2001)](https://web.cs.wpi.edu/~claypool/courses/4513-B03/papers/games/bernier.pdf) · Valve · Having clients report results works only when clients can be trusted; authoritative servers exist because of cheating concerns - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · Server-authoritative model: the server never trusts the client’s view of the game state - [Distributed authority topologies (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/terms-concepts/distributed-authority.html) · Unity · Splitting authority across clients makes cheating easier and removes the single simulation that governs every entity #### sy-lockstep · Lockstep waiting on the slowest player · Lockstep waits for the slowest peer When everyone computes the same turn together, one player’s late input makes everyone wait. - Why → Effect → On screen: Each turn can be computed only after every player’s input has arrived → One player’s input arrives late because of jitter or packet loss → Everyone hitches at the same time, and in bad cases a “Waiting for players” window appears - Symptoms: Freeze, Stutter, Input lag / Factors: Jitter, Packet loss, Stall - Who: Specific zone/channel / When: Randomly - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: adjust input delay automatically to ping, drop only the lagging player for a moment so everyone else keeps going without waiting. Client: apply the agreed input delay, in P2P without a relay server have the host client also handle input delay adjustment and lagging players. - Ballpark numbers: If you set input delay shorter than “time for an input to reach the other player + jitter,” freezes become frequent. That time is half the ping when players exchange inputs directly, and about half the sum of both players’ pings through a relay server. - On the graph: Random spikes (Turn wait time, input arrival delay per player) - Where to look: Per turn, each player’s input arrival time and how long the turn stalled waiting, and whose input each stalled turn was waiting for. With a relay server, also visible in a server-side packet capture as the arrival interval of each player’s input packets - Confirmed if: In every stalled turn, the same one player’s input arrived later than the input delay, and that player’s jitter/loss spiked at the same time - Ruled out if: All inputs arrived on time but it still stalls: points to the slowest PC’s computation time or server processing. No stalls but results differ between two screens: a desync, so check “Pathfinding mismatch in command sync” - Check with: Game server or client logs and metrics - Sources: - [Deterministic Lockstep](https://gafferongames.com/post/deterministic_lockstep/) · Gaffer On Games · Frame n can be computed only after all its inputs arrive, so a late input means waiting; a small playout delay buffer for absorbing jitter causes hitches - [1500 Archers on a 28.8: Network Programming in Age of Empires and Beyond](https://www.gamedeveloper.com/programming/1500-archers-on-a-28-8-network-programming-in-age-of-empires-and-beyond) · Game Developer · Commands are scheduled to run two turns later, and turn length is adjusted to the slowest computer and ping (Speed Control) - [A Survey and Taxonomy of Latency Compensation Techniques for Network Computer Games (ACM Computing Surveys, 2022)](https://web.cs.wpi.edu/~claypool/papers/lag-taxonomy/paper.pdf) · ACM · Input delay (incoming delay) is set to the “A→server + server→B” latency so everyone applies inputs at the same moment #### sy-rollback · Rollback netcode misprediction · Rollback misprediction The game predicts the opponent’s input and shows it early, then rewinds and recomputes if the guess was wrong. The higher the ping, the further it rewinds. - Why → Effect → On screen: The opponent changes their input (different from the prediction) → The real input arrives half a ping late, so the game rewinds that far and recomputes → The opponent’s animation skips a few frames or changes suddenly - Symptoms: Teleporting / Factors: Latency, Jitter - Who: Just me / When: During specific actions, Randomly - Primary owner: Game team (Client development) - Game team action items: Mix in 1–3 frames of input delay to shorten rollbacks, cap the rollback length. - Ballpark numbers: At 100 ms ping (50 ms one way), that’s about 3 frames of rollback at 60 FPS. With 2 frames of input delay, it drops to 1 frame. - On the graph: Random spikes (Frames rolled back, RTT (ping)) - Where to look: Client log of the frame count for each rollback, the RTT at that time, the input delay setting, and the time spent rolling back and recomputing - Confirmed if: Rollback frame count is large at the moments the opponent’s animation jumped, and the average rollback is roughly (one-way latency − input delay) ÷ frame time, growing with ping - Ruled out if: Stutter even with short rollbacks: a performance problem where recomputation takes longer than one frame. Results still differ between the two screens after rollback: a desync - Check with: Game server or client logs and metrics - Sources: - [GGPO Rollback Networking SDK](https://www.ggpo.net/) · GGPO · The game predicts the opponent’s input and moves ahead; if the real input differs, it recomputes from the point where they diverged up to the present - [8 Frames in 16ms: Rollback Networking in 'Mortal Kombat' and 'Injustice 2'](https://www.gdcvault.com/play/1025471/8-Frames-in-16ms-Rollback) · GDC · Rollback removes lockstep’s local input delay, rewinding and recomputing up to 8 frames within 16 ms #### sy-no-timestamp · Events played on arrival without timestamps · Events played on arrival (no timestamps) If server events carry no timestamp and play as soon as they arrive, network jitter carries straight through into uneven animation timing. - Why → Effect → On screen: “Attack start” and “play effect” events run as soon as they arrive → Each packet arrives at a different time, so the intervals are uneven → Chained attack animations speed up and slow down, and boss pattern timing differs every time - Symptoms: Stutter, Fast-forward / Factors: Jitter - Who: Just me / When: Always - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: play events at the time attached to them (scheduled events, interpolation buffer). Server: attach the time the event happened (server time) when sending. - On the graph: Random spikes (Event playback interval, packet arrival interval) - Where to look: Event times from server logs and arrival and playback times from client logs, matched by event number, comparing intervals. Reproduce in a development build by adding jitter (the jitter value in tc netem, min/max latency in Unreal network emulation) - Confirmed if: Events happen at steady intervals on the server, but playback intervals follow the uneven arrival intervals exactly - Ruled out if: Arrival intervals are even but playback is uneven: client frame problem (frame time spikes). Intervals are already uneven on the server: tick overrun - Check with: Game server or client logs and metrics - Sources: - [Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization (Yahn W. Bernier, GDC 2001)](https://web.cs.wpi.edu/~claypool/courses/4513-B03/papers/games/bernier.pdf) · Valve · Stamp every update with server time and draw at the position for a target time of the current time minus the interpolation time (100 ms) - [Snapshot Interpolation](https://gafferongames.com/post/snapshot_interpolation/) · Gaffer On Games · Drawing received snapshots immediately stutters because of jitter; holding them briefly in an interpolation buffer before drawing makes motion smooth - [NetworkTime and ticks (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/advanced-topics/networktime-ticks.html) · Unity · Example of putting the send time in an RPC so the receiver plays the effect in step with server time - [tc-netem(8) — Linux manual page](https://man7.org/linux/man-pages/man8/tc-netem.8.html) · iproute2 · A test tool that adds delay and jitter (delay TIME JITTER) and loss (loss random PERCENT) to outgoing packets to mimic a real network - [Using Network Emulation in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/using-network-emulation-in-unreal-engine) · Epic Games · Test with minimum and maximum latency and a packet loss percentage on server and client; set from the console, e.g., NetEmulation.PktLag #### sy-double-tick · Double tick wait · Double tick quantization If requests wait for the next tick to be processed and the results wait for the tick after that to be sent, the tick interval is added twice. - Why → Effect → On screen: Received requests are processed on the next tick → Results are also batched and sent on the next send tick → Ping on the connection is low, but responses are consistently late by about 1.5 times the tick interval. On a 10-tick server, 0.15 s on average and 0.2 s at worst - Symptoms: Input lag / Factors: Latency - Who: Whole server / When: Always - Primary owner: Game team (Server development) - Game team action items: Send the response on the same tick that processes the request, raise the tick rate, send important responses immediately. - Ballpark numbers: On a 10-tick server one tick is 100 ms, so the tick waits alone add 150 ms on average and 200 ms at worst. With a single wait, it’s 50 ms on average. - On the graph: Always high (Time from request arrival to response send) - Where to look: Server-side packet capture while a test account repeats the same action (e.g., using an item), measuring the gap between the request packet’s arrival and the response packet’s departure. With server logs, request arrival time, the tick number that processed it, and the time the response was sent - Confirmed if: Time spent inside the server averages about 1.5 times the tick interval, up to about 2 times, and stays constant regardless of RTT - Ruled out if: Time inside the server is around half the tick interval on average: only one tick wait. Longer than the tick interval and uneven: check tick overrun - Check with: Infra tools (no game code needed) - Sources: - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · An arriving input waits up to one tick for the tick boundary, and applying and sending take another frame; this shrinks as the tick rate goes up - [VALORANT's 128-Tick Servers](https://www.riotgames.com/en/news/valorants-128-tick-servers) · Riot Games · Part of the latency comes from the network, part from the server tick rate - [NetworkTime and ticks (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/advanced-topics/networktime-ticks.html) · Unity · NetworkVariable changes aren’t sent immediately; they’re batched and sent on each network tick #### sy-strict-check · Overly strict server validation · Over-strict server validation If the server checks movement speed, cooldowns, and range too strictly, it rejects even valid inputs that arrive bunched together because of jitter. - Why → Effect → On screen: Strict rules such as “max distance per tick” or “0 ms cooldown tolerance” → When jitter makes two commands arrive in the same tick, they’re judged as rule violations → Rubber-banding, skills rejected even though the cooldown is up - Symptoms: Rubber-banding, Dropped action / rollback / Factors: Jitter - Who: Just me / When: Randomly, While moving or changing zones - Primary owner: Game team (Server development) - Game team action items: Check against an accumulated allowance (token bucket), leave slack for ping and jitter. - On the graph: Random spikes (Server validation rejections and position corrections) - Where to look: Server log for each validation rejection or position correction with the reason, how many commands from that player arrived in that tick, and the arrival gap from the previous command - Confirmed if: Rejections and corrections cluster at moments when 2 or more commands arrived in one tick, while movement and use counts summed over a few seconds stay within the rules - Ruled out if: Still over the limit even when summed over a few seconds: real speed hacking or cheating is possible. Rejections concentrated on one ISP in the evening: points to validation false positives concentrated on one ISP’s players - Check with: Game server or client logs and metrics - Sources: - [Source SDK 2013: player.cpp](https://raw.githubusercontent.com/ValveSoftware/source-sdk-2013/master/src/game/server/player.cpp) · Valve · A per-tick command processing budget that accumulates (up to sv_maxusrcmdprocessticks, 24 ticks) lets bunched-up commands through; a developer comment says stricter restrictions caused stutter even for legitimate players - [RFC 2697: A Single Rate Three Color Marker](https://www.rfc-editor.org/rfc/rfc2697) · IETF · Token bucket: judged by average rate (CIR) and the burst size allowed at once (CBS) #### sy-host · Player-hosted server (host) · Listen server / host advantage When one player’s PC acts as the server, that player’s connection and PC performance decide how the game feels for everyone. - Why → Effect → On screen: The host’s PC acts as the server (P2P, listen server) → If the host’s connection or PC is slow, everyone feels it, while the host has zero ping → Only the host has the advantage, and when the host leaves everyone freezes or disconnects - Symptoms: Stutter, Freeze, Disconnect / Factors: Latency, Stall - Who: Specific zone/channel / When: Randomly - Primary owner: Game team (Server development) / Also: Game team (Client development), Infra team (Server infrastructure) - Game team action items: Server: move to dedicated servers that make the calls, until then pick a player with a good connection and PC as host during matchmaking. Client: support host migration, measure and report ping to other participants, upload speed, and PC performance during matchmaking. - Infra team action items: Secure server machines or instances for dedicated servers, place them close to regions with many players. - On the graph: Outliers only (Lag reports and disconnects per host) - Where to look: Match logs with the host’s upload speed, each participant’s RTT to the host, the host PC’s frame time, and when the host left, with lag and disconnect reports grouped by host. Players can also confirm it by playing again with the same group and only the host changed - Confirmed if: Lag and disconnects concentrate in one host’s matches, all participants get worse together when that host’s upload speed is low or frame time is long, and without host migration everyone disconnects the moment the host leaves - Ruled out if: Only participants in the same region are affected, regardless of host: connection or route problem. With dedicated servers, this isn’t the cause - Check with: Game server or client logs and metrics - Sources: - [Networking Overview for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/networking-overview-for-unreal-engine) · Epic Games · A listen server’s host has an advantage over other clients and carries a heavy load, running both the server and rendering - [Distributed authority topologies (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/terms-concepts/distributed-authority.html) · Unity · When the session owner leaves, a new owner is elected automatically from the remaining clients - [A Survey and Taxonomy of Latency Compensation Techniques for Network Computer Games (ACM Computing Surveys, 2022)](https://web.cs.wpi.edu/~claypool/papers/lag-taxonomy/paper.pdf) · ACM · P2P games can also be seen as client-server, with the host doubling as the server #### sy-optimistic-reject · Server rejects after client-side feedback · Client-side feedback rejected by server When the server later refuses a hit or skill your screen already showed, the result you clearly saw never happened. - Why → Effect → On screen: Hit effects and skill animations play before the server confirms (client-side feedback) → The server rechecks range, target position, cooldown, and resources and rejects the action → Blood sprays but there’s no damage, the skill animation plays with no effect, only the cooldown runs - Symptoms: Dropped action / rollback, Rubber-banding / Factors: Latency - Who: Just me, One feature only / When: During specific actions - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: show only what needs confirming (damage numbers, deaths, rewards) from server results, check common rejection reasons locally first, when rejected refund the cooldown and resources and show the reason. Server: leave slack for ping in range and target position checks, include the reason in rejection responses, collect rejection rate per skill as a metric. - Ballpark numbers: The rejection arrives ping + tick wait after the press. At 150 ms ping, you think you hit for about 0.2 s. - On the graph: Outliers only (Server rejection rate per skill (by ping bracket)) - Where to look: Server-side rejection rate and rejection reasons (range, target position, cooldown, resources) per skill, broken down by player RTT bracket. On the client, a count of actions shown with client-side feedback that were then rejected - Confirmed if: Rejections concentrate on specific skills with range or target position as the reason, and the rejection rate rises with ping - Ruled out if: Rejection reasons are cooldown or resources and unrelated to ping: check whether client and server data values (cooldown, cost) differ. No rejections but feedback starts only after the server responds: points to “Feedback only after the server responds (request-response)” - Check with: Game server or client logs and metrics - Learn more: Client-side feedback is the best way to hide ping. However, the more the information the client and server use to decide (enemy position, remaining resources) differs, the more often rejections happen. Collecting rejection rate per skill as a metric makes it easy to find where the calls disagree. - Sources: - [Using Gameplay Abilities in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/using-gameplay-abilities-in-unreal-engine) · Epic Games · Local Predicted abilities run immediately on the client, but the server makes the final decision and can reverse the result - [Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization (Yahn W. Bernier, GDC 2001)](https://web.cs.wpi.edu/~claypool/courses/4513-B03/papers/games/bernier.pdf) · Valve · Weapon fire is predicted on the client and its effects play first; server results then correct prediction errors #### sy-path-mismatch · Pathfinding mismatch in command sync · Command sync with divergent pathing When only “go here” is exchanged and each side computes the path itself, even a small difference in the calculation sends a character or monster down a different path until it gets pulled back into place. - Why → Effect → On screen: With click-to-move and monster chasing, only the destination is sent and the client computes the path on its own → Differences in terrain data, collisions with other characters, or calculation order make it move along a different path from the server’s → Monsters walk through walls and then snap to another spot, clicked characters change direction as if sliding - Symptoms: Teleporting, Rubber-banding / Factors: Latency - Who: Specific zone/channel, Just me / When: While moving or changing zones - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: send the path’s intermediate points (waypoints) as well, sync positions periodically. Client: converge smoothly when positions diverge, use the same terrain data as the server. - On the graph: Random spikes (Position correction count and distance per entity) - Where to look: Per entity, the difference between the position the server sent and the position the client computed, with the coordinates of corrections plotted on the map. Periodically comparing checksums of both sides’ path results or positions pinpoints when they started to diverge - Confirmed if: Corrections cluster at specific terrain (ledges, narrow passages, slopes) or crowded spots and repeat at the same place even for players with normal network metrics - Ruled out if: Corrections happen only when loss or jitter spikes, regardless of place: connection problem. A single monster jumping on several players’ screens at once: check whether a slow client has control of the monster - Check with: Game server or client logs and metrics - Learn more: This approach is one reason click-to-move and tab-target games are less sensitive to ping. The trade-off is that nothing guarantees both sides get the same result, so a mechanism that syncs positions now and then is essential. Floating-point results can differ slightly depending on CPU type, compiler, and optimization settings (including debug versus release builds). In designs that exchange only inputs and assume both sides compute identical results, such as lockstep and rollback, these tiny differences can accumulate until the game state on the two screens splits apart (desync). - Sources: - [Deterministic Lockstep](https://gafferongames.com/post/deterministic_lockstep/) · Gaffer On Games · Even if deterministic on the same machine, floating-point results can differ across compilers, OSes, and CPUs - [State Synchronization](https://gafferongames.com/post/state_synchronization/) · Gaffer On Games · Sending state along with inputs keeps both sides in sync without perfect determinism - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · Packet loss, or two characters trying to move to the same spot, makes server and client simulations diverge and need correction - [1500 Archers on a 28.8: Network Programming in Age of Empires and Beyond](https://www.gamedeveloper.com/programming/1500-archers-on-a-28-8-network-programming-in-age-of-empires-and-beyond) · Game Developer · Tiny differences grow over time until a worker’s path drifts off; the world, entities, and pathfinding are compared by checksum to catch out-of-sync states - [Floating Point Determinism](https://gafferongames.com/post/floating_point_determinism/) · Gaffer On Games · The same floating-point code can give different results depending on compiler, CPU architecture, and debug or release build; includes a case where AMD and Intel CPUs returned slightly different values from transcendental functions - [/fp (Specify floating-point behavior)](https://learn.microsoft.com/en-us/cpp/build/reference/fp-specify-floating-point-behavior) · Microsoft · /fp:fast may reorder or combine floating-point operations and give results different from other /fp settings, and operations fused with FMA can also differ from a separate multiply and add #### sy-low-send-rate · Low snapshot send rate · Low snapshot / update rate If the server sends position updates (snapshots) only a few times a second, the interpolation buffer has to be that much longer, and you see other characters further in the past. - Why → Effect → On screen: Position updates sent only 5–10 times a second to save bandwidth → Smooth rendering needs a buffer of twice the packet interval (200–400 ms); with a shorter buffer, a single missed packet causes a freeze → Opponents’ direction changes show up late and disagree with hit registration. With a short buffer: stutter, plus teleporting on packet loss - Symptoms: Stutter, Teleporting, Dropped action / rollback / Factors: Latency, Packet loss - Who: Whole server / When: Always - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: send nearby or in-combat targets often and distant ones rarely, send only what changed (delta compression) to shrink each update and raise the rate. Client: adjust the interpolation buffer length automatically to the packet interval. - Ballpark numbers: At 10 per second, the packet interval is 100 ms and the buffer 200 ms. Add the 75 ms one-way latency of a 150 ms ping, and you see opponents about 0.3 s in the past. - On the graph: Always high (Packet arrival interval per client, interpolation buffer length) - Where to look: Server-side packet capture filtered to the flow going to one player, with packets per second and intervals in Wireshark I/O Graphs. With game logs, update interval per entity together with the client’s interpolation buffer headroom (time left until the next snapshot arrives) - Confirmed if: Position updates are always sparse at 5–10 per second (100–200 ms apart), and the interpolation buffer is set above 200 ms or its headroom often hits 0 - Ruled out if: Updates go out densely but only the arrival intervals wobble: points to jitter or loss. Only distant entities arrive rarely when it’s crowded: per-connection send budget and priority - Check with: Infra tools (no game code needed) - Sources: - [Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization (Yahn W. Bernier, GDC 2001)](https://web.cs.wpi.edu/~claypool/courses/4513-B03/papers/games/bernier.pdf) · Valve · At 10 updates per second, 200 ms of interpolation survives one missed update. Half-Life defaults to 20 per second with 100 ms interpolation - [Snapshot Interpolation](https://gafferongames.com/post/snapshot_interpolation/) · Gaffer On Games · At 10 per second, surviving two consecutive losses needs 350 ms of delay; at 30 per second, it drops to 150 ms - [State Synchronization](https://gafferongames.com/post/state_synchronization/) · Gaffer On Games · Priority accumulation sends important entities more often and rotates through the rest within the bandwidth limit - [8.8. The “I/O Graphs” Window](https://www.wireshark.org/docs/wsug_html_chunked/ChStatIOGraphs.html) · Wireshark · Graphs the packet count and bytes matching a display filter per time interval ### Problems only some players hit (causes: 24) #### pt-slow-burst · A lagging player moves in bursts on others’ screens · Laggy player seen by others (bursty inputs) Inputs from a player with a bad connection reach the server unevenly, in bunches. If the server applies whatever arrived on each tick, other players see that character hitch and then cover several steps at once. - Why → Effect → On screen: The lagging player’s move commands arrive 0 at a time on some ticks and 2–3 at a time on others → The server applies them all on the tick they arrive, so that character’s position changes in steps → On other players’ screens, only that character hitches and then moves several steps at once. Everyone else looks fine - Symptoms: Fast-forward, Teleporting / Factors: Jitter, Packet loss - Who: One character looks off, Specific region/ISP / When: Always, While moving or changing zones - Primary owner: Game team (Server development) / Also: External (External) - Game team action items: Spread commands evenly with a per-player input buffer, apply them at their original spacing using input sequence numbers, lengthening only the interpolation buffer on other players’ screens isn’t enough (the server’s own position history is already stair-stepped). - External action items: Tell lagging players to use a wired connection and check their Wi-Fi and router. - Ballpark numbers: With 80 ms of jitter on a 20-tick (50 ms) server, commands per tick swing between 0 and 3. - On the graph: Outliers only (Commands applied per tick per player, jitter per player) - Where to look: Server-side packet capture filtered to packets from the reported player, counting how many arrived in each tick interval (e.g., 50 ms) and comparing with other players. With server logs, move commands applied per tick and input sequence numbers per player - Confirmed if: Only the reported player’s packets arrive in bunches, alternating between 0 and 2–3 per tick, that player’s jitter/loss is high, and other players’ packets arrive evenly. It improves when that player switches to wired - Ruled out if: Several characters move in bursts at once: server tick delay or the viewer’s own connection. Arrival and application are even but that character still looks jumpy: interpolation or display problem on the viewer’s side - Check with: Infra tools (no game code needed) - Learn more: In a server-authoritative design, this is normal behavior. One lagging player’s lag shows up to others only as “that player moving strangely” and doesn’t affect anyone else’s controls or monster movement. Anything that involves that player directly (trades, party mechanics, PvP hit registration), however, is delayed along with them. - Sources: - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · The server puts arriving inputs into a per-player move queue in tick order and fills gaps with prediction; corrections are visible only to that player, and everyone else sees smooth motion - [State Synchronization](https://gafferongames.com/post/state_synchronization/) · Gaffer On Games · Even packets sent 60 times a second arrive bunched, e.g., 2 in one frame and 0 in the next - [Source SDK 2013: player.cpp](https://raw.githubusercontent.com/ValveSoftware/source-sdk-2013/master/src/game/server/player.cpp) · Valve · Commands that arrive bunched are spread out (metered out) across server ticks #### pt-event-server · Fast-forward on servers that process on arrival · Event-driven processing of bursty inputs On a server that processes and broadcasts packets as soon as they arrive, a lagging player’s bunched-up actions run back to back immediately. - Why → Effect → On screen: A lagging player’s skill and move requests arrive in a bunch → The server runs them in order the moment they arrive and tells everyone right away → Others see that player use several skills in an instant or move as if fast-forwarding - Symptoms: Fast-forward / Factors: Jitter - Who: One character looks off / When: During specific actions, Always - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: run actions at the spacing of their attached input times (accepting those times only within an allowed range), or run bunched actions one after another spaced by a minimum interval (global cooldown) without rejecting them, don’t check cooldowns by arrival time alone (valid inputs get dropped). Client: attach the input time to each action. - On the graph: Outliers only (Action execution interval per player) - Where to look: Server log of each player’s action arrival time, execution time, and client-attached input time (if any), comparing execution intervals with input intervals. Also the arrival intervals of that player’s packets in a server-side packet capture - Confirmed if: Input intervals are normal, but server arrival and execution intervals are bunched within a few ms, and the bunches line up with the times other players reported fast-forward - Ruled out if: Intervals are already bunched in the input times: points to the client or a macro. Server execution intervals are even but look bunched only on other players’ screens: the viewer’s connection - Check with: Game server or client logs and metrics - Sources: - [Understanding Networked Movement in the Character Movement Component for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/understanding-networked-movement-in-the-character-movement-component-for-unreal-engine) · Epic Games · The server computes a move each time it receives ServerMove and sets the time step from the timestamp difference with the previous move; if the difference from server time is too large, it discards the move - [Deterministic Lockstep](https://gafferongames.com/post/deterministic_lockstep/) · Gaffer On Games · Applying inputs as they arrive gives uneven results even when they’re sent at 60 Hz, because the spacing isn’t even #### pt-input-buffer · Per-player input buffer size · Per-player server input buffer (jitter buffer) If the server holds a few inputs per player and takes out one per tick, others see smooth motion, but your own actions are confirmed on the server that much later. - Why → Effect → On screen: The server collects a lagging player’s inputs in a buffer and applies one per tick → A small buffer often runs empty, so the character stands still or the server guesses from the last input; a large buffer confirms the player’s own inputs late → Too small: others see hitches. Too large: your own skill results come late (input lag) - Symptoms: Stutter, Input lag / Factors: Jitter - Who: One character looks off, Just me / When: Always - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: size each player’s buffer automatically to their connection, take two inputs at a time to catch up when behind, tell clients whose buffers often run empty to send inputs earlier. Client: send inputs slightly earlier as the server instructs (client time adjustment). - Ballpark numbers: It varies by game, but 1–3 ticks’ worth is typical. VALORANT keeps the server buffer shorter, averaging half a frame (about 4 ms) on 128-tick servers. Adaptive buffers that grow only for players with high jitter are common. - On the graph: Outliers only (Input buffer length and empty count per player) - Where to look: Server-side, per player per tick: inputs left in the input buffer, times the buffer ran empty and was filled with a guess from the last input, and time from input arrival to application - Confirmed if: Players with small buffers run empty often and briefly freeze on others’ screens at those moments, and players with large buffers have input-to-application times longer by the buffer length - Ruled out if: The buffer almost never runs empty but others see stutter: interpolation problem on the viewer’s side. Short buffer but still high input lag: RTT itself or a double tick wait - Check with: Game server or client logs and metrics - Sources: - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · The server adjusts the client’s time base so the input queue stays just long enough to absorb uneven arrivals with minimal latency; the server buffering target averages half a frame - [NetworkTimeSystem class (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/api/Unity.Netcode.NetworkTimeSystem.html) · Unity · LocalBufferSec: how long the server buffers client messages; client time is moved ahead so messages reach the server earlier - [NetworkTime and ticks (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/advanced-topics/networktime-ticks.html) · Unity · Players with bad connections can raise the buffer value #### pt-isp-validation · Validation false positives concentrated on one ISP · Anti-cheat / movement validation false positives on bad ISPs Players on high-jitter connections have their inputs arrive in bunches, so they often trip the server’s speed and cooldown checks. - Why → Effect → On screen: Jitter on a specific ISP’s or region’s connections rises in the evening → The server judges valid inputs that arrived in a bunch as speeding or cooldown violations → Only that ISP’s players get rubber-banding and rejected skills, and in bad cases the server kicks them (disconnect) - Symptoms: Rubber-banding, Dropped action / rollback, Disconnect / Factors: Jitter - Who: Specific region/ISP, Just me / When: Evening peak hours, While moving or changing zones - Primary owner: Game team (Server development) / Also: Infra team (Network infrastructure) - Game team action items: Check against an allowance accumulated over several seconds, loosen thresholds based on connection quality (ping, jitter), add a warning stage before a forced disconnect, spread bunched inputs evenly across ticks with a per-player input buffer to cut false positives at the source. - Infra team action items: Check loss rate and jitter distribution per ISP by time of day and share them with the game team, check the path through that ISP (bidirectional mtr), reroute or escalate to the ISP if needed. - On the graph: High at certain hours (Validation rejections and forced disconnects per ISP (ASN), jitter per ISP) - Where to look: Server logs of validation rejections, corrections, and forced disconnects tagged with the ISP (ASN) of the connecting IP and the time, counted by ISP and time of day. The infra team runs bidirectional mtr toward that ISP at the same time to check jitter and loss - Confirmed if: Rejections and forced disconnects concentrate on one ISP and rise in the evening, that ISP’s jitter is high at the same time, and movement summed over a few seconds stays within the rules - Ruled out if: Only specific accounts repeat regardless of ISP: real cheating possible. Rising across all ISPs together: a server-side cause where lagging server ticks apply commands in bunches (tick overrun) - Check with: Game server or client logs and metrics - Sources: - [Source SDK 2013: player.cpp](https://raw.githubusercontent.com/ValveSoftware/source-sdk-2013/master/src/game/server/player.cpp) · Valve · A per-tick command budget stops speed hacks; a developer comment says stricter limits cause stutter for legitimate players too - [RFC 2697: A Single Rate Three Color Marker](https://www.rfc-editor.org/rfc/rfc2697) · IETF · Token bucket: judged by average rate and allowed burst size - [Understanding Networked Movement in the Character Movement Component for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/understanding-networked-movement-in-the-character-movement-component-for-unreal-engine) · Epic Games · When client and server timestamps differ by too much, the move is discarded or handled by a time discrepancy resolution procedure; computing with server time prevents speed hacks #### pt-raid-member · One lagging party member and boss mechanics · One laggy member in a synchronized mechanic In raid mechanics where everyone has to react together at a set moment, one lagging player’s late reaction fails the whole party. - Why → Effect → On screen: Group mechanics such as “everyone spread out at once” or “one player presses the button” → The lagging player sees the telegraph late, and their input also arrives late → One player causes a wipe, and the rest of the party feels it happened “because of the laggy player” - Symptoms: Dropped action / rollback, Input lag / Factors: Latency - Who: Specific zone/channel, One character looks off / When: When crowds gather, During specific actions - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: give mechanic timing windows slack for ping, send telegraphs ahead with server time, design mechanics so one player’s failure doesn’t wipe the party. Client: play received telegraphs in step with server time. - On the graph: Outliers only (RTT of each player who caused a mechanic failure) - Where to look: Server mechanic log with the player who caused the failure, their input arrival time, the timing window, and their RTT and loss - Confirmed if: Most of the inputs that caused wipes come from the same one player, whose RTT is clearly higher than the party average and whose inputs arrive just after the timing window - Ruled out if: Failures are spread evenly across party members: the window itself is too short (“Short timing windows eaten up by ping”). The lagging player’s input arrived inside the window and still failed: the server’s window-checking code - Check with: Game server or client logs and metrics - Sources: - [NetworkTime and ticks (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/advanced-topics/networktime-ticks.html) · Unity · Scheduling events at a server time so everyone plays them at the same moment - [Factors influencing the latency of simple reaction time (Frontiers in Human Neuroscience, 2015)](https://doi.org/10.3389/fnhum.2015.00131) · Frontiers · Average simple reaction time about 231 ms #### pt-mob-control · A lagging client controls the monster · Monster movement delegated to a player client Some games hand monster movement to one nearby player’s client to reduce server load. If that player’s connection is bad, the monster moves strangely on everyone’s screen. - Why → Effect → On screen: The server hands monster movement to the client of the nearest (or first-arriving) player → That player’s reports reach the server late or in bunches → Only that monster hitches and then teleports on every nearby screen. It looks fine on the controlling player’s own screen - Symptoms: Teleporting, Stutter, Fast-forward / Factors: Jitter, Packet loss - Who: One character looks off, Specific zone/channel / When: Always, Randomly - Primary owner: Game team (Server development) - Game team action items: Hand control to a player with a good connection (by ping and loss), have the server take control back immediately when reports stop, have the server compute important monsters such as bosses itself. - On the graph: Outliers only (Position report interval per monster (by controlling client)) - Where to look: Server-side, per monster: the client with control and that client’s report interval, RTT, and loss. The arrival interval of that client’s packets is also visible in a server-side packet capture - Confirmed if: Every monster that moves strangely is controlled by the same one player, that player’s reports are uneven or stop, and handing control to someone else fixes it right away - Ruled out if: Monsters the server computes itself jump the same way: server tick delay or the viewer’s connection. Still jumping after control moves: pathfinding mismatch in command sync - Check with: Game server or client logs and metrics - Learn more: The controlling player sees nothing wrong, so reports only say “the monster is acting weird.” If everyone except one player sees the same monster moving strangely, first check who has control of that monster. - Sources: - [Authority (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/terms-concepts/authority.html) · Unity · In a distributed authority model, each game instance (client) takes authority over some network objects and simulates them - [Distributed authority topologies (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/terms-concepts/distributed-authority.html) · Unity · Splitting authority across clients leaves no single simulation and makes cheating easier #### pt-heavy-char · Bloated data on one character · One character with oversized data (inventory, mail, buffs) A character with thousands of items or mails piled up, or an unusually large friend list, block list, or set of buffs, has several times more to load, save, and announce to nearby players than others. It’s slow only on that character, regardless of connection. - Why → Effect → On screen: Thousands of items or event rewards pile up in the inventory and mailbox of a long-played character → Every login, zone move, and save reads and writes that much from the DB, and the equipment and buff data sent to nearby players is large too → Only that character has long loading screens and hitches when opening the inventory or mail. On a server where the game thread waits on saves, nearby players freeze briefly too - Symptoms: Can’t connect / infinite loading, Input lag, Freeze / Factors: Stall - Who: Just me, One feature only / When: Right after login or maintenance, During specific actions, While moving or changing zones - Primary owner: Game team (Server development) / Also: Infra team (DB infrastructure) - Game team action items: Cap inventory and mail storage and auto-clean old mail, load only the parts you need in pieces, save only what changed and do it off the game thread. - Infra team action items: Find slow queries that repeat for the same character in the slow query log and pass them to the game team, provide a top list of characters with the most item and mail rows. - Ballpark numbers: If one item is one DB row, a character with 5,000 items reads 5,000 rows on every login. That’s tens of times more than an ordinary character. - On the graph: Outliers only (Login and save time per character, DB rows read per character) - Where to look: Slow reads and saves that repeat for the same character ID in the DB slow query log (MySQL slow query log, PostgreSQL log_min_duration_statement), plus a top list of per-character row counts in the item and mail tables - Confirmed if: Slow queries concentrate on a few character IDs, those characters have tens of times the average number of item and mail rows, and it’s just as slow from another PC or connection - Ruled out if: Other characters on the same account or other players are slow too: DB hardware or locking. The character is fine on another PC: the player’s environment - Check with: Infra tools (no game code needed) - Learn more: If the same character is just as slow from a different PC and connection while other characters on the same account are fine, suspect the character data. That’s why reports need the character name. - Sources: - [Extraneous Fetching antipattern](https://learn.microsoft.com/en-us/azure/architecture/antipatterns/extraneous-fetching/) · Microsoft Azure · Fetching more data than needed increases I/O load and slows responses - [PostgreSQL Documentation: Error Reporting and Logging](https://www.postgresql.org/docs/current/runtime-config-logging.html) · PostgreSQL · log_min_duration_statement: logs SQL statements that run longer than a set time, for tracking slow queries - [The Slow Query Log](https://dev.mysql.com/doc/refman/8.4/en/slow-query-log.html) · MySQL · Slow query log that records queries exceeding long_query_time #### pt-phase · Different channel, instance, or phase · Different channel / instance / phase If two characters are in different channels or instances, or in different “phases” where the visible NPCs depend on quest progress, they see different worlds. - Why → Effect → On screen: The second character is assigned to a different channel, or its quest stage differs → The server doesn’t send that NPC to that character (working as intended) → The NPC is missing on one side only. It looks like a bug but is by design - Symptoms: Invisible / ghost entities / Factors: Packet loss - Who: One client on the same PC, Just me / When: Always, Right after login or maintenance - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: send channel and phase info to the client, add “check both characters’ channel and quest stage” to the QA checklist. Client: show the channel and phase on screen. - On the graph: Outliers only (Nearby entity count per client, channel/phase) - Where to look: The two characters’ channel numbers and progress on the relevant quest compared in the game, then checked again with both on the same channel and stage. With server entity send logs, why that NPC wasn’t sent to that character (channel, phase) - Confirmed if: The two characters differ in channel or quest stage, and the NPC shows up once they match - Ruled out if: Same channel and stage but the NPC is missing on one side only: points to spawn messages dropped during loading, spawn data lost in the burst after entering, or an AOI registration race - Check with: The player’s own environment - Learn more: Also check whether quest progress is saved per account or per character. With two characters on the same account, one character’s progress can change the other’s phase. - Sources: - [Actor Relevancy in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-relevancy-in-unreal-engine) · Epic Games · The server replicates only relevant actors to each connection and doesn’t send irrelevant ones - [Object visibility (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/basics/object-visibility.html) · Unity · CheckObjectVisibility decides which objects each client can see, and hidden objects aren’t sent to that client #### pt-loading-drop · Spawn messages dropped during loading · Spawn messages dropped before the client is ready Right after you enter a zone, the server sends spawn messages for nearby NPCs, but the client is still loading the map and throws them away. - Why → Effect → On screen: The server sends spawn messages for nearby entities right after processing the entry → The client is still loading and has no message handler yet, so it drops the messages → The server treats them as sent and never resends them. The NPC stays invisible until it leaves view and comes back - Symptoms: Invisible / ghost entities / Factors: Packet loss - Who: One client on the same PC, Just me / When: Right after login or maintenance, While moving or changing zones - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: send “ready” when loading finishes, or hold packets received during loading and process them afterward. Server: send nearby information only after receiving “ready.” - Ballpark numbers: If two clients load at the same time on the same PC, or the loading one is a background window, they share CPU and disk and processing is throttled, so that client’s loading can take several times longer. The same bug also surfaces when the server gets faster at processing entries. - On the graph: Outliers only (Loading time per client, messages dropped during loading) - Where to look: Count and type of messages the client received and dropped during loading and the time loading finished, compared with the time the server sent the spawn messages. Easy to reproduce by loading two clients at once on the same PC or leaving the loading one as a background window - Confirmed if: The server sent the spawn message for the missing NPC, it arrived before loading finished, and the dropped-message count rose at that time. Happens only on the client with the longer loading - Ruled out if: Spawn message arrived after loading finished but the NPC is still invisible: lost baseline snapshot or entity ID reuse mix-up. The server never sent that NPC’s message at all: AOI registration race or a channel/phase difference - Check with: Game server or client logs and metrics - Sources: - [NetworkConfig class (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/api/Unity.Netcode.NetworkConfig.html) · Unity · SpawnTimeout: messages for objects not yet spawned are held, then dropped if the object isn’t spawned in time - [Actor Relevancy in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-relevancy-in-unreal-engine) · Epic Games · Actors that are no longer relevant are destroyed on the client and replicated anew when they become relevant again #### pt-aoi-race · AOI registration race · Interest-management race on enter/leave If a character registers in the AOI grid at the same moment an NPC moves between grid cells, that NPC’s spawn message can be missed. - Why → Effect → On screen: Processing an entry, channel change, or teleport happens at the same instant an NPC moves → That NPC is left out of the “newly visible entities” calculation → A few specific NPCs are invisible, or NPCs that already left are still there - Symptoms: Invisible / ghost entities / Factors: Packet loss - Who: One client on the same PC, Just me / When: While moving or changing zones, Randomly - Primary owner: Game team (Server development) - Game team action items: Handle AOI updates on one thread in one order, periodically resync the entire “visible list.” - On the graph: Random spikes (Mismatches between the server’s visible list and the client’s entity list) - Where to look: Server log of AOI grid registration, entity cell moves, and spawn/despawn message sends with tick numbers, plus periodic comparison of the server’s “visible list” with the list the client holds - Confirmed if: The missing NPC changed cells in the same tick as that character’s entry or teleport, and there’s no record of a spawn message sent for that NPC - Ruled out if: The spawn message was sent but the client didn’t receive it or dropped it: delivery side (spawn data lost in the burst after entering, spawn messages dropped during loading). Always the same NPC missing: phase or display settings difference - Check with: Game server or client logs and metrics - Sources: - [Replication Graph in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/replication-graph-in-unreal-engine) · Epic Games · MMORPGs and similar games split the world into a grid with an actor list per cell and send based on the cell the client is in - [Actor Relevancy in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-relevancy-in-unreal-engine) · Epic Games · Relevancy is decided per connection, and actors that are no longer relevant are destroyed on the client #### pt-baseline · Lost baseline snapshot · Lost baseline for delta compression When the server sends “only what changed since last time,” losing the full state sent once at the start (the baseline) means later changes can’t be applied. - Why → Effect → On screen: The packet with an entity’s full state (baseline) is lost or dropped before processing → The client has nothing to apply later changes to, so it ignores them → That entity is invisible, or suddenly appears much later - Symptoms: Invisible / ghost entities, Teleporting / Factors: Packet loss - Who: One client on the same PC, Just me / When: Randomly - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: always resend the baseline until an acknowledgment (ACK) arrives, build changes only against a baseline the client has acknowledged. Client: send the ACK for a baseline only after actually applying it, ask the server again when changes arrive for an unknown entity. - On the graph: Random spikes (Changes received for unknown entities) - Where to look: Times the client dropped changes that arrived without a baseline, with entity IDs, matched against when the server sent that entity’s baseline and when the ACK came back. Reproduce in development by adding loss (loss in tc netem, packet loss percentage in Unreal network emulation) - Confirmed if: For the invisible entity, the server sent a baseline and never got an ACK, yet kept sending only changes, which the client dropped - Ruled out if: The baseline was ACKed and applied on the client but the entity is still invisible: missed despawn message or entity ID reuse mix-up - Check with: Game server or client logs and metrics - Sources: - [Snapshot Compression](https://gafferongames.com/post/snapshot_compression/) · Gaffer On Games · Changes must be built only against a baseline the other side has acknowledged (acked), and the initial state is sent separately - [Quake III Arena source: code/server/sv_snapshot.c](https://raw.githubusercontent.com/id-Software/Quake-III-Arena/master/code/server/sv_snapshot.c) · id Software · Delta compression uses the snapshot the client acknowledged as its baseline, and a full snapshot is sent when the baseline gets too old - [tc-netem(8) — Linux manual page](https://man7.org/linux/man-pages/man8/tc-netem.8.html) · iproute2 · A test tool that adds delay and jitter (delay TIME JITTER) and loss (loss random PERCENT) to outgoing packets to mimic a real network - [Using Network Emulation in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/using-network-emulation-in-unreal-engine) · Epic Games · Test with minimum and maximum latency and a packet loss percentage on server and client; set from the console, e.g., NetEmulation.PktLag #### pt-ghost · Missed despawn message (ghost entity) · Missed despawn (ghost entity) The reverse case: if the “it’s gone” message is missed, NPCs or players that already died or left stay on your screen only. - Why → Effect → On screen: Death, leave, or out-of-view messages are lost or arrive out of order → The client thinks the entity is still there → A monster that doesn’t react when hit, a player who already left still standing there - Symptoms: Invisible / ghost entities / Factors: Packet loss - Who: One client on the same PC, Just me / When: Randomly, When crowds gather - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: periodically send the “currently visible list.” Client: remove entities not on the list, hide entities that should be moving but haven’t updated in a long time. - On the graph: Random spikes (Entities that exist only on the client) - Where to look: The server’s “currently visible list” compared with the client’s entity list, counting entities that exist only on the client, and despawn message send and receive logs matched by entity ID - Confirmed if: The server sent the ghost entity’s despawn message but the client has no record of receiving it, or the despawn arrived before the spawn and the order is reversed - Ruled out if: The entity is still on the server’s visible list too: the server failed to clean up the entity. Right after a new entity appeared with the same ID: entity ID reuse mix-up - Check with: Game server or client logs and metrics - Sources: - [Actor Relevancy in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-relevancy-in-unreal-engine) · Epic Games · Dynamic actors that are no longer relevant are deleted on the client - [Object visibility (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/basics/object-visibility.html) · Unity · When an object is hidden, that client despawns and destroys it #### pt-spawn-burst · Spawn data lost in the burst after entering · Initial spawn burst lost (unreliable channel, receive buffer, fragmentation) The moment you enter a zone, the server sends spawn data for tens to hundreds of nearby entities all at once. If it goes over an unreliable channel, or the receive buffer overflows while the client is loading and can’t read the socket, part of it disappears and never comes back. - Why → Effect → On screen: Spawn data arrives in a short burst right after entering → A loading client reads the socket late and the OS receive buffer overflows, or a large UDP packet is fragmented and losing a single fragment loses the whole packet. On an unreliable channel, nothing is resent either → A few NPCs are missing only on the slower-loading client. They show up after leaving view and coming back - Symptoms: Invisible / ghost entities / Factors: Packet loss - Who: One client on the same PC, Just me / When: Right after login or maintenance, While moving or changing zones - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: always send spawn and despawn messages over a reliable channel with guaranteed retransmission, send initial data in chunks. Client: receive on a thread separate from loading, increase the receive buffer size. - Ballpark numbers: The default UDP receive buffer on a PC varies by OS but is usually tens to hundreds of KB. If the entry data for a crowded town is bigger than that, even a brief pause in reading the socket during loading overflows it. - On the graph: Surge after opening (Data received right after entering, missed spawn messages) - Where to look: Spawn messages the server sent right after entry compared with the number the client received, and which channel (reliable or unreliable) they went over. In a server-side packet capture, the volume sent to that player right after entry and fragmented packets (Wireshark filter ip.flags.mf == 1 || ip.frag_offset > 0) - Confirmed if: The client received fewer than were sent, the missing ones cluster in the burst right after entry, and they went over an unreliable channel or large packets were fragmented. Happens more often on the slower-loading client - Ruled out if: Sent and received counts match but entities are still invisible: dropped after receipt (spawn messages dropped during loading) or an AOI calculation problem. Missing at random times unrelated to entry: packet loss on the connection - Check with: Game server or client logs and metrics - Sources: - [RFC 8085: UDP Usage Guidelines](https://www.rfc-editor.org/rfc/rfc8085) · IETF · Losing just one fragment of a fragmented packet loses the whole packet - [UDP vs. TCP](https://gafferongames.com/post/udp_vs_tcp/) · Gaffer On Games · UDP doesn’t guarantee delivery or order, so the application must detect and resend lost packets itself - [Socket.ReceiveBufferSize Property](https://learn.microsoft.com/en-us/dotnet/api/system.net.sockets.socket.receivebuffersize) · Microsoft · The default socket receive buffer size varies by OS - [Display Filter Reference: Internet Protocol Version 4](https://www.wireshark.org/docs/dfref/i/ip.html) · Wireshark · Filters fragmented IP packets with ip.flags.mf (More fragments) and ip.frag_offset (Fragment Offset) #### pt-id-reuse · Entity ID reuse mix-up · Entity ID reused without a generation counter If the server reuses the same entity ID when a dead NPC respawns, a client that missed the despawn message in between mistakes the new NPC for the old one. - Why → Effect → On screen: An NPC dies and respawns with the same entity ID → A client that missed the despawn message ignores the spawn message as an “already known entity,” or leaves the entity in its dead state → The NPC is missing on one screen only or appears lying dead, and sometimes shows up looking like a different NPC - Symptoms: Invisible / ghost entities / Factors: Packet loss - Who: One client on the same PC, Just me / When: Randomly, When crowds gather - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: add a generation number to entity IDs to tell reuses apart. Client: when a spawn message arrives for a known ID, delete the existing entity and create a new one. - On the graph: Random spikes (Spawn messages for already known IDs) - Where to look: Server log of creation and deletion times per entity ID (with generation number if any), plus the count of spawn messages the client received for known IDs and of delete-and-recreate cases treated as “no change” during AOI updates - Confirmed if: The invisible or dead-looking NPC has the same ID as an NPC that just died, and in between the client didn’t receive the despawn message or the server sent neither the despawn nor the spawn message - Ruled out if: IDs carry a generation number that is also used in comparisons: not this cause. Invisible even though the ID wasn’t reused: points to a lost spawn message - Check with: Game server or client logs and metrics - Learn more: It happens on the server side too. If the list of entities in view is compared by ID only, an NPC that died and respawned with the same ID between two AOI updates looks like “no change,” and neither the despawn nor the spawn message is sent. If AOI update timing differs from player to player, only the clients whose update lands on that moment are affected. - Sources: - [Entity struct (Entities 1.3)](https://docs.unity3d.com/Packages/com.unity.entities@1.3/api/Unity.Entities.Entity.html) · Unity · An Entity consists of an Index and a generation number (Version), which tells whether a reused Index is still valid - [NetworkConfig class (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/api/Unity.Netcode.NetworkConfig.html) · Unity · RecycleNetworkIds and NetworkIdRecycleDelay: network IDs are reused only after sitting unused for a set time #### pt-port-collision · Fixed UDP port collision · Two clients bound to the same local UDP port If the client is built to use a fixed local port, a second client on the same PC either can’t get the port or ends up splitting packets with the first. - Why → Effect → On screen: Two clients try to open the same local UDP port (forcing a share with a reuse option) → The OS delivers incoming packets to only one socket, or doesn’t guarantee which one gets them. The router and server also see both clients as the same address → One client misses world packets, so NPCs and other players are invisible, or it disconnects - Symptoms: Invisible / ghost entities, Disconnect, Can’t connect / infinite loading / Factors: Packet loss - Who: One client on the same PC / When: Right after login or maintenance, Always - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: let the OS pick the local port automatically (bind to port 0). Server: tell connections apart by a session token issued per connection. - On the graph: Outliers only (Packets received per client) - Where to look: On the player’s PC with both clients running, netstat -ano -p udp in a command prompt to see the local UDP ports each game process (PID) has open. On the server side, whether the two sessions come in from the same public IP and the same port - Confirmed if: Both game processes are bound to the same local port, or the server sees both sessions as the same IP and port. Fine with only one running - Ruled out if: The two clients use different local ports but one still misbehaves: sessions keyed by IP or device, or a multi-client restriction - Check with: The player’s own environment - Sources: - [Using SO_REUSEADDR and SO_EXCLUSIVEADDRUSE](https://learn.microsoft.com/en-us/windows/win32/winsock/using-so-reuseaddr-and-so-exclusiveaddruse) · Microsoft · A second bind to the same port with SO_REUSEADDR hijacks the port, and which socket receives packets is undefined - [bind function (winsock.h)](https://learn.microsoft.com/en-us/windows/win32/api/winsock/nf-winsock-bind) · Microsoft · Binding to port 0 assigns a unique port from the dynamic port range (49152–65535) - [netstat](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/netstat) · Microsoft · -a shows TCP and UDP ports, -n numeric addresses, -o the process ID (PID), and -p udp UDP only #### pt-session-key · Sessions keyed by IP or device (bug) · Session keyed by IP or machine ID If the server or an intermediate server identifies connections by IP or device ID, it treats two clients on the same PC (same public IP) as one person. - Why → Effect → On screen: The session table is keyed by IP, or by IP + device ID → The second client’s data overwrites or gets mixed into the first session → One side can’t see NPCs, and the other disconnects or receives someone else’s data - Symptoms: Invisible / ghost entities, Disconnect / Factors: Packet loss - Who: One client on the same PC, Same household, Specific region/ISP / When: Right after login or maintenance - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: identify every connection by a unique session token on both the server and intermediate servers, and be sure to fix this, because several people in the same home (behind router NAT) and players on mobile networks where the carrier shares one IP among many subscribers (CGNAT) hit the same problem. Client: use the session token each running client received separately. - On the graph: Outliers only (Concurrent sessions from the same public IP, session overwrites) - Where to look: Server and intermediate server logs of the key used to look up the session, the session token, and the client IP and port, checking whether the existing session changed the moment a second connection came in from the same IP. Reproduces by starting two clients one after another on the same PC - Confirmed if: The moment the second client connects, the first session’s address or character data changes, and the same disconnects appear for other players behind the same router or mobile network (CGNAT) - Ruled out if: Two sessions from the same IP are kept apart with different tokens: not this cause. Two processes use the same local port: fixed UDP port collision - Check with: Game server or client logs and metrics - Sources: - [RFC 6269: Issues with IP Address Sharing](https://www.rfc-editor.org/rfc/rfc6269) · IETF · When many subscribers share one IPv4 address through NAT or CGN, IP alone can’t tell users apart #### pt-multiclient · Multi-client restriction · Multi-client restriction policy If an anti-cheat module or server policy limits multiple clients on one PC, the second client is blocked from launching or connecting, or the first one gets disconnected. Some games only block features on the extra client. - Why → Effect → On screen: The anti-cheat module detects a duplicate launch, or the server limits extra connections from the same device → The second launch or connection is refused, or one side is disconnected. Rarely, only some features on the extra client are blocked → Can’t connect, or one side disconnects. In games that only block features, NPCs or shops are invisible on one side only - Symptoms: Can’t connect / infinite loading, Invisible / ghost entities, Disconnect / Factors: Packet loss - Who: One client on the same PC / When: Right after login or maintenance - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: if you restrict it, show a clear message, add a QA exception to the anti-cheat module. Server: add a QA exception to the same-device connection limit too. - On the graph: Outliers only (Connection refusals and disconnects by reason (duplicate login)) - Where to look: The message shown when the second client launches and the disconnect message on the first. Whether the server’s connection refusal and forced disconnect logs record reason codes such as duplicate login or same device - Confirmed if: A refusal message appears the moment the second client launches or connects, or the first one is disconnected with a duplicate login reason, and there’s no problem with only one client running - Ruled out if: Both connect without any refusal or disconnect reason but NPCs are invisible on one side: fixed UDP port collision, sessions keyed by IP or device, or a loading or display cause - Check with: The player’s own environment - Sources: - [CreateMutexW function (synchapi.h)](https://learn.microsoft.com/en-us/windows/win32/api/synchapi/nf-synchapi-createmutexw) · Microsoft · If a named mutex already exists, ERROR_ALREADY_EXISTS is returned, which is used to detect duplicate launches and allow only a single instance #### pt-background · Background window throttling · Background window throttling For a client in a background window, the game, engine, and OS cut its frame rate and processing. Received packets aren’t processed in time, so they back up or overflow. - Why → Effect → On screen: Background frame limits in game options or the graphics driver (e.g., the NVIDIA driver lets you pick 20–200 per second), power saving, the engine’s background pause setting. The OS also gives CPU and GPU priority to the window in front (foreground) → Fewer packets are processed per frame, so the queue builds up, and packets are dropped when the receive buffer overflows → When the window comes to the front, things appear all at once, or some NPCs never show up - Symptoms: Invisible / ghost entities, Fast-forward, Disconnect / Factors: Stall, Packet loss - Who: One client on the same PC / When: After sitting idle, Always - Primary owner: Game team (Client development) / Also: External (External) - Game team action items: Keep network receiving going on a thread separate from the game loop, guarantee a minimum processing rate in the background, turn on the engine’s run-in-background setting (runInBackground in Unity). - External action items: Tell players to turn off the graphics driver’s background frame limit and the PC’s power-saving mode. - Ballpark numbers: In Unity, with runInBackground off, the game loop stops the moment the window loses focus. If receiving happens only in that loop, no packets are processed at all in the meantime. - On the graph: Gap then burst (Client frame interval, packets processed per frame) - Where to look: On the same PC, one window in front and the other behind, swapping roles and comparing. Frame intervals of both processes measured with PresentMon, and with game logs, window focus state and packets processed per frame - Confirmed if: Frame intervals grow sharply (with a driver limit, flattening at the interval that matches the configured frame rate) or processing stops only while the window is in the background, and switching windows moves the problem to the other client - Ruled out if: It happens the same way in the foreground window: not background throttling. Always the same client misbehaving regardless of window position: display settings or version difference - Check with: The player’s own environment - Sources: - [Application.runInBackground](https://docs.unity3d.com/ScriptReference/Application-runInBackground.html) · Unity · runInBackground defaults to false, in which case the app pauses in the background - [Manage 3D Settings (reference) — NVIDIA Control Panel Help](https://www.nvidia.com/content/Control-Panel-Help/vLatest/en-us/mergedProjects/nv3d/Manage_3D_Settings_(reference).htm) · NVIDIA · Background Application Max Frame Rate: caps a background game’s frame rate at 20–200 per second - [Priority Boosts](https://learn.microsoft.com/en-us/windows/win32/procthread/priority-boosts) · Microsoft · Windows raises the priority of the foreground window’s process above background processes - [PresentMon README](https://raw.githubusercontent.com/GameTechDev/PresentMon/main/README.md) · Intel · A tool that collects CPU, GPU, and display frame times per app for Windows graphics apps #### pt-asset-lock · Simultaneous access to cache or asset files · Shared cache / asset file lock conflicts If two clients write to the same cache folder at the same time or lock its files, one of them can’t load NPC models or textures. - Why → Effect → On screen: Two clients write the cache and patch files in the same install folder at the same time → Loading fails because a file lock failed or a half-written file was read → A name tag with no character model, or a transparent NPC - Symptoms: Invisible / ghost entities / Factors: Stall - Who: One client on the same PC / When: Right after login or maintenance, While moving or changing zones - Primary owner: Game team (Client development) - Game team action items: Use a separate cache folder per client, retry when a file lock fails, show at least a default model when loading fails. - On the graph: Outliers only (Asset loading failures per client) - Where to look: On the player’s PC, Process Monitor filtered to the game install and cache folder paths, showing the file open and write results of both game processes. With client logs, asset loading failures and the file open error code (ERROR_SHARING_VIOLATION) - Confirmed if: Opening the missing model’s file ended in a sharing violation or lock failure while the other client was writing that file. It goes away with only one client running or with separate install and cache folders - Ruled out if: The same model is missing even with only one client running: file corruption or a client version/data mismatch. Files open fine but nothing is drawn: memory/VRAM shortage - Check with: The player’s own environment - Sources: - [Creating and Opening Files](https://learn.microsoft.com/en-us/windows/win32/fileio/creating-and-opening-files) · Microsoft · A file opened without a sharing mode can’t be opened by another process, which gets ERROR_SHARING_VIOLATION - [Process Monitor](https://learn.microsoft.com/en-us/sysinternals/downloads/procmon) · Microsoft · Records file system, registry, and process activity in real time and can filter on any field, such as path #### pt-vram · Streaming failure from memory or VRAM shortage · Memory / VRAM exhaustion When two clients share graphics memory, there’s no room to load newly needed models and textures, and some of them don’t get drawn. - Why → Effect → On screen: Two clients share VRAM and RAM. The OS may also shrink a background window’s graphics memory allowance first → The engine can’t load new models and textures, or keeps evicting and reloading them → NPCs appear late, look blurry, or are invisible, and the game stutters - Symptoms: Invisible / ghost entities, Stutter / Factors: Stall - Who: One client on the same PC / When: While moving or changing zones, When crowds gather - Primary owner: Game team (Client development) / Also: External (External) - Game team action items: Adjust quality automatically to fit the memory budget, show a fallback model when loading fails. - External action items: Tell players who run two clients at once to lower graphics quality or use low-spec mode, publish recommended VRAM and RAM specs. - On the graph: Hits a ceiling (Dedicated GPU memory usage per process) - Where to look: In Task Manager on the player’s PC, the dedicated GPU memory column added to the “Details” tab, with the two clients’ combined usage compared against the graphics card’s VRAM. On the game side, the budget (Budget) and current usage (CurrentUsage) reported by DXGI’s QueryVideoMemoryInfo, logged - Confirmed if: The two clients’ combined usage flattens near VRAM capacity, and model and texture loading failures cluster when current usage goes over budget. It goes away with lower quality or only one client running - Ruled out if: Invisible even with VRAM to spare: simultaneous access to cache or asset files, or a display settings difference - Check with: The player’s own environment - Sources: - [Residency (Direct3D 12)](https://learn.microsoft.com/en-us/windows/win32/direct3d12/residency) · Microsoft · The video memory budget can shrink a lot when you switch to another app, and going over budget causes stalls or failed allocations; outside the foreground, even reservations aren’t guaranteed - [GPUs in the task manager](https://devblogs.microsoft.com/directx/gpus-in-the-task-manager/) · Microsoft · Adding columns to Task Manager’s Details tab shows dedicated and shared GPU memory usage per process; dedicated GPU memory is the graphics card’s VRAM - [DXGI_QUERY_VIDEO_MEMORY_INFO structure (dxgi1_4.h)](https://learn.microsoft.com/en-us/windows/win32/api/dxgi1_4/ns-dxgi1_4-dxgi_query_video_memory_info) · Microsoft · Budget (the video memory budget set by the OS) and CurrentUsage (the app’s current usage); usage over budget can cause stutter #### pt-display-option · Different display settings · Different display settings If settings such as a visible player limit, hidden NPC name tags or models, or low-spec mode differ between two clients, they see different things. - Why → Effect → On screen: Only one client has a “limit nearby characters shown” setting or low-spec mode on → Distant or low-priority NPCs aren’t drawn (working as intended) → The NPC is missing on one side only - Symptoms: Invisible / ghost entities / Factors: Stall - Who: One client on the same PC, Just me / When: When crowds gather, Always - Primary owner: Game team (Client development) - Game team action items: Show when an entity is hidden by a setting, keep settings files separate per client so they don’t get mixed up. - On the graph: Outliers only (Entities drawn on screen per client) - Where to look: The two clients’ visible player limit, name tag/model hiding, and low-spec mode settings side by side, with one set to match the other. Also whether the two clients share one settings file and overwrite each other - Confirmed if: With matching settings both screens show the same thing, and the missing NPCs were distant entities beyond the display limit or low-priority entities - Ruled out if: Still missing on one side with identical settings: points to a channel/phase difference or a lost spawn message - Check with: The player’s own environment - Sources: - [Changing the Quantity of Characters Displayed On-screen (FINAL FANTASY XIV UI Guide)](https://na.finalfantasyxiv.com/uiguide/faq/faq-other/setting_ch_quantity.html) · Square Enix · The display limit setting (Character and Object Quantity) controls how many characters and objects are drawn on screen #### pt-version · Client version or data mismatch · Client version / data table mismatch If the second client is a different install or isn’t fully patched, it doesn’t recognize new NPC IDs the server sends and silently ignores them. - Why → Effect → On screen: An install in a different folder, or a client launched mid-patch → Unknown NPC IDs or model IDs are skipped → Only newly added NPCs are invisible on one side - Symptoms: Invisible / ghost entities / Factors: Packet loss - Who: One client on the same PC / When: Right after login or maintenance - Primary owner: Game team (Client development) / Also: Game team (Server development) - Game team action items: Client: send the data version on connect, when an unknown ID arrives log it and show a placeholder. Server: check the data version on connect and, if it differs, refuse the connection and prompt a patch. - On the graph: Outliers only (Unknown IDs received per client version) - Where to look: The two clients’ executable paths and the client and data versions shown on screen or in logs. On the game side, the data version sent on connect and how many times unknown NPC or model IDs were received and skipped, logged - Confirmed if: The two clients differ in version or install folder, the invisible NPCs were added in a recent patch, and they show up in the fully patched install - Ruled out if: Same version and install folder but missing on one side only: points to a channel/phase difference or a loading or delivery cause - Check with: The player’s own environment - Sources: - [NetworkConfig class (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/api/Unity.Netcode.NetworkConfig.html) · Unity · Clients with different ProtocolVersion values don’t communicate, and ForceSamePrefabs checks for prefab list differences on connect #### pt-priority · Per-connection send budget and priority · Per-connection bandwidth budget and priority If the server caps how much it sends per connection and sends the nearest things first, a connection with a low cap gets distant NPCs late or not at all. - Why → Effect → On screen: In crowded places, the server sends in order of importance within each connection’s send cap → A connection with a low bandwidth estimate (e.g., a background window that’s slow to acknowledge) keeps pushing back entities further down the list → Distant NPCs appear late or not at all on one side only - Symptoms: Invisible / ghost entities, Input lag / Factors: Latency - Who: One client on the same PC, Specific zone/channel / When: When crowds gather - Primary owner: Game team (Server development) / Also: Game team (Client development) - Game team action items: Server: raise the priority of deferred entities the longer they wait (prevent starvation), guarantee a minimum update interval. Client: send acknowledgments on time even in the background so the bandwidth estimate doesn’t drop. - On the graph: Rises with load (Deferred entities per connection, bytes sent per connection) - Where to look: Server-side, per connection: bytes sent per tick, send cap (estimated bandwidth), entities deferred for lack of room, and time since each entity was last sent. In Unreal, Networking Insights shows packet sizes per connection and the replicated objects inside them - Confirmed if: The invisible NPC is an entity deferred for a long time on that connection, that connection’s cap is lower than the others’, and deferred entities grow as it gets more crowded - Ruled out if: Nothing deferred and that NPC was sent on time: a stage after sending (receive buffer, loading, display settings). Every connection is at its cap: a server-wide send volume or AOI design problem - Check with: Game server or client logs and metrics - Sources: - [Actor Priority in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-priority-in-unreal-engine) · Epic Games · When bandwidth is saturated, actors to replicate are chosen by priority (distance, line of sight, time since last replication); not every actor is replicated every time - [State Synchronization](https://gafferongames.com/post/state_synchronization/) · Gaffer On Games · Priority accumulation: entities that didn’t fit in this packet go first in the next one, and the bandwidth limit is adjusted in real time - [Networking Insights in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/networking-insights-in-unreal-engine) · Epic Games · Shows the size of packets sent and received per connection and the replicated objects and properties inside them #### pt-clock-hold · Entities held back by clock estimate error · Clock estimate error holds or discards entities If the client’s estimate of the server time is wrong, it holds back freshly arrived entity data as “still in the future” or discards it as “too old.” - Why → Effect → On screen: One client’s estimate of the server time is far off (measured during loading, or after waking from sleep) → The interpolation reference time and the entity data’s timestamp don’t match → Entities appear late or look frozen - Symptoms: Invisible / ghost entities, Stutter / Factors: Latency - Who: One client on the same PC / When: After sitting idle, Right after login or maintenance - Primary owner: Game team (Client development) - Game team action items: Redo time sync periodically and reset immediately when the gap is large, don’t use values measured during loading or right after waking from sleep. - On the graph: Outliers only (Server time estimate error per client) - Where to look: Client log of the estimated server time, RTT, when time sync was redone, and how many times entity data was held back or discarded. Try to reproduce right after loading or right after waking from sleep - Confirmed if: Only the affected client’s estimate error exceeds the reset threshold (hardResetThresholdSec in Unity, 0.2 s by default), there are records of entity data held back as future or discarded as past, and redoing time sync fixes it right away - Ruled out if: Estimate error is small but entities still appear late: points to per-connection send budget and priority, or loading - Check with: Game server or client logs and metrics - Sources: - [NetworkTimeSystem class (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/api/Unity.Netcode.NetworkTimeSystem.html) · Unity · If the time gap exceeds hardResetThresholdSec (0.2 s by default), the clock is forced into sync; otherwise adjustmentRatio speeds it up or slows it down a little at a time - [NetworkTime and ticks (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/advanced-topics/networktime-ticks.html) · Unity · LocalTime runs ahead of the server and ServerTime runs behind; a message that arrives late can have a negative wait time ### Root causes of TCP retransmission (causes: 20) #### rt-wireless · Wireless link loss · Wi-Fi / cellular link loss Wi-Fi and mobile networks retransmit a few times on the wireless link and drop the packet if that still fails. TCP resends the dropped packet only much later. - Why → Effect → On screen: A weak signal or heavy interference makes wireless transmissions fail several times in a row → Once the wireless device’s retry limit (usually a few to ten-odd attempts) is exceeded, the packet is dropped → Freeze for as long as the TCP retransmission wait, while later packets sit in the receive buffer and then fast-forward - Symptoms: Freeze, Fast-forward, Teleporting / Factors: Packet loss, Jitter - Who: Just me, Same household / When: Randomly, While moving or changing zones - Primary owner: External (External) / Also: Infra team (Server infrastructure), Game team (Server development), Game team (Client development) - Game team action items: Server: turn on TCP_NODELAY (with Nagle on, RACK has no following packets to use for detecting loss), and while retransmission has the connection blocked, keep only the latest state update pending (cap what queues in the kernel with TCP_NOTSENT_LOWAT). Client: turn on TCP_NODELAY (the client OS recovers losses in the input direction), show network status on screen when losses cluster or ping spikes. - Infra team action items: Speed up loss recovery with RACK-TLP (the server can’t prevent wireless loss; the most it can do is recover faster), confirm the current Linux defaults net.ipv4.tcp_recovery=1 (RACK) and net.ipv4.tcp_early_retrans=3 (TLP) haven’t been changed. - External action items: Tell players to use a wired connection or 5 GHz/6 GHz Wi-Fi, and to move the router or change its channel. - Ballpark numbers: At 1% wireless loss, 1 in every 100 game packets disappears. Receiving 10 a second, that’s a hitch about once every 10 seconds. Without RACK-TLP, each loss freezes the game for an RTO (ping + 200 ms or more). - On the graph: Outliers only (Per-connection retransmission rate, per-connection RTT (ping)) - Where to look: From the player’s PC, ping both the router (gateway) and the game server a few hundred times and compare loss and latency spread, then measure again on a wired connection or mobile data. On the server, check that player’s connection in ss -ti for retrans and rtt (mean/deviation) - Confirmed if: Ping to the router already shows loss or erratic latency, and it goes away on a wired connection. From the server, only that player’s connection shows high retrans and RTT deviation - Ruled out if: Clean up to the router with loss starting beyond it: points to the ISP or route (“Bottleneck queue overflow (congestion loss),” “Route change / bad ECMP path”). If several players on the same ISP get worse at the same time, start with the ISP segment - Check with: The player’s own environment - Learn more: Wireless retries add jitter (a few ms per retry), and only packets that exceed the retry limit become losses. So as wireless quality degrades, symptoms grow in this order: “jitter → occasional freezes → frequent freezes.” While a device moves between access points (roaming), it can lose packets in a row for tens of milliseconds to several seconds. Mobile networks retransmit heavily on the radio link to the cell tower, so trouble there more often shows up as latency spikes of hundreds of ms than as loss. - Real incidents: ffxiv-2021 - Sources: - [net/wireless/core.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/wireless/core.c?h=v6.12) · Linux kernel · Default retry limits in the Linux wireless stack: 7 for short frames, 4 for long frames (dot11ShortRetryLimit, dot11LongRetryLimit) - [RFC 3481: TCP over Second (2.5G) and Third (3G) Generation Wireless Networks](https://www.rfc-editor.org/rfc/rfc3481) · IETF · Link-layer retransmission keeps IP loss low on mobile networks, but that recovery shows up as jitter and latency spikes - [Wi-Fi roaming support in Apple devices](https://support.apple.com/guide/deployment/wi-fi-roaming-support-dep98f116c0f/web) · Apple · When a device moves to a new AP, it can’t send data until authentication with the new AP completes, which can take a few seconds with 802.1X - [RFC 8985: The RACK-TLP Loss Detection Algorithm for TCP](https://www.rfc-editor.org/rfc/rfc8985) · IETF · Definitions of RACK (time-based loss detection) and TLP (retransmitting the last packet) - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_recovery defaults to 0x1 (RACK), tcp_early_retrans defaults to 3 (TLP on); TCP_NOTSENT_LOWAT and tcp_notsent_lowat cap the amount of data not yet sent - [tcp(7) — Linux manual page](https://man7.org/linux/man-pages/man7/tcp.7.html) · Linux man-pages · TCP_NODELAY turns off the Nagle algorithm so even small data goes out immediately - [include/net/tcp.h](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/net/tcp.h?h=v6.12) · Linux kernel · Minimum RTO: TCP_RTO_MIN = 200 ms - [misc/ss.c](https://git.kernel.org/pub/scm/network/iproute2/iproute2.git/tree/misc/ss.c?h=v6.12.0) · iproute2 · ss -ti shows retrans:in-progress/total retransmissions and rtt:RTT/RTT deviation (rttvar) #### rt-queue-drop · Bottleneck queue overflow (congestion loss) · Tail drop at a congested bottleneck When the queue at the narrowest point fills up, such as a home router, a link between ISPs, or a data center uplink, newly arriving packets are dropped. - Why → Effect → On screen: Video, downloads, and other users’ traffic fill up the bottleneck → While the queue is full, newly arriving packets are dropped one after another (tail drop). Packets that get in wait at the back of the full queue → Several packets vanish at once, causing a long freeze then fast-forward; common in the evening - Symptoms: Freeze, Fast-forward, Rubber-banding / Factors: Packet loss, Latency - Who: Same household, Specific region/ISP, Whole server / When: Evening peak hours, When crowds gather - Primary owner: Infra team (Network infrastructure) / Also: External (External), Game team (Client development) - Game team action items: Show network status on screen when losses cluster or ping spikes (mention that a large transfer on the same connection may be the cause). - Infra team action items: Keep headroom on data center links, check output drop counters on our links and switch ports, route around a congested ISP segment through another link or peering. - External action items: Tell players to use SQM (fq_codel, CAKE) and ECN on their router (so senders slow down before the queue overflows), ask the ISP to add capacity at the bottleneck. - Ballpark numbers: When a queue overflows, a large share of incoming packets can vanish at once over tens of milliseconds. Losing several packets in a row, and even losing the retransmissions, is common, so recovery often has to wait for an RTO. - On the graph: High at certain hours (Retransmission rate, RTT (ping)) - Where to look: Server retransmission rate (TcpRetransSegs ÷ TcpOutSegs deltas from nstat run every minute) and per-connection RTT, split by region, ISP, and time of day, alongside output discards (ifOutDiscards) on our links and switch ports. Compare mtr runs to the affected region at peak and off-peak hours - Confirmed if: Retransmission rate rises only at evening peak, and RTT climbs just before the loss (the queue filling up). mtr shows loss and latency growing together from one hop to the end, only at peak hours - Ruled out if: No RTT rise before the loss points to “Policer drops excess traffic.” Similar loss at every hour points to “Physical errors (bad cable, optics, connectors)” or “Route change / bad ECMP path” - Check with: Infra tools (no game code needed) - Real incidents: riot-direct-2015 - Sources: - [RFC 7567: IETF Recommendations Regarding Active Queue Management](https://www.rfc-editor.org/rfc/rfc7567) · IETF · Tail drop keeps queues full for long periods, adding delay and causing clustered loss; recommends AQM - [RFC 8290: The Flow Queue CoDel Packet Scheduler and Active Queue Management Algorithm](https://www.rfc-editor.org/rfc/rfc8290) · IETF · fq_codel: per-flow queues plus AQM keep queues short and reduce bufferbloat - [RFC 3168: The Addition of Explicit Congestion Notification (ECN) to IP](https://www.rfc-editor.org/rfc/rfc3168) · IETF · ECN: signals congestion with a mark in the IP header, without dropping packets - [Smart Queue Management](https://www.bufferbloat.net/projects/cerowrt/wiki/Smart_Queue_Management/) · Bufferbloat.net · SQM: combines per-flow scheduling, queue length management (AQM), and shaping - [Cake](https://www.bufferbloat.net/projects/codel/wiki/Cake/) · Bufferbloat.net · CAKE: router SQM that combines a shaper with fq_codel-style queue management - [net/ipv4/proc.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/proc.c?h=v6.12) · Linux kernel · TcpRetransSegs and TcpOutSegs as shown by nstat (RetransSegs and OutSegs under Tcp) - [nstat(8) — Linux manual page](https://man7.org/linux/man-pages/man8/nstat.8.html) · iproute2 · By default, nstat shows the increase since its previous run - [RFC 2863: The Interfaces Group MIB](https://www.rfc-editor.org/rfc/rfc2863) · IETF · ifOutDiscards: outbound packets discarded even though no error was detected, for example to free up buffer space - [An Internet-Wide Analysis of Traffic Policing](https://research.google/pubs/an-internet-wide-analysis-of-traffic-policing/) · Google · Queue overflow raises queuing delay and RTT before the loss, while policing drops the excess with no RTT increase (SIGCOMM 2016) #### rt-burst · Send bursts overflow shallow buffers · Sender bursts overflow shallow buffers When a server sends a whole tick of updates for thousands of players in one instant, a switch’s small buffer or a cloud instance’s short-term limit overflows in under 1 ms and some packets are dropped. - Why → Effect → On screen: At the start of each tick, the server sends everyone’s packets all at once → A switch port buffer where traffic from many servers converges (hundreds of KB to a few MB per port) or a cloud instance limit overflows for an instant (average utilization stays low) → Many players teleport or hitch at the same time; averaged metrics don’t reveal the cause - Symptoms: Teleporting, Freeze, Fast-forward / Factors: Packet loss - Who: Specific zone/channel, Whole server / When: When crowds gather - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure), Infra team (Network infrastructure) - Game team action items: Spread each tick’s sends across the tick (per-connection pacing does little for thousands of connections that all send at tick start), stagger tick start times across servers, cap the rate of connections that send large data with SO_MAX_PACING_RATE. - Infra team action items: Servers/OS: cap the whole server’s send rate (a shaper in the server OS, Linux tc), smooth out a single connection’s bursts with pacing (Linux fq qdisc, BBR). Network: use deep-buffer switches, check switch port output drop counters at short intervals (average utilization won’t show them). - Ballpark numbers: A 10 Gbps port can send about 1.25 MB in 1 ms. When several servers’ ticks line up and converge on one port, the buffer fills almost instantly. - On the graph: Rises with load (Switch port output drops, retransmission rate) - Where to look: Output discards (ifOutDiscards) collected every few seconds on the switch port the server connects to and on the port above it; in the cloud, bw_out_allowance_exceeded and pps_allowance_exceeded from ethtool -S. Line these up against retransmissions collected at the same moments with bcc tcpretrans - Confirmed if: Output discards or allowance overruns grow while per-minute average utilization stays low, and they scale with concurrent users and crowding in one spot. Retransmissions hit many connections on that server at the same moment, with no concentration on particular player IP ranges (ISP or region) - Ruled out if: CRC and input errors rising on the same port point to “Physical errors (bad cable, optics, connectors).” NIC drop counters or softnet dropped rising on the receiving server point to “Packet drops on the receiving host” - Check with: Infra tools (no game code needed) - Learn more: Pacing works per connection. When thousands of connections each send one or two packets at tick start, per-connection pacing does little to spread them out, so the game server has to split up its send timing itself. By contrast, when one connection sends a large amount of data, the NIC cuts tens of KB into packet-sized pieces and sends them back to back (TSO), and pacing spreads that kind of burst out well. - Sources: - [High-Resolution Measurement of Data Center Microbursts](https://research.facebook.com/publications/high-resolution-measurement-of-data-center-microbursts/) · Meta · Over 70% of bursts at data center rack switches end within tens of µs, and per-minute average utilization correlates only weakly with drops (IMC 2017) - [tc-fq(8) — Linux manual page](https://man7.org/linux/man-pages/man8/tc-fq.8.html) · iproute2 · The fq qdisc paces each socket (connection), and SO_MAX_PACING_RATE sets a per-connection maximum rate - [net/ipv4/tcp_bbr.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_bbr.c?h=v6.12) · Linux kernel · BBR sets pacing_rate from its estimated bottleneck bandwidth - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · TCP sizes TSO frames to the flow’s rate (up to 64 KB, tcp_min_tso_segs) - [Monitor network performance for ENA settings on your EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-network-performance-ena.html) · AWS · bw_out_allowance_exceeded, pps_allowance_exceeded: packets queued or dropped because the instance exceeded its bandwidth or packets-per-second limit - [RFC 2863: The Interfaces Group MIB](https://www.rfc-editor.org/rfc/rfc2863) · IETF · ifOutDiscards: outbound packets discarded even though no error was detected, for example to free up buffer space - [Demonstrations of tcpretrans, the Linux eBPF/bcc version](https://raw.githubusercontent.com/iovisor/bcc/master/tools/tcpretrans_example.txt) · IO Visor · Shows one line per retransmission with the remote address, port, and connection state #### rt-policer · Policer drops excess traffic · Traffic policing ISP plans, cloud instance limits, and DDoS protection devices sometimes drop packets over a set rate right away, without queuing them. - Why → Effect → On screen: Momentary send volume exceeds the allowed rate or allowed burst → Packets over the limit are dropped immediately, with no queue (policing) → Each large burst loses several packets, causing a freeze then fast-forward, while the average rate looks below the limit - Symptoms: Freeze, Fast-forward, Teleporting / Factors: Packet loss - Who: Whole server, Specific region/ISP, Just me / When: When crowds gather, Evening peak hours - Primary owner: Infra team (Network infrastructure) / Also: Infra team (Server infrastructure), Game team (Server development) - Game team action items: Spread each tick’s burst across the tick to keep momentary send volume under the allowed burst, and if you hit a packets-per-second limit, combine one tick’s messages into one packet. - Infra team action items: Network: check the device’s policer exceed counters, replace the policer with a shaper, raise the allowed burst. Servers/OS: check cloud limit-exceeded metrics (on AWS, bw_out_allowance_exceeded and pps_allowance_exceeded in ethtool -S), move to a larger instance, pace on the server (Linux fq qdisc). - Ballpark numbers: A shaper (queues packets and delays them) adds latency; a policer (drops them immediately) adds loss. A TCP game connection can freeze for hundreds of ms on a single loss, so when traffic only briefly goes over the limit, the policer usually does more damage. - On the graph: Hits a ceiling (Send volume at short intervals, policer and allowance exceed counters) - Where to look: Exceed and drop counters on the device doing the policing; in the cloud, bw_out_allowance_exceeded and pps_allowance_exceeded from ethtool -S. For connections that lost packets, RTT just before the loss from ss -ti rtt or a packet capture - Confirmed if: Exceed counters rise, and send volume at short intervals looks flat, as if cut off at a fixed value. RTT doesn’t rise before the loss, and several packets vanish at once only during large bursts - Ruled out if: RTT rising first, before the loss, points to queue overflow (“Bottleneck queue overflow (congestion loss),” “Send bursts overflow shallow buffers”). Exceed counters unchanged: a different cause - Check with: Infra tools (no game code needed) - Sources: - [RFC 2475: An Architecture for Differentiated Services](https://www.rfc-editor.org/rfc/rfc2475) · IETF · Definitions: shaping delays packets to fit a traffic profile, and policing drops packets that exceed the profile - [An Internet-Wide Analysis of Traffic Policing](https://research.google/pubs/an-internet-wide-analysis-of-traffic-policing/) · Google · Policed transfers see loss rates 6 times higher on average, and pacing or shaping can achieve the same goal. Policing drops the excess with no RTT increase, while queue overflow raises RTT before the loss (SIGCOMM 2016) - [Monitor network performance for ENA settings on your EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-network-performance-ena.html) · AWS · bw_out_allowance_exceeded, pps_allowance_exceeded in ethtool -S: packets queued or dropped for exceeding instance limits - [tc-fq(8) — Linux manual page](https://man7.org/linux/man-pages/man8/tc-fq.8.html) · iproute2 · Per-connection pacing in the Linux fq qdisc - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · rtt (average round-trip time) in ss -i #### rt-physical · Physical errors (bad cable, optics, connectors) · Bit errors: bad cable, optics, dirty fiber Damaged cables, dirty fiber connectors, and worn-out optics cause bit errors, and network equipment silently drops the corrupted packets. - Why → Effect → On screen: A bad cable, optic, or connector flips bits → Equipment drops packets whose checksum (CRC) doesn’t match → Only players whose traffic takes that path keep getting short hitches followed by fast-forward, at any time of day - Symptoms: Freeze, Fast-forward, Teleporting / Factors: Packet loss - Who: Specific zone/channel, Same household / When: Always - Primary owner: Infra team (Network infrastructure) / Also: Infra team (Server infrastructure), External (External) - Infra team action items: CRC errors accumulate on the receiving end of the bad direction, so check both ends. Network: check CRC and input error counters on device ports, check optical power (the switch’s optics info), clean fiber connectors, replace cables and optics. Servers/OS: check rx_crc_errors in ethtool -S on the server (the name varies slightly by driver), check optical power (ethtool -m), replace the server-side cable or NIC. - External action items: If the problem is in the player’s home, tell them to replace the Ethernet cable or router; if it’s on the ISP’s line, ask the ISP to inspect the line. - Ballpark numbers: Even 0.1% loss is one in every 1,000 game packets. With dozens of players on that path, someone hitches every few seconds. Bit errors hit larger packets more often. - On the graph: Outliers only (CRC errors per port, retransmission rate per server and per port) - Where to look: CRC counters at both ends of the link: on servers, rx_crc_errors in ethtool -S or crc in ip -s -s link; on switches, the port’s FCS errors (dot3StatsFCSErrors) and input errors (ifInErrors). For fiber links, received optical power from ethtool -m and the switch’s optics info - Confirmed if: CRC errors on one port climb steadily at all hours, and only servers and connections through that port have high retransmission rates. Received optical power is lower than on other links of the same type - Ruled out if: CRC flat while only output discards rise points to queue overflow (“Send bursts overflow shallow buffers,” “Bottleneck queue overflow (congestion loss)”). Late collisions on one side rising together with CRC errors on the other point to “Duplex mismatch” - Check with: Infra tools (no game code needed) - Sources: - [Interface statistics](https://docs.kernel.org/networking/statistics.html) · Linux kernel · rx_crc_errors counts packets the receiving interface flagged with CRC errors; check with ip -s -s link and ethtool -S - [ethtool(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ethtool.8.html) · ethtool · -S for NIC and driver statistics, -m for optic module (SFP+, QSFP) EEPROM and optical diagnostics - [RFC 3635: Definitions of Managed Objects for the Ethernet-like Interface Types](https://www.rfc-editor.org/rfc/rfc3635) · IETF · Switch port FCS errors (dot3StatsFCSErrors), which are also counted in input errors (ifInErrors) #### rt-duplex · Duplex mismatch · Duplex mismatch If one end autonegotiates while the other has speed and duplex hard-set, one side runs half duplex and loses packets to collisions whenever load picks up. - Why → Effect → On screen: Speed and duplex hard-set on only one of the two devices → One side runs full duplex and the other half duplex, causing collisions and late collisions → Fine normally, but as traffic grows, everyone going through that device freezes then fast-forwards - Symptoms: Freeze, Fast-forward / Factors: Packet loss - Who: Whole server, Specific zone/channel / When: When crowds gather, Evening peak hours - Primary owner: Infra team (Network infrastructure) / Also: Infra team (Server infrastructure) - Infra team action items: Set both ends to autonegotiate, or hard-set both ends to the same values. Network: check speed and duplex in the switch port status, and in the port counters check whether late collisions grow on the half-duplex side and CRC errors and runts (frames that are too short) grow on the full-duplex side. Servers/OS: check speed and duplex with ethtool. - Ballpark numbers: Autonegotiation is mandatory on 1 Gbps copper, and half duplex doesn’t exist at all at 10 Gbps and above. So these days it mostly happens on old gear at 100 Mbps or below, on management ports, and on some carrier circuit hand-offs. - On the graph: Rises with load (Port late collisions and CRC errors, retransmission rate) - Where to look: Actual speed and duplex on both ends of the link: on servers, ethtool run with just the interface name; on switches, port status or dot3StatsDuplexStatus via SNMP. Late collisions (tx_window_errors on servers, dot3StatsLateCollisions on switches) and CRC errors alongside - Confirmed if: One side reports half duplex and the other full duplex. Every time traffic grows, late collisions rise on the half-duplex side and CRC errors rise on the full-duplex side - Ruled out if: Speed and duplex match on both ends and only CRC errors rise: “Physical errors (bad cable, optics, connectors).” Links at 10 Gbps and above have no half duplex, so rule this cause out for them - Check with: Infra tools (no game code needed) - Sources: - [Linux Base Driver for Intel(R) Ethernet Network Connection](https://docs.kernel.org/networking/device_drivers/ethernet/intel/e1000.html) · Linux kernel · The 1000BASE-T standard requires autonegotiation - [IEEE 802.3ae 10 Gigabit Ethernet: HSSG Objectives](https://www.ieee802.org/3/ae/objectives.pdf) · IEEE · 10 Gigabit Ethernet supports full duplex only - [IEEE P802.3ba Objectives](https://www.ieee802.org/3/ba/PAR/P802.3ba_Objectives_0709.pdf) · IEEE · 40 and 100 Gigabit Ethernet also support full duplex only - [Interface statistics](https://docs.kernel.org/networking/statistics.html) · Linux kernel · tx_window_errors counts transmissions that failed from late collisions; rx_crc_errors counts packets received with CRC errors - [ethtool(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ethtool.8.html) · ethtool · ethtool -s sets speed, duplex, and autonegotiation (speed, duplex, autoneg); given only the interface name, ethtool shows the current settings - [RFC 3635: Definitions of Managed Objects for the Ethernet-like Interface Types](https://www.rfc-editor.org/rfc/rfc3635) · IETF · dot3StatsDuplexStatus (current duplex as halfDuplex or fullDuplex), dot3StatsLateCollisions (late collision count) #### rt-host-drop · Packet drops on the receiving host · Receiver host drops (ring, softirq, CPU) Packets reach the server but get dropped, because the NIC’s ring buffer (which briefly holds arriving packets) overflows or the kernel cores that handle receive processing are saturated. - Why → Effect → On screen: A surge of players, interrupts piled on one core, CPU steal on a virtual machine, or an overloaded virtual switch → Drops at the ring buffer (rx_missed_errors and similar; the name varies by driver) or at the kernel receive queue (softnet dropped) → When crowds gather, input registers late and hitches hit the whole server at once - Symptoms: Input lag, Freeze, Fast-forward, Teleporting / Factors: Packet loss, Stall - Who: Whole server / When: When crowds gather - Primary owner: Infra team (Server infrastructure) - Infra team action items: Add drop counters such as rx_missed_errors in ethtool -S and dropped in /proc/net/softnet_stat to monitoring, enlarge ring buffers (ethtool -G), spread RSS and interrupts across multiple cores, keep the game thread and receive-processing cores separate, keep CPU headroom, and on virtual machines check CPU steal and virtual switch load. - Ballpark numbers: Packets the server drops on receive are resent by the client. So they barely show up in the server’s retransmission metrics and appear first in drop counters such as rx_missed_errors in ethtool -S (the name varies by driver) and in dropped in /proc/net/softnet_stat. - On the graph: Hits a ceiling (softirq utilization per core, NIC drop counters) - Where to look: Drop counters in ethtool -S (rx_missed_errors and similar; on mlx5, rx_out_of_buffer and rx_discards_phy), missed in ip -s -s link, and the 2nd column (dropped) and 3rd column (time_squeeze) of /proc/net/softnet_stat, plus %soft (softirq processing) per core from mpstat -P ALL. On virtual machines, %steal too - Confirmed if: Drop counters or softnet dropped rise when crowds gather, and %soft on the receive-processing cores sits near 100% and can’t go higher. Input slows on every connection to that server at the same time - Ruled out if: Server drop counters flat while retransmissions concentrate on connections from a specific region or ISP: loss on the path. When the path loses packets the server sent, the server’s nstat TcpRetransSegs rises while these counters stay flat - Check with: Infra tools (no game code needed) - Sources: - [Interface statistics](https://docs.kernel.org/networking/statistics.html) · Linux kernel · rx_missed_errors counts packets the host missed because no buffer was available (counted into drop in /proc/net/dev); check with ip -s -s link - [drivers/net/ethernet/intel/igb/igb_ethtool.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/drivers/net/ethernet/intel/igb/igb_ethtool.c?h=v6.12) · Linux kernel · ethtool -S counter names are set by the driver (e.g., igb’s rx_missed_errors and rx_no_buffer_count) - [Ethtool counters](https://docs.kernel.org/networking/device_drivers/ethernet/mellanox/mlx5/counters.html) · Linux kernel · The mlx5 driver’s rx_out_of_buffer (no buffer in the receive queue) and rx_discards_phy (dropped for lack of port buffer) - [net/core/net-procfs.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/core/net-procfs.c?h=v6.12) · Linux kernel · /proc/net/softnet_stat has one line per CPU, in hex; the 2nd column is dropped and the 3rd is time_squeeze - [Documentation for /proc/sys/net/](https://docs.kernel.org/admin-guide/sysctl/net.html) · Linux kernel · netdev_max_backlog: limit of the receive queue that holds packets arriving faster than the kernel can process them - [Scaling in the Linux Networking Stack](https://docs.kernel.org/networking/scaling.html) · Linux kernel · RSS: the NIC splits traffic across multiple receive queues handled by multiple CPUs - [ethtool(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ethtool.8.html) · ethtool · -G (--set-ring) changes the ring buffer size - [proc_stat(5) — Linux manual page](https://man7.org/linux/man-pages/man5/proc_stat.5.html) · Linux man-pages · steal in /proc/stat: time spent in other operating systems when running in a virtualized environment - [mpstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/mpstat.1.html) · sysstat · mpstat -P ALL shows per-core utilization; %soft is time spent servicing softirqs, and %steal is time spent waiting while the hypervisor serviced another virtual CPU - [net/ipv4/proc.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/proc.c?h=v6.12) · Linux kernel · TcpRetransSegs as shown by nstat (RetransSegs under Tcp) #### rt-stateful-fw · Firewall and connection tracking drops · Stateful firewall / conntrack drops Firewalls and Linux connection tracking (conntrack, which records passing connections in a table) drop packets when the table is full or when they decide a packet doesn’t match the connection’s state. - Why → Effect → On screen: The connection tracking table is full (table full), or traffic takes a different path each way so only one direction passes through the firewall (asymmetric routing) → The firewall treats the packets as belonging to an “unknown connection” or carrying a “sequence number outside the window” and drops them → A full table blocks new connections; a path mismatch makes only players on that path disconnect after repeated retransmissions - Symptoms: Freeze, Disconnect, Can’t connect / infinite loading / Factors: Packet loss - Who: Whole server, Specific region/ISP / When: When crowds gather, Right after login or maintenance, Randomly - Primary owner: Infra team (Network infrastructure) / Also: Infra team (Server infrastructure), Game team (Server development), Game team (Client development) - Game team action items: Server: in case the table fills up, throttle connection surges with a login queue, reuse connections so you don’t keep opening short-lived ones (including server-to-server calls), proactively close connections whose heartbeats have stopped. Client: when connecting fails or the connection drops, retry at growing, randomized intervals (so players don’t all pile back in at once while the table is full). - Infra team action items: Network: enlarge the firewall’s connection tracking table, exempt game ports from connection tracking, align routing so both directions pass through the same firewall, check the firewall’s TCP window checking settings. Servers/OS: enlarge the Linux table (nf_conntrack_max), exempt game ports from connection tracking (NOTRACK), check the TCP window checking setting (nf_conntrack_tcp_be_liberal), and on AWS also check conntrack_allowance_exceeded. - Ballpark numbers: The default Linux conntrack limit (nf_conntrack_max) ranges from tens of thousands to hundreds of thousands of entries depending on memory. When the current count (nf_conntrack_count) reaches the limit, the log shows “nf_conntrack: table full, dropping packet”. - On the graph: Hits a ceiling (conntrack entry count (nf_conntrack_count), failed new connections) - Where to look: On Linux servers, nf_conntrack_count and nf_conntrack_max, “nf_conntrack: table full, dropping packet” in dmesg, and drop and invalid in /proc/net/stat/nf_conntrack (one line per core, in hex). On firewalls, session table usage and drop logs; on AWS, conntrack_allowance_exceeded in ethtool -S - Confirmed if: Entry count flattens at the limit, and table full logs and drop, or conntrack_allowance_exceeded, rise at the same moment. With asymmetric routing, the table has headroom, but invalid and firewall drop logs rise for connections on a specific path - Ruled out if: Entry count far from the limit, with invalid and drop logs flat: a different cause. Table has headroom but the firewall’s CPU or packets per second is maxed out: “Middlebox over capacity (firewall, IPS, DDoS protection)” - Check with: Infra tools (no game code needed) - Sources: - [Netfilter Conntrack Sysfs variables](https://docs.kernel.org/networking/nf_conntrack-sysctl.html) · Linux kernel · nf_conntrack_max defaults to the number of hash buckets (memory ÷ 16384, 1,024–262,144), the current count is nf_conntrack_count, and nf_conntrack_tcp_be_liberal marks only out-of-window RSTs as INVALID - [net/netfilter/nf_conntrack_core.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/netfilter/nf_conntrack_core.c?h=v6.12) · Linux kernel · When the table is full, logs “nf_conntrack: table full, dropping packet” and drops the packet (the drop stat increases); packets that don’t match the connection state increase the invalid stat - [iptables-extensions(8) — Linux manual page](https://man7.org/linux/man-pages/man8/iptables-extensions.8.html) · netfilter · CT --notrack in the raw table exempts traffic from connection tracking - [Amazon EC2 security group connection tracking](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/security-group-connection-tracking.html) · AWS · When an instance exceeds its tracked-connection limit, packets are dropped, visible in conntrack_allowance_exceeded; recommends avoiding asymmetric routing - [net/netfilter/nf_conntrack_standalone.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/netfilter/nf_conntrack_standalone.c?h=v6.12) · Linux kernel · /proc/net/stat/nf_conntrack has one line per core, in hex, with columns such as entries, invalid, insert_failed, drop, and early_drop #### rt-appliance-pps · Middlebox over capacity (firewall, IPS, DDoS protection) · Inline appliance PPS / CPU overload Firewalls, intrusion prevention systems (IPS), and DDoS protection devices inspect every packet passing through. The moment traffic exceeds their inspection capacity, they drop the packets they can’t process. - Why → Effect → On screen: Hundreds of thousands or more small game packets per second at peak hours or events, or heavy inspection rules → The device maxes out its CPU or packets-per-second limit and drops packets. False positives block legitimate packets too → Freezes and teleporting hit every server behind that device at once, getting worse only when crowds gather - Symptoms: Freeze, Fast-forward, Teleporting, Disconnect / Factors: Packet loss, Latency - Who: Whole server, Specific region/ISP / When: Evening peak hours, When crowds gather - Primary owner: Infra team (Network infrastructure) / Also: Game team (Server development) - Game team action items: Share the game’s traffic pattern (ports, packet sizes, packets per second) with the infra team, batch the small messages for one tick and send them together to cut the packet count. - Infra team action items: Watch the device’s CPU, packets per second, and drop counters alongside game metrics, size devices for small packets, exempt game ports from heavy inspection, tune DDoS protection rules to the game’s traffic pattern. - Ballpark numbers: A “10 Gbps” rating on a spec sheet is often based on large 1,500-byte packets. Game packets are around 100 bytes, so the same bandwidth means more than 10 times as many packets, and the packets-per-second limit fills up first even when the link looks idle. - On the graph: Hits a ceiling (Device packets per second and CPU utilization, device drops) - Where to look: The device’s CPU, packets per second, and drop counters, plus packet counts on the switch ports in front of and behind it, compared at the same interval. Overlaid on one screen with concurrent users and server retransmission rate - Confirmed if: At peaks and events, the device’s packets per second or CPU stops at one value and can’t go higher, fewer packets leave the device than enter it, and at the same moment retransmission rates rise across every server behind it - Ruled out if: Same packet counts in front of and behind the device and no device drops: a different cause. NIC drop counters or softnet dropped rising on the server: “Packet drops on the receiving host” - Check with: Infra tools (no game code needed) - Sources: - [RFC 2544: Benchmarking Methodology for Network Interconnect Devices](https://www.rfc-editor.org/rfc/rfc2544) · IETF · Device performance must be tested at multiple frame sizes, including the minimum and maximum (throughput varies with packet size) #### rt-mtu · MTU black hole (only large packets keep getting lost) · PMTU black hole If the largest packet size a link along the way can carry shrinks and the “too big” notice (ICMP) is blocked, large packets keep vanishing no matter how many times they’re resent. - Why → Effect → On screen: The maximum size shrinks on a VPN or tunnel segment, and a firewall blocks the “too big” notices → The sender, with no idea why, keeps retransmitting the same large packet, and the RTO doubles each time → Fine normally, but when large data moves (inventory, crowded areas, loading into a zone), everything stops, including the small packets behind it, ending in a disconnect or infinite loading - Symptoms: Freeze, Disconnect, Can’t connect / infinite loading / Factors: Packet loss - Who: Specific region/ISP, Just me / When: During specific actions, Right after login or maintenance - Primary owner: Infra team (Network infrastructure) / Also: Infra team (Server infrastructure), Game team (Server development) - Game team action items: To lower it from the server side, set the socket’s maximum segment size (TCP_MAXSEG); just splitting messages into smaller pieces in game code won’t prevent it (TCP re-packs the outgoing data into MSS-sized segments). - Infra team action items: Network: clamp the MSS on edge devices, allow the “too big” ICMP (type 3 code 4, fragmentation needed) through firewalls and cloud network ACLs. Servers/OS: set the path MTU, make sure the server firewall and cloud security groups don’t block “too big” ICMP either, and as a last safety net set Linux tcp_mtu_probing=1. - Ballpark numbers: Usually 1,500 bytes; around 1,400 through a tunnel. If the same packet is retransmitted 5–6 times, the freeze exceeds 10 seconds. - On the graph: Outliers only (Per-connection RTO and backoff, disconnects per region and ISP) - Where to look: Retransmissions on the problem connection from a server-side packet capture or bcc tcpretrans -s (shows sequence numbers), plus mss, pmtu, and backoff for that connection from ss -ti. From the server to that player’s address, a small ping compared with a 1,500-byte ping with DF set (ping -M do -s 1472) - Confirmed if: Full MSS-sized packets are retransmitted over and over with the same sequence number at doubling intervals, while smaller packets get through. No “too big” ICMP arrives (Wireshark filter icmp.type == 3 and icmp.code == 4), and the small ping gets replies while only the large DF ping vanishes without one - Ruled out if: Small packets vanishing too points to loss unrelated to size (“Bottleneck queue overflow (congestion loss),” “Route change / bad ECMP path”). If “too big” ICMP arrives and pmtu in ss -ti drops, path MTU discovery is working properly - Check with: Infra tools (no game code needed) - Learn more: tcp_mtu_probing=1 declares a black hole and lowers the MSS to 1,024 bytes only after retransmission timeouts have gone on for a few seconds (the equivalent of tcp_retries1=3). The connection is frozen until then, so keep it as a last safety net and put MSS clamping, which prevents the problem up front, first. - Sources: - [RFC 1191: Path MTU discovery](https://www.rfc-editor.org/rfc/rfc1191) · IETF · Path MTU discovery: a packet that is too large triggers ICMP “fragmentation needed and DF set” (type 3 code 4) - [RFC 2923: TCP Problems with Path MTU Discovery](https://www.rfc-editor.org/rfc/rfc2923) · IETF · The PMTU black hole problem, where blocked ICMP makes only large packets keep vanishing - [RFC 4821: Packetization Layer Path MTU Discovery](https://www.rfc-editor.org/rfc/rfc4821) · IETF · A way for the transport layer to discover packet size without ICMP (the basis of Linux tcp_mtu_probing) - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_mtu_probing: 0 off, 1 only when a black hole is detected, 2 always (starting MSS is tcp_base_mss). tcp_retries1 defaults to 3 - [net/ipv4/tcp_timer.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_timer.c?h=v6.12) · Linux kernel · When RTO retransmissions go on for tcp_retries1 times, treats it as a detected black hole, turns on MTU probing, and lowers the MSS - [include/net/tcp.h](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/net/tcp.h?h=v6.12) · Linux kernel · TCP_BASE_MSS = 1,024 bytes - [RFC 6298: Computing TCP's Retransmission Timer](https://www.rfc-editor.org/rfc/rfc6298) · IETF · Doubles the RTO each time the retransmission timer expires - [iptables-extensions(8) — Linux manual page](https://man7.org/linux/man-pages/man8/iptables-extensions.8.html) · netfilter · TCPMSS --clamp-mss-to-pmtu: works around large packets stalling on segments that block ICMP by adjusting the MSS in the SYN - [tcp(7) — Linux manual page](https://man7.org/linux/man-pages/man7/tcp.7.html) · Linux man-pages · TCP_MAXSEG: maximum segment size for outgoing packets; set before connecting, it also changes the MSS advertised to the peer - [Network maximum transmission unit (MTU) for your EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/network_mtu.html) · AWS · Traffic through internet gateways and VPNs uses an MTU of 1,500; PMTUD needs ICMP type 3 code 4, which doesn’t get through if security groups or network ACLs block it - [MTU considerations | Cloud VPN](https://cloud.google.com/network-connectivity/docs/vpn/concepts/mtu-considerations) · Google Cloud · Cloud VPN gateway MTU is 1,460 bytes, and the payload MTU of an IPv4 tunnel is 1,406 bytes (around 1,400 through a tunnel) - [Demonstrations of tcpretrans, the Linux eBPF/bcc version](https://raw.githubusercontent.com/iovisor/bcc/master/tools/tcpretrans_example.txt) · IO Visor · Shows one line per retransmission; -s also shows the sequence numbers of retransmitted packets - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · mss, pmtu (path MTU), and backoff (how many times the RTO has doubled) in ss -i - [ping(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ping.8.html) · iputils · -M do sets DF and won’t send packets larger than the path MTU the kernel knows; -s is the data size (the 8-byte ICMP header comes on top) - [Display Filter Reference: Internet Control Message Protocol](https://www.wireshark.org/docs/dfref/i/icmp.html) · Wireshark · icmp.type and icmp.code display filters #### rt-mapping · NAT or load balancer mapping expires mid-connection · NAT / load balancer mapping expired mid-connection If a device along the way deletes the mapping for an idle connection (the entry that records where to forward that connection), the next packet sent can’t be delivered. The connection either retransmits over and over until it disconnects, or the device sends back a reset (RST) and it disconnects right away. - Why → Effect → On screen: A connection with no packets going either way for a while (AFK, lobby) → A home router’s NAT, the ISP’s CGNAT, a firewall, a load balancer, or a cloud security group deletes the idle mapping → When the player moves again, retransmissions go on until a disconnect, or the disconnect is immediate - Symptoms: Disconnect, Freeze / Factors: Packet loss - Who: Just me, Specific region/ISP / When: After sitting idle - Primary owner: Game team (Client development) / Also: Game team (Server development), Infra team (Network infrastructure), Infra team (Server infrastructure) - Game team action items: Client: send heartbeats at no more than half the shortest idle timeout (mappings in players’ routers and ISP CGNAT are reliably refreshed only by outbound packets, and we can’t change their timeouts, so the client sends them), reconnect automatically after a disconnect. Server: answer heartbeats, and close the connection first when none arrive for a set time (shorten the TCP keepalive interval with socket options such as TCP_KEEPIDLE, detect failures quickly with TCP_USER_TIMEOUT), resume sessions with a session token. - Infra team action items: Network: collect the idle timeouts of firewalls and load balancers on the path and share them with the game team, extend them on our own firewalls and load balancers if needed. Servers/OS: check the cloud security group’s connection tracking timeout and share it with the game team. - Ballpark numbers: How long devices keep a TCP mapping varies widely, from a few minutes to several hours. When a cloud security group tracks connections, AWS Nitro v6 instance types delete the tracking entry after 350 seconds by default (other types after 5 days; see “Cloud security group connection tracking expiry”). The Linux TCP keepalive default is “probe after 2 hours idle,” which is later than most devices. - On the graph: Mass disconnect (Disconnects, idle time before disconnect) - Where to look: The last few minutes of a dropped connection in a server-side packet capture; for live connections, idle time from lastsnd and lastrcv in ss -ti (ms since the last send and receive). nstat TcpExtTCPAbortOnTimeout (connections abandoned when the timer ran out) alongside - Confirmed if: Each dropped connection had just been idle past a similar value (the idle timeout of a device on the path, e.g., 350 s for security groups on AWS Nitro v6 instances), and from the first packet after the idle period, retransmissions go on with no ACK until the connection gives up, or an RST comes back right away - Ruled out if: Disconnects during play regardless of idle time point to a different cause (“Route change / bad ECMP path,” “Firewall and connection tracking drops”). Rule this cause out for connections that exchange heartbeats at no more than half the shortest idle timeout - Check with: Infra tools (no game code needed) - Sources: - [RFC 5382: NAT Behavioral Requirements for TCP](https://www.rfc-editor.org/rfc/rfc5382) · IETF · Recommends a TCP NAT established-connection idle timeout of at least 2 hours 4 minutes (on the premise that devices may delete idle sessions earlier) - [RFC 4787: Network Address Translation (NAT) Behavioral Requirements for Unicast UDP](https://www.rfc-editor.org/rfc/rfc4787) · IETF · NAT mappings must be refreshed by outbound packets (REQ-6); refresh by inbound packets is optional (for UDP) - [Amazon EC2 security group connection tracking](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/security-group-connection-tracking.html) · AWS · The default TCP idle tracking timeout is 350 seconds on Nitro v6 instance types and 432,000 seconds (5 days) on other types; recommends keepalives at intervals shorter than 5 minutes - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_keepalive_time defaults to 2 hours - [tcp(7) — Linux manual page](https://man7.org/linux/man-pages/man7/tcp.7.html) · Linux man-pages · TCP_KEEPIDLE (idle time before keepalive starts), TCP_USER_TIMEOUT (how long to wait for unacknowledged data before closing the connection) - [RFC 5482: TCP User Timeout Option](https://www.rfc-editor.org/rfc/rfc5482) · IETF · TCP user timeout: how long sent data can go unacknowledged before the connection is closed - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · lastsnd and lastrcv in ss -i: time since the last send and receive (ms) - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · TcpExtTCPAbortOnTimeout: connections abandoned without an RST because a TCP timer ran out #### rt-path · Route change / bad ECMP path · Route change / bad ECMP member Packets vanish for a few seconds while an internet route changes, or steadily on connections assigned to a faulty path among several ECMP paths. - Why → Effect → On screen: BGP route recalculation, or faulty equipment or a bad link on one of several paths (ECMP, LAG) → Temporary loss during the route switch, or steady loss only on connections using that path → A sudden freeze of a few seconds then fast-forward, or “it gets better after reconnecting” (assigned to a different path) - Symptoms: Freeze, Fast-forward, Teleporting / Factors: Packet loss - Who: Specific region/ISP / When: Randomly - Primary owner: Infra team (Network infrastructure) / Also: Game team (Server development), External (External) - Game team action items: Log per-connection retransmission stats (TCP_INFO) so you can pull the IP, port, and time for affected players, don’t drop connections that stall for a few seconds right away. - Infra team action items: Monitor retransmission rate per region and ISP, check whether reconnecting changes the path, secure links from multiple ISPs, check our ECMP and LAG paths for bad links, run route measurements on the same TCP port as the game (mtr --tcp --port; the path is chosen by address and port, so an ordinary ping can take a different path and look fine). - External action items: Report the bad path to the ISP with route measurements taken on the same TCP port and a comparison from before and after reconnecting. - On the graph: Step change (RTT (ping), retransmission rate per region and ISP) - Where to look: Retransmissions grouped per connection with bcc tcpretrans -c to pull the affected players’ addresses and ports; mtr on the same TCP port as the game (mtr -T -P PORT) from the server toward the player and from the player toward the server, compared. Results before and after reconnecting compared too - Confirmed if: From a certain moment, RTT for one region or ISP shifts like a step and a few seconds of loss cluster, or even within one ISP only some connections (address and port combinations) keep retransmitting and get better after reconnecting. A plain ping can look fine while only TCP mtr shows loss - Ruled out if: All connections on that ISP getting worse together at evening peak: “Bottleneck queue overflow (congestion loss).” Only one player affected, with loss already on the ping to their router: “Wireless link loss” - Check with: Infra tools (no game code needed) - Real incidents: cloudflare-2020 - Sources: - [RFC 2991: Multipath Issues in Unicast and Multicast Next-Hop Selection](https://www.rfc-editor.org/rfc/rfc2991) · IETF · Diagnostic tools such as ping and traceroute are hard to trust over multiple paths; describes pinning each flow to one path by hashing it - [RFC 2992: Analysis of an Equal-Cost Multi-Path Algorithm](https://www.rfc-editor.org/rfc/rfc2992) · IETF · ECMP picks the next hop from a hash of the header fields that identify a flow (the same flow takes the same path) - [tcp(7) — Linux manual page](https://man7.org/linux/man-pages/man7/tcp.7.html) · Linux man-pages · TCP_INFO: query per-socket state (struct tcp_info) - [mtr(8) manual page source](https://raw.githubusercontent.com/traviscross/mtr/master/man/mtr.8.in) · mtr · -T (--tcp) uses TCP SYN in place of ICMP, -P (--port) sets the target port - [Demonstrations of tcpretrans, the Linux eBPF/bcc version](https://raw.githubusercontent.com/iovisor/bcc/master/tools/tcpretrans_example.txt) · IO Visor · Shows one line per retransmission with the remote address and port; -c counts retransmissions per flow #### rt-spurious-delay · Spurious retransmission from latency spikes · Spurious RTO from delay spikes A packet that isn’t lost, just very late for a moment, still gets retransmitted if the delay is longer than the RTO, because the sender treats it as lost. - Why → Effect → On screen: Bufferbloat, Wi-Fi power saving, mobile radio state changes, or a virtual machine pause cause momentary delays of hundreds of ms → The RTO expires first and the packet is retransmitted; the original arrives soon after (the receiver gets a duplicate) → The freeze and fast-forward come from the latency spike itself. The spurious retransmission barely lengthens the freeze; it only pushes up retransmission metrics, which get mistaken for loss - Symptoms: Freeze, Fast-forward, Input lag / Factors: Latency, Jitter - Who: Just me, Whole server / When: Randomly, After sitting idle - Primary owner: External (External) / Also: Infra team (Server infrastructure), Game team (Client development) - Game team action items: On Android 10 and later clients, request low-latency Wi-Fi mode during play (the WIFI_MODE_FULL_LOW_LATENCY Wi-Fi lock, which applies only while the screen is on and the game is in the foreground) to cut latency spikes caused by power saving. - Infra team action items: Avoid burstable instances, don’t set the RTO minimum too low, keep F-RTO and timestamps on (tcp_frto, tcp_timestamps), read retransmission metrics together with nstat TCPSpuriousRTOs and TCPDSACKRecv so they aren’t mistaken for loss. - External action items: To reduce the latency spikes themselves, tell players to use SQM on their router and turn off Wi-Fi power saving. - Ballpark numbers: Linux uses F-RTO to detect spurious RTOs and can undo the cut in sending rate. Check with nstat’s TCPSpuriousRTOs (times an RTO was judged spurious) and TCPDSACKRecv (times the receiver reported “already got this”). - On the graph: Random spikes (RTT (ping), spurious RTOs) - Where to look: Deltas of TcpExtTCPTimeouts (RTO expirations), TcpExtTCPSpuriousRTOs, TcpExtTCPDSACKRecv, and TcpExtTCPLostRetransmit together, from nstat run every minute. With a packet capture, the Wireshark filter tcp.analysis.spurious_retransmission - Confirmed if: When RTOs rise, TcpExtTCPSpuriousRTOs or TcpExtTCPDSACKRecv rise with them, and RTT jumps to hundreds of ms at the same moment. The receiver-side capture shows both the original and the retransmission arriving - Ruled out if: TcpExtTCPSpuriousRTOs and DSACK flat while TcpExtTCPLostRetransmit (even the resent packet lost again) rises: real loss. RTT not jumping but DSACK steadily high: “Spurious fast retransmit from reordering” - Check with: Infra tools (no game code needed) - Sources: - [RFC 5682: Forward RTO-Recovery (F-RTO): An Algorithm for Detecting Spurious Retransmission Timeouts with TCP](https://www.rfc-editor.org/rfc/rfc5682) · IETF · F-RTO: uses the ACKs that arrive after an RTO to tell whether the RTO was spurious - [RFC 3481: TCP over Second (2.5G) and Third (3G) Generation Wireless Networks](https://www.rfc-editor.org/rfc/rfc3481) · IETF · Latency spikes on mobile networks (handovers, link-layer recovery, and so on) cause spurious TCP timeouts and retransmissions and shrink the congestion window - [RFC 2883: An Extension to the Selective Acknowledgement (SACK) Option for TCP](https://www.rfc-editor.org/rfc/rfc2883) · IETF · DSACK: the receiver reports duplicates, so the sender can tell a retransmission was unnecessary - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · TcpExtTCPSpuriousRTOs (spurious RTOs detected by F-RTO), TcpExtTCPDSACKRecv (DSACKs received), TcpExtTCPLostRetransmit (SACK reported a retransmitted packet lost again) - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_frto is on by default (helps on wireless networks with fluctuating RTT), tcp_timestamps defaults to 1 - [RFC 6298: Computing TCP's Retransmission Timer](https://www.rfc-editor.org/rfc/rfc6298) · IETF · Evidence that avoiding spurious retransmissions needs a large minimum RTO (recommends at least 1 second) - [WifiManager](https://developer.android.com/reference/android/net/wifi/WifiManager#WIFI_MODE_FULL_LOW_LATENCY) · Android (Google) · WIFI_MODE_FULL_LOW_LATENCY (API 29, Android 10): a low-latency Wi-Fi lock that applies only while connected to an AP, with the screen on and the app in the foreground - [net/ipv4/proc.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/proc.c?h=v6.12) · Linux kernel · Counter names as shown by nstat: TCPTimeouts, TCPSpuriousRTOs, TCPDSACKRecv, TCPLostRetransmit - [net/ipv4/tcp_timer.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_timer.c?h=v6.12) · Linux kernel · TCPTimeouts increments when the retransmission timer (RTO) expires - [nstat(8) — Linux manual page](https://man7.org/linux/man-pages/man8/nstat.8.html) · iproute2 · By default, nstat shows the increase since its previous run - [Display Filter Reference: Transmission Control Protocol](https://www.wireshark.org/docs/dfref/t/tcp.html) · Wireshark · tcp.analysis.spurious_retransmission display filter #### rt-reorder · Spurious fast retransmit from reordering · Reordering triggers spurious fast retransmit When packets get out of order crossing multiple paths or bundled links, the receiver signals “a packet is missing” with duplicate ACKs, and the sender resends a packet that arrived fine. - Why → Effect → On screen: Devices that split traffic across paths per packet, LAGs (link bundles) that spread traffic per packet, and route changes shuffle packet order → Later packets arrive first and three duplicate ACKs pile up → fast retransmit → Sparse game packets are barely affected. Large updates in crowded areas and patch downloads slow down, with occasional stutter - Symptoms: Stutter, Input lag / Factors: Jitter - Who: Specific region/ISP, Whole server / When: Always, When crowds gather - Primary owner: Infra team (Network infrastructure) / Also: Infra team (Server infrastructure) - Infra team action items: Network: switch from per-packet to per-connection load balancing (hash ECMP and LAG on address and port). Servers/OS: use RACK (time-based loss detection that tolerates reordering; when DSACK reveals spurious retransmissions, it automatically widens its reordering allowance), check the reordering degree Linux estimates automatically for each connection (the reordering value in ss -ti, starting from tcp_reordering=3). - On the graph: Always high (Reordering detections, DSACKs received) - Where to look: nstat TcpExtTCPSACKReorder and TcpExtTCPTSReorder (reordering detections) and TcpExtTCPDSACKRecv; per connection, reordering (shown when not 3) and reord_seen in ss -ti. In a packet capture, the Wireshark filter tcp.analysis.out_of_order - Confirmed if: Reordering counters and DSACK climb steadily at all hours, and connections through a specific path or device show a reordering value above 3. The receiver-side capture shows later packets arriving first, with the earlier ones following shortly - Ruled out if: Reordering counters flat while TcpExtTCPLostRetransmit rises: real loss. DSACK rising only at moments when RTT jumps: “Spurious retransmission from latency spikes” - Check with: Infra tools (no game code needed) - Sources: - [RFC 2991: Multipath Issues in Unicast and Multicast Next-Hop Selection](https://www.rfc-editor.org/rfc/rfc2991) · IETF · Splitting paths per packet reorders packets, and when 3 or more later packets arrive first, TCP performs a spurious fast retransmit - [RFC 2992: Analysis of an Equal-Cost Multi-Path Algorithm](https://www.rfc-editor.org/rfc/rfc2992) · IETF · ECMP that picks a path from a hash of the header fields identifying a flow (per-flow load balancing) - [RFC 5681: TCP Congestion Control](https://www.rfc-editor.org/rfc/rfc5681) · IETF · Fast retransmit on the third duplicate ACK - [RFC 8985: The RACK-TLP Loss Detection Algorithm for TCP](https://www.rfc-editor.org/rfc/rfc8985) · IETF · RACK detects loss by time, so it tolerates reordering, and it widens its reordering window (reo_wnd) when it receives DSACKs - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_reordering starts at 3 (adjusted automatically per connection up to tcp_max_reordering), RACK setting in tcp_recovery - [misc/ss.c](https://git.kernel.org/pub/scm/network/iproute2/iproute2.git/tree/misc/ss.c?h=v6.12.0) · iproute2 · ss -ti shows reordering:value when a connection’s reordering differs from the default of 3, and reord_seen:count once the connection has seen reordering - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · TcpExtTCPSACKReorder, TcpExtTCPTSReorder (reordering detected), TcpExtTCPDSACKRecv (DSACKs received), TcpExtTCPLostRetransmit (a retransmitted packet lost again) - [include/uapi/linux/tcp.h](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/uapi/linux/tcp.h?h=v6.12) · Linux kernel · tcpi_reord_seen in tcp_info: how many times the connection has seen reordering - [Display Filter Reference: Transmission Control Protocol](https://www.wireshark.org/docs/dfref/t/tcp.html) · Wireshark · tcp.analysis.out_of_order display filter #### rt-ack-path · Late or lost ACKs (saturated upload) · ACK path congestion on asymmetric links Data arrives fine, but if the “got it” ACK is delayed or dropped in a full upload queue, the sender treats the data as lost and retransmits. - Why → Effect → On screen: Video uploads or cloud backups at home saturate the upload → ACKs sit in the router’s queue for hundreds of ms or get dropped when it overflows → Game packets from the server mostly arrive on time. Your inputs, stuck in the same upload queue, go out late, causing input lag and rubber-banding, with occasional spurious retransmissions - Symptoms: Input lag, Rubber-banding / Factors: Latency, Packet loss - Who: Same household / When: Randomly, Evening peak hours - Primary owner: External (External) / Also: Game team (Client development) - Game team action items: Show network status on screen when ping spikes, display a hint to “check for programs that are uploading.” - External action items: Tell players to keep the upload queue short with SQM on their router, prioritize small packets (ACKs), and cap upload speed (video uploads, cloud backups). - Ballpark numbers: A later ACK confirms everything an earlier one did, so losing a few ACKs is usually fine. The real trouble is ACKs delayed in the queue. - On the graph: Outliers only (Per-connection RTT (ping)) - Where to look: Ping to the game server from the player’s PC with an upload (video upload, cloud backup) running and stopped, compared. On the server, rtt for that player’s connection in ss -ti - Confirmed if: Ping climbs to hundreds of ms only during the upload, with input lag and rubber-banding, and recovers soon after the upload stops. From the server, that connection’s rtt rises at the same time - Ruled out if: Loss and latency regardless of uploads: “Wireless link loss” or a path-side cause. Only the server-to-player direction slow, with no link to uploads: “Bottleneck queue overflow (congestion loss)” - Check with: The player’s own environment - Sources: - [RFC 3449: TCP Performance Implications of Network Path Asymmetry](https://www.rfc-editor.org/rfc/rfc3449) · IETF · On asymmetric links with a narrow upload, delayed or lost ACKs hurt TCP performance; ACKs are cumulative, so a later ACK covers for lost ones; remedies such as ACK-prioritizing scheduling - [Smart Queue Management](https://www.bufferbloat.net/projects/cerowrt/wiki/Smart_Queue_Management/) · Bufferbloat.net · Keeping router queues short with queue management and shaping - [tc-cake(8) — Linux manual page](https://man7.org/linux/man-pages/man8/tc-cake.8.html) · iproute2 · CAKE separates flows and minimizes latency for flows that send sparsely (sparse flows) - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · rtt (average round-trip time) and rttvar (deviation) in ss -i #### rt-rto-setting · RTO settings that don’t fit the environment · RTO min too low or too high Set the RTO minimum too low and even small delays cause spurious retransmissions; leave the default (200 ms) and it’s too long for games, so every loss means a long freeze. - Why → Effect → On screen: RTO minimum lowered sharply for data center use, or the default left as is on internet paths → Too low: retransmission storms on momentary delays. Too high: a long wait on every loss → At the default, each loss means a freeze of hundreds of ms then fast-forward; set too low, freezes get shorter, but spurious retransmissions surge and waste bandwidth - Symptoms: Freeze, Fast-forward, Input lag / Factors: Latency - Who: Whole server / When: Always - Primary owner: Infra team (Server infrastructure) / Also: Game team (Server development) - Game team action items: On Linux 6.15 and later, consider lowering the RTO cap for game connections with TCP_RTO_MAX_MS (this also shortens the time until the connection gives up, so set the disconnect detection time with TCP_USER_TIMEOUT alongside it), lower the RTO minimum only on internal server-to-server connections with the TCP_RTO_MIN_US socket option (6.15 and later), consider the TCP_THIN_LINEAR_TIMEOUTS socket option so that consecutive RTOs don’t double on game connections. - Infra team action items: Lower rto_min per route only for internal server-to-server connections, keep the default on internet paths and compensate with RACK-TLP and the thin stream setting (tcp_thin_linear_timeouts). - Ballpark numbers: Linux RTO = round-trip time + max(200 ms, RTT deviation × 4). It doubles on every failure, up to 120 seconds. On Linux 6.15 and later, TCP_RTO_MAX_MS can lower this cap to as little as 1 second. - On the graph: Always high (Per-connection RTO, spurious RTOs) - Where to look: The server’s RTO minimum setting (rto_min in ip route show; on Linux 6.11 and later, sysctl net.ipv4.tcp_rto_min_us), rto and rtt in ss -ti, and deltas of nstat TcpExtTCPSpuriousRTOs - Confirmed if: On a server with a lowered minimum, rto on internet connections hugs rtt and TcpExtTCPSpuriousRTOs climbs sharply. At the default, rto on game connections sits 200 ms or more above rtt, and every loss freezes the game for that long - Ruled out if: rto follows the default calculation (about rtt + 200 ms) and spurious RTOs are few, yet freezes are unusually long: points to consecutive losses or the recovery method (“Slow recovery on thin streams,” “Middlebox strips TCP options”) - Check with: Infra tools (no game code needed) - Sources: - [RFC 6298: Computing TCP's Retransmission Timer](https://www.rfc-editor.org/rfc/rfc6298) · IETF · RTO = SRTT + max(G, 4·RTTVAR), recommended minimum of 1 second, doubles on every failure, any maximum must be at least 60 seconds - [include/net/tcp.h](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/net/tcp.h?h=v6.12) · Linux kernel · Linux TCP_RTO_MIN 200 ms, TCP_RTO_MAX 120 seconds - [net/ipv4/tcp_input.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_input.c?h=v6.12) · Linux kernel · Linux RTO is SRTT + rttvar, and rttvar never goes below the RTO minimum (default 200 ms) - [tcp: add the ability to control max RTO](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=54a378f43425085d0684679d99735696b69165bc) · Linux kernel · Adds the TCP_RTO_MAX_MS socket option (1–120 seconds), from Linux 6.15 - [tcp: add sysctl_tcp_rto_min_us](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=f086edef71be7174a16c1ed67ac65a085cda28b1) · Linux kernel · Adds tcp_rto_min_us, the server-wide default RTO minimum, from Linux 6.11 - [tcp: support TCP_RTO_MIN_US for set/getsockopt use](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=f38805c5d26fe4af97837c10d58074a7496638bf) · Linux kernel · Adds the TCP_RTO_MIN_US socket option to set the RTO minimum per socket, from Linux 6.15 - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_rto_min_us defaults to 200000 (the route option rto_min and the socket option TCP_RTO_MIN_US take precedence), tcp_rto_max_ms 1,000–120,000 (default 120,000), tcp_thin_linear_timeouts - [ip-route(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ip-route.8.html) · iproute2 · Per-route rto_min option: the RTO minimum used when communicating with that destination - [Thin-streams and TCP](https://docs.kernel.org/networking/tcp-thin.html) · Linux kernel · TCP_THIN_LINEAR_TIMEOUTS can turn off exponential backoff for thin stream connections only - [tcp(7) — Linux manual page](https://man7.org/linux/man-pages/man7/tcp.7.html) · Linux man-pages · TCP_USER_TIMEOUT: how long to wait for unacknowledged data before closing the connection - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · rto (ms) and rtt in ss -i - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · TcpExtTCPSpuriousRTOs: spurious RTOs detected by F-RTO #### rt-thin · Slow recovery on thin streams · Thin streams fall back to RTO When a game sends small packets sparsely, the RTO fires before “three following packets” can pile up. The same loss causes a much longer freeze than it would on a bulk transfer. - Why → Effect → On screen: Packets go out about 100 ms apart, so only a few packets are ever in flight (not yet ACKed) → Collecting three duplicate ACKs takes over 300 ms, so the RTO (ping + 200 ms) fires first, doubling on consecutive losses → Each loss freezes the game for about 0.3 seconds; if the retransmission is lost too, the freeze lasts close to 1 second, then fast-forward - Symptoms: Freeze, Fast-forward / Factors: Packet loss, Stall - Who: Just me, Whole server / When: Randomly - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure), Game team (Client development) - Game team action items: Server: turn on TCP_NODELAY (with Nagle on, RACK has no following packets to base its decision on), send real-time packets over UDP with your own retransmission. Client: turn on TCP_NODELAY, send real-time packets the same way as the server (UDP). - Infra team action items: Use RACK-TLP (the default on current Linux), use tcp_thin_linear_timeouts so consecutive RTOs don’t double. - Ballpark numbers: With 100 ms between packets and a 60 ms ping, fast retransmit takes about 360 ms (until three later packets arrive and their acknowledgments come back), while the RTO is about 260 ms. With RACK, the packet is resent right away at about 160 ms, when the acknowledgment for the next packet comes back. Once packets are more than 200 ms apart, RACK is no faster than the RTO either. - On the graph: Gap then burst (Per-connection receive volume, RTO expirations) - Where to look: Deltas of nstat TcpExtTCPTimeouts (RTO expirations), TcpExtTCPFastRetrans (fast retransmits), and TcpExtTCPLossProbes and TcpExtTCPLossProbeRecovery (TLP), compared, plus rto and backoff on game connections in ss -ti. The server’s net.ipv4.tcp_recovery, tcp_early_retrans, and tcp_sack values too - Confirmed if: Among retransmissions, RTO expirations outnumber fast retransmits, and game connections often show backoff above 0 (in the middle of an RTO). Received volume sits at 0 during the freeze and then arrives all at once on recovery - Ruled out if: Bulk transfers on the same server stalling just as long: a loss problem unrelated to the connection’s traffic pattern. Concentrated on connections missing SACK or timestamps: “Middlebox strips TCP options” - Check with: Infra tools (no game code needed) - Learn more: Linux once had an option to retransmit on a single duplicate ACK for thin streams (tcp_thin_dupack), but it was removed in 2017, and RACK now fills that role. With Nagle on (TCP_NODELAY off), no new packets go out while the sender waits for the lost packet’s acknowledgment, so RACK has no following packets to base its decision on and the connection ends up waiting for the RTO. - Sources: - [Thin-streams and TCP](https://docs.kernel.org/networking/tcp-thin.html) · Linux kernel · Thin streams that send sparsely, like games, don’t trigger fast retransmit well and rely on long timeouts; the threshold is fewer than 4 packets in flight (not yet ACKed) - [RFC 5681: TCP Congestion Control](https://www.rfc-editor.org/rfc/rfc5681) · IETF · Fast retransmit on the third duplicate ACK - [RFC 8985: The RACK-TLP Loss Detection Algorithm for TCP](https://www.rfc-editor.org/rfc/rfc8985) · IETF · RACK detects loss from the delivery of packets sent later; the TLP wait is 2·SRTT (plus delayed ACK slack when only one packet is unacknowledged) - [include/net/tcp.h](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/net/tcp.h?h=v6.12) · Linux kernel · TCP_RTO_MIN 200 ms, thin stream detection (fewer than 4 packets in flight), and 6 linear retries - [tcp: remove thin_dupack feature](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=4a7f6009441144783e5925551c72e3f2e1b0839b) · Linux kernel · thin_dupack removed in January 2017 (Linux 4.11), with an explanation that RACK now fills that role - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_thin_linear_timeouts: for thin streams, doesn’t double the RTO for up to 6 retries (off by default) - [tcp(7) — Linux manual page](https://man7.org/linux/man-pages/man7/tcp.7.html) · Linux man-pages · TCP_NODELAY turns off the Nagle algorithm - [net/ipv4/proc.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/proc.c?h=v6.12) · Linux kernel · Counter names as shown by nstat: TCPTimeouts, TCPFastRetrans, TCPLossProbes, TCPLossProbeRecovery - [net/ipv4/tcp_timer.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_timer.c?h=v6.12) · Linux kernel · TCPTimeouts increments when the retransmission timer (RTO) expires - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · TcpExtTCPFastRetrans (retransmissions outside the Loss state), TcpExtTCPLossProbes (TLPs sent), TcpExtTCPLossProbeRecovery (losses recovered by TLP) - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · rto (ms) and backoff (how many times the RTO has doubled) in ss -i #### rt-sack-stripped · Middlebox strips TCP options · Middlebox strips TCP options When some firewalls or accelerators remove or rewrite TCP options, multiple losses get recovered only one per round trip, or the window (how much can be sent at once) shrinks, and everything slows down. - Why → Effect → On screen: A firewall’s “TCP normalization” or an old accelerator strips the SACK, timestamp, and window scale options → With several packets lost, recovery goes one packet per round trip, and the window is capped at 64 KB → Every loss causes a much longer freeze (without SACK, RACK-TLP can’t be used either), then fast-forward when it clears. Bulk transfers such as patches are slow too - Symptoms: Freeze, Fast-forward / Factors: Stall, Latency - Who: Specific region/ISP, Whole server / When: Always - Primary owner: Infra team (Network infrastructure) / Also: Infra team (Server infrastructure) - Infra team action items: Network: turn off TCP normalization on the device in question, check the firewall’s sequence number randomization too, compare the options in the SYN with packet captures at both ends. Servers/OS: check in ss -ti whether connections missing sack or wscale concentrate on a specific path (Windows PCs may not use ts depending on their settings, so ts alone missing can be normal), check that the server’s net.ipv4.tcp_sack is 1. - On the graph: Always high (Recoveries started without SACK (TcpExtTCPRenoRecovery)) - Where to look: Whether each connection shows sack and wscale in ss -ti, the ratio of nstat TcpExtTCPRenoRecovery (recovery started without SACK) to TcpExtTCPSackRecovery, and TcpExtTCPSACKDiscard (SACK blocks discarded as inconsistent). On a suspect path, SYNs captured at both ends with their options compared (Wireshark tcp.options.sack_perm and similar) - Confirmed if: Only connections through a specific path or device lack sack and wscale, and TcpExtTCPRenoRecovery makes up a large share. The SACK-permitted option present in the SYN as sent is missing from the SYN as received. When sequence number randomization is the cause, the options survive but TcpExtTCPSACKDiscard rises - Ruled out if: sack missing on every connection: check the server’s net.ipv4.tcp_sack value first. Options intact and TcpExtTCPSACKDiscard flat: the slow recovery has another cause (“Slow recovery on thin streams”) - Check with: Infra tools (no game code needed) - Learn more: SACK can break even when the options survive. If a firewall’s sequence number randomization rewrites only the sequence numbers in the header and leaves the numbers inside SACK blocks untouched, the sender discards the inconsistent SACKs. A server where tcp_sack=0 was set during the 2019 SACK security issue and then forgotten ends up the same way. - Sources: - [RFC 2018: TCP Selective Acknowledgment Options](https://www.rfc-editor.org/rfc/rfc2018) · IETF · Without SACK, cumulative ACKs alone reveal only one lost packet per round trip - [RFC 7323: TCP Extensions for High Performance](https://www.rfc-editor.org/rfc/rfc7323) · IETF · Without the window scale option, the window is at most 2^16 = 64 KiB - [RFC 8985: The RACK-TLP Loss Detection Algorithm for TCP](https://www.rfc-editor.org/rfc/rfc8985) · IETF · RACK-TLP requires SACK - [net/ipv4/tcp_output.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_output.c?h=v6.12) · Linux kernel · Linux schedules TLP only on connections that use SACK - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_sack defaults to 1 (on) - [misc/ss.c](https://git.kernel.org/pub/scm/network/iproute2/iproute2.git/tree/misc/ss.c?h=v6.12.0) · iproute2 · ss -ti shows ts, sack, and wscale:send,receive depending on the options a connection uses - [tcp: limit payload size of sacked skbs](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=3b4929f65b0d8249f19a50245cd88ed1a2f78cff) · Linux kernel · Commit fixing the 2019 SACK handling vulnerability (CVE-2019-11477) - [Linux and FreeBSD Kernel: Multiple TCP-based remote denial of service vulnerabilities (NFLX-2019-001)](https://raw.githubusercontent.com/Netflix/security-bulletins/master/advisories/third-party/2019-001.md) · Netflix · At the time, tcp_sack=0 (turning off SACK processing) was advised as a temporary mitigation - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · TcpExtTCPRenoRecovery (recovery started without SACK), TcpExtTCPSackRecovery (recovery started with SACK), TcpExtTCPSACKDiscard (invalid SACK blocks) - [Display Filter Reference: Transmission Control Protocol](https://www.wireshark.org/docs/dfref/t/tcp.html) · Wireshark · tcp.options.sack_perm (SACK-permitted option in the SYN) display filter #### rt-zero-window · Zero window (a stall that looks like retransmission) · Zero window, often mistaken for retransmission When the receiving program doesn’t read its socket in time and the buffer fills up, the sender stops sending and sends only zero window probes. The network itself is fine. - Why → Effect → On screen: A client frame freeze or a blocked server thread keeps the socket from being read → The receive window drops to 0, so the sender stops sending and sends only probes (at growing intervals) → Freeze then fast-forward. A packet capture shows “ZeroWindow” and no loss - Symptoms: Freeze, Fast-forward / Factors: Stall - Who: Just me, Whole server / When: When crowds gather, Randomly - Primary owner: Game team (Client development) / Also: Game team (Server development), Infra team (Server infrastructure) - Game team action items: Start with the side that sent ZeroWindow in the packet capture (the side that can’t read its socket), keep reading network input on a separate thread, size the receive buffer appropriately. Client: fix the causes of frame freezes such as loading and GC. Server: fix whatever blocks the thread that reads the socket. - Infra team action items: Add the server’s nstat TcpExtTCPToZeroWindowAdv (times the server advertised a zero receive window) to monitoring (if it rises, the problem is on the server side, so pass it to server development), provide server-side packet captures. - On the graph: Gap then burst (Per-connection receive volume, zero window count) - Where to look: In a packet capture, the side that advertised window 0, found with the Wireshark filter tcp.analysis.zero_window. In the server’s nstat, TcpExtTCPToZeroWindowAdv (the server advertised window 0) and TcpExtTCPWinProbe (probes sent in response to the peer’s window 0) viewed separately, plus Recv-Q on the server socket (bytes in ss the program hasn’t read yet) - Confirmed if: No retransmissions during the freeze, only zero window and probe packets. TcpExtTCPToZeroWindowAdv or server socket Recv-Q rising: the server isn’t reading in time. TcpExtTCPWinProbe rising: the client isn’t reading in time - Ruled out if: No zero window in the capture and the same data being resent: points to loss or spurious retransmission - Check with: Infra tools (no game code needed) - Real incidents: roblox-2021 - Sources: - [RFC 9293: Transmission Control Protocol (TCP)](https://www.rfc-editor.org/rfc/rfc9293) · IETF · When the receive window is 0, the sender sends zero window probes and backs off the probe interval exponentially - [7.5. TCP Analysis](https://www.wireshark.org/docs/wsug_html_chunked/ChAdvTCPAnalysis.html) · Wireshark · TCP ZeroWindow: a packet in which the receiver advertises window 0, telling the sender to stop sending - [Display Filter Reference: Transmission Control Protocol](https://www.wireshark.org/docs/dfref/t/tcp.html) · Wireshark · tcp.analysis.zero_window, tcp.analysis.zero_window_probe display filters - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · TcpExtTCPToZeroWindowAdv: times the receive window was advertised as 0 after being nonzero - [net/ipv4/proc.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/proc.c?h=v6.12) · Linux kernel · Counter names as shown by nstat: TCPToZeroWindowAdv, TCPWinProbe - [net/ipv4/tcp_output.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_output.c?h=v6.12) · Linux kernel · TCPWinProbe: increments for each probe sent while the peer’s receive window is 0 (tcp_send_probe0) - [net/ipv4/tcp_diag.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_diag.c?h=v6.12) · Linux kernel · For a connected socket, Recv-Q in ss is the number of bytes received that the program hasn’t read yet #### rt-syn · Connection request (SYN) retransmission · SYN retransmission on connect If a connection request is lost because the connection queue (backlog) overflows or a firewall blocks it, the client OS resends it starting 1 second later, at growing intervals. - Why → Effect → On screen: A connection surge right after maintenance overflows the server’s connection queue, or a firewall or DDoS protection drops the SYN → The client OS retransmits the SYN starting 1 second later, at set intervals (older Linux: 1 s → 2 s → 4 s) → After pressing Connect, delays come in whole seconds, such as 1 or 3 seconds; if it keeps failing: can’t connect / infinite loading - Symptoms: Can’t connect / infinite loading / Factors: Packet loss - Who: Whole server, Specific region/ISP / When: Right after login or maintenance - Primary owner: Game team (Server development) / Also: Infra team (Server infrastructure), Infra team (Network infrastructure), Game team (Client development) - Game team action items: Server: raise the backlog argument to listen (together with somaxconn), make sure the game server calls accept promptly, use a login queue. Client: lengthen connection retry intervals (randomized to spread retries out). - Infra team action items: Servers/OS: confirm server connection queue overflow with nstat TcpExtListenOverflows and TcpExtListenDrops and the “Possible SYN flooding” warning in the logs, raise somaxconn (together with the listen argument), use SYN cookies. Network: relax SYN limits on firewalls and DDoS protection. - Ballpark numbers: On Linux (Android included), the first SYN retransmission comes after 1 second. Older kernels then double the interval each time and resend at 1, 3, 7, 15 seconds …, while 6.5 and later (tcp_syn_linear_timeouts=4) resend five times at 1, 2, 3, 4, and 5 seconds, then double (7, 11, 19 seconds …). Android phones often keep their launch kernel even after OS updates, so behavior can differ between devices on the same Android version. Either way, if every attempt fails, the OS gives up after about 2 minutes. Windows starts at 1 or 3 seconds depending on version and settings and resends 2–4 times, so it gives up after 20–30 seconds (check that PC’s value with Max SYN Retransmissions in netsh int tcp show global). - On the graph: Surge after opening (Connection attempts, connection queue overflows) - Where to look: In the server’s nstat, TcpExtListenOverflows and TcpExtListenDrops, plus the “Possible SYN flooding on port” warning in dmesg; in ss -lnt, whether a listening socket’s Recv-Q (connections waiting for accept) reaches its Send-Q (backlog limit). In a server-side capture, whether SYNs arrive and whether SYN-ACKs go back - Confirmed if: During the connection surge right after maintenance, TcpExtListenOverflows rises and Recv-Q sits at Send-Q. The capture shows the same client’s SYN coming back at whole-second intervals with no server response - Ruled out if: SYNs never reach the server and server counters stay flat: a firewall or DDoS protection in front dropped them, so check that device’s SYN limits and drop logs. Server sent SYN-ACKs but connecting is still slow: loss in the return direction - Check with: Infra tools (no game code needed) - Sources: - [include/net/tcp.h](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/net/tcp.h?h=v6.12) · Linux kernel · Initial RTO TCP_TIMEOUT_INIT = 1 second (the initial value from RFC 6298) - [RFC 6298: Computing TCP's Retransmission Timer](https://www.rfc-editor.org/rfc/rfc6298) · IETF · Initial RTO of 1 second, doubling on every retransmission - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_syn_retries defaults to 6, tcp_syn_linear_timeouts defaults to 4 (SYN RTO 1, 1, 1, 1, 1, 2, 4 …), last retransmission at 67 seconds and giving up at 131 seconds, somaxconn defaults to 4096, tcp_syncookies defaults to 1 - [tcp: make the first N SYN RTO backoffs linear](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=ccce324dabfe2143519daf50ed8b1ef1d0c542f7) · Linux kernel · Commit that made the first SYN retransmissions use a fixed interval, from Linux 6.5 (the default of 4 follows the macOS and iOS behavior) - [Android common kernels](https://source.android.com/docs/core/architecture/kernel/android-common) · Android (Google) · Common kernels 5.10 through 6.18 are supported side by side, and a kernel for an earlier platform (e.g., android14-6.1) can be used to launch or upgrade new Android devices - [TcpMaxConnectRetransmissions](https://learn.microsoft.com/en-us/previous-versions/windows/it-pro/windows-2000-server/cc938209(v=technet.10)) · Microsoft · Older Windows defaults: 2 SYN retransmissions, starting with a 3-second wait and doubling, then waiting double again after the last one before giving up (3+6+12=21 seconds) - [TCP/IP connectivity issues troubleshooting](https://learn.microsoft.com/en-us/troubleshoot/windows-client/networking/tcp-ip-connectivity-issues-troubleshooting) · Microsoft · The number of SYN retransmissions varies by OS; check it with Max SYN Retransmissions in netsh int tcp show global - [listen(2) — Linux manual page](https://man7.org/linux/man-pages/man2/listen.2.html) · Linux man-pages · The backlog argument to listen is truncated to somaxconn (default 4096 since Linux 5.4, 128 before that) - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · When the accept queue is full, SYNs are dropped and TcpExtListenOverflows and TcpExtListenDrops rise together; TcpExtTCPSynRetrans - [net/ipv4/tcp_input.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_input.c?h=v6.12) · Linux kernel · “Possible SYN flooding on port …” log message - [net/ipv4/tcp_diag.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_diag.c?h=v6.12) · Linux kernel · For a listening socket, Recv-Q in ss is the number of connections waiting for accept, and Send-Q is the backlog limit ## Sources by chapter ### #l-client-game - [Slow rendering](https://developer.android.com/topic/performance/vitals/render) · Android (Google) · Hitting 60 FPS means drawing each frame within 16 ms; late frames get skipped and show up as stutter - [Interpolation and extrapolation (Netcode for Entities 6.5)](https://docs.unity3d.com/Packages/com.unity.netcode@6.5/manual/interpolation.html) · Unity · Interpolation, which draws the in-between states of snapshots that arrive at intervals, and extrapolation, which keeps going in the same direction and speed when data is late, plus its limit - [Introduction to prediction (Netcode for Entities 6.5)](https://docs.unity3d.com/Packages/com.unity.netcode@6.5/manual/intro-to-prediction.html) · Unity · Prediction, where the client moves on its own input without waiting for the server’s result, and correction when it differs from the server - [Peeking into VALORANT's Netcode](https://technology.riotgames.com/news/peeking-valorants-netcode) · Riot Games · Buffering unevenly arriving data smooths it out but adds that much latency, and wrong guesses make characters jump or slide - [Garbage collection modes](https://docs.unity3d.com/Manual/performance-incremental-garbage-collection.html) · Unity · Simulation’s GC spike: with incremental GC off, the main thread stops while the whole heap is scanned, blowing past the 16 ms frame budget - [Scripting.GarbageCollector.CollectIncremental](https://docs.unity3d.com/ScriptReference/Scripting.GarbageCollector.CollectIncremental.html) · Unity · Simulation’s 3 ms incremental GC time slice: incrementalTimeSliceNanoseconds defaults to 3 ms - [Shader loading](https://docs.unity3d.com/Manual/shader-loading.html) · Unity · Simulation’s new-area loading: the first time a shader variant is used, the game can stall while the driver builds it for the GPU - [Frame Pacing library](https://developer.android.com/games/sdk/frame-pacing) · Android (Google) · Simulation’s V-Sync boundary: a 60 Hz display shows the previous frame again when no new frame is ready - [Set fixed timestep to optimize physics simulation frequency](https://docs.unity3d.com/Manual/physics-optimization-cpu-frequency.html) · Unity · Simulation’s fixed-step catch-up: when a frame takes longer than the step interval, the step runs several times in one frame, adding load - [Handling variation in time](https://docs.unity3d.com/Manual/time-handling-variations.html) · Unity · Simulation’s catch-up cap (the simulation allows up to 5 per frame): Unity caps one frame’s game time at 1/3 second to stop a catch-up spiral, and the game clock falls behind by the excess ### #l-client-os - [Multitasking](https://learn.microsoft.com/en-us/windows/win32/procthread/multitasking) · Microsoft · Preemptive multitasking: each thread gets a time slice (about 20 ms, varying by OS and CPU), and when it’s used up the CPU moves to the next thread - [Scheduling Priorities](https://learn.microsoft.com/en-us/windows/win32/procthread/scheduling-priorities) · Microsoft · Among runnable threads, those with the highest priority take turns getting time slices (round robin) - [Priority Boosts](https://learn.microsoft.com/en-us/windows/win32/procthread/priority-boosts) · Microsoft · Raises the priority of the foreground window’s process to at least that of background processes - [socket(7) — Linux manual page](https://man7.org/linux/man-pages/man7/socket.7.html) · Linux man-pages · Each socket has a receive buffer (SO_RCVBUF) whose default and maximum sizes are set by system settings - [Wi-Fi low-latency mode](https://source.android.com/docs/core/connect/wifi-low-latency) · Android (Google) · In Wi-Fi low-latency mode on Android 10 and later, the framework explicitly turns off Wi-Fi power saving (doze) while the app is in the foreground and the screen is on - [Cached apps freezer](https://source.android.com/docs/core/perf/cached-apps-freezer) · Android (Google) · Android 14 and later freeze cached apps after 10 seconds so they can’t use the CPU - [Extending your app’s background execution time](https://developer.apple.com/documentation/uikit/extending-your-app-s-background-execution-time) · Apple · iOS suspends an app a few seconds after it moves to the background - [_WDF_TIMER_CONFIG (wdftimer.h)](https://learn.microsoft.com/en-us/windows-hardware/drivers/ddi/wdftimer/ns-wdftimer-_wdf_timer_config) · Microsoft · Simulation’s 15.6 ms timer: the Windows system clock tick has a default interval of 15.6 ms - [timeBeginPeriod function (timeapi.h)](https://learn.microsoft.com/en-us/windows/win32/api/timeapi/nf-timeapi-timebeginperiod) · Microsoft · Simulation’s 1 ms timer: a program can request a higher timer resolution with timeBeginPeriod - [Customize the Windows performance power slider](https://learn.microsoft.com/en-us/windows-hardware/customize/desktop/customize-power-slider) · Microsoft · Simulation’s power-saving mode: Windows power modes change power and CPU settings to extend battery life at the cost of performance - [Thermal API](https://developer.android.com/games/optimize/adpf/thermal) · Android (Google) · Simulation’s phone heat: devices sustain high performance only for a limited time before thermal throttling kicks in ### #l-memory - [Designs, Lessons and Advice from Building Large Distributed Systems (LADIS 2009 keynote)](https://www.cs.cornell.edu/projects/ladis2009/talks/dean-keynote-ladis2009.pdf) · Google · Basis for the latency numbers table: L1 cache 0.5 ns, L2 7 ns, main memory 100 ns, round trip within the same data center 0.5 ms, disk seek 10 ms (as of 2009; the table’s cache figures are slightly different approximations) - [Solidigm™ D7-P5520 and D7-P5620 Product Brief](https://www.solidigm.com/products/data-center/product-briefs/d7-p5520-p5620-product-brief.html) · Solidigm · Four-nines (99.99%) latency of 130 µs for a server NVMe SSD: basis for the table’s SSD read and swap read-back of around 100 µs - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_rto_min_us defaults to 200,000 µs: Linux TCP’s minimum retransmission wait of 200 ms (the table’s TCP retransmission row) - [What is NUMA?](https://docs.kernel.org/mm/numa.html) · Linux kernel · Memory attached to another CPU (remote) is slower to access and has lower bandwidth than local memory (the table’s NUMA row) - [JEP 439: Generational ZGC](https://openjdk.org/jeps/439) · OpenJDK · G1 pauses range from a few ms to a few seconds, and ZGC pauses are 1 ms or less (the body’s full-heap GC of hundreds of ms to a few seconds, the table’s 1-second large-heap GC) - [Available Collectors](https://docs.oracle.com/en/java/javase/25/gctuning/available-collectors.html) · Oracle · ZGC gives up a little throughput to keep maximum pauses under 1 ms, independent of heap size - [Garbage Collector Implementation](https://docs.oracle.com/en/java/javase/25/gctuning/garbage-collector-implementation.html) · Oracle · Generational collection: minor collections that cover only the young generation are short, and major collections that cover the whole heap take much longer (the GC simulation’s generational mode) - [The Z Garbage Collector](https://docs.oracle.com/en/java/javase/25/gctuning/z-garbage-collector1.html) · Oracle · ZGC does its expensive work concurrently and doesn’t pause for more than 1 ms, but if reclamation falls behind, the application can stall waiting for GC (the GC simulation’s concurrent mode) - [A Guide to the Go Garbage Collector](https://go.dev/doc/gc-guide) · Go · The GC mark phase uses 25% of the CPU, slowing the program meanwhile, and with heavy allocation, goroutines get delayed helping the GC (assist) (the GC simulation’s concurrent mode) - [Go 1.8 Release Notes](https://go.dev/doc/go1.8) · Go · Go GC pauses are usually under 100 µs - [Debug a memory leak in .NET](https://learn.microsoft.com/en-us/dotnet/core/diagnostics/debug-memory-leak) · .NET · Even with GC, leaks happen when code keeps referencing objects it no longer needs - [Concepts overview](https://docs.kernel.org/admin-guide/mm/concepts.html) · Linux kernel · When memory runs short, the kernel reclaims page cache and swappable pages, and if that isn’t enough, the OOM killer force-kills a process (leak simulation) ### #l-disk - [Exos X18 Data Sheet](https://www.seagate.com/www-content/datasheets/pdfs/exos-x18-mango-DS2045-1N-2007US-en_US.pdf) · Seagate · 4K random reads at 170 IOPS and 4.16 ms average rotational latency for a 7,200 rpm HDD (the body’s HDD figure of a little over 150, the disk simulation’s HDD) - [D3-S4520 SSD](https://www.solidigm.com/products/data-center/d3/s4520.html) · Solidigm · SATA SSD 4 KB random read/write up to 92K/48K IOPS (the body’s tens of thousands for SSDs) - [Solidigm™ D7-P5520 and D7-P5620 Product Brief](https://www.solidigm.com/products/data-center/product-briefs/d7-p5520-p5620-product-brief.html) · Solidigm · NVMe SSD random read/write of 1,000K/200K IOPS (the body’s hundreds of thousands) - [Amazon EBS General Purpose SSD volumes](https://docs.aws.amazon.com/ebs/latest/userguide/general-purpose.html) · AWS · gp3 has a baseline of 3,000 IOPS; gp2 bursts to 3,000 IOPS on I/O credits and drops to baseline performance when the credits run out - [Amazon EBS-optimized instance types](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ebs-optimized.html) · AWS · Smaller instances deliver maximum EBS performance for only 30 minutes once every 24 hours, then return to baseline (e.g., t4g.2xlarge with a baseline of 4,000 and a maximum of 15,700 IOPS; the disk simulation’s burstable cloud assumption) - [Managed disk bursting](https://learn.microsoft.com/en-us/azure/virtual-machines/disk-bursting) · Microsoft Azure · Small disks and VMs can burst on credits for up to 30 minutes - [Concepts overview](https://docs.kernel.org/admin-guide/mm/concepts.html) · Linux kernel · Data written to a file lands in the page cache first, is marked dirty, and is written to disk later - [Documentation for /proc/sys/vm/](https://docs.kernel.org/admin-guide/sysctl/vm.html) · Linux kernel · When pending writes reach dirty_ratio, the writing process has to do the disk writes itself (the disk simulation’s OS write limit) - [fsync(2) — Linux manual page](https://man7.org/linux/man-pages/man2/fsync.2.html) · Linux man-pages · fsync blocks until the device reports that the write is complete ### #l-db - [How MySQL Uses Indexes](https://dev.mysql.com/doc/refman/8.4/en/mysql-indexes.html) · MySQL · Without an index, the whole table is read starting from the first row (full table scan) - [InnoDB Locking](https://dev.mysql.com/doc/refman/8.4/en/innodb-locking.html) · MySQL · Requests to modify a row wait until the transaction holding its row lock finishes (hot row) - [Number Of Database Connections](https://wiki.postgresql.org/wiki/Number_Of_Database_Connections) · PostgreSQL · Once DB resources are used up, adding connections lowers throughput (basis for the DB simulation, where a bigger pool slows everything down when CPU cores are short) - [WAL Configuration (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/wal-configuration.html) · PostgreSQL · A checkpoint is an expensive operation that writes out dirty pages in bulk every 5 minutes or 1 GB of WAL by default; spreading out the writes avoids I/O spikes - [Semisynchronous Replication](https://dev.mysql.com/doc/refman/8.4/en/replication-semisync.html) · MySQL · With asynchronous replication, if the primary dies, committed transactions may be missing on the standby (rolled back after failover) - [Log-Shipping Standby Servers (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/warm-standby.html) · PostgreSQL · Streaming replication is asynchronous by default, so there’s a delay between a commit and the replica reflecting it (data just written doesn’t show up on the replica) - [Asynchronous Commit (PostgreSQL Documentation)](https://www.postgresql.org/docs/current/wal-async-commit.html) · PostgreSQL · The trade-off of batching writes and flushing them late: faster, but recent changes are lost in a failure (the same structure as a game server that saves every few minutes) ### #judge - [Service Level Objectives (Site Reliability Engineering, ch. 4)](https://sre.google/sre-book/service-level-objectives/) · Google · Percentiles (50th, 95th, 99th) over averages to see the shape and tail of the latency distribution - [The Tail at Scale](https://research.google/pubs/the-tail-at-scale/) · Google · Occasional long delays (tail latency) increasingly dominate how the whole service feels as scale grows - [RFC 3550: RTP, A Transport Protocol for Real-Time Applications](https://www.rfc-editor.org/rfc/rfc3550) · IETF · Definition and calculation of interarrival jitter - [RFC 1812: Requirements for IP Version 4 Routers](https://www.rfc-editor.org/rfc/rfc1812) · IETF · Routers must be able to rate-limit ICMP error messages such as Time Exceeded, and may also limit Echo Replies (read mtr and ping results with care) - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · icmp_ratelimit, icmp_ratemask: Linux rate-limits ICMP responses such as Time Exceeded and Destination Unreachable by default - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · The rtt (average round-trip time)/rttvar fields in ss -ti - [pidstat(1) — Linux manual page](https://man7.org/linux/man-pages/man1/pidstat.1.html) · sysstat · pidstat -t: CPU utilization per thread - [bcc tools: runqlat examples](https://raw.githubusercontent.com/iovisor/bcc/master/tools/runqlat_example.txt) · IO Visor · Measuring the distribution of run queue latency (time a thread spends waiting for a CPU) - [RIPE Atlas documentation](https://atlas.ripe.net/docs/) · RIPE NCC · A public synthetic monitoring tool that runs ping and traceroute from probes around the world - [GeoLite2 Free Geolocation Data](https://dev.maxmind.com/geoip/geolite2-free-geolocation-data) · MaxMind · A public database that maps IP addresses to countries and ASNs - [View CloudWatch metrics for your instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/viewing_metrics_with_cloudwatch.html) · AWS · Basic monitoring at 5-minute intervals, detailed monitoring at 1-minute intervals (the aggregation interval hides short spikes) ### #l-nic - [NAPI](https://docs.kernel.org/networking/napi.html) · Linux kernel · When a device signals new packets with an interrupt, the kernel pulls and processes them through NAPI; interrupt coalescing is usually done by the device - [Scaling in the Linux Networking Stack](https://docs.kernel.org/networking/scaling.html) · Linux kernel · How RSS spreads multiple receive queues across multiple cores, with an interrupt per queue and a hash to pick the queue - [Interface statistics](https://docs.kernel.org/networking/statistics.html) · Linux kernel · Packets the device dropped for lack of buffers (rx_missed_errors) and per-driver statistics in ethtool -S - [ethtool(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ethtool.8.html) · ethtool · Setting and checking ring buffers (-G), interrupt coalescing (-C), receive hashing (-N), and statistics (-S) - [How to receive a million packets per second](https://blog.cloudflare.com/how-to-receive-a-million-packets/) · Cloudflare · Measurements showing one queue on one core tops out at about 350,000–430,000 packets per second, and receiving 1 million pps takes more queues and cores (the simulation assumes a more generous 700,000 pps per core) - [Monitor network performance for ENA settings on your EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-network-performance-ena.html) · AWS · When a cloud instance exceeds its bandwidth, PPS, or connection tracking limits, traffic is queued outside the instance and then dropped; limit-exceeded counters ### #l-server-os - [listen(2) — Linux manual page](https://man7.org/linux/man-pages/man2/listen.2.html) · Linux man-pages · The connection queue (backlog) and the somaxconn cap (default 4,096 since 5.4, 128 before that) - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · When the accept queue is full, Linux drops connection requests (SYN) and increments TcpExtListenOverflows - [listen function (winsock2.h)](https://learn.microsoft.com/en-us/windows/win32/api/winsock2/nf-winsock2-listen) · Microsoft · On Windows, when the queue is full, the client gets WSAECONNREFUSED - [getrlimit(2) — Linux manual page](https://man7.org/linux/man-pages/man2/getrlimit.2.html) · Linux man-pages · RLIMIT_NOFILE: the limit on how many fds a process can open; going over it returns EMFILE - [systemd-system.conf(5) — Linux manual page](https://man7.org/linux/man-pages/man5/systemd-system.conf.5.html) · systemd · The default fd limit for services is 1024:524288 (the simulation’s misconfigured fd limit of 1,024) - [The /proc Filesystem](https://docs.kernel.org/filesystems/proc.html) · Linux kernel · The OOM killer picks the process to kill by a score (badness) based on its share of memory use, adjustable with oom_score_adj - [Control Group v2](https://docs.kernel.org/admin-guide/cgroup-v2.html) · Linux kernel · When memory.max is reached and memory can’t be reclaimed, the OOM killer runs inside that cgroup; cpu.max sets the CPU limit - [proc_stat(5) — Linux manual page](https://man7.org/linux/man-pages/man5/proc_stat.5.html) · Linux man-pages · steal: CPU time taken by other operating systems in a virtualized environment - [Exponential Backoff And Jitter](https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/) · AWS · Exponential backoff alone still bunches retries together; adding randomness (jitter) reduces contention (the simulation’s retry method) ### #l-socket - [RFC 9293: Transmission Control Protocol (TCP)](https://www.rfc-editor.org/rfc/rfc9293) · IETF · TCP is a reliable, in-order byte stream; definitions of Nagle and delayed ACK - [RFC 768: User Datagram Protocol](https://www.rfc-editor.org/rfc/rfc768) · IETF · UDP doesn’t guarantee delivery or protect against duplicates - [RFC 8085: UDP Usage Guidelines](https://www.rfc-editor.org/rfc/rfc8085) · IETF · UDP applications that need reliability or ordering have to implement it themselves - [tcp(7) — Linux manual page](https://man7.org/linux/man-pages/man7/tcp.7.html) · Linux man-pages · TCP socket options such as TCP_NODELAY, TCP_USER_TIMEOUT, and keepalive - [socket(7) — Linux manual page](https://man7.org/linux/man-pages/man7/socket.7.html) · Linux man-pages · SO_SNDBUF, SO_RCVBUF, SO_KEEPALIVE, SO_LINGER socket options - [RFC 5681: TCP Congestion Control](https://www.rfc-editor.org/rfc/rfc5681) · IETF · Fast retransmit on 3 duplicate ACKs, and a congestion window of 1 segment after a timer-based retransmission (the simulation’s textbook rules) - [RFC 8985: The RACK-TLP Loss Detection Algorithm for TCP](https://www.rfc-editor.org/rfc/rfc8985) · IETF · RACK, which detects loss from send times without counting duplicate ACKs - [RFC 6298: Computing TCP's Retransmission Timer](https://www.rfc-editor.org/rfc/rfc6298) · IETF · Recommended minimum RTO of 1 second, with backoff that doubles on each expiry - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · Linux tcp_rto_min_us defaults to 200 ms; loss detection is RACK (tcp_recovery) - [include/net/tcp.h (Linux v6.12)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/net/tcp.h?h=v6.12) · Linux kernel · The simulation’s Linux delayed ACK of 40 ms: TCP_DELACK_MIN (HZ/25 = 40 ms), maximum TCP_DELACK_MAX (HZ/5 = 200 ms) - [Design issues - Sending small data segments over TCP with Winsock](https://learn.microsoft.com/en-us/previous-versions/troubleshoot/windows/win32/data-segment-tcp-winsock) · Microsoft · The simulation’s Windows delayed ACK of 200 ms: receiving data starts a 200 ms delayed ACK timer, and combined with Nagle, small packets wait for the ACK - [RFC 896: Congestion Control in IP/TCP Internetworks](https://www.rfc-editor.org/rfc/rfc896) · IETF · The original purpose of the Nagle rule: the remote terminal problem where every 1-byte keystroke sent a 41-byte packet - [send(2) — Linux manual page](https://man7.org/linux/man-pages/man2/send.2.html) · Linux man-pages · When the send buffer is full, a blocking send() doesn’t return ### #journey - [ITU-T G.114: One-way transmission time](https://www.itu.int/rec/T-REC-G.114-200305-I/en) · ITU · Fiber propagation delay of 5 µs/km (5 ms one way over 1,000 km); a one-way delay of 150 ms or less is nearly transparent to most applications, but highly interactive tasks can be affected even below 100 ms - [Azure network round-trip latency statistics](https://learn.microsoft.com/en-us/azure/networking/azure-network-latency) · Microsoft Azure · Scale of the internet path by server location in the simulation: from Seoul, the Busan region 8 ms, Tokyo 30 ms, Singapore 68 ms, US West 124–136 ms, Europe 234–244 ms (median round trip) - [The Internet at the Speed of Light (HotNets 2014)](https://conferences.sigcomm.org/hotnets/2014/papers/hotnets-XIII-final111.pdf) · ACM · The simulation’s path multiplier of 1.5: real router paths are about 1.5 times the straight-line fiber distance at the median - [Report ITU-R M.2134: Requirements related to technical performance for IMT-Advanced radio interface(s)](https://www.itu.int/pub/R-REP-M.2134-2008) · ITU · LTE-Advanced one-way radio latency requirement under 10 ms (unloaded, small packets; the simulation’s LTE value assumes load and scheduling waits on top of this) - [Report ITU-R M.2410: Minimum requirements related to technical performance for IMT-2020 radio interface(s)](https://www.itu.int/pub/R-REP-M.2410-2017) · ITU · 5G (IMT-2020) one-way radio latency requirement of 4 ms (eMBB, unloaded) - [An Experimental Study of Home Gateway Characteristics (IMC 2010)](https://conferences.sigcomm.org/imc/2010/papers/p260.pdf) · ACM · The simulation’s router queue: when uploads and downloads overlap, home router queuing delay can grow to hundreds of ms (up to about 400 ms) - [Maximum transmission unit and maximum segment size](https://developers.cloudflare.com/magic-transit/reference/mtu-mss/) · Cloudflare · The simulation’s route through DDoS protection: only incoming traffic passes through the protection network, and server responses go straight out to the internet (DSR) ### #l-home - [RFC 7567: IETF Recommendations Regarding Active Queue Management](https://www.rfc-editor.org/rfc/rfc7567) · IETF · Queues building up in network equipment are a major cause of internet latency; recommends queue management (AQM) by default - [RFC 8289: Controlled Delay Active Queue Management](https://www.rfc-editor.org/rfc/rfc8289) · IETF · CoDel’s 5 ms target queuing delay and 100 ms interval (the simulation’s SQM queue of about 5 ms) - [RFC 8290: The Flow Queue CoDel Packet Scheduler and Active Queue Management Algorithm](https://www.rfc-editor.org/rfc/rfc8290) · IETF · FQ-CoDel: splits traffic into per-flow queues by address and port and sends small flows that don’t build a queue first - [What Can I Do About Bufferbloat?](https://www.bufferbloat.net/projects/bloat/wiki/What_can_I_do_about_Bufferbloat/) · Bufferbloat.net · Use a router that supports SQM such as cake or fq_codel, and tune the SQM rate while measuring latency under load - [An Experimental Study of Home Gateway Characteristics (IMC 2010)](https://conferences.sigcomm.org/imc/2010/papers/p260.pdf) · ACM · Measurements of 34 home routers: queuing delay of up to about 400 ms with overlapping uploads and downloads, median UDP mapping lifetime of 90 seconds - [Ending the Anomaly: Achieving Low Latency and Airtime Fairness in WiFi (USENIX ATC 2017)](https://www.usenix.org/system/files/conference/atc17/atc17-hoiland-jorgensen.pdf) · USENIX · The Wi-Fi radio queue also adds hundreds of ms of latency under load, and slow devices eat into other devices’ airtime - [RFC 4787: Network Address Translation (NAT) Behavioral Requirements for Unicast UDP](https://www.rfc-editor.org/rfc/rfc4787) · IETF · NAT UDP mapping timer requirements (at least 2 minutes, with 5 minutes or more recommended as the default) and refresh by packets going out from inside - [RFC 8325: Mapping Diffserv to IEEE 802.11](https://www.rfc-editor.org/rfc/rfc8325) · IETF · Wi-Fi uses CSMA/CA: it defers while the channel is busy and transmits after a random backoff - [Resolve Wi-Fi and Bluetooth issues caused by wireless interference](https://support.apple.com/en-us/102319) · Apple · The simulation’s microwave interference: microwave ovens, Bluetooth, and similar devices interfere with 2.4 GHz Wi-Fi, and moving to 5 GHz helps - [RFC 9438: CUBIC for Fast and Long-Distance Networks](https://www.rfc-editor.org/rfc/rfc9438) · IETF · The simulation’s upload rate control: CUBIC shrinks its sending window to 0.7 times (about a 30% cut) after a loss, then grows it again ### #l-isp - [ITU-T G.114: One-way transmission time](https://www.itu.int/rec/T-REC-G.114-200305-I/en) · ITU · Fiber propagation delay of 5 µs/km: light travels through fiber at about 200,000 km per second, so a 1,000 km round trip takes at least 10 ms - [The Internet at the Speed of Light (HotNets 2014)](https://conferences.sigcomm.org/hotnets/2014/papers/hotnets-XIII-final111.pdf) · ACM · Real router paths are about 1.5 times the straight-line fiber distance at the median, and minimum ping is 3.2 times the speed-of-light bound (about 2 times the straight-line fiber distance) - [Azure network round-trip latency statistics](https://learn.microsoft.com/en-us/azure/networking/azure-network-latency) · Microsoft Azure · Measured median round trips from Seoul: Tokyo 30 ms, US West 124–136 ms, Europe 234–244 ms (Korea–Europe is about 2.8 times the straight-line fiber distance) - [AAE-1 & SMW5 cable cuts impact millions of users across multiple countries](https://blog.cloudflare.com/aae-1-smw5-cable-cuts/) · Cloudflare · Europe–Asia traffic mostly passes through Egypt, and submarine cable repairs take days to weeks - [Quantifying the Causes of Path Inflation (SIGCOMM 2003)](https://conferences.sigcomm.org/sigcomm/2003/papers/p113-spring.pdf) · ACM · Peering policies between ISPs and interdomain routing inflate paths significantly - [Inferring Persistent Interdomain Congestion (SIGCOMM 2018)](https://www.caida.org/catalog/papers/2018_inferring_persistent_interdomain_congestion/inferring_persistent_interdomain_congestion.pdf) · ACM · Some interconnects between ISPs show recurring congestion, with latency and loss rising at peak hours every day - [RFC 4271: A Border Gateway Protocol 4 (BGP-4)](https://www.rfc-editor.org/rfc/rfc4271) · IETF · BGP, which exchanges internet routing information, with a suggested default hold time of 90 seconds - [BGP updates in 2024](https://blog.apnic.net/2025/01/07/bgp-updates-in-2024/) · APNIC · After a path change, routing takes a daily average of 25–35 seconds to settle again on IPv4 and 40–50 seconds on IPv6 - [Delayed Internet Routing Convergence (SIGCOMM 2000)](https://conferences.sigcomm.org/sigcomm/2000/conf/paper/sigcomm2000-5-2.pdf) · ACM · Convergence after a route failure can take up to several minutes, with higher loss and latency in the meantime (measured in 2000) ### #l-dc-net - [High-Resolution Measurement of Data Center Microbursts (IMC 2017)](https://conferences.sigcomm.org/imc/2017/papers/imc17-final60.pdf) · ACM · End-to-end latency in data center networks is under 1 ms, and over 70% of bursts end within tens of µs, invisible in average utilization - [Data Center TCP (DCTCP) (SIGCOMM 2010)](https://conferences.sigcomm.org/sigcomm/2010/papers/sigcomm/p63.pdf) · ACM · Commodity switches share shallow buffers across ports, and when many flows converge on one port for a brief moment, the buffer overflows and packets are lost - [Netfilter Conntrack Sysfs variables](https://docs.kernel.org/networking/nf_conntrack-sysctl.html) · Linux kernel · Defaults for the connection tracking table’s maximum entries and timeouts (UDP 30 seconds, stream 120 seconds, established TCP 5 days) - [Edit attributes for your Application Load Balancer](https://docs.aws.amazon.com/elasticloadbalancing/latest/application/edit-load-balancer-attributes.html) · AWS · ALB idle timeout defaults to 60 seconds; when it expires, the load balancer closes the connection - [Network Load Balancers](https://docs.aws.amazon.com/elasticloadbalancing/latest/network/network-load-balancers.html) · AWS · The simulation’s load balancer defaults: NLB TCP 350 seconds, UDP 120 seconds (fixed); after the idle timeout, the NLB silently stops tracking the flow - [Configure load balancer TCP reset and idle timeout](https://learn.microsoft.com/en-us/azure/load-balancer/load-balancer-tcp-idle-timeout) · Microsoft Azure · Azure Load Balancer idle timeout defaults to 4 minutes - [Amazon EC2 security group connection tracking](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/security-group-connection-tracking.html) · AWS · Security group connection tracking defaults, and a note that TCP idle timeouts on load balancers and firewalls are commonly 60–90 minutes - [Flow-Based Sessions](https://www.juniper.net/documentation/us/en/software/junos/flow-packet-processing/topics/topic-map/security-flow-based-session-for-srx-series-devices.html) · Juniper Networks · The simulation’s corporate firewall: SRX firewall default session timeouts of 1,800 seconds (30 minutes) for TCP and 60 seconds for UDP - [An Experimental Study of Home Gateway Characteristics (IMC 2010)](https://conferences.sigcomm.org/imc/2010/papers/p260.pdf) · ACM · The simulation’s 1-hour home router TCP timeout: median TCP mapping of about 60 minutes. UDP mappings ranged from 30 to 691 seconds by device, with medians of 90 seconds one-way and about 180 seconds bidirectional - [A Multi-perspective Analysis of Carrier-Grade NAT Deployment (IMC 2016)](https://www.icir.org/vern/papers/cgn-imc16.pdf) · ACM · The simulation’s 30-second CGNAT UDP and 1-minute home router UDP timeouts: median CGN UDP mapping of 35 seconds on fixed networks and 65 seconds on mobile networks, 74% of measured NATs at 1 minute or less, most home router (CPE) NATs at 65 seconds - [tcp(7) — Linux manual page](https://man7.org/linux/man-pages/man7/tcp.7.html) · Linux man-pages · The simulation’s TCP keepalive: by default, after 2 hours (7,200 seconds) idle, probes 9 times at 75-second intervals and drops the connection if there’s no response - [Cached apps freezer](https://source.android.com/docs/core/perf/cached-apps-freezer) · Android (Google) · The simulation’s mobile background: Android 14 and later freeze app processes that enter the cached state after 10 seconds - [RFC 5880: Bidirectional Forwarding Detection (BFD)](https://www.rfc-editor.org/rfc/rfc5880) · IETF · BFD, which detects path failures faster than routing protocols’ Hellos, which work on the order of seconds - [RFC 2923: TCP Problems with Path MTU Discovery](https://www.rfc-editor.org/rfc/rfc2923) · IETF · Path MTU black holes, where blocked ICMP makes only large packets vanish ### #basics - [Teletraffic Engineering Handbook (ITU-D Study Group 2 Question 16/2)](https://www.itu.int/dms_pub/itu-d/opb/stg/D-STG-SG02.16.2.1-2002-PDF-E.pdf) · ITU · M/M/1 mean wait W = A·s/(1−A): at 50%, 80%, and 90% utilization, 1, 4, and 9 times the service time. At the same utilization, more workers (servers) and more even arrivals mean shorter waits - [CloudWatch metrics that are available for your instances](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/viewing_metrics_with_cloudwatch.html) · AWS · EC2 CPUUtilization is an instance-wide value, aggregated every 5 minutes by default or every 1 minute with detailed monitoring - [Designs, Lessons and Advice from Building Large Distributed Systems (Jeff Dean, LADIS 2009 keynote)](https://www.cs.cornell.edu/projects/ladis2009/talks/dean-keynote-ladis2009.pdf) · Google · Main memory reference 100 ns, round trip within the same data center 500,000 ns (0.5 ms) - [ITU-T G.114: One-way transmission time](https://www.itu.int/rec/T-REC-G.114-200305-I/en) · ITU · Fiber propagation delay of 5 µs/km (basis for the speed-of-light limit calculation) - [Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization (Yahn W. Bernier, GDC 2001)](https://web.cs.wpi.edu/~claypool/courses/4513-B03/papers/games/bernier.pdf) · Valve · Half-Life defaults: 20 updates per second, 100 ms interpolation. At 10 per second, 200 ms of interpolation survives one missed update - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_rto_min_us defaults to 200000 (200 ms) - [Factors influencing the latency of simple reaction time (Frontiers in Human Neuroscience, 2015)](https://doi.org/10.3389/fnhum.2015.00131) · Frontiers · Mean simple reaction time of about 231 ms (213 ms after correcting for equipment latency) - [A Survey and Taxonomy of Latency Compensation Techniques for Network Computer Games (ACM Computing Surveys, 2022)](https://web.cs.wpi.edu/~claypool/papers/lag-taxonomy/paper.pdf) · ACM · Latency under 100 ms still affects performance on game tasks, and for tasks like dragging, about 10 ms is noticeable - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · Skilled players notice differences of about 10 ms in blind tests - [Latency and Player Actions in Online Games (Communications of the ACM, 2006)](https://web.cs.wpi.edu/~claypool/papers/precision-deadline/final.pdf) · ACM · Latency tolerance by genre: about 100 ms for first-person, about 500 ms for third-person (RPG, MMO), about 1,000 ms for RTS - [JEP 333: ZGC: A Scalable Low-Latency Garbage Collector](https://openjdk.org/jeps/333) · OpenJDK · G1 on a 128 GB heap averages 157 ms pauses with a maximum of 544 ms, while ZGC stays around 1–2 ms regardless of heap and live data size - [CommonNetworkParametersExtensions (Unity Transport 2.5)](https://docs.unity3d.com/Packages/com.unity.transport@2.5/api/Unity.Networking.Transport.CommonNetworkParametersExtensions.html) · Unity · disconnectTimeoutMS: disconnects if nothing is received for this long (default 30,000 ms) ### #lab - [Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization (Yahn W. Bernier, GDC 2001)](https://web.cs.wpi.edu/~claypool/courses/4513-B03/papers/games/bernier.pdf) · Valve · Interpolation is smooth but shows the past; extrapolation can’t anticipate direction changes and jumps when it guesses wrong. Prediction errors are corrected with the server’s result - [UDP vs. TCP](https://gafferongames.com/post/udp_vs_tcp/) · Gaffer On Games · When TCP loses a packet, it holds back new data that has arrived until the retransmission comes in (usually 2×RTT or more) - [Deterministic Lockstep](https://gafferongames.com/post/deterministic_lockstep/) · Gaffer On Games · Sending unacknowledged inputs redundantly in every packet avoids waiting for retransmission (up to 2 seconds’ worth in the worst case) - [Snapshot Interpolation](https://gafferongames.com/post/snapshot_interpolation/) · Gaffer On Games · Drawing updates the moment they arrive stutters with jitter; an interpolation buffer adds a little latency but makes motion smooth - [Understanding Networked Movement in the Character Movement Component for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/understanding-networked-movement-in-the-character-movement-component-for-unreal-engine) · Epic Games · When client movement goes missing or goes wrong because of connection problems, the server corrects the position (rubber-banding) - [Source SDK 2013: player.cpp](https://raw.githubusercontent.com/ValveSoftware/source-sdk-2013/master/src/game/server/player.cpp) · Valve · A command budget that accrues each tick limits commands that arrive in a bunch. Too strict, and even legitimate players stutter - [Introducing Time Dilation (TiDi)](https://www.eveonline.com/news/view/introducing-time-dilation-tidi) · CCP Games · A design that slows the game clock when the server is overloaded (Time Dilation) so everything runs slower - [CommonNetworkParametersExtensions (Unity Transport 2.5)](https://docs.unity3d.com/Packages/com.unity.transport@2.5/api/Unity.Networking.Transport.CommonNetworkParametersExtensions.html) · Unity · An inactivity timeout that disconnects after nothing is received for a set time ### #symptoms - [Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization (Yahn W. Bernier, GDC 2001)](https://web.cs.wpi.edu/~claypool/courses/4513-B03/papers/games/bernier.pdf) · Valve · When updates go missing, entities stop at their last position (stutter) or extrapolate and then jump (teleporting). Extrapolation time must be capped - [Understanding Networked Movement in the Character Movement Component for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/understanding-networked-movement-in-the-character-movement-component-for-unreal-engine) · Epic Games · When client movement goes missing or differs from the server’s calculation, the server sends a correction that snaps the position back (rubber-banding) - [UDP vs. TCP](https://gafferongames.com/post/udp_vs_tcp/) · Gaffer On Games · TCP holds later data until the lost packet is retransmitted, then delivers it all at once (fast-forward) - [Deterministic Lockstep](https://gafferongames.com/post/deterministic_lockstep/) · Gaffer On Games · When delayed inputs arrive all at once, several frames are computed in a batch to catch up - [Introducing Time Dilation (TiDi)](https://www.eveonline.com/news/view/introducing-time-dilation-tidi) · CCP Games · Time Dilation (TiDi), which slows the game clock when the server is overloaded, and actions lagging by several seconds under overload - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · Buffering time depends on the server tick rate and the client’s render frame rate - [Using Gameplay Abilities in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/using-gameplay-abilities-in-unreal-engine) · Epic Games · The server can overturn abilities that were run on prediction (dropped action / rollback) - [Actor Relevancy in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-relevancy-in-unreal-engine) · Epic Games · Actors the server deems irrelevant aren’t replicated, or are removed on the client (invisible / ghost entities) - [CommonNetworkParametersExtensions (Unity Transport 2.5)](https://docs.unity3d.com/Packages/com.unity.transport@2.5/api/Unity.Networking.Transport.CommonNetworkParametersExtensions.html) · Unity · Disconnects once the inactivity timeout passes (disconnect) ### #sync - [Latency Compensating Methods in Client/Server In-game Protocol Design and Optimization (Yahn W. Bernier, GDC 2001)](https://web.cs.wpi.edu/~claypool/courses/4513-B03/papers/games/bernier.pdf) · Valve · Principles of authoritative servers, client-side prediction and reconciliation, interpolation, and lag compensation, and trade-offs such as “getting shot around the corner” - [What Every Programmer Needs To Know About Game Networking](https://gafferongames.com/post/what_every_programmer_needs_to_know_about_game_networking/) · Gaffer On Games · How netcode evolved from P2P lockstep to client/server and then to client-side prediction - [Using Gameplay Abilities in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/using-gameplay-abilities-in-unreal-engine) · Epic Games · The responsiveness versus accuracy trade-off between Local Predicted and Server Initiated - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · Server authority, prediction, buffering, rewind-based hit registration, and rewind limits - [NetworkTime and ticks (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/advanced-topics/networktime-ticks.html) · Unity · Scheduling events to play back at a given server time - [1500 Archers on a 28.8: Network Programming in Age of Empires and Beyond](https://www.gamedeveloper.com/programming/1500-archers-on-a-28-8-network-programming-in-age-of-empires-and-beyond) · Game Developer · Lockstep: commands scheduled two turns ahead, with turn length set by the slowest computer; a steady 500 ms latency is fine, but erratic latency feels bad - [Deterministic Lockstep](https://gafferongames.com/post/deterministic_lockstep/) · Gaffer On Games · Lockstep advances only when all inputs have arrived, and a playback delay buffer absorbs jitter - [GGPO Rollback Networking SDK](https://www.ggpo.net/) · GGPO · Rollback: predicts the opponent’s input and keeps going, rewinding and recomputing when the prediction was wrong - [8 Frames in 16ms: Rollback Networking in 'Mortal Kombat' and 'Injustice 2'](https://www.gdcvault.com/play/1025471/8-Frames-in-16ms-Rollback) · GDC · Rollback removes lockstep’s local input delay, re-simulating up to 8 frames - [Latency and Player Actions in Online Games (Communications of the ACM, 2006)](https://web.cs.wpi.edu/~claypool/papers/precision-deadline/final.pdf) · ACM · The same latency affects play differently depending on the precision and deadline of the action and the perspective (first-person, third-person, omnipresent) - [Networking Overview for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/networking-overview-for-unreal-engine) · Epic Games · The advantages and load of a listen server host - [Factors influencing the latency of simple reaction time (Frontiers in Human Neuroscience, 2015)](https://doi.org/10.3389/fnhum.2015.00131) · Frontiers · Human simple reaction time of about 0.23 seconds ### #partial - [Peeking into VALORANT's Netcode](https://www.riotgames.com/en/news/peeking-valorants-netcode) · Riot Games · When one player is out of sync, only that player gets corrected, and the other nine see smooth play. Rewind-based hit registration has a limit - [State Synchronization](https://gafferongames.com/post/state_synchronization/) · Gaffer On Games · Packets arrive bunched up, 2 in one frame and 0 in the next, and a jitter buffer evens them out - [Source SDK 2013: player.cpp](https://raw.githubusercontent.com/ValveSoftware/source-sdk-2013/master/src/game/server/player.cpp) · Valve · A command budget that spreads bunched-up commands across ticks, and the side effects of strict limits - [Source SDK 2013: player_lagcompensation.cpp](https://raw.githubusercontent.com/ValveSoftware/source-sdk-2013/master/src/game/server/player_lagcompensation.cpp) · Valve · The Source engine’s rewind cap sv_maxunlag defaults to 1 second - [A Survey and Taxonomy of Latency Compensation Techniques for Network Computer Games (ACM Computing Surveys, 2022)](https://web.cs.wpi.edu/~claypool/papers/lag-taxonomy/paper.pdf) · ACM · Lag compensation’s “getting shot around the corner,” and input delay that lines up everyone’s inputs - [send(2) — Linux manual page](https://man7.org/linux/man-pages/man2/send.2.html) · Linux man-pages · When the send buffer is full, send() waits in blocking mode and returns EAGAIN right away in nonblocking mode - [Networking Overview for Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/networking-overview-for-unreal-engine) · Epic Games · A listen server host has an advantage over other players - [Authority (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/manual/terms-concepts/authority.html) · Unity · Distributed authority: each client owns and computes some of the entities - [Actor Relevancy in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-relevancy-in-unreal-engine) · Epic Games · The server sends each connection only relevant actors and removes them on the client once they’re no longer relevant - [NetworkConfig class (Netcode for GameObjects 2.5)](https://docs.unity3d.com/Packages/com.unity.netcode.gameobjects@2.5/api/Unity.Netcode.NetworkConfig.html) · Unity · Messages for entities not yet spawned are held and then dropped after a timeout (SpawnTimeout) - [RFC 8085: UDP Usage Guidelines](https://www.rfc-editor.org/rfc/rfc8085) · IETF · A fragmented UDP packet is lost entirely if even one fragment is lost - [Using SO_REUSEADDR and SO_EXCLUSIVEADDRUSE](https://learn.microsoft.com/en-us/windows/win32/winsock/using-so-reuseaddr-and-so-exclusiveaddruse) · Microsoft · When two sockets share the same port, which one receives a given packet is unpredictable - [RFC 6269: Issues with IP Address Sharing](https://www.rfc-editor.org/rfc/rfc6269) · IETF · When multiple subscribers share one IP address, users can’t be told apart by IP alone - [Application.runInBackground](https://docs.unity3d.com/ScriptReference/Application-runInBackground.html) · Unity · Default behavior of apps pausing in the background - [Actor Priority in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/actor-priority-in-unreal-engine) · Epic Games · When bandwidth is saturated, only some actors are replicated, by priority ### #retrans - [RFC 6298: Computing TCP's Retransmission Timer](https://www.rfc-editor.org/rfc/rfc6298) · IETF · RTO = SRTT + max(G, 4·RTTVAR), initial value 1 second, recommended minimum 1 second, doubles on every expiry, any maximum must be at least 60 seconds, retransmitted packets are excluded from RTT samples (Karn) - [RFC 5681: TCP Congestion Control](https://www.rfc-editor.org/rfc/rfc5681) · IETF · Fast retransmit on the third duplicate ACK; after an RTO, the congestion window restarts from 1 segment (the loss window) - [RFC 6675: A Conservative Loss Recovery Algorithm Based on Selective Acknowledgment (SACK) for TCP](https://www.rfc-editor.org/rfc/rfc6675) · IETF · Loss recovery that uses SACK information to identify missing packets (duplicate ACK and SACK signals for fast retransmit) - [RFC 8985: The RACK-TLP Loss Detection Algorithm for TCP](https://www.rfc-editor.org/rfc/rfc8985) · IETF · RACK’s reordering allowance (min_RTT/4) and its adjustment based on DSACK, TLP wait of 2·SRTT (plus delayed ACK slack when only one packet is unacknowledged), RTO reset after sending a TLP, SACK required - [RFC 2018: TCP Selective Acknowledgment Options](https://www.rfc-editor.org/rfc/rfc2018) · IETF · SACK: the receiver reports the gaps in the middle - [RFC 2883: An Extension to the Selective Acknowledgement (SACK) Option for TCP](https://www.rfc-editor.org/rfc/rfc2883) · IETF · DSACK: reports receiving something already received, exposing spurious retransmissions - [RFC 5682: Forward RTO-Recovery (F-RTO): An Algorithm for Detecting Spurious Retransmission Timeouts with TCP](https://www.rfc-editor.org/rfc/rfc5682) · IETF · F-RTO: detecting spurious RTOs - [RFC 3522: The Eifel Detection Algorithm for TCP](https://www.rfc-editor.org/rfc/rfc3522) · IETF · Using timestamps to tell after the fact whether a recovery was unnecessary (the simulation’s congestion window undo) - [RFC 7323: TCP Extensions for High Performance](https://www.rfc-editor.org/rfc/rfc7323) · IETF · Timestamp and window scale options; without scaling, the window is at most 64 KiB - [RFC 6937: Proportional Rate Reduction for TCP](https://www.rfc-editor.org/rfc/rfc6937) · IETF · PRR: during recovery, scales the amount sent to the amount newly delivered (the simulation’s send limit during recovery) - [RFC 9438: CUBIC for Fast and Long-Distance Networks](https://www.rfc-editor.org/rfc/rfc9438) · IETF · CUBIC cuts the congestion window to 0.7 times on loss (the simulation’s 30% cut) - [RFC 9293: Transmission Control Protocol (TCP)](https://www.rfc-editor.org/rfc/rfc9293) · IETF · Zero window probe: probes are sent even when the window is 0, at exponentially increasing intervals - [RFC 9000: QUIC: A UDP-Based Multiplexed and Secure Transport](https://www.rfc-editor.org/rfc/rfc9000) · IETF · In QUIC, a loss blocks only the streams carried in that packet, and other streams keep going (stream separation) - [RFC 2475: An Architecture for Differentiated Services](https://www.rfc-editor.org/rfc/rfc2475) · IETF · Definitions: shaping delays packets, and policing drops the excess - [RFC 3168: The Addition of Explicit Congestion Notification (ECN) to IP](https://www.rfc-editor.org/rfc/rfc3168) · IETF · ECN: signals congestion without dropping packets - [Smart Queue Management](https://www.bufferbloat.net/projects/cerowrt/wiki/Smart_Queue_Management/) · Bufferbloat.net · Router SQM: per-flow scheduling, AQM, and shaping reduce queue overflow and bufferbloat - [RFC 1191: Path MTU discovery](https://www.rfc-editor.org/rfc/rfc1191) · IETF · Finding the path MTU with the “too big” ICMP (type 3 code 4) - [RFC 1812: Requirements for IP Version 4 Routers](https://www.rfc-editor.org/rfc/rfc1812) · IETF · Routers may rate-limit how often they generate ICMP error messages (mtr results where only one hop in the middle appears to lose packets) - [RFC 2991: Multipath Issues in Unicast and Multicast Next-Hop Selection](https://www.rfc-editor.org/rfc/rfc2991) · IETF · Diagnostic results such as ping and traceroute are hard to trust over multiple paths; pinning each flow to one path with a flow hash - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · tcp_recovery (default 0x1 RACK; from 6.17, RACK is the only loss detection, so setting 0 has no effect), tcp_early_retrans (default 3, 0 disables TLP), tcp_sack, tcp_dsack, tcp_timestamps (on by default), tcp_thin_linear_timeouts (fewer than 4 packets in flight, up to 6 linear retries), tcp_rto_max_ms, tcp_mtu_probing and tcp_base_mss, tcp_rto_min_us, tcp_retries2 (default 15, about 924.6 seconds), tcp_syn_linear_timeouts, tcp_notsent_lowat - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · TcpOutSegs excludes retransmissions; meanings of TCPLossProbes, TCPLossProbeRecovery, TCPLostRetransmit, TCPSpuriousRTOs, TCPDSACKRecv, and TCPSynRetrans - [net/ipv4/proc.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/proc.c?h=v6.12) · Linux kernel · Counter names as shown by nstat (RetransSegs, TCPTimeouts, TCPLossProbes, TCPSpuriousRTOs, TCPDSACKRecv, and so on) - [Thin-streams and TCP](https://docs.kernel.org/networking/tcp-thin.html) · Linux kernel · Thin streams such as games don’t trigger fast retransmit well and rely on long timeouts; TCP_THIN_LINEAR_TIMEOUTS - [include/net/tcp.h](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/net/tcp.h?h=v6.12) · Linux kernel · TCP_RTO_MIN 200 ms, TCP_RTO_MAX 120 seconds, TCP_TIMEOUT_INIT 1 second, delayed ACK 40–200 ms (TCP_DELACK_MIN and MAX), TCP_BASE_MSS 1,024, thin stream threshold (fewer than 4 packets in flight) - [net/ipv4/tcp_input.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_input.c?h=v6.12) · Linux kernel · RTO = SRTT + rttvar (no lower than the RTO minimum), RACK applies only to SACK connections, PRR reduces the congestion window during recovery - [net/ipv4/tcp_output.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_output.c?h=v6.12) · Linux kernel · TLP only on SACK connections, with a wait of 2·RTT, plus the RTO minimum when only one packet is unacknowledged - [net/ipv4/tcp_timer.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_timer.c?h=v6.12) · Linux kernel · When RTOs go on for tcp_retries1 times, MTU probing starts on black hole detection; linear timeouts for thin streams and SYNs - [net/ipv4/tcp_recovery.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_recovery.c?h=v6.12) · Linux kernel · RACK reordering window = min(min_RTT/4 × step count, SRTT); retransmissions acknowledged faster than the minimum RTT are excluded from the reference - [net/ipv4/tcp_ipv4.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_ipv4.c?h=v6.12) · Linux kernel · Default initialization: tcp_early_retrans 3, tcp_recovery RACK, tcp_syn_linear_timeouts 4, tcp_base_mss 1,024 (tcp_mtu_probing isn’t set explicitly, so it’s 0) - [net/ipv4/tcp_bbr.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/ipv4/tcp_bbr.c?h=v6.12) · Linux kernel · BBR’s pacing_rate = pacing_gain × bottleneck bandwidth (spreads sends evenly through pacing) - [tcp: use RACK to detect losses](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=4f41b1c58a32537542f14c1150099131613a5e8a) · Linux kernel · Introduction of RACK and tcp_recovery (Linux 4.4), initially working as a supplement to the existing method - [tcp: disable RFC6675 loss detection](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=b38a51fec1c1f693f03b1aa19d0622123634d4b7) · Linux kernel · Made RACK the default loss detection (2018, Linux 4.18) - [tcp: remove obsolete and unused RFC3517/RFC6675 loss recovery code](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=1c120191dcec510cc17d587ece48a7ae875a90c5) · Linux kernel · Removal of the RFC6675 loss recovery code (Linux 6.17), noting that RACK-TLP has been the default since 2018 - [tcp: make the first N SYN RTO backoffs linear](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=ccce324dabfe2143519daf50ed8b1ef1d0c542f7) · Linux kernel · The first 4 SYN RTO backoffs no longer double (after the first 1-second RTO, four more 1-second waits, then 2, 4 seconds …, Linux 6.5) - [tcp: add sysctl_tcp_rto_min_us](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=f086edef71be7174a16c1ed67ac65a085cda28b1) · Linux kernel · Server-wide RTO minimum tcp_rto_min_us (Linux 6.11) - [tcp: add the ability to control max RTO](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=54a378f43425085d0684679d99735696b69165bc) · Linux kernel · TCP_RTO_MAX_MS socket option, 1–120 seconds (Linux 6.15) - [tcp: support TCP_RTO_MIN_US for set/getsockopt use](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=f38805c5d26fe4af97837c10d58074a7496638bf) · Linux kernel · TCP_RTO_MIN_US socket option to set the RTO minimum per connection (Linux 6.15) - [Android common kernels](https://source.android.com/docs/core/architecture/kernel/android-common) · Android (Google) · Supported Android common kernels are 5.10 and later (newer than 4.18, where RACK became the default) - [tcp(7) — Linux manual page](https://man7.org/linux/man-pages/man7/tcp.7.html) · Linux man-pages · TCP_NODELAY (turns off Nagle), TCP_USER_TIMEOUT (sets only when to give up, leaving retransmission timing unchanged), TCP_KEEPIDLE, TCP_INFO - [ip-route(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ip-route.8.html) · iproute2 · Per-route rto_min option - [tc-fq(8) — Linux manual page](https://man7.org/linux/man-pages/man8/tc-fq.8.html) · iproute2 · Per-connection pacing in the fq qdisc and SO_MAX_PACING_RATE - [iptables-extensions(8) — Linux manual page](https://man7.org/linux/man-pages/man8/iptables-extensions.8.html) · netfilter · TCPMSS --clamp-mss-to-pmtu: works around segments that block ICMP by adjusting the MSS in the SYN - [ss(8) — Linux manual page](https://man7.org/linux/man-pages/man8/ss.8.html) · iproute2 · rto (ms), backoff (number of exponential backoffs), rtt/rttvar, and cwnd in ss -i - [misc/ss.c](https://git.kernel.org/pub/scm/network/iproute2/iproute2.git/tree/misc/ss.c?h=v6.12.0) · iproute2 · ss -ti output for retrans:current/total, lost, reordering, bytes_sent, and bytes_retrans - [nstat(8) — Linux manual page](https://man7.org/linux/man-pages/man8/nstat.8.html) · iproute2 · By default, nstat shows the increase since its previous run - [Demonstrations of tcpretrans, the Linux eBPF/bcc version](https://raw.githubusercontent.com/iovisor/bcc/master/tools/tcpretrans_example.txt) · IO Visor · One line per retransmission with address, port, and state; -c aggregates per flow, -l includes TLP attempts - [Interface statistics](https://docs.kernel.org/networking/statistics.html) · Linux kernel · Checking detailed errors with ip -s -s link, the meanings of rx_missed_errors and rx_crc_errors, and per-driver statistics in ethtool -S - [drivers/net/ethernet/intel/igb/igb_ethtool.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/drivers/net/ethernet/intel/igb/igb_ethtool.c?h=v6.12) · Linux kernel · Examples of per-driver counter names: rx_missed_errors, rx_no_buffer_count, rx_crc_errors - [Ethtool counters](https://docs.kernel.org/networking/device_drivers/ethernet/mellanox/mlx5/counters.html) · Linux kernel · mlx5’s rx_out_of_buffer and rx_discards_phy - [net/core/net-procfs.c](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/net/core/net-procfs.c?h=v6.12) · Linux kernel · /proc/net/softnet_stat has one line per CPU, in hex; the 2nd column is dropped and the 3rd is time_squeeze - [Monitor network performance for ENA settings on your EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-network-performance-ena.html) · AWS · Meanings of bw_in, bw_out, pps, conntrack, and linklocal_allowance_exceeded; viewing them in CloudWatch requires installing the CloudWatch agent - [High-Resolution Measurement of Data Center Microbursts](https://research.facebook.com/publications/high-resolution-measurement-of-data-center-microbursts/) · Meta · Most bursts end within tens of µs, so per-minute average utilization doesn’t reveal the cause of drops (IMC 2017) - [mtr(8) manual page source](https://raw.githubusercontent.com/traviscross/mtr/master/man/mtr.8.in) · mtr · -T (--tcp) for TCP SYN, -P (--port) to set the target port - [pathping](https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/pathping) · Microsoft · A Windows path diagnostic command that sends many probes and computes loss and latency per hop - [Display Filter Reference: Transmission Control Protocol](https://www.wireshark.org/docs/dfref/t/tcp.html) · Wireshark · tcp.analysis.retransmission, fast_retransmission, spurious_retransmission, duplicate_ack, lost_segment, and zero_window display filters - [7.5. TCP Analysis](https://www.wireshark.org/docs/wsug_html_chunked/ChAdvTCPAnalysis.html) · Wireshark · The conditions Wireshark uses to flag retransmissions, fast retransmissions, spurious retransmissions, and ZeroWindow - [Network-Related Performance Counters](https://learn.microsoft.com/en-us/windows-server/networking/technologies/network-subsystem/net-sub-performance-counters) · Microsoft · TCPv4 Segments Retransmitted/sec·Segments Sent/sec, Network Interface Packets Received Discarded - [TCP/IP connectivity issues troubleshooting](https://learn.microsoft.com/en-us/troubleshoot/windows-client/networking/tcp-ip-connectivity-issues-troubleshooting) · Microsoft · Checking global TCP settings (Max SYN Retransmissions) with netsh int tcp show global, and confirming loss along the way with simultaneous captures at both ends - [Packet Monitor (Pktmon)](https://learn.microsoft.com/en-us/windows-server/networking/technologies/pktmon/pktmon) · Microsoft · A built-in tool that shows where and why packets are dropped at multiple points in the Windows network stack - [Pktmon command formatting](https://learn.microsoft.com/en-us/windows-server/networking/technologies/pktmon/pktmon-syntax) · Microsoft · Built into Windows 10 and Windows Server 2019 (1809 and later) as pktmon.exe - [TCP improvements in the Windows network stack (IETF 98 TCPM)](https://datatracker.ietf.org/meeting/98/materials/slides-98-tcpm-tcp-improvements-in-windows-01) · Microsoft · Windows 10 Anniversary Update (1607) and Server 2016 turned on TLP and RACK by default (for connections with RTT over 10 ms); when only one packet remains, TLP allows for the 200 ms delayed ACK - [Algorithmic improvements boost TCP performance on the Internet](https://techcommunity.microsoft.com/blog/networkingblog/algorithmic-improvements-boost-tcp-performance-on-the-internet/2347061) · Microsoft · TLP on by default since Windows Server 2016, the new RACK that also recovers lost retransmissions ships in Server 2022, PRR on by default since Windows 10 1903 - [TcpMaxConnectRetransmissions](https://learn.microsoft.com/en-us/previous-versions/windows/it-pro/windows-2000-server/cc938209(v=technet.10)) · Microsoft · Older Windows SYN retransmission: 2 attempts, starting at 3 seconds and doubling ### #owners - [RFC 4787: Network Address Translation (NAT) Behavioral Requirements for Unicast UDP](https://www.rfc-editor.org/rfc/rfc4787) · IETF · NAT mappings are always refreshed by outbound packets (REQ-6), while refresh by inbound packets is optional (for UDP). That’s why the client sends the heartbeats - [RFC 5382: NAT Behavioral Requirements for TCP](https://www.rfc-editor.org/rfc/rfc5382) · IETF · NATs may delete idle TCP sessions; the recommended idle timeout is at least 2 hours 4 minutes (settings can vary by device) - [Amazon EC2 security group connection tracking](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/security-group-connection-tracking.html) · AWS · Security group connection tracking TCP idle timeout (350 seconds on Nitro v6 instance types, 5 days on others, adjustable from 60 seconds to 5 days), recommends keepalives shorter than 5 minutes, 350 seconds for TCP through an NLB - [Control subnet traffic with network access control lists](https://docs.aws.amazon.com/vpc/latest/userguide/vpc-network-acls.html) · AWS · Network ACLs are stateless (no connection tracking), so response traffic must be allowed by its own rules - [Monitor network performance for ENA settings on your EC2 instance](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-network-performance-ena.html) · AWS · Instance counters for exceeding connection tracking and packets-per-second limits (conntrack_allowance_exceeded, pps_allowance_exceeded) - [listen(2) — Linux manual page](https://man7.org/linux/man-pages/man2/listen.2.html) · Linux man-pages · The backlog argument to listen is the actual queue size, truncated to somaxconn if larger (default 4096 since 5.4) - [IP Sysctl](https://docs.kernel.org/networking/ip-sysctl.html) · Linux kernel · somaxconn (listen backlog cap), tcp_max_syn_backlog, tcp_syncookies (default 1, the fallback used when the SYN queue overflows), tcp_keepalive_time (default 2 hours) - [SNMP counter](https://docs.kernel.org/networking/snmp_counter.html) · Linux kernel · When the accept queue is full, SYNs are dropped and TcpExtListenOverflows and TcpExtListenDrops rise - [Netfilter Conntrack Sysfs variables](https://docs.kernel.org/networking/nf_conntrack-sysctl.html) · Linux kernel · Calculating conntrack utilization from nf_conntrack_count and nf_conntrack_max - [Control Group v2](https://docs.kernel.org/admin-guide/cgroup-v2.html) · Linux kernel · nr_throttled in cpu.stat: how many times the container was throttled by its CPU limit - [proc_stat(5) — Linux manual page](https://man7.org/linux/man-pages/man5/proc_stat.5.html) · Linux man-pages · steal in /proc/stat: time spent in other operating systems when running in a virtualized environment - [RFC 2992: Analysis of an Equal-Cost Multi-Path Algorithm](https://www.rfc-editor.org/rfc/rfc2992) · IETF · ECMP picks a path from a hash of the header fields identifying a flow (the same flow takes the same path; different flows may take different paths) - [mtr(8) manual page source](https://raw.githubusercontent.com/traviscross/mtr/master/man/mtr.8.in) · mtr · Measuring the path with the same protocol and port as the game (-T, -u, -P) ### #l-server-proc - [VALORANT's 128-Tick Servers](https://www.riotgames.com/en/news/valorants-128-tick-servers) · Riot Games · Tick budget: a 128-tick server must finish each frame within 7.8125 ms; frame time is measured per subsystem and the budget is split among them - [Introducing Time Dilation (TiDi)](https://www.eveonline.com/news/view/introducing-time-dilation-tidi) · CCP Games · EVE Online’s physics simulation updates once a second, and under overload the game clock slows down to reduce time-bound load proportionally - [HED-GP Technical Retrospective: What a HED-ache](https://www.eveonline.com/news/view/what-a-hed-ache) · CCP Games · Time Dilation floor of 10%; O(n²) messaging, where each of n players’ actions goes out to n players, is the limiting factor in large battles - [Handling variation in time](https://docs.unity3d.com/Manual/time-handling-variations.html) · Unity · The tick simulation’s model: when fixed-interval updates fall behind, catch-up steps run in a burst (fast-forward), and time beyond the cap is dropped, so game time slows down (slow motion) - [Comparing Interest Management Algorithms for Massively Multiplayer Games](https://www.sable.mcgill.ca/~clump/papers/boulanger-06-comparing.pdf) · ACM · NetGames 2006 paper (author’s copy). Comparing distances between every pair doesn’t scale as player count grows; splitting the world into a grid means checking only nearby cells - [Replication Graph in Unreal Engine](https://dev.epicgames.com/documentation/en-us/unreal-engine/replication-graph-in-unreal-engine) · Epic Games · Splitting the world into a grid and choosing recipients from per-cell lists saves server CPU even with many players and actors - [Amdahl's Law in the Multicore Era](https://research.cs.wisc.edu/multifacet/papers/ieeecomputer08_amdahl_multicore.pdf) · IEEE · The lock simulation’s throughput ceiling: if the fraction of work that can only run one at a time is 1−f, speedup can’t exceed 1/(1−f) - [Runtime locking correctness validator](https://docs.kernel.org/locking/lockdep-design.html) · Linux kernel · The lock simulation’s deadlock: taking two locks in opposite orders causes a circular wait and deadlock - [Liveness, Readiness, and Startup Probes](https://kubernetes.io/docs/concepts/workloads/pods/probes/) · Kubernetes · The lock simulation’s watchdog: a liveness check catches the deadlocked state and restarts the process; by default it checks every 10 seconds and restarts after 3 consecutive failures (about 30 seconds) - [ASP.NET Core Best Practices](https://learn.microsoft.com/en-us/aspnet/core/fundamentals/best-practices) · Microsoft · Synchronous calls: make data access and I/O calls asynchronous; blocking calls lead to thread pool exhaustion and slow responses ### #l-infra - [Site Reliability Engineering, Chapter 22: Addressing Cascading Failures](https://sre.google/sre-book/addressing-cascading-failures/) · Google · Cascading failures: how a slow backend ties up threads and resources in front of it, how retries, failing health checks, and restarts with empty caches spread the failure, and how to respond - [Circuit Breaker Pattern](https://learn.microsoft.com/en-us/azure/architecture/patterns/circuit-breaker) · Microsoft Azure · Circuit breaker closed, open, and half-open states and failure count thresholds; long timeouts tie up threads until the breaker trips - [Bulkhead Pattern](https://learn.microsoft.com/en-us/azure/architecture/patterns/bulkhead) · Microsoft Azure · Isolating resources per feature and per call target so a failure in one place doesn’t spread - [Timeouts, retries, and backoff with jitter](https://d1.awsstatic.com/builderslibrary/pdfs/timeouts-retries-and-backoff-with-jitter.pdf) · AWS · Amazon Builders’ Library. Timeouts free up resources, and retries are capped in number with jitter added - [The Unique Architecture behind Amazon Games’ Seamless MMO New World](https://aws.amazon.com/blogs/gametech/the-unique-architecture-behind-amazon-games-seamless-mmo-new-world/) · AWS · Example MMO server architecture: entry servers, per-grid simulation servers (hubs), a shared server pool for sessions, and a database for state - [Working with DB instance read replicas](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_ReadRepl.html) · AWS · The architecture simulation’s replication lag: read replicas are updated asynchronously, so reads can return stale data - [Site Reliability Engineering, Chapter 20: Load Balancing in the Datacenter](https://sre.google/sre-book/load-balancing-datacenter/) · Google · Deploys and restarts: put the server into a lame duck state to send new requests elsewhere before shutting down, and warm up right after a restart - [Amazon EC2 Auto Scaling lifecycle hooks](https://docs.aws.amazon.com/autoscaling/ec2/userguide/lifecycle-hooks.html) · AWS · During scale-out and scale-in, instances are held in a wait state to finish setup and cleanup (up to 1 hour by default) - [systemd.service(5) — Linux manual page](https://man7.org/linux/man-pages/man5/systemd.service.5.html) · systemd · The architecture simulation’s watchdog: stops a service whose heartbeat has stopped and restarts it automatically ## Graph shapes - **Periodic spikes** (`periodic`): Low most of the time, then spikes at the same interval: every few seconds, every few minutes, or on the hour. - **Random spikes** (`random`): Spikes at irregular intervals, then quickly returns to normal. - **Step change** (`step`): Steps up one level at a specific moment, such as a patch, config change, or route change, and stays there. - **Slow climb** (`ramp`): Creeps up over hours or days. Grows with uptime. - **Sawtooth** (`sawtooth`): Climbs slowly, then drops sharply at a restart or cleanup, over and over. - **High at certain hours** (`peak`): Rises like a hill at the same time of day, such as the evening peak. - **Rises with load** (`load`): Climbs as concurrent users or the crowd in one spot grows, and more steeply than the head count itself. - **Hits a ceiling** (`ceiling`): Throughput or connection count reaches a value and can’t go higher; from then on, waits and errors pile up. - **Always high** (`high`): Stays high without spiking. This comes from structure: distance, routing, or design. - **Outliers only** (`outlier`): Most are normal, while specific players, regions, ISPs, or devices are high on their own. - **Gap then burst** (`gap`): Received data drops to zero for a while, then arrives all at once. - **Mass disconnect** (`drop`): The connection count plunges, or the disconnect count shoots up in an instant. - **Surge after opening** (`surge`): Spikes right after servers open or an event starts, then slowly settles down. ## Playbooks ### Lag after a patch When lag reports increase after a specific patch or deployment. Use it when reports like “it’s been weird since this update” pile up, or when a graph steps up at some point and stays there. 1. **Pin down the start time and gather every change around it**: Find when reports first spiked and when the graph stepped up, and list every change that went out around that time. Cover client patches, server deployments, configuration changes, DB schema changes (DDL) and restarts, network and firewall work, and infrastructure swaps (instance type, kernel, drivers). If you use your monitoring tool’s annotation feature to mark every deployment with a vertical line on all graphs, this step goes quickly. If a game patch and infrastructure work went out in the same maintenance window, keep both as suspects. Who to call first: the game team and the infra team, for the changes each of them shipped. (causes: in-deploy, so-os-update, db-ddl-lock, db-plan-flip, db-cold-cache) 2. **Break down the scope: build, device, server, region**: Find the dimension where the problem clusters. Suspect the client first if only players on the new build are affected; client performance or drivers if only certain OSes, graphics cards, or devices are; the server if only certain servers, channels, or zones are; the network path if only certain countries or ISPs are; and shared resources (DB, load balancers, gateways) or the server deployment that just went out if everyone is affected at once. If client telemetry includes the build number, put ping, FPS, frame spikes, and disconnect counts for the old and new builds side by side. If ping is unchanged and only FPS got worse, client performance is a likelier suspect than the network. Who to call first: the game team (client) if it clusters by build or device; for servers or channels, the game team (server) if host metrics look normal, or the infra team (servers/OS) if they don’t; the infra team (network) if it clusters by country or ISP. (causes: cg-hitch, cg-sync-load, co-vram, cg-crash) 3. **Compare the new and old versions over the same time window**: A plain before-and-after comparison mixes in changes from the day of the week, time of day, and events, which muddies the picture. If possible, roll the new version out to a few servers first (canary), and compare tick time p50 and p99, tick overrun count, CPU, memory, and error rate side by side with old-version servers (the control group) over the same time window. If it’s already deployed everywhere, compare against the same day and time last week. A server-wide average hides problems on individual servers or zones, so break it down by server and zone. Who to call first: the game team (server). (causes: sp-tick-overrun, mem-alloc, mem-leak, sp-broadcast) 4. **Compare the traffic fingerprint before and after**: Even without knowing the server code, you can tell from values visible on the network whether the patch changed the shape of the traffic. Compare before and after: packets per second (pps) and bytes per player, average and maximum packet size, connection count, and the size of the send burst that goes out all at once each tick. If UDP packets have started to exceed the path MTU (usually 1,500 bytes), IP fragmentation occurs. Losing a single fragment loses the whole packet, and some NATs and firewalls drop fragments outright. Players whose path crosses a segment with a smaller MTU (a tunnel or VPN) lose only the large packets. If pps went up, check whether you’re hitting the cloud instance’s PPS limit or the throughput limit of firewalls or DDoS protection appliances. Who to call first: if the fingerprint changed, the game team (server), with the evidence attached; if the fingerprint is the same and only loss and retransmissions went up, the infra team (network). (causes: sp-patch-traffic, sk-fragment, rt-mtu, nic-cloud-pps, rt-appliance-pps, rt-burst) 5. **Compare DB query types and counts before and after**: If DB latency went up, start by checking whether query volume (QPS) went up with it. PostgreSQL’s pg_stat_statements and the digest summaries in MySQL Performance Schema group queries that differ only in their values into one entry and track execution count and total time. Comparing the top query lists from before and after the patch reveals new queries, queries whose count multiplied (N+1), and queries that read the whole table without an index (the SUM_NO_INDEX_USED column in MySQL). Who to call first: the game team (server) if QPS or the shape of the queries changed; the infra team (DB: query plans, IOPS, locks) if the queries are the same and only latency went up. (causes: db-no-index, db-login-storm, db-plan-flip, db-cache-stampede) 6. **Use host and server process metrics to find the layer**: Use values visible from the OS, without needing the game code, to tell problems inside the server process from problems on the host. If the receive queue (Recv-Q) on the server socket is building up, the server process isn’t reading in time (tick stalls, GC, locks). If one thread alone is at 100%, it’s a single-thread bottleneck. If pause times in the GC log went up, the memory usage pattern changed. Also check whether the build shipped with a higher log level and log writes went up. Conversely, if CPU steal, throttling, or NIC drops went up, look at infrastructure that changed at the same time (instance type, kernel, container limits). Who to call first: the game team (server) for signals inside the process; the infra team (servers/OS) for host signals. (causes: mem-gc, sp-hotzone, dk-sync-log, so-cpu-quota, so-steal, so-os-update) 7. **Roll back to confirm, then record the result**: Revert the most likely change for only some servers or some players (roll back, or turn off a feature flag), or set the configuration back to its previous value, and see whether the symptom goes away with it. If only the reverted side improves, the cause is confirmed. Reverting can itself cause a brief slowdown from restarts and cold caches, so if it isn’t urgent, do it during off-peak hours. Record the result in the incident log along with the cause ID, and add limits on packet size, query count, and tick time to the pre-deployment checklist for the next patch. Who to call first: the team that made the change. (causes: in-deploy, db-cold-cache) ### Launching in a new country or region When you launch the service in a new country or add a new region or data center. Use it both for pre-launch checks and for sorting out reports like “it’s fine back home, but players in the new country are lagging.” 1. **Measure path quality for each local ISP before launch**: For each major ISP (ASN) in the target country, measure the round-trip time (RTT) distribution, jitter, and packet loss to each candidate game server location. A single average hides the differences between ISPs, so look at the median and 95th percentile per ISP, separately for the evening peak and the early morning hours. The public measurement network RIPE Atlas lets you pick countries and ASNs and send ping and traceroute from probes around the world, or you can spin up temporary VMs in the candidate regions and measure from there. Devices along the path sometimes rate-limit ICMP responses, so when possible, also measure with the same protocol and port the game uses. If one ISP’s traffic stands out by passing through distant cities, it’s a peering or routing problem. ISPs choose lower-cost routes even when lower-latency ones exist, so even nearby destinations can end up taking long detours. Who to call first: the infra team (network); external (ISP, IX) if the route problem is on the ISP side. (causes: isp-distance, isp-routing, isp-peak, isp-cable) 2. **Compare the measurements with the limits the game design can tolerate**: Compare the measured RTT and jitter with the game’s timing windows (reaction times for dodges, parries, and so on), lag compensation limit, interpolation buffer length, and input buffer size. For example, if the parry window is 0.2 s, players on ISPs where round-trip delay plus the interpolation buffer adds up to more than that will be late even when they react in time. Widen lag compensation to make up for it, and now the players on the receiving end start reporting “I got hit behind a wall.” If many ISPs exceed the limits, the infra team should look at placing regions or edge PoPs closer, and the game team should review the timing window, interpolation, and lag compensation values. The white paper’s chapter “Same ping, different feel: netcode models” serves as the reference table. Who to call first: the game team (server and client: design limits) and the infra team (network: region and PoP locations). (causes: sy-short-window, sy-no-lagcomp, sy-lagcomp-overreach, cg-no-buffer) 3. **Check MTU and whether UDP gets through**: Check that the game’s largest packets make it through local networks intact. Measure the path MTU by sending pings of various sizes with the Don’t Fragment (DF) bit set, and look for segments smaller than 1,500 bytes, such as PPPoE, tunnels, and mobile networks. The standard for datagram transports such as UDP (RFC 8899) recommends 1,200 bytes as the base size that can cross most paths on IPv4, so if the game’s largest packet is bigger than that, decide with the game team whether to shrink it or split it up. Also check whether UDP or the game’s ports are blocked or throttled on public Wi-Fi, corporate networks, or certain ISPs, and whether there’s a fallback path (TCP, port 443) for when they are. Who to call first: the infra team (network) and the game team (server: packet size). (causes: dc-mtu, rt-mtu, sk-fragment, isp-udp-block, hn-captive, isp-shaping) 4. **Measure NAT and CGNAT idle timeouts and set the heartbeat interval to match**: Measure how long local home routers and mobile networks (CGNAT) keep the mapping for an idle UDP connection before deleting it. For each trial, send one packet from a test device to the server to create the mapping. The device then sends nothing, and the server sends a packet back to the device after a set wait (30 s, 60 s, 120 s …). The wait at which the device stops receiving that packet is the network’s idle timeout. The standard (RFC 4787) says UDP mappings must not expire in less than 2 minutes and recommends a default of 5 minutes or more, but values vary widely between devices, and some delete mappings sooner. Only outbound packets from the device reliably refresh the mapping, so have the client send the heartbeat, and check that its interval is at most half of the shortest of these: the measured value and the idle timeouts of the load balancer and cloud security groups. Who to call first: the game team (client: heartbeat interval; server: timeout values) and the infra team (load balancer and security group settings). (causes: hn-nat, isp-cgnat, rt-mapping, dc-lb-idle, dc-cloud-conntrack) 5. **Check the external services and security appliances on the local path**: Check that local platform login, payment, and identity verification respond at normal speed, that local DNS resolves the login and patch server addresses correctly, and that the CDN serves patches from locations close to that country. Make sure the new country’s IP ranges aren’t caught by country-blocking rules or rate limits in DDoS protection and firewalls. In particular, make sure CGNAT ranges, where many subscribers share one IP, don’t get blocked wholesale. Who to call first: the infra team (security appliances, DNS, CDN) and external parties (platforms, payment providers, ISPs). (causes: in-external, isp-dns, dc-ddos, isp-cgnat) 6. **After launch, break the data down by country and ASN**: Tag client IPs in connection logs and load balancer logs with country and ASN, and look at RTT, retransmissions, and disconnect counts and reasons (heartbeat timeout, RST, server kick) by country and ISP. Free databases such as MaxMind GeoLite ASN map IPs to an ASN and organization name. To comply with local privacy rules, store IPs truncated to /24 or reduced to the ASN. Look first at that ISP’s route if problems concentrate in one ASN (infra team, external); at distance and design limits if the whole new country is bad (infra team, game team); and at peering congestion if it only gets worse in the evening. If some players always have high ping, check with the game team (server) whether they’re being assigned to a distant region because of GeoIP errors, VPNs, or region assignment based on the party leader. If synthetic monitoring looks normal and only players are having problems, it points to the player’s environment or the client. (causes: isp-peak, isp-routing, rt-queue-drop, pt-isp-validation, in-region-match) 7. **Check how distant players affect everyone else**: When more players connect from far away, the damage doesn’t stop at their own screens. A slow player’s inputs arrive in bunches, so on other players’ screens that one character moves in fast-forward, and it trips the server’s speed and cooldown checks, causing rubber-banding or rejected skills. In party mechanics, one slow player’s late reaction can fail the whole party, and in lockstep, everyone waits for the slowest player. Check whether reports of “one character looks off” from existing players went up after the new country opened, and work out input buffers, validation tolerances, and separate matchmaking regions with the game team. Who to call first: the game team (server). (causes: pt-slow-burst, pt-isp-validation, pt-raid-member, sy-lockstep) ## Real-world incidents ### eve-hedgp-2014 · CCP Games 2014: Server overload in EVE Online’s massive HED-GP fleet battle - What happened: During the massive fleet battle in the HED-GP system, covered in a January 2014 retrospective, the server was badly overloaded. Even after Time Dilation (a feature that slows game time under overload) hit its 10% floor and the whole battlefield went into slow motion, load kept piling up. The backlog in processing module deactivations and repeat cycles (Dogma Lateness) peaked at 193 seconds of game time, about 32 minutes in real time. The July 2013 battle in 6VDT, of nearly the same size, peaked at 42 seconds (about 7 minutes in real time). - Cause: CCP cautioned that it couldn’t be certain, because its profiling tools add load of their own and aren’t run in situations like this, and named two likely causes. The first was unprocessed load that kept building up as the battle dragged on. The second was heavier drone use: unique drones deployed during the battle went from 21,123 in 6VDT to 38,852 in HED-GP, 84% more. Telling everyone in view about each player’s actions takes traffic that grows with the square of the player count (O(n²)), and drones generate more messages per attack. The code drones use to pick targets also often scans every attackable target on the same battlefield, so its cost grows close to n². - Lessons: When the processing load in one crowded area exceeds its limit, the whole area goes into slow motion, and the longer the battle runs, the more the backlog grows and the worse the input lag gets. Signals to check: tick time and backlog on the server (node) handling that area, plus player and entity counts. A telltale sign is that other areas stay fine. The primary owner is the game team (server), and the things to fix are how many recipients each action is sent to and the cost of AI target searches. Slowing game time can’t eliminate the overload, but it slows everyone down at the same rate, which keeps a subset of actions from falling behind indefinitely. - Related causes: sp-broadcast, sp-tick-overrun, sp-queue, sp-hotzone - Original post: [CCP Games](https://www.eveonline.com/news/view/what-a-hed-ache) ### riot-direct-2015 · Riot Games 2015: League of Legends traffic on roundabout routes, and Riot Direct - What happened: A technical post in which Riot Games explains why the internet is a poor fit for real-time games. Real traffic reported by a League of Legends player should have gone straight from San Francisco to Portland, but it went through Los Angeles, Denver, and Seattle, taking 70 ms for a trip that would take 14 ms on a direct path. Riot explained that when routers overflow and drop packets, other champions appear to jump around the screen and projectiles seem to teleport. - Cause: Riot pointed to routes and routers. Backbone providers and ISPs send traffic along the cheapest path even when a lower-latency path exists, and when the route BGP settles on takes a long detour, traffic also passes through more routers. A router’s processing load depends on the number of packets, whatever their size. Game packets are around 55 bytes, so the same amount of data takes 27 times as many packets as it would in 1,500-byte packets, and fills router input buffers that much faster. According to Riot, many routers drop UDP packets first when they’re overloaded. As a fix, Riot built its own network, Riot Direct, with routers at 10 major internet hubs in the US and direct connections (peering) with as many ISPs as possible. According to Part II, the share of players with a ping under 80 ms rose from 31% to 50% in a little over 9 months, and hit 80% overnight after the game servers moved to Chicago. - Lessons: If only customers of one ISP have unusually high ping, even within the same country, suspect the route. Signals to check: the RTT distribution per ISP (ASN) and the cities that show up as hops in traceroute. The primary owner is the infra team (network), and the fixes are direct peering with ISPs, connecting at IXs (internet exchanges), and choosing server locations. Routing policy on the ISP side has to be worked out with the external party (the ISP). The case also shows that simply moving servers closer to the center of the player base makes a big difference. - Related causes: isp-routing, isp-distance, rt-queue-drop - Original post: [Riot Games](https://www.riotgames.com/en/news/fixing-internet-real-time-applications-part-i) ### riot-edge-2020 · Riot Games 2020: Edge host overload on League of Legends servers in Europe and Brazil - What happened: In late February 2020, the League of Legends EUW, EUNE, and BR servers had several outages, and the number of new games starting dropped sharply. Backend services such as matchmaking and game servers all reported healthy, yet almost no traffic was coming in. Riot pushed the tournament mode (Clash) back a week to avoid launching it on clusters that might be unstable. The postmortem doesn’t say how long each outage lasted. - Cause: Three things came together. Requests to one service were malformed, so in certain cases they kept failing and being retried, and request volume exploded. A known compatibility problem between the container system and the OS version was leaking memory inside the OS. The OS upgrade was finished on only about 60% of Riot’s entire container environment and was still in progress on the Europe and Latin America clusters. Edge containers, which receive internet traffic, filter it, and pass it to the backend, were kept apart within a shard (server group), but nothing kept different shards apart, so in every outage edge containers from at least three shards were packed onto a single host. The retry surge landed on that host, and the memory leak brought it to a halt. - Lessons: When every backend service reports “healthy, but no traffic is coming in,” look at what sits in front of them (edge, gateways, load balancers). Signals to check: inbound connection counts skewed toward particular hosts, and the failure and retry rate of specific requests. The primary owner is the game team (server: the malformed request and the retry behavior), and the infra team (servers/OS) handles container placement rules, OS upgrades, and skew alerts. Riot fixed the request code, changed retries so they wouldn’t spike, and put skew alerts in place until it could implement spreading across shards. - Related causes: in-gateway, in-cascade - Original post: [Riot Games](https://www.leagueoflegends.com/en-gb/news/riot-games/incident-report-recent-outages-in-europe-brazil/) ### riot-euw-2021 · Riot Games 2021: League of Legends EUW 5-hour outage: one auxiliary DB halted the whole server - What happened: On January 22, 2021, the League of Legends EUW server didn’t work properly for a little over 5 hours. The metrics for logged-in players and players in a game cut out at the same moment, and between the two restarts, logins went up but almost no games started. - Cause: The primary server of a database behind a non-critical feature had a hardware failure, and that database had no automatic failover to a standby configured. Each database had its own connection pool, but all the pools shared one thread pool; work sent to the failed database never finished and held on to threads, until the whole system ran out of threads. Amid a flood of alerts, the team first suspected a recent malicious network attack and hardware work in another region, so the alert for the failed database wasn’t noticed until about 1 hour later. Because every system ran inside a single JVM, when GC paused the process for several seconds at a time under the reconnect load after the restart, metrics collection also developed large gaps. The login queue also didn’t hold to its configured limit, so players flowed in unevenly. - Lessons: Even one auxiliary database that nobody considered critical can halt everything through a shared resource such as a thread pool. Signals to check: pending requests per database, thread pool utilization, and a ratio of game starts to logins that is far too low. The owners are the game team (server: thread pool isolation, timeouts) and the infra team (DB: automatic failover). When alerts flood in, it’s easy to suspect whatever hit you recently (an attack, for example) first, so rule things out one at a time in the decision order (scope → timing → layer). After a restart, also check that the login queue actually limits inflow as configured. - Related causes: sp-threadpool, db-failover, in-cascade, mem-gc - Original post: [Riot Games](https://www.riotgames.com/en/news/keeping-legacy-software-alive-case-study) ### roblox-2021 · Roblox 2021: Roblox 73-hour outage: contention in the service discovery (Consul) cluster - What happened: It began on the afternoon of October 28, 2021 (Pacific Time) with high CPU load on one Consul server. At 16:35 the number of players online fell to half of normal, and then the entire service went down. Not until 16:45 on October 31 could all players get back in, 73 hours after the outage began. Roblox said 50 million people use it every day. - Cause: Roblox uses HashiCorp Consul for service discovery (how services find each other’s addresses), health checks, and a key-value (KV) store, and a single Consul cluster was handling several workloads at once. There were two root causes. First, the day before the outage, Consul’s new streaming feature, which had been rolled out gradually over several months, was turned on for the traffic routing service too, and that service’s node count was raised by 50%. Under very heavy read and write load, the feature caused contention on a single shared resource (a Go channel). The contention was even worse on the dual-socket (NUMA) servers with more cores that were swapped in during the outage. Second, free-page list (freelist) management in BoltDB, which Consul uses to store its Raft log, became pathologically slow and wrote 7.8 MB to disk for every append of 16 kB or less. Median KV write latency, normally under 300 ms, rose to 2 s, and zero windows (full TCP buffers) were seen on the slow leader server. Because telemetry depended on Consul, the metrics needed to find the cause disappeared along with it. - Lessons: When a foundational system that many services rely on (service discovery, configuration store, authentication) slows down, every feature stops at once. Signals to check: that system’s write latency, leader changes, and CPU, plus any configuration change made just before the outage. Ownership is shared between the game team (server) and the infra team (servers/OS). Keep monitoring separate from the systems it watches, so you can still see metrics during an outage. During recovery, caches are empty, and letting everyone in at once can knock things over again, so Roblox used DNS to control the share of players let in and raised it about 10% at a time. - Related causes: in-cascade, sp-lock, rt-zero-window, db-cold-cache - Original post: [Roblox](https://about.roblox.com/newsroom/2022/01/roblox-return-to-service-10-28-10-31-2021) ### ffxiv-2021 · Square Enix 2021: FINAL FANTASY XIV expansion launch congestion and login queue errors - What happened: From the start of early access for the Endwalker expansion in December 2021, every World was extremely congested. Login queues grew long, and Error 2002 appeared often when logging in from the character selection screen or while waiting in the queue. Some Worlds and zones also went down (Error 3001), and queues timed out (Error 4004). As of the December 11 notice, on day 8 of early access, the congestion was still going on. - Cause: Error 2002 occurs in two cases. The first is when more than 17,000 players are waiting on a logical data center. This cap exists to keep the login server from going down under an overly long queue, and when it’s hit, the client shuts down completely. On December 7, spare development hardware was added to the lobby servers to raise the cap; this error became less common, but the queues actually got longer. The second case is when a waiting player’s connection is unstable. As waits got longer, brief disconnects caused by packet loss on the internet path or unstable Wi-Fi became more common. The lobby server waits somewhere from tens of seconds to about 1 minute for a reconnect. Players who reconnect in time keep their place in the queue, but anyone who takes longer goes to the back of the line. Square Enix said most reports fell into this case. The semiconductor shortage also meant new Worlds couldn’t be added right away. - Lessons: The longer the queue, the more often a brief connection drop for a waiting player turns into a connection error. Under the same congestion, errors cluster among players on Wi-Fi or unstable connections, so it becomes a problem that hits “only some players.” Signals to check: queue length and wait time, and the share of disconnects that happen while players wait in the queue. The primary owner is the game team (server: the queue cap and the reconnect grace period), and the infra team joins in on adding lobby and World servers. A generous reconnect grace period keeps more of these brief drops from costing players their place in the queue. - Related causes: in-login-queue, hn-wifi, rt-wireless - Original post: [Square Enix](https://na.finalfantasyxiv.com/lodestone/news/detail/6a94b30182b6d963994fdc0b789264ac9f24986f) ### cloudflare-2020 · Cloudflare 2020: Traffic loss in some cities from a Cloudflare backbone configuration error - What happened: Many games rely on CDN providers for their websites, APIs, and DDoS protection, so this is the kind of infrastructure outage that affects games too. For 27 minutes on July 17, 2020, from 21:12 to 21:39 (UTC), traffic across Cloudflare’s entire network dropped by about 50%. The impact was limited to some city locations (PoPs) in the US, Europe, Russia, and Brazil that were connected to the backbone. Other locations were fine. - Cause: An outage on the Newark–Chicago backbone link congested the Atlanta–Washington link, so an engineer changed a router configuration to take some backbone traffic off Atlanta. The change was supposed to disable an entire policy term, but it disabled only the condition inside it (the prefix-list), so the Atlanta router advertised all its BGP routes across the whole backbone with a higher preference (local-preference 200). Each location gave the routes to its own servers a preference of 100, so traffic from every backbone-connected location was pulled to Atlanta. Atlanta was overloaded, and the affected locations were left with almost no traffic to handle. Service recovered once the Atlanta router was removed from the backbone. Cloudflare stated that the outage was unrelated to any attack or breach. - Lessons: If players in particular cities or regions all hit disconnects or can’t connect / infinite loading at once while everyone else is fine, first suspect a routing configuration change made just before. On the graph, CPU and traffic spike at one location only, while the affected locations actually drop to close to 0. The primary owner is the infra team (network), or external if the outage is on the provider’s side. Cloudflare decided to cap the number of routes each backbone BGP session can accept (maximum-prefix), and adjusted preferences so that one location can’t pull in traffic meant for other locations. - Related causes: isp-bgp, rt-path - Original post: [Cloudflare](https://blog.cloudflare.com/cloudflare-outage-on-july-17-2020/) ### fastly-2021 · Fastly 2021: Worldwide Fastly CDN errors - What happened: Many games deliver patch files, launchers, and web pages through a CDN, so this is the kind of infrastructure outage that affects games too. Starting at 09:47 (UTC) on June 8, 2021, 85% of Fastly’s network returned errors. Within 49 minutes, 95% of the network was back to normal, and the incident was resolved at 12:35. - Cause: A software deployment that began on May 12 contained a bug that would trigger when a specific customer configuration met specific conditions. On June 8, a customer pushed a valid configuration change that met those conditions. Fastly detected the problem within 1 minute, and recovery began once it identified and disabled the customer configuration that triggered it. Deployment of the bug fix began at 17:25 the same day. - Lessons: Code deployed weeks earlier can still cause a global outage in an instant when it meets a rare condition. On the game side, the signals are HTTP error rates for patch, launcher, and web requests rising in every region at the same time, and the CDN provider’s status page. A telltale sign is that game connections already in progress stay fine if they don’t go through the CDN, and only new connections, patch downloads, and web logins are blocked. The primary owner is external (the CDN provider). The game and infra teams should have a fallback ready, such as a second CDN or a path that fetches directly from the origin server. - Related causes: in-external - Original post: [Fastly](https://www.fastly.com/blog/summary-of-june-8-outage) ### meta-2021 · Meta 2021: Facebook outage: one backbone command took DNS down with it - What happened: An infrastructure outage whose lessons apply directly to a game company’s own network and DNS. On October 4, 2021, Facebook (now Meta) services were unreachable worldwide. The backbone linking its data centers went down completely, and Facebook’s DNS servers could no longer be found from the internet. The postmortem doesn’t say how long the outage lasted. - Cause: During routine maintenance, a command issued to assess global backbone capacity unintentionally took down every connection in the backbone, and an audit tool meant to block commands like this failed to stop it because of a bug. DNS servers at smaller locations are designed to mark themselves unhealthy and withdraw their BGP advertisements when they can’t talk to the data centers, so the DNS servers became unreachable from the internet even though they were still running. The normal access paths and out-of-band access were both down, and internal tools had lost DNS as well, so engineers had to be sent to the data centers in person, and security procedures slowed that down further. By the time of recovery, power draw at each data center had dropped by tens of MW, and bringing everything back at once could put everything from electrical systems to caches at risk, so load was raised in stages. - Lessons: If can’t connect / infinite loading hits every region and every ISP at the same time, check DNS and BGP routes before the game servers. You can confirm this from outside the company with external DNS lookups and public BGP route data. The primary owner is the infra team (network). Check in advance that your out-of-band access path and the internal tools you’d use in an outage don’t depend on the same DNS and network, and during recovery, raise load in stages so reconnects don’t all arrive at once. - Related causes: isp-bgp, isp-dns - Original post: [Meta](https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/) ### aws-2021 · AWS 2021: AWS us-east-1 internal network congestion - What happened: Many games run their servers, login, and data on public clouds, so this is the kind of infrastructure outage that affects games too. At 7:30 AM PST on December 7, 2021, the internal network in the Northern Virginia region (us-east-1) became congested. From 7:33 AM, EC2 API errors and latency rose, making it hard to launch new instances (instance launches recovered at 2:40 PM), followed by console login failures, Route 53 configuration changes being blocked, and delayed and partly lost CloudWatch metrics. The network devices fully recovered at 2:22 PM. EC2 instances that were already running and existing DNS responses were not affected. - Cause: An automated activity to scale capacity for a service in the main network triggered unexpected behavior from a large number of clients in the internal network, causing a surge in connection attempts. The devices linking the internal network to the main network were overwhelmed and communication slowed down. The delays in turn drove more connection attempts and retries, so the congestion persisted. The clients had backoff behavior that spaces out requests during congestion like this, but a latent defect kept it from working properly. Internal monitoring depended on the same network, so the operations team had to respond using logs, without real-time metrics. - Lessons: When retries can’t back off, a brief bout of congestion turns into an outage that lasts hours. From the game’s point of view, game servers that are already running may be fine, but new server capacity (autoscaling), any login, matchmaking, or payment flow that calls cloud APIs, and monitoring can all be blocked at the same time. Signals to check: the cloud provider’s status page, cloud API error rates, and instance launch failures. The primary owner is external (the cloud provider). The game team should give every retry exponential backoff with randomized delays and a retry limit, and the infra team should keep enough spare capacity to ride out blocked scale-out, plus a fallback in another region. - Related causes: in-cascade, in-autoscale, in-external - Original post: [AWS](https://aws.amazon.com/message/12721/) ### cloudflare-dns-2025 · Cloudflare 2025: Cloudflare public DNS 1.1.1.1 outage - What happened: An outage of a public DNS resolver that players set up themselves on their devices or routers. This type of outage blocks every game and service at once, but only for players using that setting. For 62 minutes on July 14, 2025, from 21:52 to 22:54 (UTC), the 1.1.1.1 resolver stopped responding worldwide. Cloudflare said that for many users, this meant being unable to use essentially any internet service. Queries over UDP, TCP, and DNS over TLS were affected, while DNS over HTTPS, which connects by domain name, stayed relatively stable. - Cause: On June 6, while a service topology (the configuration that decides which locations advertise which IP ranges) was being prepared for a different service not yet in use, the 1.1.1.1 resolver’s IP ranges were mistakenly included in it. When that service’s configuration was changed on July 14, the locations advertising the resolver ranges shrank from every location to a single offline one, and the BGP routes were withdrawn worldwide. The change skipped canary deployment and went straight to every data center. Reverting the configuration at 22:20 brought traffic back to about 77%, but in the meantime about 23% of edge servers had lost required IP configuration, and re-applying it meant service wasn’t back to normal until 22:54. Cloudflare said it was an internal configuration error unrelated to any attack or BGP hijack. - Lessons: If the game servers and other players are fine but some players get can’t connect / infinite loading on the login or patch servers, suspect the DNS those players use. A telltale sign is that existing sessions stay connected and only new connections fail. Having those players change their DNS settings or look up the server address directly settles it right away. The primary owner is external (the DNS operator or ISP). If the game team (client) reports name resolution failures separately from other errors, customer support can make the call right away. - Related causes: isp-dns, isp-bgp - Original post: [Cloudflare](https://blog.cloudflare.com/cloudflare-1-1-1-1-incident-on-july-14-2025/) ### aws-2025 · AWS 2025: AWS us-east-1 DynamoDB DNS outage and long recovery - What happened: Many games run their servers, login, and data on public clouds, so this is the kind of infrastructure outage that affects games too. From 11:48 PM on October 19, 2025, to 2:20 PM on October 20 (PDT), the Northern Virginia region was affected in three stages. DynamoDB API errors were elevated until 2:40 AM on the 20th; new EC2 instance launches failed from 2:25 AM to 10:36 AM (connection problems on some new instances cleared up at 1:50 PM); and some Network Load Balancers (NLB) saw more connection errors from 5:30 AM to 2:09 PM. - Cause: The automation that manages DynamoDB’s DNS had a latent race condition. Among the processes that apply DNS plans in different Availability Zones (DNS Enactors), one that was running unusually late overwrote a newer plan with an old one. Right after that, another Enactor’s cleanup job deleted that old plan, leaving the DNS record for the regional endpoint (dynamodb.us-east-1.amazonaws.com) empty. The automation couldn’t repair this, so it had to be fixed by hand. EC2’s system for managing physical servers depends on DynamoDB, so in the meantime the lease it kept for each physical server expired. After DynamoDB came back, there were so many physical servers that re-establishing the leases timed out before finishing, and retries piled up again, pushing the system into “congestive collapse.” Network configuration for newly launched instances propagated slowly, so NLB health checks flapped between passing and failing, and even healthy nodes were repeatedly removed from DNS and added back. - Lessons: A DNS record error in one place spreads to the other services that depend on that service, and even after the cause is fixed, backlogged work and flapping health checks stretch recovery out by several more hours. From the game’s side, servers already running may hold up, but new servers can’t launch, so autoscaling stalls, and flapping health checks can make the load balancer pull healthy servers out of rotation. Signals to check: the cloud status page, managed service API error rates, instance launch failures, and the load balancer’s healthy target count. The primary owner is external (the cloud provider). The infra team should cap how many servers can drop out at once on failed health checks and have a fallback in another region ready. - Related causes: in-external, in-cascade, in-autoscale, dc-lb-imbalance, isp-dns - Original post: [AWS](https://aws.amazon.com/message/101925/) ## Glossary - **Ping** (Ping, RTT): The time it takes for a signal you send to reach the server and come back (round trip). The ping a game displays sometimes also includes time spent waiting for server processing. - **Latency** (Latency): The time a packet takes to get from where it was sent to where it arrives. It often refers to one direction only, so it’s roughly half the ping. - **Jitter** (Jitter): Variation in packet arrival intervals. Even at the same average ping, high jitter makes the screen stutter. - **Packet** (Packet): A chunk of data sent over the network in one go. Usually 1,500 bytes at most; game updates are tens to hundreds of bytes. - **Packet loss** (Packet loss): Packets that are sent but never arrive. In a game that uses TCP, even 1% loss causes a noticeable hitch every few seconds to ten-odd seconds, while a UDP game with interpolation and redundant input sending can sometimes hide a few percent. - **Bandwidth** (Bandwidth): The maximum amount of data a connection can carry per second (Mbps). It’s a different concept from how quickly data arrives (latency). - **Tick** (Tick): One step in which the server computes the game state. A 20-tick server computes 20 times per second, every 50 ms. - **Tick rate** (Tick rate): How many ticks the server runs per second. Higher means faster response, but server cost and bandwidth go up. To save bandwidth, some games send packets less often than the tick rate. - **Tick budget** (Tick budget): The time limit for finishing one tick. Go over it and the next tick starts late, stretching the tick interval. - **FPS** (Frames per second): How many times per second the screen is drawn. At 60 FPS, each frame gets 16.7 ms. - **Frame time** (Frame time): How long it took to draw one frame. Occasional spiky frames matter more to how a game feels than average FPS. - **Snapshot** (Snapshot): A summary of “the game state right now” that the server sends every tick: position, health, status, and so on. It usually carries only what changed relative to what the receiver already has (delta compression). - **Interpolation** (Interpolation): Drawing smooth motion between two received snapshots. The trade-off is that it shows a moment slightly in the past. - **Interpolation buffer** (Interpolation buffer): Time the client deliberately waits before drawing, so it has something to interpolate. It’s the slack that absorbs jitter and one or two lost packets. It’s usually twice the packet interval (100 ms at 20 updates per second), and some games grow it automatically when jitter increases. - **Extrapolation** (Extrapolation, Dead reckoning): When no new packet has arrived, guessing where an entity is headed from its last velocity and drawing it there. A wrong guess looks like teleporting, so many games extrapolate for only about 0.25 s and then stop (the Source engine default is 0.25 s). - **Client-side prediction** (Client-side prediction): Moving your own character on screen right away, without waiting for server confirmation. - **Server reconciliation** (Reconciliation): When the server’s result arrives, comparing it with the prediction and correcting your character’s position. The client starts from the position the server confirmed and replays the inputs the server hasn’t confirmed yet. A large correction looks like rubber-banding. - **Lag compensation** (Lag compensation): When the server judges an attack, it rewinds to the moment the attacker was seeing and checks whether the attack hit. To keep things fair for the player being hit, the rewind is capped: 0.2–0.25 s is common in competitive shooters, and some games rewind as far as 1 s, like the Source engine default. - **Authoritative server** (Authoritative server): A design in which only the server makes final decisions. It stops cheating, but every outcome needs a round trip to the server, so prediction and client-side feedback hide the wait. - **Lockstep** (Deterministic lockstep): Everyone exchanges only inputs and runs the exact same computation on the same turn. Inputs get a fixed delay, and if anyone’s input is late, everyone waits. - **Server-side input buffer** (Server-side input buffer): A per-player buffer in which the server collects a few inputs and consumes one per tick. Even a player with high jitter looks smooth to others, but that player’s actions are confirmed on the server correspondingly later. - **Listen server** (Listen server): One player’s PC plays the game and acts as the server at the same time. The host has zero ping, but if the host’s connection or PC is slow, everyone lags. - **Phasing** (Phasing): Showing different NPCs and terrain in the same place depending on quest progress. If two characters are at different points in a quest, it’s normal for an NPC to be missing for one of them. - **Rollback netcode** (Rollback netcode (GGPO)): Predicting the opponent’s input to keep the game moving, then rewinding to a past frame and recomputing when the real input turns out different. Common in fighting games. Unrelated to a database rollback. - **Input buffering** (Input buffer, spell queue): Accepting the next input pressed shortly before a cooldown or animation ends, and executing it the moment it ends. This keeps a round trip from sneaking in between combo steps. - **Client-side feedback** (Client-side feedback): Playing animations, sounds, and effects right away, without waiting for server confirmation. Only results that need confirmation, like damage and rewards, wait for the server’s reply. If the server rejects the action, what was shown has to be reverted. - **TCP** (Transmission Control Protocol): A protocol that delivers data in order, with nothing missing. It holds later packets back from the game until a lost packet has been received again. - **UDP** (User Datagram Protocol): A protocol that delivers whatever is sent, with no guarantees. There’s no waiting, but the game has to handle loss and ordering itself. - **Reliable UDP** (Reliable UDP (KCP, ENet…)): An approach that implements just as much retransmission and ordering as needed on top of UDP. - **HOL blocking** (Head-of-line blocking): One stuck item at the front makes everything behind it wait. This is what causes fast-forward on TCP. - **RTO** (Retransmission timeout): The retransmission timer: how long TCP waits before deciding a packet is lost and sending it again. On Linux it’s ping + 200 ms or more, doubling after each failure. - **Nagle’s algorithm** (Nagle’s algorithm): A TCP feature that holds small pieces of data until the acknowledgment (ACK) for earlier data arrives, then sends them together to save packets. Games should usually turn it off. - **TCP_NODELAY** (TCP_NODELAY): The socket option that turns off Nagle’s algorithm. Small messages go out immediately. - **Delayed ACK** (Delayed ACK): Sending the acknowledgment of receipt a little later, bundled with other data. Linux typically waits 40 ms (up to 200 ms); older Windows versions wait 200 ms and current ones 40 ms. - **Socket buffer** (SO_SNDBUF / SO_RCVBUF): The size of the send and receive queues the OS keeps for each socket. Too small and they overflow; too large and stale data piles up and waits. - **keepalive** (SO_KEEPALIVE): A TCP feature that checks whether an idle connection is still alive. It’s off by default, and even when turned on, the default is to check only after 2 hours. - **RST** (TCP reset): A TCP signal that kills a connection on the spot. Any data not yet sent is discarded. - **Heartbeat** (Heartbeat): An “I’m alive” signal the game itself sends periodically. It’s used to detect dead connections and to keep connections open on devices along the path. - **Timeout** (Timeout): How long to wait for a response before treating it as a failure. Too short causes false alarms; too long delays detection. - **NAT** (Network Address Translation): A router feature that sends traffic from several devices at home out through one public IP, recording each connection in the NAT table. - **CGNAT** (Carrier-grade NAT): Large-scale NAT in which an ISP shares one IP address among many subscribers. - **MTU** (Maximum Transmission Unit): The largest packet that can be sent in one piece. Usually 1,500 bytes, and smaller on VPN and PPPoE segments. - **Bufferbloat** (Bufferbloat): Network equipment letting its queues grow so large that latency climbs to hundreds of ms. - **SQM** (Smart Queue Management (fq_codel, CAKE)): A router feature that keeps queues short and sends traffic out fairly per flow. The fix for bufferbloat. - **QoS** (Quality of Service): A feature that gives important traffic priority so it goes out first. - **Peering** (Peering): The points where ISPs connect their networks to each other. They tend to get congested in the evening. - **BGP** (Border Gateway Protocol): The protocol ISPs use to tell each other which routes to take across the internet. When routes change, the path and ping change too. - **DDoS** (Distributed Denial of Service): An attack that floods a service with traffic from many sources to knock it offline. - **Scrubbing center** (DDoS scrubbing center): A DDoS mitigation provider’s site that receives traffic headed for your servers during an attack, filters out the attack, and forwards only legitimate traffic. A distant site makes the route longer. - **Firewall** (Firewall): A device or program that lets through only permitted connections. It tracks connections in a session table. - **Load balancer** (Load balancer): A device that spreads incoming connections across several servers. - **Session table** (Session table, conntrack): The table a device or OS uses to track current connections. Its size is limited. - **Microburst** (Microburst): Traffic piling up in very short windows of 1 ms or less, even though the average is low. - **NIC** (Network Interface Card): A server’s network card. - **Ring buffer** (Ring buffer): The buffer that holds packets the NIC has received until the CPU picks them up. It cycles through a fixed number of slots, and when every slot is full, new packets are dropped. - **Interrupt** (Interrupt): A signal a device sends to tell the CPU that there’s work to handle. - **RSS** (Receive Side Scaling): A NIC feature that spreads incoming packets across several receive queues so multiple CPU cores can process them. - **PPS** (Packets per second): Packets per second. Game servers often hit this limit before they run out of bandwidth. - **Kernel** (Kernel): The core of the operating system. It manages networking, memory, and how CPU time is shared out. - **backlog** (Listen backlog): The queue where new connection requests wait until the server accepts them. When it’s full, Linux silently drops new requests and Windows sends a rejection. - **TIME_WAIT** (TIME_WAIT): The state in which the side that closed a connection first keeps that port pair reserved for a while (60 seconds on Linux) in case late packets arrive. - **CPU steal** (Steal time): Time a virtual machine spent waiting for CPU because the physical host was giving it to other virtual machines. Shown as the st value in top. - **CPU throttling** (CFS throttling): Forcibly pausing a container until the next period once it has used up its CPU quota within a set period (CFS period, usually 100 ms). - **File descriptor** (File descriptor): The number (fd) attached to each file or connection a process opens. There’s a limit on how many a process can have. - **Thread** (Thread): A unit of work that runs independently inside a program. Several threads can run at the same time. - **Context switching** (Context switch): The CPU swapping out the running thread for another one. It has a cost. - **Lock** (Lock, Mutex): A guard that lets only one thread at a time use shared data. - **Deadlock** (Deadlock): A state in which threads each wait for a lock the other holds and stay stuck forever. - **Thread pool** (Thread pool): A set of worker threads created in advance. When all of them are busy, new work waits. - **Asynchronous I/O** (epoll, IOCP, io_uring): A model in which a thread keeps doing other work while I/O is in progress and gets notified when it completes. - **AOI** (Area of Interest): The range each player “can see.” Only changes inside it are sent, which cuts bandwidth. To reduce the cost of working out who’s in range, the map is usually divided into a grid and only nearby cells are checked. - **Broadcast** (Broadcast, fan-out): Sending one change to everyone who can see it. If everyone in a crowd can see everyone else, the amount to send grows with the square of the player count. - **GC** (Garbage collection): Automatically reclaiming memory that’s no longer in use. Programs can pause during GC. - **Heap** (Heap): The memory area a program allocates from whenever it needs memory while running. - **Memory leak** (Memory leak): A bug in which memory that’s no longer needed is never released, so usage keeps growing. It happens even with GC if something still holds a reference to objects the program is done with. - **Swap** (Swap, paging): Moving part of memory to disk when RAM runs short. Using that memory again is more than 1,000 times slower than RAM. - **OOM killer** (Out-of-memory killer): A Linux feature that, when memory runs out, picks the process using the most memory and forcibly kills it. In a container it kicks in as soon as the container’s memory limit is hit. - **Cache miss** (Cache miss): The data isn’t in the cache close to the CPU, so it has to come from slower memory. - **IOPS** (I/O operations per second): How many reads and writes a disk can handle per second. On cloud disks, the limit depends on what you pay for. - **fsync** (fsync): A call that waits until data is safely written to disk. Normal writes land in OS memory first and reach the disk later, so they can be lost if the server loses power in between. fsync is safe but slow. - **Burst credits** (Burst credits): A balance that cloud disks and servers build up so they can briefly run above their baseline performance. When it runs out, performance drops back to baseline. - **Index** (Index): A lookup structure in a database. Without one, the database has to read the whole table. - **Full table scan** (Full table scan): A query that checks every row of a table without using an index. - **Query plan** (Query plan): The database’s plan for executing a query: in what order and with which indexes. Even with unchanged code, the same query can suddenly slow down if the database switches plans. - **Transaction** (Transaction): A group of DB operations that either all succeed or all fail. Trades must always be handled in a transaction. Rows it modifies stay locked until it finishes, so shorter is better. - **Connection pool** (Connection pool): A set of DB connections opened in advance. When all of them are in use, new requests wait. - **Hot row** (Hot row): A single row that many requests try to modify at the same time. A cause of lock contention. - **Replication lag** (Replication lag): How far a replica database has fallen behind the primary. - **Rollback** (Rollback): A save is canceled and data returns to its previous state. Players experience it as “my item disappeared.” - **Cache** (Cache (Redis etc.)): A copy of frequently used data kept somewhere fast. It reduces database load. - **Checkpoint** (Checkpoint): The database periodically writing the changes it has collected in memory out to disk in one batch. Saves and queries can slow down briefly while it happens. - **Failover** (Failover): Switching over to a standby when the primary server or DB goes down. Writes stop briefly during the switch, and if replication was lagging, the most recent data can be lost. - **MVCC** (Multi-version concurrency control): A way for a database to keep old row versions for a while so readers and writers don’t block each other. A long-open transaction lets old versions pile up and slows things down. - **Cache stampede** (Cache stampede): Cache entries empty out all at once and requests flood the origin (DB). - **Gateway** (Gateway): An intermediate server that accepts client connections and forwards them to the game servers behind it. - **Circuit breaker** (Circuit breaker): A mechanism that temporarily stops calling a service that keeps failing and returns a failure right away, which prevents cascading failures. After a while it tries a call or two, and if the service has recovered, it lets calls through again. - **Cascading failure** (Cascading failure): A failure in one place spreading to other services along the call chain. - **Autoscaling** (Autoscaling): Automatically adding and removing servers based on load. Adding servers takes time. - **Watchdog** (Watchdog): A timer that watches whether the server has frozen. If the game loop stalls longer than a set time (a few seconds to tens of seconds), it writes a state dump and forcibly kills the server so it restarts. - **Utilization** (Utilization): The fraction of time a worker (anything that processes requests, such as a CPU core, thread, or DB connection) is busy. Above 80–90%, waiting time climbs steeply. - **p99** (99th percentile): The value that 99 out of 100 measurements beat, with about 1 in 100 slower. It reflects the lag players feel better than the average does. - **V-Sync** (Vertical sync): Sending frames in step with the display’s refresh cycle. It eliminates screen tearing but adds input lag, and when FPS drops below the refresh rate, it can bounce between 60 and 30 and stutter. - **Variable refresh rate** (VRR, G-Sync, FreeSync): The monitor refreshes whenever a frame is ready. It reduces the stutter and input lag caused by V-Sync bouncing between 60 and 30. - **Anti-cheat** (Anti-cheat): A security module that blocks game hacks. When its periodic scans or its heartbeats to the server fail, it can cause stutter or disconnects. - **Overlay** (Overlay): A feature of chat, recording, or FPS-counter programs that draws on top of the game screen. It hooks into the game’s rendering and can cause stutter. - **Shader compilation** (Shader compilation): Converting graphics effect programs into code for the GPU. If it isn’t done ahead of time, the screen hitches the first time an effect appears, and updating the graphics driver invalidates the saved results, so it happens all over again. - **Main thread** (Main thread, Game thread): The game’s central thread, which runs game logic and prepares each frame in turn. If any one task on it takes long, the screen freezes for that long. - **Timer resolution** (Timer resolution): The shortest interval at which the OS can wake a sleeping program. The Windows default is 15.6 ms, so unless a program changes it, even “wake me in 1 ms” wakes up late. - **Thermal throttling** (Thermal throttling): A protective feature that lowers CPU and GPU speed when a device gets hot. On phones it’s common after a few to tens of minutes of play. - **VRAM** (Video memory): Dedicated memory on the graphics card. Textures and models are loaded here for rendering. When it runs short, data shuttles to and from PC memory over a slower path and the game stutters. - **Net graph** (Net graph): A dev and debug overlay that shows ping, packet loss, FPS, and tick as live graphs on the game screen. Having it in a lag report video makes finding the cause much easier. - **Retransmission rate** (Retransmission rate): The share of sent TCP packets that had to be sent again. There’s no official threshold, but a server-wide average below 0.1% is generally healthy, and above 1% many players are likely to feel lag. Also check how many times higher it is than the usual level. - **SACK** (Selective ACK): A TCP feature in which the receiver reports in detail: “I got this range; only this part is missing.” It can recover several losses at once. - **RACK-TLP** (Recent ACK, Tail Loss Probe): A TCP feature that detects loss based on time and, when no ACK arrives for a while, resends the last packet to speed up recovery. It’s the default on current Linux and Android. On Windows, TLP and RACK are on by default from Windows 10 (1607) and Server 2016, and the newer RACK that also recovers lost retransmissions arrived with Server 2022. It works only on connections with SACK enabled. - **Spurious retransmission** (Spurious retransmission): Resending a packet that wasn’t lost: it arrived late or out of order, and the sender mistook it for lost. It wastes bandwidth and needlessly cuts the sending rate. - **Zero window** (Zero window): The receiver’s buffer is full, so it has told the sender “stop sending for now.” It looks like retransmission, yet the network is fine; the receiving program just didn’t read the data in time. - **thin stream** (Thin stream): A connection that sends small packets sparsely, like a game. Fast retransmit signals rarely build up, so losses cause long stalls. - **Policer** (Policer): A rate limiter that drops packets exceeding a set rate right away, without queuing them. A limiter that queues them and releases them slowly is called a shaper. - **Pacing** (Pacing): Spreading outgoing packets evenly over time so they don’t all go out in one burst. It keeps small buffers from overflowing. - **ECN** (Explicit Congestion Notification): Under congestion, marking packets with a “congested” flag, without dropping them, so the sender slows down. It signals congestion with no loss. It only works if both endpoints and the equipment on the congested segment all support it. - **MSS** (Maximum Segment Size): The maximum amount of data TCP puts in one packet. Usually 1,460 bytes; lowering it to fit tunnel segments prevents MTU black holes. - **Handover** (Handover): A moving phone switching the cell tower it’s connected to. - **Percentile** (Percentile (p50, p95, p99)): The value at a given percentage position when all values are sorted from smallest to largest. p50 is the median; p99 is the value near the slowest 1 in 100. It reveals the spikes an average hides. - **Tail latency** (Tail latency): Long delays that happen occasionally even though most requests are fast. They barely show up in the average, but they’re what players remember as lag. - **Synthetic monitoring** (Synthetic monitoring): Measuring path quality with dedicated probes or servers that periodically run ping, traceroute, and similar tests from fixed locations, standing in for real players. RIPE Atlas is the best-known public tool. - **Aggregation interval** (Aggregation interval): How many seconds or minutes of data one point on a graph combines. The longer the interval, the more short spikes get averaged away. - **Postmortem** (Postmortem): A write-up after an incident covering what happened, why it happened, and what will change. It is written blamelessly, with the goal of preventing a repeat. - **C-state** (CPU idle state): Power-saving states a CPU enters when idle. Deeper states save more power but take longer to wake from. - **Live migration** (Live migration): A cloud provider moving a running virtual machine to another host, for example for host maintenance. The VM can pause briefly during the move. - **SNAT** (Source NAT): NAT that rewrites the source address of outgoing packets to a public address. Each public address has a limited number of usable ports, and new connections fail once they run out. - **NAT gateway** (NAT gateway): A cloud device that lets servers on a private network share one public address to reach the internet. It limits concurrent connections per destination. - **LEO satellite internet** (LEO satellite internet): Internet service through satellites orbiting hundreds to thousands of km up. Latency is far lower than with geostationary satellites, but it can spike when the connection switches to another satellite. - **GeoIP** (IP geolocation): A database that estimates country, city, and ISP from an IP address. Wrong or outdated entries can get players assigned to a distant server. - **TLS certificate** (TLS certificate): A digital document proving a server is who it claims to be. It has an expiry date, and once it expires, encrypted connections fail and players can’t connect. - **Frame generation** (Frame generation): A technology in which the graphics card inserts predicted frames between the frames it actually rendered to raise FPS. Motion looks smoother, but input-to-screen latency can increase.