Game Lag White Paper
English

Game Lag
White Paper

This white paper explains the causes of stutter, teleporting, and disconnects across 13 layers, from your game screen to the server’s database. Each cause comes with how to confirm it and the team that owns it, and you can change conditions in the simulations to check for yourself. The examples center on MMOs, but most of it applies to online games in general, regardless of genre.

00Getting started

The four factors that cause lag

There are more than a hundred causes, but the factors that create lag fall into four groups: packets arrive late, arrive unevenly, never arrive, or something stops processing. Games use many techniques to hide these factors, and whatever shows through when the hiding fails is the “shape” of the lag we see.

In an MMO, the world on your screen is redrawn from packets the server sends. It varies by game, but a server typically computes the game state 10–30 times a second (each pass is called a tick), then picks out only the changes around each player and sends them as packets. Your PC reads the packets as they arrive and draws the screen. So most lag comes down to how the game shows “packets that didn’t arrive on time.”

Some lag falls outside the four factors. If the server and your PC compute the same thing differently (different movement rules, a bug), you get rubber-banding or invisible entities even when the connection is fine. The clue for this kind of lag is that it repeats at the same place or with the same action, regardless of ping.

From factor to symptom

Analogy

Think of package delivery. If delivery always takes 3 days, that’s latency. If one package takes a day and another takes five, that’s jitter. A box that disappears is packet loss, and a distribution center that shuts its doors is a stall. Games get by with tricks like “if a box is late or missing, fill in its contents by guessing from the boxes before and after it (interpolation, extrapolation)” and “if it never comes, ask for it to be sent again (retransmission).”

Getting a feel for time scales

Most lag talk happens in milliseconds (ms, 1/1000 of a second). Remember just a few numbers from the table below and what the game team and infra team say will make much more sense.

Reference pointTimeWhat it means

Queue behavior common to every layer

CPU, disk, database, router, ISP link. The layers differ, but the structure is the same. There are workers that process requests (CPU cores, threads, DB connections, and so on), and a queue forms in front of them. When the workers are idle the queue stays empty, but once utilization (how busy they are) passes 80–90%, the queue grows sharply. With a single worker and requests arriving at random, the average wait equals the processing time at 50% utilization, 4 times that at 80%, and 9 times at 90%. That’s the answer to “The CPU still has 10% left, so why is it lagging?” On top of that, the CPU number on a monitoring dashboard is usually averaged across several cores and over 1–5 minutes, so it hides a single core pinned at 100% or a spike lasting only a few seconds.

What this white paper covers and what it doesn’t

This white paper covers the causes of lag during online gameplay, from the PC or phone drawing your game screen through your home network, the ISP, and the data center to the servers and databases. It does not cover cloud gaming, which streams the game screen to you as video, voice chat, which runs as a separate service, or patch and download speed, because they are built differently. The network causes inside them (Wi-Fi, bufferbloat, congested links, and so on) are the same as the ones here, though.

01The full map

The packet’s path: from input to the server database

When you press a skill button, the signal passes through your PC, your home network, the ISP, and the data center to reach the server. The result, processed by several layers inside the server, travels back through the same layers and is drawn on your screen. There are 13 layers in all, and a blockage in any one of them means lag. Click a layer on the map below to jump to its chapter.

02Try it yourself

Lag lab

A small simulation with one server, one connection, and one PC. Break the conditions one at a time and watch how stutter, teleporting, rubber-banding, fast-forward, slow motion, input lag, freezes, and disconnects each come about. The packet timeline draws a line from when each packet was sent to when it arrived. The more a line slants, the longer that packet took, and an × marks a lost packet.

03Find it by what you see

Symptom guide

Players just say “it’s lagging,” but the shape of the lag tells you a lot about the cause. The small picture for each symptom traces the path of a character on screen. Overlapping dots mean it stopped, and widening gaps mean it sped up or skipped ahead.

Stutter

Movement isn’t smooth: it keeps pausing briefly and moving again. If ping looks normal, the likely cause is frame timing on your PC (client or OS); if ping jumps around, jitter on Wi-Fi or the connection is more likely. Keep in mind that most in-game ping readouts are measured inside the game loop, which runs once per frame, so a frame spike can make the ping number spike too.

Teleporting

A character moves to a distant spot in one step, with no movement in between. It usually means packets stopped arriving for a while. Check for packet loss, brief connection drops, server stalls, and extrapolation failures. If one player jumps around while everyone else is fine, suspect that player’s connection first.

Rubber-banding

Your character runs forward, then gets dragged back to where it just was. Your screen (prediction) and the server’s call disagree. Your input never reached the server (packet loss), the server’s movement validation cut it off, or the two sides compute movement differently.

Fast-forward

A frozen screen starts moving again, and the backlog of movement, hits, and damage plays out all at once at high speed. Packets piled up somewhere and were released all at once. Typical causes are waiting on a TCP retransmission, the server catching up, and a processing backlog on the client.

Slow motion

Everything moves slowly. Skill casts and monster movement look stretched out. Depending on the server design, speed can stay normal while it shows up as stutter or teleporting. The server can’t finish its ticks on time. The network is fine, so ping measured outside the game stays the same; in-game ping can rise slightly if it includes time waiting for server processing. Check for player surges, AOI calculation, broadcast, and memory shortage.

Input lag

It takes a while from pressing a button to seeing the result. The screen itself can still be smooth. Round-trip time (ping) is long, or a queue is building up somewhere. Check distance, router queues, Nagle (a TCP feature that collects small packets and sends them together), and server queues. If it always feels sluggish even with low ping, look at your PC (V-Sync, low FPS) or at a design that waits for server confirmation on every action (see the netcode models chapter).

Freeze

Everything on screen stops for a moment (0.5 s to a few seconds), then moves again. The whole server stopped (GC, deadlock, blocking call), the connection dropped briefly, or your PC froze.

Dropped action / rollback

Something you definitely did never happened, or its result gets reversed much later. The request was lost (packet loss, queue overflow), the server ruled differently from what your screen showed (timing difference, rejection after client-side feedback), or saving failed partway (DB lock or outage, server crash).

Disconnect

The connection drops mid-game and you land back at the login screen or a reconnect dialog. Not a single packet arrived within the timeout. Check for long connection drops, idle timeouts, server crashes or restarts, and a server or PC that stalled longer than the timeout (long loading). If the game just closed with no message, look at a forced client shutdown (crash, out of memory) before the connection.

Can’t connect / infinite loading

You can’t get into the game, or you’re stuck on a loading or entry screen. Whatever accepts new connections (the server’s connection queue, firewall, login server, DB) is full. It’s especially common right after maintenance.

Invisible / ghost entities

NPCs, monsters, or players that should be there are missing only on your screen, or entities that are already gone remain only on your screen. This usually has little to do with speed: one packet went missing, or drawing the entity failed. Check for channel or phasing differences, lost spawn and despawn messages, messages discarded during loading, and asset loading failures. The decisive clue is whether it shows up after you leave view range and come back.

04Netcode design

Netcode models and feel

Some games feel fine at 150 ms of ping, while others feel sluggish at 60 ms. That holds even outside action games. On the same connection, the difference usually comes from how the client and server have agreed on what gets decided, when, and by whom, in other words the netcode design. Some of it is a deliberate choice, and some of it really is badly built.

Every networked game solves the same problem. There is always a time gap between the server and your PC, and one side has to decide how to handle things that “aren’t confirmed yet.” There are four broad options.

  • Wait: show nothing until the server confirms. It’s accurate, but your ping becomes your response time.
  • Show it first, fix it later: play your own actions immediately and correct them if the server’s result differs. It’s fast, but you occasionally see rubber-banding or canceled actions.
  • Schedule it in advance: announce the event along with a future time, like “ground slam in 1.5 seconds.” If the lead-up lasts longer than the ping, the ping doesn’t show at all.
  • Everyone runs the same computation: exchange only inputs and have every machine compute the same thing (lockstep, rollback). It sends very little data, but one player’s delay spreads to everyone.

So sensitivity to ping depends less on genre and more on two questions: how many server round trips a single core action waits for, and whether the game rules allow comfortably more time than “ping + human reaction time”.

Common netcode models

ModelHow it worksCommon inWhat it looks like at 150 ms pingWeak spots
Request-response
Shown after the server confirms
Press a button, ask the server, and play the result only when the answer arrives.Turn-based, card, and idle games; shop, trade, and crafting UIs; skill and item use in older MMOsEvery action starts about 0.2 s late. Barely noticeable in turn-based gamesChained actions, UIs with several round trips on one screen
State sync + interpolation
Server-authoritative
The server sends the game state every tick, and the client draws the motion between two states.Showing other players and monsters in most MMOsOther players appear about 0.2 s in the past. Usually barely noticeableJitter (variation in packet arrival times) and loss → teleporting, low tick rates
Client-side prediction + server reconciliationYour input takes effect immediately, and when the server’s result arrives, the client compares the two and corrects.FPS games, action MMOs, movement in most MMOsYour controls respond instantly. Occasional brief rubber-bandingFrequent corrections if the client and server compute differently
Lag compensation
Server rewinds for hit registration
The server rewinds to the past moment the attacker was seeing and registers the hit there.FPS games, non-target actionThe shooter feels it’s fair, but the target gets “I was behind cover and still got hit”Feels unfair to the target. The higher the attacker’s ping, the further the rewind and the worse it gets
Command/destination syncSend only the intent, like “go here” or “attack this target,” and both sides work out the rest.Click-to-move MMOs, tab-target combat, some MOBAsOnly the start is slightly late. Movement and attacks stay smoothNeeds correction if paths or results diverge
Scheduled events
Based on server time
Announce the event along with a future time, like “start at server time T,” and every client plays it at that time.Raid boss patterns, cutscenes, on-the-hour eventsPractically no effect if the warning is longer than the pingIf the message arrives after the scheduled time, the beginning gets skipped
Deterministic lockstepGather everyone’s inputs and compute the same turn identically on every machine. Inputs get a fixed delay.RTS (StarCraft-style), some co-op and puzzle gamesAll input is consistently late (hidden by instant sound and visual cues on press). High jitter makes everyone freezeJitter, loss, the slowest player
Rollback
Predict, then rewind
Predict the opponent’s input and run ahead. If the guess was wrong, rewind and recompute.Fighting games (GGPO-style), some action and sports gamesControls feel almost instant (usually 1–3 frames of input delay). The opponent’s animation occasionally jumps a few framesWith high ping the rewinds get bigger and look like teleporting
Client-authoritativeEach client decides its own results, and the server only relays and records them.Some mobile and casual games, P2P and relay setupsYour own screen feels smooth. Results disagree with what other players seeCheating, “I hit them but it didn’t count”

Real games mix these models. It’s normal to pick one per action: prediction for movement, client-side feedback followed by server confirmation for skills, scheduled events for boss patterns, and request-response for trades.

What games that feel smooth even at 150 ms have in common

1. No round trip sits inside a single action. Pressing a button starts the animation, sound, and effects immediately (client-side feedback), and the server’s result is used only for parts where a delay doesn’t show, like damage numbers.

2. The game rules allow comfortably more time than the ping. If a boss telegraphs an attack 1–2 seconds ahead, you can still dodge easily even when packets arrive about 0.2 s late and your reaction takes 0.25 s. If your own skill has a cast time, the server confirmation finishes while the cast bar fills, so the wait hides inside the cast. This is the biggest reason tab-target MMOs are insensitive to ping. Conversely, a short telegraph of around 0.5 s is hard to see and dodge with as little as 150 ms of ping (see the timing window simulation below).

3. Chained actions are accepted in advance. With input buffering (a skill queue) that accepts your next skill even if you press it before the cooldown ends, no round trip gets squeezed in between combo steps.

4. Jitter gets absorbed. An interpolation buffer and playback based on server time turn packets that take “usually 150 ms, sometimes 250 ms” into a steady stream that is always about 250 ms late. You see a little further into the past, but motion stays smooth. People get used to a constant delay quickly, but have a hard time getting used to an uneven one. Well-made games grow and shrink the buffer on their own as jitter rises and falls.

5. The server’s call matches what you saw. Dodges and hits are judged from the moment the player saw them (lag compensation), or the rules don’t depend on position in the first place (tab targeting).

6. One player’s delay doesn’t make everyone else wait. With a server-authoritative setup, other players are fine even when your ping is bad. With lockstep or a host-based setup, the slowest player sets the feel for everyone.

Telling intended design apart from flawed design

Sometimes a game is sensitive to ping by design, and sometimes because it was built wrong.

May be a deliberate choice
  • Short timing windows: games where a short window is the fun itself, like a 0.2 s parry or a perfect dodge. Ping inevitably eats into your reaction time and jitter throws off the timing, so these games soften it with lag compensation or regional servers.
  • Server confirmation to prevent cheating: for results that must never be faked, like currency, items, and rankings, waiting for the server to confirm is the right call.
  • Lockstep: the most practical structure for syncing hundreds of units through inputs alone. The trade-off is that input delay has to be tuned to the ping.
  • Fairness: some games deliberately weaken lag compensation so that the player getting hit isn’t punished because someone else has high ping.
Signs it was likely built wrong
  • A game you control directly with keyboard or gamepad waits for server confirmation even for movement and basic attacks. Without prediction, ping becomes the feel of the controls. Command-style controls like click-to-move hide a server wait much better, so some games, MOBAs among them, choose that on purpose.
  • Several round trips for one UI action: if opening a window → fetching a list → confirming → buying are each a separate round trip, the whole thing takes 0.7–0.8 s at 150 ms of ping. These can be bundled into one.
  • Consistently sluggish even though connection ping is low: suspect Nagle (TCP’s default behavior of holding small packets to send them together, turned off with TCP_NODELAY), a double wait where requests are held until the next tick and results go out on the tick after that, or a design where each action responds only after its DB save finishes. V-Sync or a low FPS on your PC gives the same feel.
  • “Confirm, then accept the next input” with no skill queue: a round trip lands between every combo step, so DPS drops in proportion to ping.
  • Playing things the moment they arrive: if the client plays animations in the order received, with no interpolation buffer or server time, jitter turns straight into stuttering animation.

Causes of lag in netcode design

Feedback only after the server responds (request-response) Request-response (no client-side feedback)

Press a button and there’s no animation or sound until the server answers. Your ping becomes your response time.

Why: Skills, movement, and item pickups play only after the server confirms them → Effect: From the moment you press, nothing happens for a full round trip plus the tick wait → On screen: At 150 ms ping, every action feels 0.2 s sluggish

Symptoms: Input lag · Primary owner Game team (Client development) · Also Game team (Server development)

Chatty protocol (many sequential round trips) Chatty protocol / sequential round trips

If one action needs several server round trips one after another, your ping is multiplied by that many.

Why: Open shop → request list → check price → buy → refresh inventory, each as a separate request → Effect: Each request is sent only after the answer to the previous one arrives → On screen: At 150 ms ping, a single purchase takes close to 1 s. Loading takes unusually long

Symptoms: Input lag, Can’t connect / infinite loading · Primary owner Game team (Server development) · Also Game team (Client development)

No skill input buffering No input/spell queue

If you can’t press the next skill until the server confirms the previous one has finished, a round trip gets inserted between every skill in a rotation.

Why: The next skill input is accepted only “after the previous skill is confirmed” → Effect: A gap as long as your ping opens between every skill → On screen: Gaps between skills in a rotation, and the higher the ping, the lower the DPS

Symptoms: Input lag, Dropped action / rollback · Primary owner Game team (Client development) · Also Game team (Server development)

Short timing windows eaten up by ping Timing window too short for latency + reaction

When the time you have to react is short, as with dodges, parries, and guards, ping eats up that time and some attacks become impossible to avoid.

Why: Short timing windows, such as a 0.5 s boss attack telegraph or a 0.2 s parry window → Effect: You see the telegraph late (downstream latency + interpolation), and your input also arrives late (upstream latency + tick wait) → On screen: You get hit even though you clearly dodged, parries don’t go off

Symptoms: Dropped action / rollback, Input lag · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)

Hit registration without lag compensation Server-now hit validation

If the server checks hits only against where targets are on the server right now, what you saw on your screen and the server’s call disagree.

Why: The opponent on your screen is at a position about 0.2 s in the past (at 150 ms ping and 100 ms interpolation) → Effect: The server checks against the current position, so the target has already left the spot you aimed at → On screen: A clear hit counts as a miss. You have to lead moving targets

Symptoms: Dropped action / rollback · Primary owner Game team (Server development) · Also Game team (Client development)

Too much lag compensation Excessive lag compensation

If the server rewinds too far in the attacker’s favor, the target gets hit even after they’ve already taken cover.

Why: The server rewinds a long way to check hits for a high-ping attacker → Effect: On the target’s screen, they were already behind cover → On screen: “I got shot behind a wall,” high-ping players have the advantage

Symptoms: Dropped action / rollback · Primary owner Game team (Server development)

Client authority Client-authoritative results

When each client decides its own results, your own screen feels responsive, but results disagree with other players’ screens and the game is easy to hack.

Why: The client decides position and hits, and the server only relays them → Effect: Two players each claim they hit first, and the server can’t verify either claim → On screen: Opponents teleport or pass through walls, “I hit them but it didn’t count”

Symptoms: Teleporting, Dropped action / rollback · Primary owner Game team (Server development) · Also Game team (Client development)

Lockstep waiting on the slowest player Lockstep waits for the slowest peer

When everyone computes the same turn together, one player’s late input makes everyone wait.

Why: Each turn can be computed only after every player’s input has arrived → Effect: One player’s input arrives late because of jitter or packet loss → On screen: Everyone hitches at the same time, and in bad cases a “Waiting for players” window appears

Symptoms: Freeze, Stutter, Input lag · Primary owner Game team (Server development) · Also Game team (Client development)

Rollback netcode misprediction Rollback misprediction

The game predicts the opponent’s input and shows it early, then rewinds and recomputes if the guess was wrong. The higher the ping, the further it rewinds.

Why: The opponent changes their input (different from the prediction) → Effect: The real input arrives half a ping late, so the game rewinds that far and recomputes → On screen: The opponent’s animation skips a few frames or changes suddenly

Symptoms: Teleporting · Primary owner Game team (Client development)

Events played on arrival without timestamps Events played on arrival (no timestamps)

If server events carry no timestamp and play as soon as they arrive, network jitter carries straight through into uneven animation timing.

Why: “Attack start” and “play effect” events run as soon as they arrive → Effect: Each packet arrives at a different time, so the intervals are uneven → On screen: Chained attack animations speed up and slow down, and boss pattern timing differs every time

Symptoms: Stutter, Fast-forward · Primary owner Game team (Client development) · Also Game team (Server development)

Double tick wait Double tick quantization

If requests wait for the next tick to be processed and the results wait for the tick after that to be sent, the tick interval is added twice.

Why: Received requests are processed on the next tick → Effect: Results are also batched and sent on the next send tick → On screen: Ping on the connection is low, but responses are consistently late by about 1.5 times the tick interval. On a 10-tick server, 0.15 s on average and 0.2 s at worst

Symptoms: Input lag · Primary owner Game team (Server development)

Overly strict server validation Over-strict server validation

If the server checks movement speed, cooldowns, and range too strictly, it rejects even valid inputs that arrive bunched together because of jitter.

Why: Strict rules such as “max distance per tick” or “0 ms cooldown tolerance” → Effect: When jitter makes two commands arrive in the same tick, they’re judged as rule violations → On screen: Rubber-banding, skills rejected even though the cooldown is up

Symptoms: Rubber-banding, Dropped action / rollback · Primary owner Game team (Server development)

Player-hosted server (host) Listen server / host advantage

When one player’s PC acts as the server, that player’s connection and PC performance decide how the game feels for everyone.

Why: The host’s PC acts as the server (P2P, listen server) → Effect: If the host’s connection or PC is slow, everyone feels it, while the host has zero ping → On screen: Only the host has the advantage, and when the host leaves everyone freezes or disconnects

Symptoms: Stutter, Freeze, Disconnect · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)

Server rejects after client-side feedback Client-side feedback rejected by server

When the server later refuses a hit or skill your screen already showed, the result you clearly saw never happened.

Why: Hit effects and skill animations play before the server confirms (client-side feedback) → Effect: The server rechecks range, target position, cooldown, and resources and rejects the action → On screen: Blood sprays but there’s no damage, the skill animation plays with no effect, only the cooldown runs

Symptoms: Dropped action / rollback, Rubber-banding · Primary owner Game team (Client development) · Also Game team (Server development)

Pathfinding mismatch in command sync Command sync with divergent pathing

When only “go here” is exchanged and each side computes the path itself, even a small difference in the calculation sends a character or monster down a different path until it gets pulled back into place.

Why: With click-to-move and monster chasing, only the destination is sent and the client computes the path on its own → Effect: Differences in terrain data, collisions with other characters, or calculation order make it move along a different path from the server’s → On screen: Monsters walk through walls and then snap to another spot, clicked characters change direction as if sliding

Symptoms: Teleporting, Rubber-banding · Primary owner Game team (Server development) · Also Game team (Client development)

Low snapshot send rate Low snapshot / update rate

If the server sends position updates (snapshots) only a few times a second, the interpolation buffer has to be that much longer, and you see other characters further in the past.

Why: Position updates sent only 5–10 times a second to save bandwidth → Effect: Smooth rendering needs a buffer of twice the packet interval (200–400 ms); with a shorter buffer, a single missed packet causes a freeze → On screen: Opponents’ direction changes show up late and disagree with hit registration. With a short buffer: stutter, plus teleporting on packet loss

Symptoms: Stutter, Teleporting, Dropped action / rollback · Primary owner Game team (Server development) · Also Game team (Client development)

05Scope of impact

When only one player or one client is affected

The most confusing lag reports are the ones where only a few people, or only one side, are affected. How one slow player looks to everyone else, and whether that player slows others down too, depends entirely on how the server processes input and on the netcode model. This chapter also covers two clients on the same PC where only one of them can’t see NPCs.

When only certain players or certain connections are slow

Most MMOs today are server-authoritative. The server decides every outcome, and the client draws the results it receives. In this setup, lag mostly shows up only for the slow player.

  • The slow player gets input lag: actions that need server confirmation, like skills and picking up items, run late by their ping. Movement shows up right away thanks to prediction, but with high jitter (variation in packet arrival times) they also get rubber-banding and see other players teleport.
  • Everyone else only sees the slow player’s character hitch and then fast-forward, or teleport. Their own controls and the monsters behave normally. If the slow player has high ping but no jitter or loss, their character just moves smoothly a little behind where it really is. What makes a player look laggy to others is jitter and loss far more than ping.
  • If only a specific ISP or region has a bad connection, all of those players hit the symptoms above at once. From the server’s point of view only their inputs arrive unevenly, so false positives from movement validation or cheat detection also pile up on those players.
  • If it’s slow only on one specific character, suspect that character’s data before the connection. A character with thousands of items or mail messages piled up reads and writes several times more than others every time it logs in or saves. To tell the two apart, log in with the same character from a different PC or connection and see whether it’s just as slow.

Some structures, though, let one slow player slow everyone down. What they have in common is that “someone is waiting for that player.”

  • Everyone waits for the same turn: lockstep (RTS) and co-op content that advances turn by turn. If one player’s input is late, everyone stops. Even with no jitter, simply being late means everyone’s input takes effect as late as the slowest player’s ping.
  • The server waits while sending to the slow player: blocking sends (a send call that stops and waits until the send buffer has room) and synchronous processing. Everyone handled by that server thread slows down. The usual design gives each player a separate send queue and never waits. Then the lag shows up only for the slow player, and if that queue grows too long, only that player gets disconnected.
  • The slow player plays a central role: P2P (players connect directly with no server) or a listen server (a player’s PC doubles as the server) with that player’s PC as the host, and events that advance only on the party leader’s authority. In games that offload monster movement calculations to a nearby player’s client to reduce server load, the monsters that player handles stutter on everyone’s screen.
  • The server rewinds to the slow player’s view when registering hits: lag compensation. The slow player hits fairly, but the target feels cheated: “I was already behind cover and still got hit.” So games cap how far they rewind. The cap varies by game and is roughly 0.2–1 s (the Source engine default is 1 s).

How it looks depends on how the server processes input

How the server processes inputWhat the slow player experiencesHow the slow player looks to othersOther players’ own game
Batched per tick
Fixed tick, all received inputs at once
Skill results are late by ping plus the wait for the next tick (input lag). Rubber-banding if movement validation is strictHitches, then several steps at once (fast-forward, teleporting). Smooth if ping is high but there’s no jitterNo effect
Processed on arrival
Event-driven, applied and sent as soon as received
Input lag equal to ping. Faster only by the tick wait it skipsMovement speeds up and slows down (mild fast-forward). Several skills that arrived together go off in one instantNo effect
Per-player input buffer
Buffered per player, one input per tick
Confirmation is late by the buffer lengthFairly smooth. Stands still briefly when the buffer runs emptyNo effect
Lag-compensated hit registration
Rewinds to what the attacker saw
Hits land where you aimed (within the rewind cap)Their attacks still land after you take coverUnfair hits (spreads)
Lockstep / turn waitsInput lag. Freezes if input is lateEveryone freezesFreeze. Input lag even when merely late (spreads to everyone)
Blocking sends / synchronous processing
The server waits for that player
Freeze, then fast-forwardEveryone on that server thread slows downSlow motion / freeze (spreads to the players on that thread)
Slow player is the host
P2P, listen server
Their own ping is 0Stutter on everyone’s screenEveryone lags
Slow player controls monsters
Monster movement computed on a client
Monsters look fine on their own screenThe monsters they handle hitch, then teleportEveryone fighting those monsters (spreads)

Two clients on the same PC, only one can’t see NPCs

If one person runs two clients on the same PC and only one of them can’t see NPCs, the connection is almost never the cause. Both clients use the same router and the same connection. The difference comes from one of three places.

  1. The server never sent it to that client: a different channel, instance, or quest phase (a feature that splits which NPCs you see by quest progress), an AOI (visibility) registration race, a per-connection send limit, a session bug that treats the same PC or same IP as one person, or a multi-client restriction.
  2. It was sent, but the client threw it away: spawn messages discarded because they arrived during loading, a burst of spawn data right after entering lost to receive buffer overflow or an unreliable channel (a channel that doesn’t resend lost data), a lost baseline snapshot (the full state that deltas build on), a new NPC with a reused ID mistaken for the old one, another client grabbing the packets because of a fixed UDP port collision, a background window falling behind on processing until its receive buffer overflows, or display held back because the server time estimate is off.
  3. It arrived, but couldn’t be drawn: both clients writing the same cache file at once so model loading fails, running out of graphics memory (VRAM), different settings such as a cap on displayed characters, or version and data mismatches.

The three strongest clues: is the name tag there but the character model missing (the server sent it, drawing failed), does it appear after you leave view range and come back (one spawn message was missed), and does it improve when you bring the affected window to the front (background window throttling). Conversely, a “ghost entity,” such as a monster that’s already dead but still standing on your screen alone, means a despawn message was missed.

Problems only some players hit

A lagging player moves in bursts on others’ screens Laggy player seen by others (bursty inputs)

Inputs from a player with a bad connection reach the server unevenly, in bunches. If the server applies whatever arrived on each tick, other players see that character hitch and then cover several steps at once.

Why: The lagging player’s move commands arrive 0 at a time on some ticks and 2–3 at a time on others → Effect: The server applies them all on the tick they arrive, so that character’s position changes in steps → On screen: On other players’ screens, only that character hitches and then moves several steps at once. Everyone else looks fine

Symptoms: Fast-forward, Teleporting · Primary owner Game team (Server development) · Also External (External)

Fast-forward on servers that process on arrival Event-driven processing of bursty inputs

On a server that processes and broadcasts packets as soon as they arrive, a lagging player’s bunched-up actions run back to back immediately.

Why: A lagging player’s skill and move requests arrive in a bunch → Effect: The server runs them in order the moment they arrive and tells everyone right away → On screen: Others see that player use several skills in an instant or move as if fast-forwarding

Symptoms: Fast-forward · Primary owner Game team (Server development) · Also Game team (Client development)

Per-player input buffer size Per-player server input buffer (jitter buffer)

If the server holds a few inputs per player and takes out one per tick, others see smooth motion, but your own actions are confirmed on the server that much later.

Why: The server collects a lagging player’s inputs in a buffer and applies one per tick → Effect: A small buffer often runs empty, so the character stands still or the server guesses from the last input; a large buffer confirms the player’s own inputs late → On screen: Too small: others see hitches. Too large: your own skill results come late (input lag)

Symptoms: Stutter, Input lag · Primary owner Game team (Server development) · Also Game team (Client development)

Validation false positives concentrated on one ISP Anti-cheat / movement validation false positives on bad ISPs

Players on high-jitter connections have their inputs arrive in bunches, so they often trip the server’s speed and cooldown checks.

Why: Jitter on a specific ISP’s or region’s connections rises in the evening → Effect: The server judges valid inputs that arrived in a bunch as speeding or cooldown violations → On screen: Only that ISP’s players get rubber-banding and rejected skills, and in bad cases the server kicks them (disconnect)

Symptoms: Rubber-banding, Dropped action / rollback, Disconnect · Primary owner Game team (Server development) · Also Infra team (Network infrastructure)

One lagging party member and boss mechanics One laggy member in a synchronized mechanic

In raid mechanics where everyone has to react together at a set moment, one lagging player’s late reaction fails the whole party.

Why: Group mechanics such as “everyone spread out at once” or “one player presses the button” → Effect: The lagging player sees the telegraph late, and their input also arrives late → On screen: One player causes a wipe, and the rest of the party feels it happened “because of the laggy player”

Symptoms: Dropped action / rollback, Input lag · Primary owner Game team (Server development) · Also Game team (Client development)

A lagging client controls the monster Monster movement delegated to a player client

Some games hand monster movement to one nearby player’s client to reduce server load. If that player’s connection is bad, the monster moves strangely on everyone’s screen.

Why: The server hands monster movement to the client of the nearest (or first-arriving) player → Effect: That player’s reports reach the server late or in bunches → On screen: Only that monster hitches and then teleports on every nearby screen. It looks fine on the controlling player’s own screen

Symptoms: Teleporting, Stutter, Fast-forward · Primary owner Game team (Server development)

Bloated data on one character One character with oversized data (inventory, mail, buffs)

A character with thousands of items or mails piled up, or an unusually large friend list, block list, or set of buffs, has several times more to load, save, and announce to nearby players than others. It’s slow only on that character, regardless of connection.

Why: Thousands of items or event rewards pile up in the inventory and mailbox of a long-played character → Effect: Every login, zone move, and save reads and writes that much from the DB, and the equipment and buff data sent to nearby players is large too → On screen: Only that character has long loading screens and hitches when opening the inventory or mail. On a server where the game thread waits on saves, nearby players freeze briefly too

Symptoms: Can’t connect / infinite loading, Input lag, Freeze · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

Different channel, instance, or phase Different channel / instance / phase

If two characters are in different channels or instances, or in different “phases” where the visible NPCs depend on quest progress, they see different worlds.

Why: The second character is assigned to a different channel, or its quest stage differs → Effect: The server doesn’t send that NPC to that character (working as intended) → On screen: The NPC is missing on one side only. It looks like a bug but is by design

Symptoms: Invisible / ghost entities · Primary owner Game team (Server development) · Also Game team (Client development)

Spawn messages dropped during loading Spawn messages dropped before the client is ready

Right after you enter a zone, the server sends spawn messages for nearby NPCs, but the client is still loading the map and throws them away.

Why: The server sends spawn messages for nearby entities right after processing the entry → Effect: The client is still loading and has no message handler yet, so it drops the messages → On screen: The server treats them as sent and never resends them. The NPC stays invisible until it leaves view and comes back

Symptoms: Invisible / ghost entities · Primary owner Game team (Client development) · Also Game team (Server development)

AOI registration race Interest-management race on enter/leave

If a character registers in the AOI grid at the same moment an NPC moves between grid cells, that NPC’s spawn message can be missed.

Why: Processing an entry, channel change, or teleport happens at the same instant an NPC moves → Effect: That NPC is left out of the “newly visible entities” calculation → On screen: A few specific NPCs are invisible, or NPCs that already left are still there

Symptoms: Invisible / ghost entities · Primary owner Game team (Server development)

Lost baseline snapshot Lost baseline for delta compression

When the server sends “only what changed since last time,” losing the full state sent once at the start (the baseline) means later changes can’t be applied.

Why: The packet with an entity’s full state (baseline) is lost or dropped before processing → Effect: The client has nothing to apply later changes to, so it ignores them → On screen: That entity is invisible, or suddenly appears much later

Symptoms: Invisible / ghost entities, Teleporting · Primary owner Game team (Server development) · Also Game team (Client development)

Missed despawn message (ghost entity) Missed despawn (ghost entity)

The reverse case: if the “it’s gone” message is missed, NPCs or players that already died or left stay on your screen only.

Why: Death, leave, or out-of-view messages are lost or arrive out of order → Effect: The client thinks the entity is still there → On screen: A monster that doesn’t react when hit, a player who already left still standing there

Symptoms: Invisible / ghost entities · Primary owner Game team (Server development) · Also Game team (Client development)

Spawn data lost in the burst after entering Initial spawn burst lost (unreliable channel, receive buffer, fragmentation)

The moment you enter a zone, the server sends spawn data for tens to hundreds of nearby entities all at once. If it goes over an unreliable channel, or the receive buffer overflows while the client is loading and can’t read the socket, part of it disappears and never comes back.

Why: Spawn data arrives in a short burst right after entering → Effect: A loading client reads the socket late and the OS receive buffer overflows, or a large UDP packet is fragmented and losing a single fragment loses the whole packet. On an unreliable channel, nothing is resent either → On screen: A few NPCs are missing only on the slower-loading client. They show up after leaving view and coming back

Symptoms: Invisible / ghost entities · Primary owner Game team (Server development) · Also Game team (Client development)

Entity ID reuse mix-up Entity ID reused without a generation counter

If the server reuses the same entity ID when a dead NPC respawns, a client that missed the despawn message in between mistakes the new NPC for the old one.

Why: An NPC dies and respawns with the same entity ID → Effect: A client that missed the despawn message ignores the spawn message as an “already known entity,” or leaves the entity in its dead state → On screen: The NPC is missing on one screen only or appears lying dead, and sometimes shows up looking like a different NPC

Symptoms: Invisible / ghost entities · Primary owner Game team (Server development) · Also Game team (Client development)

Fixed UDP port collision Two clients bound to the same local UDP port

If the client is built to use a fixed local port, a second client on the same PC either can’t get the port or ends up splitting packets with the first.

Why: Two clients try to open the same local UDP port (forcing a share with a reuse option) → Effect: The OS delivers incoming packets to only one socket, or doesn’t guarantee which one gets them. The router and server also see both clients as the same address → On screen: One client misses world packets, so NPCs and other players are invisible, or it disconnects

Symptoms: Invisible / ghost entities, Disconnect, Can’t connect / infinite loading · Primary owner Game team (Client development) · Also Game team (Server development)

Sessions keyed by IP or device (bug) Session keyed by IP or machine ID

If the server or an intermediate server identifies connections by IP or device ID, it treats two clients on the same PC (same public IP) as one person.

Why: The session table is keyed by IP, or by IP + device ID → Effect: The second client’s data overwrites or gets mixed into the first session → On screen: One side can’t see NPCs, and the other disconnects or receives someone else’s data

Symptoms: Invisible / ghost entities, Disconnect · Primary owner Game team (Server development) · Also Game team (Client development)

Multi-client restriction Multi-client restriction policy

If an anti-cheat module or server policy limits multiple clients on one PC, the second client is blocked from launching or connecting, or the first one gets disconnected. Some games only block features on the extra client.

Why: The anti-cheat module detects a duplicate launch, or the server limits extra connections from the same device → Effect: The second launch or connection is refused, or one side is disconnected. Rarely, only some features on the extra client are blocked → On screen: Can’t connect, or one side disconnects. In games that only block features, NPCs or shops are invisible on one side only

Symptoms: Can’t connect / infinite loading, Invisible / ghost entities, Disconnect · Primary owner Game team (Client development) · Also Game team (Server development)

Background window throttling Background window throttling

For a client in a background window, the game, engine, and OS cut its frame rate and processing. Received packets aren’t processed in time, so they back up or overflow.

Why: Background frame limits in game options or the graphics driver (e.g., the NVIDIA driver lets you pick 20–200 per second), power saving, the engine’s background pause setting. The OS also gives CPU and GPU priority to the window in front (foreground) → Effect: Fewer packets are processed per frame, so the queue builds up, and packets are dropped when the receive buffer overflows → On screen: When the window comes to the front, things appear all at once, or some NPCs never show up

Symptoms: Invisible / ghost entities, Fast-forward, Disconnect · Primary owner Game team (Client development) · Also External (External)

Simultaneous access to cache or asset files Shared cache / asset file lock conflicts

If two clients write to the same cache folder at the same time or lock its files, one of them can’t load NPC models or textures.

Why: Two clients write the cache and patch files in the same install folder at the same time → Effect: Loading fails because a file lock failed or a half-written file was read → On screen: A name tag with no character model, or a transparent NPC

Symptoms: Invisible / ghost entities · Primary owner Game team (Client development)

Streaming failure from memory or VRAM shortage Memory / VRAM exhaustion

When two clients share graphics memory, there’s no room to load newly needed models and textures, and some of them don’t get drawn.

Why: Two clients share VRAM and RAM. The OS may also shrink a background window’s graphics memory allowance first → Effect: The engine can’t load new models and textures, or keeps evicting and reloading them → On screen: NPCs appear late, look blurry, or are invisible, and the game stutters

Symptoms: Invisible / ghost entities, Stutter · Primary owner Game team (Client development) · Also External (External)

Different display settings Different display settings

If settings such as a visible player limit, hidden NPC name tags or models, or low-spec mode differ between two clients, they see different things.

Why: Only one client has a “limit nearby characters shown” setting or low-spec mode on → Effect: Distant or low-priority NPCs aren’t drawn (working as intended) → On screen: The NPC is missing on one side only

Symptoms: Invisible / ghost entities · Primary owner Game team (Client development)

Client version or data mismatch Client version / data table mismatch

If the second client is a different install or isn’t fully patched, it doesn’t recognize new NPC IDs the server sends and silently ignores them.

Why: An install in a different folder, or a client launched mid-patch → Effect: Unknown NPC IDs or model IDs are skipped → On screen: Only newly added NPCs are invisible on one side

Symptoms: Invisible / ghost entities · Primary owner Game team (Client development) · Also Game team (Server development)

Per-connection send budget and priority Per-connection bandwidth budget and priority

If the server caps how much it sends per connection and sends the nearest things first, a connection with a low cap gets distant NPCs late or not at all.

Why: In crowded places, the server sends in order of importance within each connection’s send cap → Effect: A connection with a low bandwidth estimate (e.g., a background window that’s slow to acknowledge) keeps pushing back entities further down the list → On screen: Distant NPCs appear late or not at all on one side only

Symptoms: Invisible / ghost entities, Input lag · Primary owner Game team (Server development) · Also Game team (Client development)

Entities held back by clock estimate error Clock estimate error holds or discards entities

If the client’s estimate of the server time is wrong, it holds back freshly arrived entity data as “still in the future” or discards it as “too old.”

Why: One client’s estimate of the server time is far off (measured during loading, or after waking from sleep) → Effect: The interpolation reference time and the entity data’s timestamp don’t match → On screen: Entities appear late or look frozen

Symptoms: Invisible / ghost entities, Stutter · Primary owner Game team (Client development)

06A common cause

TCP retransmission: causes and why the delay grows

When “TCP retransmissions” climb in server metrics, lag reports often climb with them. A retransmission is a sign that “a packet was lost” or that the sender “wrongly concluded it was lost.” The causes can sit anywhere on the path, from Wi-Fi to the server’s network card, and on a connection that sends small packets at sparse intervals, like a game’s, losing a single packet can turn into a freeze of hundreds of milliseconds. This chapter covers the root causes of retransmission, how to find them, and how to fix them.

Server sendsGame receives123×Lost456124563Waiting for the resend (depends on the recovery method)3Nothing reaches the game: freeze3–6 at once: fast-forward
A connection that sends every 50 ms loses just packet 3. TCP hands data over only in order, so even when 4, 5, and 6 arrive, it holds them back from the game until it receives 3 again. That’s how a single loss becomes a freeze followed by fast-forward. When the resend happens depends on the recovery method, anywhere from about one round trip to “round-trip time + at least 200 ms” with the retransmission timer (see “Types of retransmission” below).

Four reasons retransmission slows games down

  1. Waiting for order (head-of-line blocking): TCP won’t hand later packets to the game until it receives the lost one again. Lose one, and everything behind it stops together, then releases all at once (a freeze, then fast-forward).
  2. Retransmission wait time: the sender resends only when the retransmission timer (RTO) expires. On Linux that’s “round-trip time + at least 200 ms.” If the resent packet is lost too, the wait doubles each time (0.3 s → 0.6 s → 1.2 s …).
  3. Thin streams (connections that send small packets at sparse intervals): fast retransmit is triggered when the receiver signals that “3 later packets have arrived” (3 duplicate ACKs; an ACK is a “got it” acknowledgment). Game packets go out one every 50–200 ms, so the RTO often fires before that signal builds up. That’s why large downloads hold up fine while games alone freeze. RACK in recent Linux can decide as soon as a single later packet arrives, which narrows this gap a lot, but when packets are around 200 ms apart, even RACK is no faster than the RTO.
  4. Cutting the sending rate: TCP treats loss as a sign of congestion and shrinks how much it sends at once (the congestion window). Once it reaches an RTO, it has to ramp back up from sending one packet at a time. Meanwhile, newly generated packets pile up on the server, and big updates in crowded areas back up one after another.

Types of retransmission

TypeWhen it happensTime to recoverWhat it looks like in the game
Fast retransmit
Triggered by duplicate ACKs or SACK
Later packets arrive first and the receiver reports “a gap in the middle” (duplicate ACKs, SACK)Round-trip time + the time for 3 later packets to arriveBrief hitch. The denser the packets, the faster
RACK/TLP
Time-based loss detection, tail packet resend
When a later packet has arrived but an earlier one still hasn’t shown up after a set time. Or, when no ACK comes for a while, the last packet is sent once moreSoon after a later packet is acknowledged (RACK; it waits about 1/4 of a round trip longer in case the packets were merely reordered). With no later packets, about 2× the round-trip time (TLP), plus an extra 200 ms to allow for delayed ACK when only one packet is still unacknowledgedA relatively brief hitch even on thin streams. Default in recent Linux
RTO retransmission
Retransmission timeout
When the wait runs out with no signal at allRound-trip time + at least 200 ms, doubling on each failureA freeze of hundreds of ms to several seconds, then fast-forward, and a disconnect if it drags on
SYN retransmissionWhen the connection request itself is lost, such as a connection queue (listen backlog) overflow or a firewall block1 s, 2 s, 4 s, 8 s … (Linux 6.5 and later resends up to five times at 1 s intervals before doubling; older Windows starts at 3 s)After you press connect, the delay lands on whole seconds like 1 s or 3 s, and if it keeps failing, can’t connect / infinite loading
Spurious retransmission
Resent without real loss
A packet that wasn’t lost arrives late or out of order, so the sender assumes it was lost and resends itNothing to recover, yet the sending rate still drops (Linux may undo the cut when it detects the case through timestamps or DSACK, the receiver’s “already got it” notice)Wasted bandwidth, slower bulk transfers. Only the retransmission rate in the metrics looks high
Zero window probe
Often mistaken for retransmission
A packet the sender uses to check in while the receiver’s buffer is full and the receiver has said “stop sending for now”Until the receiver starts readingFreeze. The connection is fine; the receiving program didn’t read in time

How to find where packets are lost

The retransmission rate is “the share of sent packets that had to be resent.” A cumulative count since boot buries recent changes under the long-run total, so compute it from the increase over a fixed interval, such as 1 minute. There’s no official threshold, but a rough feel for the server-wide average is: below 0.1% is healthy, 0.1–1% means some players hitch now and then, above 1% is noticeable to many players, and above 3% is severe. Games with many mobile or overseas players run a higher baseline. So look at how many times higher it is than normal alongside the number itself. The average gets dragged around by a few bad connections, so breaking it down by region, ISP, server, and time of day is the shortcut to the cause. If a device terminates connections in the middle and opens new ones to the server (a proxy, some load balancers and gateways), the game server’s metrics only capture the segment between that device and the server. Look for player-side retransmissions on that device.

Where to lookWhat to look atWhat it tells you
Whole server (Linux)Increase between two nstat runs 1 minute apart: TcpRetransSegs ÷ TcpOutSegs, plus the TcpExt counters TCPTimeouts, TCPLossProbes, TCPLossProbeRecovery, TCPLostRetransmit, TCPSpuriousRTOs, TCPDSACKRecv, TCPSynRetransRetransmission rate, how many times it went as far as an RTO, how many TLPs were sent and how many of those actually repaired a loss, how many times even a resent packet was lost again, connection request retransmissions. High DSACK or Spurious counts mean “resent without real loss.” Linux OutSegs excludes retransmitted segments, so the strict ratio is RetransSegs ÷ (OutSegs + RetransSegs), but around 1% the difference is small
Per connection (Linux)retrans (currently recovering/cumulative), rto, backoff, rtt, cwnd, lost, reordering, bytes_retrans from ss -tiWhether only certain players or regions retransmit a lot, and how far the RTO has grown (backoff is how many times in a row the RTO has doubled). bytes_retrans ÷ bytes_sent is that connection’s retransmission rate
Each retransmission (Linux)The eBPF tool tcpretrans (bcc). -c counts per connection, -l includes TLPPrints one line per retransmission with the remote IP, port, and connection state. A lightweight way, without packet capture, to see which player address ranges or servers they cluster on
Server network carddropped, missed, crc from ip -s -s link; rx_missed_errors, rx_no_buffer_count, rx_crc_errors, and so on from ethtool -S (names vary by driver; mlx5 uses rx_out_of_buffer and rx_discards_phy); column 2 (dropped) and column 3 (time_squeeze) of /proc/net/softnet_statWhether the server’s network card dropped packets on arrival (ring buffer, CPU), or whether a cable or optic is bad (CRC). softnet_stat has one line per CPU, in hexadecimal. If time_squeeze keeps climbing, the core handling receive processing can’t finish its work in time
Cloud networkOn AWS ENA, bw_in_allowance_exceeded, bw_out_allowance_exceeded, pps_allowance_exceeded, conntrack_allowance_exceeded, linklocal_allowance_exceeded from ethtool -S. They aren’t in the default CloudWatch view, so collect them separately with the CloudWatch agentWhether packets were silently dropped at an instance limit. If the values are rising, you’re over the limit. Other clouds also have bandwidth and connection-count limits per VM size
Switches, routers, firewallsPort CRC and input errors, output drops, policer exceeds, session table usage, drop logsWhether packets were dropped in the data center gear. Output drops rising while 5-minute average utilization is low mean microbursts (traffic crowding in for a very short moment)
PathLoss that carries through to the final hop in mtr or pathping. You need hundreds of probes or more before loss around 1% shows up, and probing the same TCP port as the game (mtr -T -P PORT) is more accurateWhich hop the loss starts at. If only one hop in the middle shows loss and the hops after it are clean, that device is just rate-limiting its replies to the probes (ICMP). The path can differ in each direction, so also measure from the server toward the player
Packet capture (both ends)Wireshark filter tcp.analysis.retransmission, plus fast_retransmission, spurious_retransmission, duplicate_ack, lost_segment, zero_window from the same tcp.analysis. familyIf the original packet is in the sender’s capture but not in the receiver’s, it was lost in between. If the receiver has it too, it’s a spurious retransmission, or the ACK was delayed or lost on the way back. Packets dropped in the receiving server’s ring buffer also look “lost in between” in a capture, so check them together with the network card counters
Windows ServerTCPv4\Segments Retransmitted/sec ÷ Segments Sent/sec in Performance Monitor, Network Interface\Packets Received Discarded, netsh int tcp show global, pktmon (built into Windows 10 1809 and Windows Server 2019 or later)Retransmission rate trend, whether the network card dropped packets on arrival, TCP settings, where inside Windows the packets were dropped

Order of checks: when working through it with the infra team, this order is fastest.

  1. When and for whom: find when the retransmission rate started rising and whether it clusters on a specific region, ISP, server, or time of day.
  2. Real loss or not: if TCPSpuriousRTOs and DSACK rise along with it, first suspect spurious retransmissions, where late packets were mistaken for lost ones.
  3. Server receive path: if network card, softnet, or cloud limit counters rose at the same moment, the server side dropped the packets.
  4. Data center gear: check the drop and CRC counters and the session tables on switches and firewalls.
  5. Outside path: run mtr in both directions, from the affected player’s side and from the server’s side, to find the hop where loss starts.
  6. If you still can’t tell: capture packets at both ends at the same time and compare.

If a report includes the time (to the second), the player’s ISP and region, the server they were on, and the symptom name, the infra team can follow this order right away.

How to fix it

1. Stop losing packets (the real fix)

  • Wired over Wi-Fi, 5 GHz or 6 GHz, SQM and ECN on the router to reduce queue overflow
  • Keep the server from firing a whole tick’s updates at once by spreading them across the tick. Smooth out bursts from a single connection with pacing (fq, BBR, a sending rate cap)
  • Enlarge ring buffers, spread interrupts across multiple cores, check cloud limits
  • Replace cables or optics that show CRC errors, match duplex settings
  • Shapers (queue the excess and send it out slowly) over policers (drop the excess immediately), raise the allowed burst
  • Leave headroom in firewall and connection tracking (conntrack) tables and in the packets per second of middleboxes, and route both directions through the same firewall
  • Prevent MTU black holes by adjusting the MSS (the maximum data one packet carries) and allowing ICMP “packet too big” notices, with MTU probing as the last safety net. Keep NAT and LB mappings alive with heartbeats sent by the client

2. Recover faster

  • Make sure SACK and timestamps aren’t turned off in the server settings or stripped by middleboxes (without SACK, RACK-TLP doesn’t work either)
  • Use RACK-TLP (the default in recent Linux and Android). Traffic the client sends, such as your input, is recovered by the client OS, so server settings don’t affect it (Windows has had TLP and RACK on by default since Windows 10 version 1607 and Windows Server 2016, and the newer RACK that also recovers lost retransmissions since Windows Server 2022)
  • For thin streams, tcp_thin_linear_timeouts; on Linux 6.15 and later, lower the RTO cap with TCP_RTO_MAX_MS
  • Keep TCP_NODELAY on for game connections (if Nagle holds new packets back, RACK loses the later packets it relies on)
  • For server-to-server connections on the internal network, lower the per-route minimum RTO (ip route … rto_min)
  • Use TCP_USER_TIMEOUT and game heartbeats to drop dead connections quickly and reconnect

3. Make the game less sensitive to retransmission (architecture)

  • Send real-time position and combat data over UDP and resend only what’s needed (an old position isn’t worth resending). If each packet also repeats the last few inputs, the next packet fills the gap when one is lost
  • Split things that need ordering, like chat and trades, and real-time packets into separate streams (QUIC streams, separate TCP connections, and so on). Loss on one no longer blocks the other
  • If you stay on TCP, don’t let old positions pile up in the send buffer; overwrite them with the latest state (TCP_NOTSENT_LOWAT and similar). The fast-forward after a long freeze gets shorter
  • Hide brief hitches on screen with an interpolation buffer and prediction. Hiding an RTO freeze of hundreds of ms is hard

Setting names: what to enable and what is easy to confuse

Most settings that affect retransmission recovery are operating system (kernel) settings, and only a few are socket options you can turn on for game connections alone. TCP_NODELAY, often confused with them because of its name, doesn’t speed up recovery. If you leave it off (Nagle on), though, new packets get delayed even more during recovery. The table below is for Linux; Windows uses different names and supports a different set.

SettingWhereWhat it changesCaution
TCP_NODELAYSocket optionTurns off Nagle. Sends small messages right away without batching themRemoves the 40–200 ms wait that happens even without loss. The retransmission timer (RTO) itself stays the same. With Nagle on, though, new packets are also held while recovery is pending, so they wait one more round trip after recovery, and the later packets that fast retransmit and RACK rely on never go out, which makes an RTO more likely. Games usually turn it on
net.ipv4.tcp_recovery (RACK)Kernel settingTime-based loss detection. Robust against reordering, and recovers thin streams quickly tooDefault 1 (on). Added in Linux 4.4 and reached its current form around 4.18. From 6.17, RACK is the only loss detection method, so setting it to 0 has no effect. Doesn’t work on connections without SACK
net.ipv4.tcp_early_retrans (TLP)Kernel settingIf no ACK comes for a while (about 2× the round-trip time), resends the last packet once to detect tail loss (loss of the final packets) quicklyDefault 3 (on); 0 turns it off. Requires SACK. If only one packet is in flight (not yet acknowledged), it waits an extra 200 ms, which makes it about as slow as an RTO
net.ipv4.tcp_sack, tcp_dsack, tcp_timestampsKernel settingSelective ACK (SACK, reports gaps), duplicate receipt notices (DSACK), round-trip time measurement (timestamps)All on by default. Some servers still have them off from the 2019 SACK security issue. With SACK off, RACK and TLP stop working too
net.ipv4.tcp_thin_linear_timeouts / TCP_THIN_LINEAR_TIMEOUTSKernel setting / socket optionFor connections with fewer than 4 packets in flight (not yet acknowledged), the RTO isn’t doubled for the first 6 timeoutsOff by default. A socket option can turn it on for game connections only. Doesn’t shorten the first RTO
TCP_RTO_MAX_MS / net.ipv4.tcp_rto_max_msSocket option / kernel setting (Linux 6.15 and later)Lowers the cap on the doubling RTO (default 120 s). Minimum 1 sKeeps the RTO from growing to tens of seconds after repeated losses. Dead connections get detected sooner too
net.ipv4.tcp_mtu_probingKernel settingIf large packets keep disappearing, shrinks the packet size to get through an MTU black holeDefault 0 (off). 1 = shrinks only when retransmissions have gone on for about 3 s and a black hole is suspected (frozen until then). 2 = starts at 1,024 bytes from the beginning and gradually probes larger sizes
ip route … rto_minRoute settingLowers the minimum RTO for that route (default 200 ms)Only for internal networks between servers. Lowering it on internet paths increases spurious retransmissions. In Linux 6.11 and later, net.ipv4.tcp_rto_min_us is a server-wide value, so it changes internet connections too. From 6.15, the socket option TCP_RTO_MIN_US can lower it for internal connections only
TCP_USER_TIMEOUTSocket optionHow long to keep retransmitting before giving up on the connectionDoesn’t make recovery faster. Drops dead connections quickly so the client can reconnect. If unset, Linux keeps retransmitting about 15 times, roughly 15 minutes, before dropping the connection (tcp_retries2)
SO_KEEPALIVE + TCP_KEEPIDLE and related optionsSocket optionChecks whether an idle connection is still aliveUnrelated to retransmission. Keeps NAT/LB mappings alive and detects dead connections
fq qdisc + SO_MAX_PACING_RATE, BBRQdisc setting / socket option / kernel settingSpreads packets out evenly to cut loss caused by bursts (sending a lot at once)“Prevents” loss. Unrelated to recovery speed

Root causes of TCP retransmission

Wireless link loss Wi-Fi / cellular link loss

Wi-Fi and mobile networks retransmit a few times on the wireless link and drop the packet if that still fails. TCP resends the dropped packet only much later.

Why: A weak signal or heavy interference makes wireless transmissions fail several times in a row → Effect: Once the wireless device’s retry limit (usually a few to ten-odd attempts) is exceeded, the packet is dropped → On screen: Freeze for as long as the TCP retransmission wait, while later packets sit in the receive buffer and then fast-forward

Symptoms: Freeze, Fast-forward, Teleporting · Primary owner External (External) · Also Infra team (Server infrastructure), Game team (Server development), Game team (Client development)

Bottleneck queue overflow (congestion loss) Tail drop at a congested bottleneck

When the queue at the narrowest point fills up, such as a home router, a link between ISPs, or a data center uplink, newly arriving packets are dropped.

Why: Video, downloads, and other users’ traffic fill up the bottleneck → Effect: While the queue is full, newly arriving packets are dropped one after another (tail drop). Packets that get in wait at the back of the full queue → On screen: Several packets vanish at once, causing a long freeze then fast-forward; common in the evening

Symptoms: Freeze, Fast-forward, Rubber-banding · Primary owner Infra team (Network infrastructure) · Also External (External), Game team (Client development)

Send bursts overflow shallow buffers Sender bursts overflow shallow buffers

When a server sends a whole tick of updates for thousands of players in one instant, a switch’s small buffer or a cloud instance’s short-term limit overflows in under 1 ms and some packets are dropped.

Why: At the start of each tick, the server sends everyone’s packets all at once → Effect: A switch port buffer where traffic from many servers converges (hundreds of KB to a few MB per port) or a cloud instance limit overflows for an instant (average utilization stays low) → On screen: Many players teleport or hitch at the same time; averaged metrics don’t reveal the cause

Symptoms: Teleporting, Freeze, Fast-forward · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure)

Policer drops excess traffic Traffic policing

ISP plans, cloud instance limits, and DDoS protection devices sometimes drop packets over a set rate right away, without queuing them.

Why: Momentary send volume exceeds the allowed rate or allowed burst → Effect: Packets over the limit are dropped immediately, with no queue (policing) → On screen: Each large burst loses several packets, causing a freeze then fast-forward, while the average rate looks below the limit

Symptoms: Freeze, Fast-forward, Teleporting · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)

Physical errors (bad cable, optics, connectors) Bit errors: bad cable, optics, dirty fiber

Damaged cables, dirty fiber connectors, and worn-out optics cause bit errors, and network equipment silently drops the corrupted packets.

Why: A bad cable, optic, or connector flips bits → Effect: Equipment drops packets whose checksum (CRC) doesn’t match → On screen: Only players whose traffic takes that path keep getting short hitches followed by fast-forward, at any time of day

Symptoms: Freeze, Fast-forward, Teleporting · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), External (External)

Duplex mismatch Duplex mismatch

If one end autonegotiates while the other has speed and duplex hard-set, one side runs half duplex and loses packets to collisions whenever load picks up.

Why: Speed and duplex hard-set on only one of the two devices → Effect: One side runs full duplex and the other half duplex, causing collisions and late collisions → On screen: Fine normally, but as traffic grows, everyone going through that device freezes then fast-forwards

Symptoms: Freeze, Fast-forward · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)

Packet drops on the receiving host Receiver host drops (ring, softirq, CPU)

Packets reach the server but get dropped, because the NIC’s ring buffer (which briefly holds arriving packets) overflows or the kernel cores that handle receive processing are saturated.

Why: A surge of players, interrupts piled on one core, CPU steal on a virtual machine, or an overloaded virtual switch → Effect: Drops at the ring buffer (rx_missed_errors and similar; the name varies by driver) or at the kernel receive queue (softnet dropped) → On screen: When crowds gather, input registers late and hitches hit the whole server at once

Symptoms: Input lag, Freeze, Fast-forward, Teleporting · Primary owner Infra team (Server infrastructure)

Firewall and connection tracking drops Stateful firewall / conntrack drops

Firewalls and Linux connection tracking (conntrack, which records passing connections in a table) drop packets when the table is full or when they decide a packet doesn’t match the connection’s state.

Why: The connection tracking table is full (table full), or traffic takes a different path each way so only one direction passes through the firewall (asymmetric routing) → Effect: The firewall treats the packets as belonging to an “unknown connection” or carrying a “sequence number outside the window” and drops them → On screen: A full table blocks new connections; a path mismatch makes only players on that path disconnect after repeated retransmissions

Symptoms: Freeze, Disconnect, Can’t connect / infinite loading · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development), Game team (Client development)

Middlebox over capacity (firewall, IPS, DDoS protection) Inline appliance PPS / CPU overload

Firewalls, intrusion prevention systems (IPS), and DDoS protection devices inspect every packet passing through. The moment traffic exceeds their inspection capacity, they drop the packets they can’t process.

Why: Hundreds of thousands or more small game packets per second at peak hours or events, or heavy inspection rules → Effect: The device maxes out its CPU or packets-per-second limit and drops packets. False positives block legitimate packets too → On screen: Freezes and teleporting hit every server behind that device at once, getting worse only when crowds gather

Symptoms: Freeze, Fast-forward, Teleporting, Disconnect · Primary owner Infra team (Network infrastructure) · Also Game team (Server development)

MTU black hole (only large packets keep getting lost) PMTU black hole

If the largest packet size a link along the way can carry shrinks and the “too big” notice (ICMP) is blocked, large packets keep vanishing no matter how many times they’re resent.

Why: The maximum size shrinks on a VPN or tunnel segment, and a firewall blocks the “too big” notices → Effect: The sender, with no idea why, keeps retransmitting the same large packet, and the RTO doubles each time → On screen: Fine normally, but when large data moves (inventory, crowded areas, loading into a zone), everything stops, including the small packets behind it, ending in a disconnect or infinite loading

Symptoms: Freeze, Disconnect, Can’t connect / infinite loading · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)

NAT or load balancer mapping expires mid-connection NAT / load balancer mapping expired mid-connection

If a device along the way deletes the mapping for an idle connection (the entry that records where to forward that connection), the next packet sent can’t be delivered. The connection either retransmits over and over until it disconnects, or the device sends back a reset (RST) and it disconnects right away.

Why: A connection with no packets going either way for a while (AFK, lobby) → Effect: A home router’s NAT, the ISP’s CGNAT, a firewall, a load balancer, or a cloud security group deletes the idle mapping → On screen: When the player moves again, retransmissions go on until a disconnect, or the disconnect is immediate

Symptoms: Disconnect, Freeze · Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Network infrastructure), Infra team (Server infrastructure)

Route change / bad ECMP path Route change / bad ECMP member

Packets vanish for a few seconds while an internet route changes, or steadily on connections assigned to a faulty path among several ECMP paths.

Why: BGP route recalculation, or faulty equipment or a bad link on one of several paths (ECMP, LAG) → Effect: Temporary loss during the route switch, or steady loss only on connections using that path → On screen: A sudden freeze of a few seconds then fast-forward, or “it gets better after reconnecting” (assigned to a different path)

Symptoms: Freeze, Fast-forward, Teleporting · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)

Spurious retransmission from latency spikes Spurious RTO from delay spikes

A packet that isn’t lost, just very late for a moment, still gets retransmitted if the delay is longer than the RTO, because the sender treats it as lost.

Why: Bufferbloat, Wi-Fi power saving, mobile radio state changes, or a virtual machine pause cause momentary delays of hundreds of ms → Effect: The RTO expires first and the packet is retransmitted; the original arrives soon after (the receiver gets a duplicate) → On screen: The freeze and fast-forward come from the latency spike itself. The spurious retransmission barely lengthens the freeze; it only pushes up retransmission metrics, which get mistaken for loss

Symptoms: Freeze, Fast-forward, Input lag · Primary owner External (External) · Also Infra team (Server infrastructure), Game team (Client development)

Spurious fast retransmit from reordering Reordering triggers spurious fast retransmit

When packets get out of order crossing multiple paths or bundled links, the receiver signals “a packet is missing” with duplicate ACKs, and the sender resends a packet that arrived fine.

Why: Devices that split traffic across paths per packet, LAGs (link bundles) that spread traffic per packet, and route changes shuffle packet order → Effect: Later packets arrive first and three duplicate ACKs pile up → fast retransmit → On screen: Sparse game packets are barely affected. Large updates in crowded areas and patch downloads slow down, with occasional stutter

Symptoms: Stutter, Input lag · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)

Late or lost ACKs (saturated upload) ACK path congestion on asymmetric links

Data arrives fine, but if the “got it” ACK is delayed or dropped in a full upload queue, the sender treats the data as lost and retransmits.

Why: Video uploads or cloud backups at home saturate the upload → Effect: ACKs sit in the router’s queue for hundreds of ms or get dropped when it overflows → On screen: Game packets from the server mostly arrive on time. Your inputs, stuck in the same upload queue, go out late, causing input lag and rubber-banding, with occasional spurious retransmissions

Symptoms: Input lag, Rubber-banding · Primary owner External (External) · Also Game team (Client development)

RTO settings that don’t fit the environment RTO min too low or too high

Set the RTO minimum too low and even small delays cause spurious retransmissions; leave the default (200 ms) and it’s too long for games, so every loss means a long freeze.

Why: RTO minimum lowered sharply for data center use, or the default left as is on internet paths → Effect: Too low: retransmission storms on momentary delays. Too high: a long wait on every loss → On screen: At the default, each loss means a freeze of hundreds of ms then fast-forward; set too low, freezes get shorter, but spurious retransmissions surge and waste bandwidth

Symptoms: Freeze, Fast-forward, Input lag · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

Slow recovery on thin streams Thin streams fall back to RTO

When a game sends small packets sparsely, the RTO fires before “three following packets” can pile up. The same loss causes a much longer freeze than it would on a bulk transfer.

Why: Packets go out about 100 ms apart, so only a few packets are ever in flight (not yet ACKed) → Effect: Collecting three duplicate ACKs takes over 300 ms, so the RTO (ping + 200 ms) fires first, doubling on consecutive losses → On screen: Each loss freezes the game for about 0.3 seconds; if the retransmission is lost too, the freeze lasts close to 1 second, then fast-forward

Symptoms: Freeze, Fast-forward · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Game team (Client development)

Middlebox strips TCP options Middlebox strips TCP options

When some firewalls or accelerators remove or rewrite TCP options, multiple losses get recovered only one per round trip, or the window (how much can be sent at once) shrinks, and everything slows down.

Why: A firewall’s “TCP normalization” or an old accelerator strips the SACK, timestamp, and window scale options → Effect: With several packets lost, recovery goes one packet per round trip, and the window is capped at 64 KB → On screen: Every loss causes a much longer freeze (without SACK, RACK-TLP can’t be used either), then fast-forward when it clears. Bulk transfers such as patches are slow too

Symptoms: Freeze, Fast-forward · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)

Zero window (a stall that looks like retransmission) Zero window, often mistaken for retransmission

When the receiving program doesn’t read its socket in time and the buffer fills up, the sender stops sending and sends only zero window probes. The network itself is fine.

Why: A client frame freeze or a blocked server thread keeps the socket from being read → Effect: The receive window drops to 0, so the sender stops sending and sends only probes (at growing intervals) → On screen: Freeze then fast-forward. A packet capture shows “ZeroWindow” and no loss

Symptoms: Freeze, Fast-forward · Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Server infrastructure)

Connection request (SYN) retransmission SYN retransmission on connect

If a connection request is lost because the connection queue (backlog) overflows or a firewall blocks it, the client OS resends it starting 1 second later, at growing intervals.

Why: A connection surge right after maintenance overflows the server’s connection queue, or a firewall or DDoS protection drops the SYN → Effect: The client OS retransmits the SYN starting 1 second later, at set intervals (older Linux: 1 s → 2 s → 4 s) → On screen: After pressing Connect, delays come in whole seconds, such as 1 or 3 seconds; if it keeps failing: can’t connect / infinite loading

Symptoms: Can’t connect / infinite loading · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure), Game team (Client development)

07Ownership

Responsibilities of the game team and the infra team

The same lag can be fixed in different places. Client and server code and netcode design belong to the game team; connections, network equipment, server hardware, and DB hosts belong to the infra team. Neither team can directly fix problems in a player’s PC and home network, the ISP’s segment, or the cloud provider, so for those they guide players, make requests, or work around them. Every cause card shows its primary owner and who else is involved, and expanding a card’s “Ballpark numbers,” “How to confirm,” and “Actions by team” section shows the action items split by team.

  1. Report / alertSymptom, time to the second, server/channel
  2. Who’s affectedOne player or household / specific ISP or region / specific server or channel / everyone
  3. Who to call firstPrimary owner of the candidate cause cards. “Start with” in the triage helper
  4. What to hand offIP and ISP, disconnect reason, related graphs, recent changes
  5. Shared workSplit it up using each card’s action items by team
The flow from an incoming ticket to action items split by team. “Who’s affected” matters most in deciding the owner. The detailed criteria are in the table below and in Diagnosing from monitoring data.
OwnerScopeTypical fixes
Game teamClientGame client code: frames, GC, loading, interpolation, extrapolation, prediction, and the client’s network handling (including sending heartbeats and auto-reconnect)Code changes, interpolation buffer and prediction tuning, changes to how loading works, heartbeat intervals and reconnect flow, client patches
Game teamServerGame server code: ticks, threads, locks, netcode design, connection handling (accept loop, listen arguments), heartbeat replies and dead-connection cleanup, socket options, query and transaction designLogic optimization, asynchronous calls, spreading load across ticks and regions, login queues, resuming sessions with session tokens, socket options (TCP_NODELAY and others), query and index design, server patches
Infra teamNetworkCircuits and data center network gear (switches, routers, firewalls, load balancers, DDoS protection); cloud network ACLs, VPC routing, and load balancers; ISPs and peeringEquipment configuration and replacement, adding link and peering capacity, route changes, escalation to ISPs, tuning idle timeouts and session limits on load balancers and firewalls
Infra teamServers/OSServer hardware and cloud instances (including security groups and connection tracking), OS and kernel settings, NICs, deployment and monitoring environmentsAdding capacity and changing instance types, kernel settings (sysctl: somaxconn, conntrack, and so on), security group setup and connection tracking timeouts, NIC ring buffers and interrupt distribution, rescheduling cron jobs and backups
Infra teamDB hostsDB servers and storage; DB configuration, replication, and backups; cache serversAdding DB capacity, securing storage IOPS, DB parameters and replication settings, tuning backups and checkpoints
ExternalPlayers/ISPs/cloudPlayers’ PCs and home networks, ISP segments (outside our contracts), cloud providersPlayer guidance (such as using a wired connection), requests to ISPs and cloud providers, workarounds and mitigations on the game side

Owners by layer and topic at a glance

The bold number is how many causes that owner is the primary owner of, and the +number is how many more causes they’re also involved in. Click a cell to see those causes and the team’s action items below.

When the boundary is unclear: the team where the cause lives is the primary owner, other teams mitigate and verify

The primary owner is where the root cause lives, or the team that can remove it. Even when a connection or a device is the cause, the game team holds the line in the meantime with designs that reduce the impact (interpolation buffers, sending inputs redundantly, reconnecting). When server code is the cause, the infra team adding hardware only postpones the problem. We settled the most commonly confused boundaries like this.

  • Disconnects after sitting idle: we can’t change the idle timeouts of players’ routers or ISP equipment, and only packets going out from the inside reliably keep those mappings alive. So the client sends heartbeats and reconnects automatically when dropped, and the server answers heartbeats, cleans up the connection first when they stop arriving, and resumes the session with a session token. The infra team shares the timeout values of our own equipment and raises them if needed.
  • Connection queue (listen backlog) overflow: the real limit is the listen argument and the accept loop in the server code, so the primary owner is Server development, while Servers/OS handles the kernel cap (somaxconn) and SYN cookies.
  • Cloud: security groups and instance connection tracking belong to Servers/OS; network ACLs, VPC routing, and cloud load balancers belong to Network.

“First” in the table is who to call first, chosen by counting the primary owners of the cause cards that match that situation. When two are listed, the first is the one that is primary owner on the most cards, and the second should be pulled in from the start as well.

SituationGame team action itemsInfra team action itemsMetrics to check first
High loss and jitter on a specific ISP or region
FirstInfra teamNetwork
Adaptive interpolation buffer, sending inputs redundantly, loss-tolerant UDP transport, pulling affected players’ IP, port, and time from per-connection loss and retransmission stats, relaxing movement validation to match connection qualityMeasure the path in both directions with the same protocol and port as the game (mtr), route around bad paths, escalate to the ISP, add peering and linksLoss rate and jitter distribution by ISP, retransmission rate
TCP retransmissions rising
FirstInfra teamNetworkGame teamServer
TCP_NODELAY, spreading a tick’s sends across the tick, reading sockets promptly (to prevent zero windows), keeping mappings alive with heartbeats, moving real-time packets to UDP or a separate connection, not piling old positions into the send buffer (TCP_NOTSENT_LOWAT)Remove loss points (cables, optics, duplex, policers, firewall connection tracking, MTU), adjust MSS, server ring buffers and interrupt distribution, kernel recovery settings (RACK, tcp_mtu_probing)Retransmission increase (nstat), zero window count, drop and CRC counters on NICs and switch ports
Tick overruns from server CPU saturation
FirstGame teamServer
Optimize AOI (visibility) calculation and broadcasts, split the tick across threads, split crowded zones and channels, match the worker thread count to the CPU limit, record tick time as a metricCPUs and instances with high single-core performance (clock speed), per-core CPU utilization alerts, check CPU steal and container CPU throttling, keep interrupt-handling cores separate from tick-thread coresTick time, per-core CPU utilization, steal, throttling count (nr_throttled)
Slow DB responses
FirstGame teamServerInfra teamDB hosts
Query, index, and transaction design (keep transactions short, lock in a consistent order), asynchronous calls off the game thread, batched lookups and caching, tuning connection pool size and wait timeoutFind slow queries, query plans, and lock waits and share them with the game team; checkpoint, replication, and statistics update settings; storage IOPS; confirm that server count × pool size stays within the max connections; add DB capacitySlow query log, lock waits, connection waits, replication lag, IOPS
Can’t connect right after maintenance
FirstGame teamServer
Keep the thread that accepts connections (accept loop) from getting stuck on other work, raise the listen backlog argument, a login queue system, batch login queries (remove N+1), have clients retry at growing, randomized intervalsKernel somaxconn and SYN cookies, firewall and load balancer session limits, server conntrack and file descriptor limits, DB cache warm-up, scaling servers out ahead of eventsListenOverflows, session table and conntrack usage, login query count and connection waits
Disconnects after sitting idle
FirstGame teamClient
Client: send heartbeats at no more than half the shortest idle timeout (so the next one gets there before the timeout even if one is late or lost), and reconnect automatically when dropped. Server: answer heartbeats, clean up the connection first if none arrive for a set time, and resume with a session token.Collect the idle timeouts of load balancers and firewalls on the path (Network) and the connection tracking timeouts of cloud security groups (Servers/OS), share them with the game team, and raise them on our own equipment if needed. Timeouts on players’ routers and ISP CGNAT can’t be changedDistribution of idle time on dropped connections (if it clusters near one value, a device with that timeout is the culprit), network type (mobile, wired)
Server freezes at set times
FirstGame teamServerInfra teamServers/OS
Randomize the timing of on-the-hour events, saves, timers, and cache expiry; break batch queries into small pieces; explicitly choose a GC with short pausesStagger cron, backup, and log compression times and lower their I/O priority; run DB backups on a replica and spread checkpoints evenly; check disk burst credits; throttle backup transfer speedFreeze times checked against job schedules (cron, backups, batches, checkpoints), GC logs
DDoS and traffic surges
FirstInfra teamNetwork
Share game traffic patterns (ports, packet sizes, packets per second) with the infra team, rate-limit requests per account and character, block malformed packets earlyDDoS protection (scrubbing) with rules tuned to game traffic, hiding server addresses, packet-per-second limits on equipment, IP-based limits that account for ISP shared IPs and PC bangs (internet cafés)Packets per second, equipment CPU and drops, connection failure rate by region and ISP (to catch false positives)
Player Wi-Fi or PC problems
FirstExternalPlayers/ISPs/cloudGame teamClient
In-game network status display (ping, loss), automatic interpolation buffer sizing based on jitter, recording network type and PC CPU usage in the logs taken when lag happens, guidance text such as recommending a wired connectionCan’t be fixed directly. If reports cluster on one ISP or region, reclassify it as a connection problemConnection and device details from reports, share of reports from the same ISP or region

What to include in a handoff

Game team → infra team

  • Exact time (to the second, with time zone) and duration, and whether it’s still happening
  • Server and channel IDs, scope (just me, specific ISP, whole server), and number of affected players (relative to concurrent users)
  • Symptom name and shape: for disconnects, the idle time before the drop; for freezes, their length and how often they repeat
  • Affected players’ IP, port, ISP, and region (if one of several paths is bad, you need the port to tell which one), the protocol the game uses (TCP, UDP), and the server port
  • Game-side metrics: tick time, ping and loss distribution, number of connections with rising retransmissions, disconnect reasons (heartbeat timeout, connection reset (RST), and so on)
  • Current heartbeat interval, the server’s no-response timeout, retry behavior
  • Any recent deploys or configuration changes, and causes already checked and ruled out

Infra team → game team

  • Equipment and link metrics for the same time window (utilization, drop and error counters, session count) and server OS metrics (ListenOverflows, conntrack usage, CPU steal)
  • Timeout and limit values of devices on the path: load balancer and firewall idle timeouts, security group connection tracking timeouts, session count and packet-per-second limits
  • Equipment and link change history and scheduled work (replacements, configuration changes, backups and cron jobs, ISP maintenance notices)
  • Ticket numbers filed with ISPs or cloud providers and when to expect an answer
  • Temporary measures (workarounds, relaxed limits) and when they’ll be rolled back
  • Where the cause was, the conclusion, and the plan to prevent a recurrence
  • Actions needed on the game side (heartbeat interval, retry behavior, connection count limits, and so on)

Both teams: name one incident lead, keep a chronological log in a single channel, and announce the next update time in advance. If primary ownership moves to another team, hand the log over too so nobody repeats the same checks. When it’s over, use the same log to correct the owner and action items on the relevant cause card.

Client game process

The game program itself, running on the player’s PC or phone. Even with a perfect network, the screen stutters if frames run late here. This is also where it’s decided how well bad network conditions get hidden.

A game repeats the same work about 60 times a second: read input, process the packets that arrived, advance the game state one step, and draw the screen. One pass of this loop is a frame, and at 60 FPS each frame gets 16.7 ms (33.3 ms for phone games running at 30 FPS). When a frame runs late, the screen stops for that long, then jumps ahead by the backlog in the next frame.

On the network side, the client’s job is “filling in missing information.” Other players’ positions arrive from the server at intervals, so the client has to draw the motion in between (interpolation). When packets stop, it has to guess and keep things moving (extrapolation). And it moves your own character right away without waiting for the server to confirm (prediction). The ways these techniques fail are exactly teleporting, rubber-banding, and stutter.

Analogy

The game client is the control room of a live broadcast. When photos arrive from the field (the server) at intervals, it stitches them together smoothly so they look like video. If a photo arrives late, there’s nothing to stitch and the picture freezes, and if the control room itself gets too busy, the broadcast cuts out too.

Causes of lag at this layer

Frame time spike Frame hitch

One frame takes several times longer than usual to compute, so the screen freezes for a moment.

Why: A burst of skill effects, a mass spawn, or a full UI refresh all land in one frame → Effect: The frame can’t finish within 16.7 ms and takes 50–300 ms → On screen: The screen hitches, then everyone moves at once on the next frame

Symptoms: Stutter, Freeze · Primary owner Game team (Client development)

Client garbage collection Client GC (Unity C#, Unreal, Lua)

The whole game freezes while it reclaims memory that was used and thrown away (garbage). The telltale sign is stutter at regular intervals.

Why: Temporary strings, arrays, and lists are created and thrown away every frame → Effect: Once garbage piles up, GC pauses the main thread to reclaim it → On screen: Regular stutter every few seconds to tens of seconds

Symptoms: Stutter, Freeze · Primary owner Game team (Client development)

Synchronous loading and shader compilation on the main thread Synchronous asset load, shader compile

The game freezes to read files and build shaders right before it draws an area, monster, or effect for the first time.

Why: Entering a new area, or a skill, piece of gear, or monster appearing for the first time → Effect: The main thread waits for file reads and shader compilation → On screen: A 0.1–1 s freeze the first time only; fine from the second time on

Symptoms: Freeze, Stutter · Primary owner Game team (Client development)

Slow storage delays asset streaming Slow storage stalls asset streaming

On slow storage such as an HDD, reading open-world textures and models can’t keep up with movement, so objects appear late or the game stutters while it waits for reads.

Why: Moving fast on a mount or by teleport, or entering a crowded area, suddenly calls for many new textures and models → Effect: Slow storage such as an HDD can’t read at the needed speed, so read requests pile up, and some loads make the main thread wait until they finish → On screen: Textures stay blurry for a while, buildings and characters pop in late, and the game stutters or freezes while it waits on reads

Symptoms: Invisible / ghost entities, Stutter, Freeze · Primary owner Game team (Client development) · Also External (External)

Rendering load from large crowds Render/animation cost of crowds

When hundreds of players fill one screen, as in a siege or a world boss fight, the cost of drawing them is more than the device can handle.

Why: Hundreds of players and effects overlap on one screen → Effect: Animation, shadow, name tag, and effect costs grow with the player count → On screen: FPS drops 60 → 15: all movement stutters, and input lag sets in

Symptoms: Stutter, Input lag · Primary owner Game team (Client development)

Packet processing bottleneck on the main thread Network processing on the main thread

If the client processes only a fixed amount of received packets per frame, a flood of packets keeps getting pushed to the next frame.

Why: Thousands of updates per second arrive in crowded places → Effect: The main thread hits its per-frame processing limit and can’t read them all → On screen: Other players’ movements show up later and later, then all at once

Symptoms: Fast-forward, Input lag · Primary owner Game team (Client development) · Also Game team (Server development)

Missing or too-short interpolation buffer Missing/short interpolation buffer

If the client draws server packets the moment they arrive, jitter (variation in packet arrival times) shows up directly on screen.

Why: Received positions are drawn immediately, or the buffer is shorter than the jitter → Effect: Characters stop for as long as a packet is late, then jump when delayed packets arrive together → On screen: Other characters move in fits and starts

Symptoms: Stutter · Primary owner Game team (Client development) · Also Game team (Server development)

Excessive extrapolation (dead reckoning) Over-extrapolation / dead reckoning

While no packets arrive, the client keeps showing characters moving at their last velocity, then snaps them back when it turns out to be wrong.

Why: Packets stop arriving, so the character keeps moving in its last direction and speed → Effect: In reality, the other player stopped or changed direction → On screen: The other character runs on for a while, then snaps to its real position or passes through walls. With erratic packet arrival intervals, it keeps overshooting and getting pulled back, so it looks like it’s shaking

Symptoms: Teleporting, Stutter · Primary owner Game team (Client development)

Client-side prediction mismatch Prediction mismatch / reconciliation

Your client shows your character moving before the server confirms it, but if the server calculates something different, your character gets pulled back.

Why: The client moves the character before the server confirms (prediction) → Effect: The server calculates collisions, movement speed, or buffs differently, or never receives the command → On screen: When the confirmation arrives, your character gets pulled back

Symptoms: Rubber-banding · Primary owner Game team (Client development) · Also Game team (Server development)

Fixed-timestep catch-up spiral Fixed-timestep catch-up / spiral of death

After one stall, the game runs its backlog of calculations all at once, and that extra work puts it behind again.

Why: The game simulation runs at a fixed interval and stalls once → Effect: The backlog of steps is computed in a single frame → On screen: A chain of long frames causes spikes, or the cap kicks in and the world slows down

Symptoms: Stutter, Fast-forward, Slow motion · Primary owner Game team (Client development)

Clock sync error Clock sync error

If the client’s estimate of the server time is wrong, interpolation timing and cooldown checks drift out of step with the server.

Why: The client syncs to server time only once when connecting and never adjusts as ping changes → Effect: The point to interpolate to and the time a cooldown ends drift away from the server’s → On screen: Opponents occasionally hitch; skills get rejected even after the cooldown has ended

Symptoms: Stutter, Dropped action / rollback · Primary owner Game team (Client development)

Float time precision loss Float time precision loss on long sessions

If the game keeps its clock in a low-precision floating-point type (float), the longer it runs, the worse its time resolution (the smallest time difference it can tell apart) becomes, and movement and effects start to shake.

Why: Time elapsed since launch is accumulated in a float or passed to shaders as is → Effect: The longer the game runs, the larger the smallest difference a float can represent → On screen: Only clients left running for days see characters, animations, and scrolling effects shake; a restart fixes it

Symptoms: Stutter · Primary owner Game team (Client development)

V-Sync and the render queue V-Sync, render queue

Input is delayed while several finished frames wait in a queue to be sent out in step with the monitor’s refresh.

Why: The graphics driver queues 1–3 frames ahead → Effect: Input takes that much longer to show up on screen → On screen: Ping is low, but controls feel heavy and sluggish

Symptoms: Input lag, Stutter · Primary owner Game team (Client development) · Also External (External)

Client memory leak Client memory leak

The longer the game stays open, the more memory it uses; it gets slower and slower until the game is eventually force-closed.

Why: Textures, UI, and effects aren’t released when moving between areas → Effect: GC runs more often, and the OS runs short of memory and starts swapping → On screen: After hours of play it stutters more and more, then gets force-closed (looks like a disconnect to the player)

Symptoms: Stutter, Disconnect · Primary owner Game team (Client development)

Client crash Client crash

An unhandled error closes the game. To the player it looks like a disconnect, but the server is fine.

Why: Null reference, out of memory, graphics driver error → Effect: The game process is forcibly terminated → On screen: Reports of “I got kicked out” while everyone else is fine at the same moment

Symptoms: Disconnect · Primary owner Game team (Client development) · Also External (External)

Anti-cheat scans Anti-cheat scan and heartbeat

The anti-cheat module that runs alongside the game to block cheats scans the system periodically. If a scan is heavy, or the heartbeat (a periodic keepalive signal) to the anti-cheat server is late, the game stutters or disconnects.

Why: The anti-cheat module periodically scans game memory, running programs, and drivers → Effect: The game thread stalls during the scan, or the heartbeat doesn’t go out on time → On screen: Hitches at regular intervals; in bad cases, a disconnect with a security error message

Symptoms: Stutter, Freeze, Disconnect · Primary owner Game team (Client development) · Also Game team (Server development)

Client OS and device

Games run on Windows, Android, and iOS, sharing CPU, memory, and network with other programs. If the operating system gives the game its CPU time late, slows things down to save battery, or suspends background apps, you get lag.

The scheduler in the operating system (OS) decides whose turn it is to use the CPU and divides CPU time among programs. The game, antivirus, browser, and updaters all wait for “their turn,” and the OS hands out cores in turns of a few to tens of milliseconds (time slices). The OS gives the game in front (the foreground app) a little more priority, but when there’s more work than cores, the game has to wait too, and that wait delays frames.

The network goes through the OS as well. Packets received by the network card or Wi-Fi chip sit in the driver’s and the OS’s receive buffer until the game picks them up. If the game is busy and picks them up late, the buffer overflows, and if it picks them all up at once, you get fast-forward. On mobile, a key point is that the OS puts the wireless connection into a power-saving state to save battery and frequently suspends the app itself.

Analogy

The OS is a head chef who makes several cooks share one kitchen. Even when the game is cooking something urgent, it has to wait if a cook called “antivirus scan” takes over the burner. And when the kitchen gets too hot (heat), the chef turns the flames down.

Causes of lag at this layer

Background processes taking up CPU Background CPU contention

When an antivirus scan, Windows Update, streaming software, or a browser video takes over CPU cores, the game thread has to wait for CPU time.

Why: Other programs hold CPU cores for a long time → Effect: The game thread waits for CPU time → On screen: Frames come late, and received packets are processed late too

Symptoms: Stutter, Fast-forward · Primary owner External (External) · Also Game team (Client development)

Power saving and thermal throttling Power saving, thermal throttling

Laptop battery mode, phone power-saving mode, and device heat slow down the CPU and GPU. With heat, the telltale sign is that the game runs fine at first and slows down only after a while.

Why: Battery or power-saving mode is on, or the device gets hot → Effect: CPU and GPU clocks drop by 30–50%, depending on the device → On screen: FPS drops and the game stutters, right from the start with power saving, or after a few minutes to about 20 minutes of play with heat

Symptoms: Stutter, Input lag · Primary owner External (External) · Also Game team (Client development)

Timer resolution Timer resolution (Windows 15.6ms)

Windows’ default timer ticks every 15.6 ms, so “sleep for just 1 ms” actually lasts until the next timer tick, up to 15.6 ms.

Why: Frame limiting and packet sending are implemented with Sleep (a short wait) → Effect: The OS wakes the thread only every 15.6 ms → On screen: Frame intervals and input send intervals become uneven

Symptoms: Stutter · Primary owner Game team (Client development)

Mobile app sent to the background App suspended in background

If you leave the game for a moment to check a notification, the OS suspends the app a few seconds later, and meanwhile the server disconnects you.

Why: The player leaves the game to read a message or take a call → Effect: The game engine pauses gameplay, and the OS soon suspends the app and its networking → On screen: Already disconnected on return, so the game reconnects

Symptoms: Disconnect · Primary owner Game team (Client development) · Also Game team (Server development)

Wi-Fi ↔ LTE/5G switching Network switch changes IP

When you walk out of the house and your phone drops Wi-Fi for LTE or 5G, your IP address changes and the existing connection stops working.

Why: The Wi-Fi signal weakens and the phone switches to the mobile network → Effect: Your IP address changes, so the connection made from the old address can’t carry any more data → On screen: A brief freeze, then a disconnect or a reconnect

Symptoms: Freeze, Disconnect · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Network infrastructure)

Packet inspection by security software Antivirus / firewall inspection

When antivirus software or a firewall inspects every packet, latency goes up, and in bad cases it mistakes the game for an attack and blocks it.

Why: Security software inspects every packet sent and received, one by one → Effect: Each packet picks up delay, and packets get dropped when inspection falls behind → On screen: Ping spikes irregularly, or connections get blocked

Symptoms: Stutter, Can’t connect / infinite loading · Primary owner External (External) · Also Game team (Client development)

Receive buffer overflow Socket receive buffer overflow

If the game is busy and pulls packets out of the socket (the network send/receive interface the OS provides) late, the OS buffer overflows.

Why: Frames fall behind and the game reads the socket late → Effect: The OS receive buffer fills up: UDP packets get dropped, and TCP shrinks the receive window so the sender stops sending → On screen: Teleporting (UDP) or fast-forward (TCP)

Symptoms: Teleporting, Fast-forward · Primary owner Game team (Client development)

Client low on memory and swapping Paging / swap on client

With dozens of browser tabs open alongside the game, the OS moves part of the game’s memory out to disk.

Why: The PC runs low on RAM overall → Effect: The OS moves game memory that isn’t in use right now to disk → On screen: The moment that memory is used again, the game freezes for tens to hundreds of ms, depending on storage

Symptoms: Freeze, Stutter · Primary owner External (External) · Also Game team (Client development)

Out of graphics memory (VRAM) VRAM over-commit

When the graphics settings need more memory than the graphics card has, the OS moves textures out to system memory and brings them back, and the game stutters.

Why: High texture settings plus all the gear and effects in a crowded place fill up graphics card memory → Effect: The OS moves textures that aren’t in use right now to system memory, then brings them back over the slow PCIe bus when needed → On screen: A hitch every time a new scene or character comes into view; textures stay blurry for a while

Symptoms: Stutter, Freeze · Primary owner Game team (Client development) · Also External (External)

Wi-Fi background scanning Periodic Wi-Fi background scan

Communication pauses briefly while the OS periodically hops across channels to look for nearby Wi-Fi networks.

Why: The OS or driver searches for nearby Wi-Fi networks on a fixed schedule → Effect: Sending and receiving pause briefly during the scan → On screen: Ping spikes at exactly regular intervals (e.g., every 60 s)

Symptoms: Stutter, Teleporting · Primary owner External (External) · Also Game team (Client development)

NIC power saving and driver issues NIC power saving, driver bugs

When a network card or Wi-Fi chip enters a power-saving state between packets, it takes time to wake back up.

Why: Network device power saving is on, or the driver is outdated → Effect: Wake-up delay, occasional device restarts → On screen: Irregular delays, occasional freezes lasting several seconds

Symptoms: Stutter, Freeze · Primary owner External (External) · Also Game team (Client development)

Other apps on the same device using up bandwidth Other apps saturating the link

When cloud sync, a large download, or a game patch runs on the same PC, game packets have to wait in a queue.

Why: Another app maxes out the upload or download → Effect: Game packets pile up in the queues on the PC and the router → On screen: Ping shoots up, input lag, fast-forward

Symptoms: Input lag, Fast-forward · Primary owner External (External) · Also Game team (Client development)

Throttling when the window is minimized or unfocused Minimized / unfocused window throttling

When you switch to another window or minimize the game, the game and Windows slow it down to save power. When you come back, the backlog of packets floods in, or you’ve already been disconnected.

Why: Switching to another window with Alt+Tab, or minimizing the game → Effect: While it isn’t visible, the game lowers FPS sharply or pauses, and Windows also lowers the priority of programs that aren’t visible → On screen: Fast-forward the moment you return; a disconnect if the game stayed minimized for a long time

Symptoms: Fast-forward, Stutter, Disconnect · Primary owner Game team (Client development)

Overlay software interference Overlays and screen hooks

Chat apps, launchers, recording tools, and FPS counters hook into the game’s rendering to draw their own UI on top of the game screen (hooking). That adds work to every frame and sometimes clashes with the game, causing hitches or crashes.

Why: Overlays from chat apps, game launchers, graphics card tools, or recording software are turned on → Effect: Every time a frame goes out to the screen, the overlay steps in and draws its own UI on top → On screen: Frames get slightly later, and when a notification pops up the game hitches, shows graphics glitches, or crashes (looks like a disconnect to the player)

Symptoms: Stutter, Freeze, Disconnect · Primary owner External (External) · Also Game team (Client development)

Display, input device, and frame generation latency Display, input device and frame generation latency

If ping is normal but controls feel heavy, a TV’s video processing, a wireless controller, or frame generation may be adding delay between your input and the screen.

Why: The TV’s game mode is off, a Bluetooth or wireless controller is in use, or frame generation (DLSS or FSR frame generation) is on → Effect: The TV sends frames out late while it processes the picture, wireless input arrives late by its polling interval plus any interference, and frame generation waits for the next frame to create an in-between frame → On screen: Ping and FPS numbers look good, but there’s a delay between pressing a button and seeing the result on screen: input lag

Symptoms: Input lag · Primary owner External (External) · Also Game team (Client development)

Home network: Wi-Fi, router, mobile network

The last few meters before a packet leaves the house. The distance is short, but a large share of lag reports start here. Wi-Fi splits the same wireless channel among many devices, and the router sends the whole household’s traffic out through a single queue.

Wi-Fi shares the same wireless channel (frequency band) with the neighbors’ routers, and the 2.4 GHz band also overlaps with Bluetooth and microwave ovens. When transmissions collide, a device waits briefly and sends again, and as these retransmissions pile up, packets arrive unevenly. Average ping can look fine while it spikes from moment to moment, which is the typical Wi-Fi pattern.

The router is the device every device in the house has to pass through to reach the internet. When you send more than the internet connection can take, a queue builds up inside the router or modem, and devices without queue management (SQM) let that queue grow to hundreds of milliseconds’ worth. Even an expensive router does the same if the feature is turned off. The moment a sibling starts uploading a video, game packets end up waiting at the back of that queue. This is called bufferbloat.

The router also records each “device inside ↔ server outside” connection in its NAT table, and deletes the entry if no packets pass for a while. This is a common cause of disconnects after sitting idle. Mobile networks add cell tower handovers, radio power-saving states, and weak signals on top of that.

Analogy

The router is the only gate of an apartment complex. If moving trucks (video uploads) are lined up, even an urgent motorcycle courier (game packets) has to wait behind them. A smart router (SQM) opens a separate lane just for couriers.

Causes of lag at this layer

Wi-Fi interference and weak signal Wi-Fi interference, weak signal

With a weak signal or interference, packets get resent several times over the wireless link, so they arrive unevenly.

Why: Walls, distance, microwaves, Bluetooth, and neighbors’ routers degrade the radio signal → Effect: Transmissions fail on the wireless link → resent several times → On screen: Packets arrive unevenly (jitter), so characters move in fits and starts; with heavy loss, they teleport

Symptoms: Stutter, Teleporting, Rubber-banding · Primary owner External (External) · Also Game team (Client development)

Congested Wi-Fi channel Crowded Wi-Fi channel

Where there are dozens of routers, as in an apartment building, they share the same channel and have to wait for a chance to transmit.

Why: Dozens of routers use the same 2.4 GHz channel → Effect: Before transmitting, a device waits until other devices finish and the channel clears → On screen: In the evening, when people get home, jitter (variation in packet arrival times) rises and the game stutters

Symptoms: Stutter, Input lag · Primary owner External (External) · Also Game team (Client development)

Bufferbloat (router queue) Bufferbloat

When someone in the household uploads a video or downloads a large file, hundreds of ms worth of packets pile up in the router’s queue, and game packets wait behind them.

Why: The connection fills up with a family member’s video upload or cloud backup, your own live stream, or a large download → Effect: The router or modem holds the overflow of packets in a large queue → On screen: Game packets wait at the back of the queue too, and ping shoots up to hundreds of ms

Symptoms: Input lag, Fast-forward, Teleporting · Primary owner External (External) · Also Game team (Client development)

NAT mapping expiry NAT mapping timeout

Routers remove idle connections that haven’t carried packets for a while from their NAT table. It’s a common reason the connection drops the moment you move after standing still.

Why: The router records the “inside device ↔ outside server” connection in its NAT table (address translation table) → Effect: If no packets pass for a while, the entry is deleted (for UDP, often after 30–120 s) → On screen: Server packets can no longer get into the home, so the connection drops

Symptoms: Disconnect · Primary owner Game team (Client development) · Also Game team (Server development)

Underpowered or overheating router Router CPU / session table exhaustion

When dozens of devices and thousands of connections pile onto a cheap router, the router itself can’t keep up.

Why: Dozens of devices, plus P2P and torrent clients opening thousands of connections → Effect: The router’s CPU and session table are saturated → On screen: Delayed and lost packets, failed new connections

Symptoms: Stutter, Can’t connect / infinite loading, Disconnect · Primary owner External (External)

Cell tower handover (while moving) Cellular handover

When you travel by bus or subway, the connection drops out while your phone switches cell towers.

Why: The phone switches to a different cell tower while on the move → Effect: Usually a gap of tens of ms, but if the signal is bad and the switch fails, it can drop out for hundreds of ms to several seconds → On screen: A freeze, then teleporting; if it lasts long, a disconnect

Symptoms: Freeze, Teleporting, Disconnect · Primary owner External (External) · Also Game team (Client development), Game team (Server development)

RRC state transition delay (mobile radio power saving) Radio state promotion (RRC)

When a phone has no traffic for a while, it drops its radio connection to a low-power state, and the next packet is delayed while it powers back up.

Why: After a short period with no traffic, the phone puts its radio connection into a power-saving state → Effect: To send the next packet, the connection has to be brought back up → On screen: Only the first action after standing idle is noticeably slow

Symptoms: Input lag · Primary owner Game team (Client development)

Weak mobile signal and dead zones Weak cellular signal

In elevators, basements, and deep inside buildings, retransmissions increase, speed drops, and eventually the connection drops.

Why: Moving into a place with weak signal → Effect: More radio retransmissions, lower speed, momentary dropouts → On screen: Jitter and loss cause stutter and teleporting, and eventually a disconnect

Symptoms: Stutter, Teleporting, Disconnect · Primary owner External (External) · Also Game team (Client development)

Frequent 5G↔LTE switching (at 5G coverage edges) 5G NSA / LTE switching

Inside buildings with weak 5G signal or at the edge of 5G coverage, the phone switches between 5G and LTE often, and each switch causes a ping spike or a brief dropout.

Why: In a place where the 5G signal comes and goes (inside a building, at the edge of 5G coverage) → Effect: The phone keeps switching between 5G and LTE, with a short gap each time → On screen: Ping spikes at random even when standing still, with occasional freezes and teleporting

Symptoms: Stutter, Teleporting, Freeze · Primary owner External (External) · Also Game team (Client development)

Public Wi-Fi and corporate network restrictions Captive portal, restrictive network

A café Wi-Fi login page or a corporate firewall blocks the game’s connection.

Why: The login page hasn’t been completed yet, or a firewall blocks the game’s ports or UDP → Effect: Connection attempts are blocked outright, or only some traffic gets through → On screen: Can’t connect, or login works but entering the game fails

Symptoms: Can’t connect / infinite loading · Primary owner External (External) · Also Game team (Client development), Game team (Server development)

Internet path: ISP networks and long-haul links

After leaving the house, a packet crosses the ISP network, the links between ISPs, and sometimes submarine cables to reach the data center where the server lives. Latency on this stretch is mostly set by distance and route selection (routing), and in many cases the game company can’t fix it directly.

Light travels about 200,000 km per second in optical fiber. For a server 1,000 km away, a round trip takes at least 10 ms, and as long as the signal travels over fiber, no server or equipment upgrade can shrink that number. Real packets don’t travel in a straight line; they detour through the points where ISPs connect (peering), so they usually take 1.5–2 times the theoretical value. On routes with almost no large cables along the straight line, such as Korea–Europe, traffic detours through Southeast Asia and Suez or through the US, reaching 2.5–3 times (about 230–270 ms round trip).

The problem is that this route changes with time and circumstances. Around 9–11 PM, when everyone is streaming video, the links between ISPs tend to get congested. When routing information (BGP) changes, packets can’t reach their destination for a few seconds to tens of seconds (rarely a few minutes). When a submarine cable is cut, traffic detours over a longer route for weeks. If the lag hits “only users of a specific ISP,” “only in the evening,” or “only from overseas,” suspect this layer first.

Analogy

The ISP network is a highway system. Even when the road from Seoul to Busan is clear, the distance still takes time. The toll gates (peering points) jam at rush hour, and after an accident the navigation app sends you on a long detour.

Causes of lag at this layer

Propagation delay (physical distance) Propagation delay

Even light travels only about 200,000 km per second in optical fiber. A distant server is slow no matter how good it is.

Why: The server is far away (an overseas server, another continent) → Effect: Round-trip time grows with distance (at least 10 ms per 1,000 km) → On screen: Constant input lag on every action and a disadvantage in hit registration

Symptoms: Input lag · Primary owner Infra team (Server infrastructure) · Also Infra team (Network infrastructure), Game team (Server development)

Satellite internet (LEO/GEO) Satellite internet (LEO, GEO)

Satellite signals have to travel to space and back. With geostationary satellites the round trip alone exceeds 0.5 seconds. Low Earth orbit satellites such as Starlink are usually fast, but latency fluctuates and the link can drop briefly at the moment routes are reassigned.

Why: Connecting from home, a ship, or a plane over GEO or LEO satellite internet, or over in-flight Wi-Fi that uses satellites → Effect: GEO satellites sit at about 36,000 km, so the distance itself is long. LEO systems reassign the terminal–satellite–ground station path at short intervals, with a brief burst of delay and loss at each reassignment → On screen: GEO: heavy input lag on every action. LEO: fine most of the time, then stutter and teleporting at regular intervals

Symptoms: Input lag, Stutter, Teleporting · Primary owner External (External) · Also Game team (Client development), Game team (Server development)

Detour routing Suboptimal routing

Because of interconnection agreements between ISPs, traffic to even a nearby server can take a long way around.

Why: Your ISP and the server’s ISP aren’t directly connected → Effect: Traffic passes through another country or city, adding distance and hops → On screen: Only players on certain ISPs have unusually high ping

Symptoms: Input lag · Primary owner Infra team (Network infrastructure) · Also External (External)

Peak-hour congestion at peering links Peak-hour congestion at peering

Around 9–11 PM, video traffic surges and the links between ISPs (peering) tend to get congested.

Why: Evening streaming and downloads pile up → Effect: Queues build and packets drop on peering links → On screen: Players on certain ISPs get stutter and teleporting only in the evening

Symptoms: Stutter, Teleporting, Rubber-banding · Primary owner Infra team (Network infrastructure) · Also External (External)

Submarine cable / international link outage Submarine cable fault

When a submarine cable is cut, traffic takes long detours for weeks (sometimes months) until it is repaired, and the remaining links get congested.

Why: Cable cut or equipment failure → Effect: Traffic crowds onto long detour routes and the remaining links → On screen: Ping surges and packet loss for overseas players that last days to weeks

Symptoms: Input lag, Teleporting · Primary owner External (External) · Also Infra team (Network infrastructure)

BGP route changes and convergence Route change / BGP convergence

When internet routing information changes, packets are lost for the few seconds to tens of seconds (rarely a few minutes) it takes to converge again.

Why: Routing information changes somewhere in an ISP’s network → Effect: For a few seconds to tens of seconds, packets vanish or switch to a new route → On screen: A sudden freeze of a few seconds, then ping settles at a different value (e.g., 40 → 70 ms)

Symptoms: Freeze, Teleporting · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)

One faulty ECMP path ECMP / link bundle member fault

ISPs and data centers keep several paths to the same destination and send each connection down one of them. If a single path fails, only the players assigned to it keep lagging.

Why: One link or device in a bundle of links is faulty or congested → Effect: The path is chosen from the address and port combination (hash), so only connections assigned to that path see loss and delay → On screen: Same region and ISP, but only some players keep teleporting. Reconnecting sometimes fixes it

Symptoms: Teleporting, Rubber-banding, Stutter · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), External (External)

ISP throttling and traffic management Traffic shaping, data caps

When you go over your data allowance, or on plans that manage certain kinds of traffic, packets get delayed or dropped.

Why: Speed throttled after the plan’s data runs out, or certain traffic restricted → Effect: Packets wait in a queue or get dropped → On screen: Lag after a certain amount of usage, especially on mobile

Symptoms: Input lag, Teleporting · Primary owner External (External) · Also Game team (Server development), Infra team (Network infrastructure)

Country- or ISP-level UDP restrictions and packet inspection UDP blocking, throttling and inspection by networks

Some networks block specific UDP addresses and ports or throttle UDP, and packet inspection equipment filters out protocols it doesn’t recognize. Games that communicate over UDP can’t connect on those networks or disconnect often.

Why: Connecting from an ISP network that throttles UDP, or from a network with country- or ISP-level traffic inspection (censorship) equipment → Effect: Specific UDP addresses and ports are blocked, UDP is throttled at busy hours, ports or protocols not on an allowlist are filtered out, or the first few packets get through before the flow is blocked → On screen: Only players in certain countries or on certain ISPs can’t connect or get infinite loading, disconnect soon after connecting, or teleport from packet loss at busy hours

Symptoms: Can’t connect / infinite loading, Disconnect, Teleporting · Primary owner External (External) · Also Game team (Client development), Game team (Server development), Infra team (Network infrastructure)

Poor line quality Faulty last-mile line / modem

Loose connectors, old wiring, or a faulty modem cause steady packet loss and periodic line drops.

Why: Damaged cable, poor contact, faulty modem or optical network terminal → Effect: Packets dropped from bit errors; now and then the line drops for a few seconds to about a minute while it reconnects → On screen: Steady low-level packet loss, occasional freezes of a few seconds or disconnects

Symptoms: Teleporting, Freeze, Disconnect · Primary owner External (External)

DNS failures and delays DNS failure / slowness

If DNS, which turns server names into addresses, is slow or fails, the game can’t find its login or patch servers.

Why: ISP DNS outage or misconfiguration → Effect: The login or patch server address can’t be resolved → On screen: A long wait after pressing Connect, or no connection at all. Players already connected are fine

Symptoms: Can’t connect / infinite loading · Primary owner External (External) · Also Game team (Client development)

Shared links saturated by DDoS DDoS saturating shared links

Massive attacks aimed at the game company, or at someone else on the same network, fill up shared links.

Why: A flood of attack traffic → Effect: Legitimate traffic on the same links gets delayed and dropped too → On screen: Many players teleport, disconnect, or can’t connect at the same time

Symptoms: Teleporting, Disconnect, Can’t connect / infinite loading · Primary owner Infra team (Network infrastructure) · Also External (External)

ISP-shared IP addresses (CGNAT) Carrier-grade NAT

Mobile networks and some ISPs have many subscribers share one IP address, and they delete the mappings of idle connections after a short time.

Why: ISP equipment manages the session table for huge numbers of subscribers → Effect: Session table limits, short idle timeouts → On screen: Disconnects after sitting idle, and false positives that block everyone sharing the same IP at once

Symptoms: Disconnect, Can’t connect / infinite loading · Primary owner Game team (Client development) · Also Game team (Server development), Infra team (Network infrastructure)

Routing through a VPN or game booster VPN / game accelerator detour

With a VPN or game booster on, packets go through that company’s relay servers. If the relay is far away or busy, the connection can actually get slower.

Why: The VPN or booster sends every game packet through its relay servers → Effect: Distance to the relay and its congestion add up, and tunnel headers shrink the MTU (the largest packet size that can be sent at once) → On screen: Higher ping and packet loss; can’t connect when the relay address gets blocked along with everyone else using it

Symptoms: Input lag, Teleporting, Can’t connect / infinite loading · Primary owner External (External) · Also Infra team (Network infrastructure), Game team (Server development)

Data center network equipment

Just before reaching the server, a packet passes through routers, DDoS protection gear, firewalls, load balancers, and switches in turn. This stretch normally takes less than 1 ms, but when a single device runs out of capacity or fails, thousands of players on the whole server are hit at the same time.

Each device has a different role. The router picks the route, the DDoS protection device filters out attack traffic, and the firewall lets only allowed connections through and tracks every connection in a session table. The load balancer spreads incoming connections across servers, and switches connect the servers to each other.

The common weak points of these devices are table size and buffer size. When a firewall’s session table fills up, it can’t accept new connections. Load balancers delete idle connections after a set time. A switch’s small buffers overflow in under 1 ms when several servers send packets to thousands of players at the same instant (a world boss spawn). And when a device fails and switches over to a standby unit (failover), everyone freezes for those few seconds.

Analogy

The data center entrance is airport security and the boarding gate. Security (the firewall) lets through only people on the list, and once the list is full, it can’t take anyone else. The gate agent (the load balancer) treats a passenger who has sat quietly for a long time as “gone” and strikes them from the list.

Causes of lag at this layer

Firewall session table full Firewall session table exhaustion

A firewall tracks every connection it lets through by recording it in a session table. Once the table is full, it can’t accept new connections.

Why: A connection surge or an attack pushes the session count to its limit → Effect: No free entry to record a new connection, so it’s refused → On screen: Players trying to get in can’t connect or get infinite loading, and some existing connections disconnect too

Symptoms: Can’t connect / infinite loading, Disconnect · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), Game team (Client development)

DDoS protection detours and false positives DDoS scrubbing latency, false positives

Diverting traffic to a scrubbing center to stop attacks makes the route longer, and legitimate players are sometimes mistaken for attackers and blocked.

Why: After an attack is detected (or all the time), inbound traffic is diverted to a scrubbing center → Effect: The route gets longer, and some legitimate packets are flagged as attack traffic → On screen: Ping rises for everyone; players in certain regions or on certain ISPs can’t connect

Symptoms: Input lag, Can’t connect / infinite loading, Teleporting · Primary owner Infra team (Network infrastructure) · Also Game team (Server development)

Load balancer idle timeout Load balancer idle timeout

A load balancer deletes idle connections after a set time. The game assumes the connection is still alive, and then the player gets disconnected.

Why: The player sends no packets for a while (chat window open, away from keyboard) → Effect: The load balancer cleans up the idle connection (common defaults are 60–350 seconds) → On screen: Disconnect the moment the player moves again

Symptoms: Disconnect · Primary owner Infra team (Network infrastructure) · Also Game team (Client development), Game team (Server development)

Cloud security group connection tracking expiry Cloud security group connection tracking timeout

The firewall attached to a cloud server (security group) also tracks connections, and tracking entries for idle connections expire after a set time. Even on servers that clients reach directly without a load balancer, players who sat idle can get disconnected.

Why: The security group is set up so that it tracks game connections (only certain addresses allowed, restricted outbound rules, traffic through an NLB, and so on) → Effect: The tracking entry for a connection that sat idle for a while expires, and the security group silently drops packets that arrive after that → On screen: After being away, the player moves again, gets no response, then disconnects. The server program doesn’t notice for a long time

Symptoms: Disconnect · Primary owner Infra team (Server infrastructure) · Also Game team (Client development), Game team (Server development)

Cloud NAT gateway connection and port limits Cloud NAT gateway connection / port limits

When servers in a private subnet connect out (platform authentication, payments, external APIs), a NAT gateway rewrites their address and port. If concurrent connections to the same destination exceed the gateway’s port limit, new connections fail.

Why: Servers open many short connections to the same external address, such as platform authentication or payments, or keep connections open for a long time → Effect: The NAT gateway can’t allocate any more source ports for that destination, so new connections fail → On screen: The game itself is fine, but only features that call external services, such as login, payments, and reward delivery, fail or slow down (can’t connect / infinite loading, dropped action / rollback)

Symptoms: Can’t connect / infinite loading, Dropped action / rollback · Primary owner Infra team (Network infrastructure) · Also Game team (Server development)

Load balancer skew and misjudged health checks LB imbalance, bad health checks

Connections pile onto one server, or players keep getting sent to a server that’s already dead.

Why: The distribution rule is a poor fit, or the health check can’t see the real state → Effect: One server alone is overloaded, or players try to connect to a dead server → On screen: Only some channels or some players get slow motion, can’t connect, or get infinite loading

Symptoms: Slow motion, Can’t connect / infinite loading · Primary owner Infra team (Network infrastructure) · Also Game team (Server development)

Switch microbursts Switch microburst drops

When several servers send packets to thousands of players at the same instant, the small buffer on the switch port where that traffic converges overflows in less than 1 ms.

Why: A world boss spawn or massive skills, or ticks on several servers lining up so they all send at once → Effect: Buffers where several ports feed into one, or where a fast port feeds a slower one (hundreds of KB to a few MB per port), fill up in an instant → On screen: Some packets dropped; many players teleport or have skills fail to go off at the same moment

Symptoms: Teleporting, Dropped action / rollback · Primary owner Game team (Server development) · Also Infra team (Network infrastructure), Infra team (Server infrastructure)

Data center link saturation Uplink saturation

When patch distribution, log shipping, or backups share a link with the game, the link fills up.

Why: Bulk transfers take over the same link → Effect: Link queues and loss grow → On screen: Higher ping and teleporting across the whole server

Symptoms: Input lag, Teleporting · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure)

Network equipment failover Network device failover

When a router or firewall fails and traffic switches to the standby unit (failover), everyone freezes for a few seconds.

Why: Switchover to standby equipment because of a failure or maintenance → Effect: The switchover takes a few seconds, and connections reset if session state isn’t synced → On screen: Every player on the server freezes at once; mass disconnects

Symptoms: Freeze, Disconnect · Primary owner Infra team (Network infrastructure) · Also Game team (Server development), Game team (Client development)

Bad cables and port errors Bad cable / optics (CRC errors)

Bad optics or a bad cable corrupt a steady share of the packets that pass through that path.

Why: Bit errors from bad optics or cables → Effect: The equipment silently drops corrupted packets → On screen: Only some servers or players using that path teleport or rubber-band from steady packet loss

Symptoms: Teleporting, Rubber-banding · Primary owner Infra team (Network infrastructure)

MTU mismatch (only large packets vanish) MTU black hole

If the MTU (the largest size that can be sent at once) shrinks somewhere along the path and the “packet too big” messages are blocked, only large packets keep vanishing.

Why: The MTU shrinks on a tunnel or VPN segment → Effect: A firewall blocks the “packet too big” messages (ICMP), so the sender never finds out → On screen: Freezes and then disconnects only when opening large screens such as the inventory or character list

Symptoms: Freeze, Disconnect, Can’t connect / infinite loading · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Server development)

Server network card (NIC)

The network card in a server takes in hundreds of thousands to millions of packets per second and hands them to the CPU. If processing falls behind here, packets are lost before the server program even knows they arrived.

The NIC puts arriving packets one after another into a ring buffer (a receive buffer that reuses a fixed number of slots in rotation) and tells the CPU “packets are here” (an interrupt). The CPU takes packets out of the ring buffer and hands them to the OS. If packets come in faster than the CPU takes them out, every slot fills up and any packet after that is dropped. Only a number in the network card stats (ethtool -S) quietly goes up, and the game server logs show no errors at all, which makes this lag hard to find.

Modern NICs have several receive queues (ring buffers) and can spread notifications across multiple CPU cores (RSS), but if that isn’t configured or traffic skews into one queue, a single core hits 100% and becomes the bottleneck. Cloud servers also have separate packets-per-second, bandwidth, and connection-count limits in front of the NIC, and whatever exceeds them is dropped before it reaches the server. This doesn’t show up in the usual metrics like CPU or ring buffers. On AWS it appears only in the ENA driver stats (pps_allowance_exceeded and others in ethtool -S).

Analogy

The NIC is an apartment building’s mailbox, the ring buffer is the number of mailbox slots, and the interrupt is the mail carrier’s doorbell. If mail pours in and only one person is taking it out, the slots overflow and letters fall on the floor. RSS means having several people take the mail out.

Causes of lag at this layer

NIC interrupts concentrated on one core Single-queue NIC / no RSS

If the NIC sends every packet-arrival interrupt to a single CPU core, that core becomes the bottleneck.

Why: A single receive queue, or RSS (which spreads packets across cores) turned off → Effect: One core hits 100% and can’t pull packets off in time → On screen: Packet loss and latency across the whole server when players crowd in (teleporting, input lag)

Symptoms: Teleporting, Rubber-banding, Input lag · Primary owner Infra team (Server infrastructure)

Ring buffer too small RX ring buffer overflow

If the NIC’s ring buffer, which briefly holds incoming packets, is small, a sudden burst overflows it and packets get dropped.

Why: The ring buffer is left at its small default (256–2,048 slots depending on the driver) → Effect: During a burst, the buffer overflows before the CPU can pull packets off → On screen: Loss only at burst moments (teleporting, skills not going off). No trace in the game server logs

Symptoms: Teleporting, Dropped action / rollback · Primary owner Infra team (Server infrastructure)

Excessive interrupt coalescing Interrupt coalescing

When the NIC collects packets and notifies the CPU once per batch to reduce CPU load, packets arrive later by the time spent collecting.

Why: The NIC collects packets for a set time or count before raising an interrupt → Effect: Packets wait while the batch fills → On screen: A small rise in latency. Usually tiny, but ms-scale if overdone

Symptoms: Input lag · Primary owner Infra team (Server infrastructure)

Cloud PPS limit exceeded Cloud PPS / bandwidth allowance

Each cloud instance type has limits on packets per second and bandwidth, and traffic over them is silently dropped.

Why: Rising CCU pushes packets per second over the instance limit → Effect: The cloud network drops the excess → On screen: Teleporting and skills not going off from unexplained packet loss. Server CPU has headroom

Symptoms: Teleporting, Dropped action / rollback · Primary owner Infra team (Server infrastructure) · Also Game team (Server development), External (External)

NIC bandwidth saturation NIC bandwidth saturation

Running a 1 Gbps or 10 Gbps card at its limit makes the transmit queue grow until packets get dropped.

Why: More broadcasts push traffic to the card’s limit → Effect: The transmit queue grows, and packets are dropped when it overflows → On screen: Latency and loss across the whole server (input lag, teleporting)

Symptoms: Input lag, Teleporting · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Virtualization overhead and noisy neighbors Noisy neighbors in virtualization

When other virtual machines on the same physical server use a lot of network or CPU, your server’s processing gets delayed at irregular times.

Why: Other VMs on the same physical server use a lot of resources → Effect: Packet processing on your VM is delayed at irregular times → On screen: Occasional jitter (variation in packet arrival times) with no obvious cause, showing up as stutter

Symptoms: Stutter · Primary owner Infra team (Server infrastructure) · Also External (External)

Cloud host maintenance and live migration Cloud host maintenance / live migration

When a cloud provider performs maintenance on a physical server (host), it moves VMs to another host (live migration) or pauses them briefly. The whole server freezes during that time, and if the pause is long, connections drop.

Why: The provider moves the VM to another host, or pauses it briefly, for host maintenance or a predicted failure → Effect: During the move, CPU, memory, and network slow down, and at the end the VM stops completely for a moment (from under 1 second to around 30 seconds, depending on the provider and method) → On screen: Everyone on the server freezes at once and then sees fast-forward and teleporting; if the freeze outlasts the timeout, mass disconnects

Symptoms: Freeze, Fast-forward, Teleporting, Disconnect · Primary owner Infra team (Server infrastructure) · Also Game team (Server development), External (External)

NIC driver and firmware problems NIC hang / reset

When a driver bug or a malfunctioning feature hangs the card, all traffic in and out stops while it restarts.

Why: Driver bug, malfunctioning offload feature → Effect: The NIC hangs and restarts (a few seconds) → On screen: Everyone on that server freezes together and then teleports or disconnects

Symptoms: Freeze, Disconnect · Primary owner Infra team (Server infrastructure)

GRO/LRO batching delay GRO/LRO batching

GRO and LRO bundle several packets into one to reduce CPU load. Depending on settings, a small game packet may wait briefly for the next packet to bundle with.

Why: The NIC and kernel bundle arriving packets together for processing → Effect: With hardware aggregation (LRO) or a batching wait time setting turned on, packets wait briefly for the next one → On screen: A small rise in latency (usually tens of µs or less)

Symptoms: Input lag · Primary owner Infra team (Server infrastructure)

Server OS (kernel)

The server’s Linux or Windows kernel accepts connections, manages socket buffers, and hands out CPU and memory to the game server program. Most defaults are conservative values meant to suit many uses, so they often don’t fit a game server that keeps tens of thousands of players connected for long stretches.

When a new connection comes in, the kernel puts the request in the connection queue (listen backlog), and the game server takes requests from it one at a time. When the queue is full, Linux silently drops new requests, while Windows sends back a refusal. On Linux, every connection needs one file descriptor (fd), the number attached to each open file or connection, and there’s a limit on how many fds one process can hold. When tens of thousands of players press the connect button right after maintenance, the connection queue and fds are the first to run out.

When memory runs short, the kernel also pushes pages out to disk (swap, if enabled), and when memory truly runs out, Linux picks the process using the most memory and kills it (the OOM killer). The game server is usually the process using the most memory on its machine, so it’s the first to go. If the container has a memory limit, the same thing happens the moment it hits that limit, even when the server as a whole has room to spare. Things that seem unrelated to the game, such as time sync (NTP), scheduled jobs, container CPU limits, and CPU steal on virtual machines (time spent waiting while other VMs use the physical CPU), can also pause the server briefly or throw its timers off.

Analogy

The server OS is an amusement park’s entrance gate and management office. When everyone rushes in at opening time (the end of maintenance), the line at the gate (backlog) overflows, and once the wristbands (file descriptors) run out, nobody else gets in.

Causes of lag at this layer

Connection queue (listen backlog) overflow Listen backlog / SYN queue overflow

When tens of thousands of players connect at once right after maintenance, the kernel’s connection queue (listen backlog) overflows and connection attempts are dropped.

Why: As maintenance ends, connections pour in faster than the game server can accept them → Effect: The kernel’s connection queue (listen backlog: the smaller of the value the server code passes to listen and the kernel cap) fills up → On screen: Connection attempts are dropped and retried again and again: can’t connect / infinite loading

Symptoms: Can’t connect / infinite loading · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Game team (Client development)

File descriptor limit File descriptor limit (ulimit)

Every connection needs a file descriptor (fd: the number the OS gives an open file or socket), and the number of fds one process can open is capped.

Why: Concurrent users reach the process’s file descriptor limit → Effect: The server can’t accept new connections (Too many open files). Opening log files and DB connections fails too → On screen: Past an exact player count, nobody gets in: can’t connect / infinite loading

Symptoms: Can’t connect / infinite loading · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

Kernel socket buffers too small Small socket buffers

With small send and receive buffers, a burst of traffic makes the kernel drop packets arriving over UDP, and TCP sends block because the buffer has no room left.

Why: SO_SNDBUF and SO_RCVBUF left at their defaults or set too small → Effect: During a burst, or while the receiving thread pauses briefly, the UDP receive buffer overflows and drops packets; TCP waits because the send buffer has no room → On screen: Teleporting (UDP loss) or fast-forward (TCP waiting)

Symptoms: Teleporting, Fast-forward · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

Too many threads and context switching Thread oversubscription, context switching

Running far more threads than there are cores makes the OS spend CPU just switching between them.

Why: Hundreds to thousands of threads, for example one thread per connection → Effect: Higher context-switching cost (swapping out the running thread) and more cache misses → On screen: CPU is busy but throughput is low and ticks are uneven: stutter, slow motion

Symptoms: Stutter, Slow motion · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

CPU steal (virtual machines) CPU steal time

While the physical server (hypervisor) briefly gives a virtual machine’s CPU time to another VM (CPU steal), the game server stalls.

Why: Other VMs on the same host use a lot of CPU → Effect: The game server’s VM loses its turn on the CPU for a few ms to tens of ms at a time → On screen: Unexplained tick-time spikes: stutter, freeze

Symptoms: Stutter, Freeze · Primary owner Infra team (Server infrastructure) · Also External (External)

Container CPU throttling (CFS quota) Container CPU throttling (CFS quota)

With a CPU limit on a container, the moment it uses up its quota within a set period (usually 100 ms), it is forced to stop for the rest of that period (throttling).

Why: A CPU limit is set on the game server container (in Kubernetes, for example) → Effect: A burst of tick work uses up the quota, and the server stops for tens of ms until the next period → On screen: Average CPU is low, yet ticks spike periodically: stutter, slow motion

Symptoms: Stutter, Slow motion · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

Latency spikes from server power management (C-states, frequency scaling) CPU power management latency (C-states, frequency scaling)

Idle CPU cores drop into deep power-saving states (C-states) and lower their frequency to save power. Waking up and raising the frequency when a packet or timer arrives takes time, which adds delay to handling small packets.

Why: The OS frequency scaling policy (governor) or the BIOS power settings allow deep C-states and low frequencies → Effect: An idle core is up to hundreds of µs late every time it wakes from a deep power-saving state, and a frequency pinned low slows the tick computation itself → On screen: Usually hard to notice, but with many server-to-server calls it adds up to input lag that gets worse when the server is quiet. With the frequency pinned low, ticks fall behind when crowds gather: slow motion

Symptoms: Input lag, Slow motion · Primary owner Infra team (Server infrastructure)

OOM killer Out-of-memory killer

When memory runs out, Linux picks the process using the most memory and kills it. Usually that’s the game server.

Why: Memory runs out from a leak or a surge in usage, or the container hits its memory limit → Effect: The kernel kills the game server process → On screen: Everyone on that server disconnects at once, and recent progress may be rolled back

Symptoms: Disconnect, Dropped action / rollback · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Stalls from memory reclaim and compaction Memory compaction / reclaim stalls (THP)

The process stalls while the OS compacts memory to build huge pages or reclaims memory to free it up.

Why: Free memory runs low, or transparent huge pages (THP) trigger memory compaction → Effect: The thread that asked for memory waits until reclaim or compaction finishes → On screen: Irregular server stalls (a few ms to hundreds of ms)

Symptoms: Freeze, Stutter · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

System clock jump (NTP step) Wall-clock jump (NTP step)

When the server clock is moved forward or back by several seconds in one step, timers that depend on the system clock fire all at once or stop.

Why: Time sync moves the clock by a large amount in one step → Effect: Timers fire in a batch or stop, and timeouts are misjudged → On screen: Buff and cooldown glitches, mass disconnects, fast-forward

Symptoms: Fast-forward, Disconnect, Dropped action / rollback · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Scheduled jobs Cron jobs (log rotation, backup, scans)

Log compression, backups, and security scans that run at the same time every day take up CPU and disk.

Why: An OS job runs at a scheduled time → Effect: It shares CPU and disk with the game server → On screen: Stutter and slow motion at a fixed time, such as 4 a.m. every day

Symptoms: Stutter, Slow motion · Primary owner Infra team (Server infrastructure)

Performance changes after OS, kernel, driver, or firmware updates Performance regression after OS / kernel / driver / firmware update

The game code hasn’t changed, but the server has been slower since an OS, kernel, driver, or firmware update. Updates can change defaults, the scheduler, CPU vulnerability mitigations, and driver behavior.

Why: A routine security patch or a new server image changes the kernel, drivers, or firmware → Effect: Changed defaults or scheduler, or newly enabled vulnerability mitigations, make the same work take more CPU time and change the order in which threads get the CPU → On screen: A server that ran fine is a little slower all the time from the day of the update: input lag, plus stutter and slow motion when crowds gather

Symptoms: Input lag, Stutter, Slow motion · Primary owner Infra team (Server infrastructure)

Server conntrack table full conntrack table full

When the connection tracking (conntrack) table, where the Linux firewall records every connection, reaches its limit, new packets are dropped.

Why: Connection surges and repeated short-lived connections pile up connection entries → Effect: The table fills up, and new connections and some packets are dropped → On screen: Can’t connect, and teleporting from unexplained packet loss

Symptoms: Can’t connect / infinite loading, Teleporting · Primary owner Infra team (Server infrastructure) · Also Game team (Server development), Game team (Client development)

Ephemeral port exhaustion on server-to-server connections Ephemeral port exhaustion (TIME_WAIT)

When a game server opens and closes short connections to the DB or other servers very often, closed connections hold their ports for a while, and new connections can’t be opened.

Why: A new connection is opened and closed for every request → Effect: The side that closes first holds the port for about 60 seconds on Linux (TIME_WAIT), and the pool of usable ports runs dry → On screen: Internal requests fail: failed saves, broken features

Symptoms: Dropped action / rollback, Can’t connect / infinite loading · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Sockets and protocols: TCP, UDP, socket options

This is the part that decides how the game sends and receives data over the network. Even on the same connection, depending on which protocol you use and how you set the socket options, losing a single packet can end as “a slight skip” or turn into “a 1-second freeze, then fast-forward.”

TCP guarantees that everything is delivered, in the order it was sent. The catch is that when one packet is lost, it holds back everything that arrived after it from the game until the lost one comes again. UDP guarantees nothing. It hands over whatever arrives with no waiting, but the game has to deal with anything lost on its own. That’s why action-heavy games build just as much reliability as they need on top of UDP (reliable UDP), while many MMOs use TCP because it’s easier to implement, and accept its weaknesses.

Socket options are the fine-grained settings for this behavior: whether to collect small packets before sending (TCP_NODELAY), how big the send and receive buffers are (SO_SNDBUF, SO_RCVBUF), when to notice a dead connection (SO_KEEPALIVE, TCP_USER_TIMEOUT), and what to do with leftover data on close (SO_LINGER). Most defaults are tuned for sending large amounts of data efficiently in few packets, which often works against games that exchange small packets frequently.

Key point

TCP hands received data to the game only in the order it was sent. If packet 17 is lost, packets 18–30 all wait until 17 comes again, even if they have already arrived (head-of-line blocking). UDP hands data over as it arrives, so even if 17 never shows up, the rest are processed on time.

Why retransmissions happen (Wi-Fi, congestion, MTU black holes, spurious retransmissions, and so on) and how to find the cause are covered cause by cause in 06 TCP retransmission.

Causes of lag at this layer

TCP head-of-line blocking Head-of-line blocking

To keep data in order, TCP holds back every packet that arrived after a lost one until the lost packet is received again.

Why: One packet goes missing → Effect: The packets behind it have arrived but wait in the receive buffer → On screen: Everything stops, then releases all at once: fast-forward

Symptoms: Freeze, Fast-forward · Primary owner Game team (Server development) · Also Game team (Client development)

TCP RTO and exponential backoff RTO and exponential backoff

Each time a retransmission fails again, the wait doubles, so a brief connection drop turns into a long stall.

Why: The connection drops briefly, and retransmissions fail one after another → Effect: The wait before the next attempt doubles each time: 0.3 → 0.6 → 1.2 → 2.4 s (at 100 ms ping) → On screen: The connection was down for 1 second, but the game stalls for over 2 seconds. A longer drop eventually ends in a disconnect

Symptoms: Freeze, Disconnect · Primary owner Game team (Server development) · Also Game team (Client development)

Nagle’s algorithm + delayed ACK Nagle + delayed ACK (TCP_NODELAY off)

Nagle’s algorithm, which batches small packets, and delayed ACK, which sends ACKs late, interact so that each message written in pieces is delayed by 40–200 ms.

Why: Small messages are written in pieces without TCP_NODELAY turned on → Effect: The sender waits for an ACK, and the receiver sends its ACK late → On screen: Ping is low, yet every action is consistently sluggish: input lag

Symptoms: Input lag · Primary owner Game team (Server development) · Also Game team (Client development)

Blocking sends caused by slow clients Blocking send on a full socket

When one player on a slow connection has a full send buffer and the server sends in blocking mode (a send call that doesn’t return until the buffer has room), the server thread waits on that one player.

Why: A slow client’s send buffer is full → Effect: With blocking sends, the server thread waits until the buffer has room → On screen: Everyone that thread handles gets a freeze or slow motion

Symptoms: Freeze, Slow motion · Primary owner Game team (Server development)

Slow client (slow consumer) handling policy Slow-consumer policy

When a client’s outgoing data keeps piling up, the server drops stale updates or disconnects it.

Why: The client’s connection can’t keep up with what the server sends → Effect: The server drops stale updates, or disconnects the client once a limit is exceeded → On screen: Just that player sees teleporting or gets a disconnect

Symptoms: Teleporting, Disconnect · Primary owner Game team (Server development)

Keepalive default of 2 hours TCP keepalive defaults

When the other side vanishes without a close signal, TCP notices only much later. Keepalive (a TCP feature that checks whether an idle connection is still alive) is off by default, and even when it’s on, checks start only after 2 hours of idle time.

Why: The client vanishes without a close signal because its power went off or its connection dropped → Effect: The server assumes the connection is still alive (keepalive default 7,200 s; if data was being sent, about 15 minutes until retransmission gives up) → On screen: A ghost character stays behind, and reconnecting fails with an “Already logged in” error

Symptoms: Can’t connect / infinite loading, Invisible / ghost entities · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)

IP fragmentation of UDP packets IP fragmentation of large UDP

A UDP packet larger than the MTU (the largest size that can be sent in one piece) is fragmented at the IP layer, and losing just one fragment throws away the whole packet.

Why: Snapshots in crowded areas exceed 1,500 bytes → Effect: They go out split into several fragments, and losing any one of them discards the whole packet → On screen: Large packets are lost several times as often. Teleporting only in crowded areas

Symptoms: Teleporting · Primary owner Game team (Server development)

Reliable UDP retransmission settings Reliable-UDP tuning (KCP, ENet…)

When the retransmission rules you built on top of UDP are too conservative, recovery is slow; when they’re too aggressive, they clog the connection even more.

Why: Retransmission interval, retry count, and window size don’t suit the connection → Effect: Slow recovery, or duplicate sends that make congestion worse → On screen: Skills not going off, fast-forward, worse lag during congestion

Symptoms: Dropped action / rollback, Fast-forward · Primary owner Game team (Server development) · Also Game team (Client development)

Slow start after idle Slow start after idle

When a connection has been idle for a while, TCP shrinks the congestion window (how much it can send at once) again, so a sudden large send goes out in several rounds.

Why: A large burst of data (on entering a town, for example) goes out over a connection that was idle → Effect: The congestion window has shrunk, so the data is spread over several round trips → On screen: Right after entering, nearby characters and NPCs appear a few round trips late (more noticeable on distant servers)

Symptoms: Input lag, Invisible / ghost entities · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

Sending rate plunges under congestion control Congestion control backoff

TCP treats loss as a sign of congestion and cuts its sending rate by 30–50%. It reacts the same way to Wi-Fi loss.

Why: A little loss on Wi-Fi or the connection while there’s a lot to send → Effect: TCP cuts its sending rate sharply and recovers slowly (CUBIC, the Linux and Windows default, cuts by 30%) → On screen: Updates fall behind in crowded areas: fast-forward, input lag

Symptoms: Fast-forward, Input lag · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

Last data lost to an abortive close (RST) SO_LINGER, abrupt RST

When the server cuts a connection abruptly, the final notice or save-complete signal it sent is lost.

Why: The server closes the connection with an abortive close (RST). This happens when SO_LINGER is set to 0 seconds, or when the socket is closed before all received data has been read → Effect: The kick reason and final data still in transit are thrown away → On screen: An unexplained “Connection closed due to an unknown error”

Symptoms: Disconnect · Primary owner Game team (Server development)

Blocking I/O design Blocking I/O model

In a design where a thread can’t do anything else while it waits on one socket, everything slows down as the player count grows.

Why: Each connection waits on its own reads and writes → Effect: A delay on one connection spreads to the other connections on the same thread → On screen: As concurrent users grow, everyone gets slow motion and input lag

Symptoms: Slow motion, Input lag · Primary owner Game team (Server development)

Uneven SO_REUSEPORT distribution SO_REUSEPORT imbalance, stuck worker

When several processes share one port, the kernel assigns each connection to a process by address hash and never reassigns it. If one of those processes stalls, only the players assigned to it wait.

Why: A gateway or login server runs several processes on one port with SO_REUSEPORT → Effect: Even when one process stalls from GC or overload, the new connections and UDP packets assigned to it don’t move to another process → On screen: Only some players can’t connect or freeze. During a restart that changes the process count, some UDP sessions drop

Symptoms: Can’t connect / infinite loading, Freeze, Disconnect · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

WSAECONNRESET errors on Windows UDP sockets WSAECONNRESET on a Windows UDP socket

When a Windows server sends UDP to a client that has already left, a “port unreachable” (ICMP) message comes back. That message makes the next receive call fail with an error, and if the server code treats the error as a failure of the socket itself, everyone using that socket is affected.

Why: UDP keeps going to the address of a client that just left, and a “port unreachable” (ICMP) message comes back → Effect: Windows fails the next receive call with WSAECONNRESET (10054), and the server code stops receiving or closes the socket → On screen: Everyone who was using that socket freezes or disconnects at once

Symptoms: Disconnect, Freeze · Primary owner Game team (Server development)

Server game process: ticks and threads

The program that actually computes the game logic. Movement, combat, monster AI, visibility calculation, and broadcasts all have to finish within one “tick.” The more people gather in one place, the more visibility calculations and outgoing packets there are, growing with the square of the player count.

The server computes the game state at a fixed tick interval. On a 20-tick server that’s once every 50 ms, and within that time it applies every player’s input, moves the monsters, works out who can see whom (visibility, AOI), and sends the changes to everyone who can see them. That 50 ms is the tick budget. Go over it, and the next tick starts late. A server that advances game time by a fixed step per tick makes time in the whole game world run slow (slow motion). A server that advances by however much real time has passed keeps the speed, but its packets come less often, so things stutter and teleport. Either way, responses get slower. If one game thread runs a whole server (channel), everyone on that server feels it; if threads are split by zone, the people in that zone feel it together.

The problem is player count. Comparing everyone with everyone means about 10,000 checks per tick for 100 players and about 1 million for 1,000. So servers divide the map into a grid and compare only nearby cells, but when everyone crowds around one cell, as at a world boss, a siege, or a town square event, the grid helps less and the computation and outgoing data explode. Add locks, where several threads wait on the same data, and blocking calls that wait for a DB response in the middle of a tick, and everyone handled by that thread freezes together while it waits.

Analogy

The server tick is a conductor’s beat. The more orchestra members (players) there are, the more sheet music has to be covered within one beat, and when a beat slips, the whole piece slows down. If someone walks off to the storeroom (DB) to find a page of music, everyone waits for them.

Causes of lag at this layer

Tick overrun Tick overrun

When one tick has more work than its budget, the server’s tick interval stretches, and the whole area slows down or stutters.

Why: One tick (e.g., 50 ms) has more work than its budget → Effect: Game state meant to update 20 times a second updates only 8 times → On screen: Slow motion across the zone (stutter on some server designs), sluggish skill response

Symptoms: Slow motion, Input lag, Stutter · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

AOI calculation blowup (N²) Area-of-interest explosion

If you compare everyone against everyone to work out who can see whom, 10 times the players means 100 times the computation.

Why: Every character’s distance is checked against every other character, or, even with a grid, hundreds of players crowd around one cell → Effect: About 10,000 comparisons for 100 players, about 1 million for 1,000 → On screen: Ticks spike where crowds gather, such as world bosses and sieges: slow motion, stutter

Symptoms: Slow motion, Stutter · Primary owner Game team (Server development)

Broadcast fan-out overload Broadcast fan-out (N×N)

Sending one player’s movement to everyone who can see them creates updates on the order of the square of the crowd size.

Why: Each player’s changes are sent to everyone who can see them → Effect: 1,000 players who all see each other means 1 million updates per tick → On screen: The send queue and bandwidth saturate, causing delay and loss (input lag, fast-forward, teleporting)

Symptoms: Input lag, Teleporting, Fast-forward · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Single-threaded zone overload (hotspot) Single-threaded hot zone

When each area runs on a single thread and players crowd into one place, only that one core hits 100%.

Why: One thread runs each area (channel) → Effect: When players crowd into one place, only that core saturates while the other cores have room to spare → On screen: Only that area lags; other areas are fine

Symptoms: Slow motion, Input lag · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Lock contention Lock contention

When several threads wait on one lock to write the same data, they run one at a time no matter how many threads you add.

Why: Several threads use shared data at once, such as the auction house or guild storage → Effect: The others wait until the thread holding the lock finishes → On screen: Only certain features are slow; in bad cases, the whole tick is delayed

Symptoms: Input lag, Freeze · Primary owner Game team (Server development)

Deadlock Deadlock

When two threads each wait for a lock the other holds, both stop forever.

Why: Thread A holds lock 1 and waits for lock 2, while B holds lock 2 and waits for lock 1 → Effect: Both stop forever, and related threads stop one after another → On screen: The whole server stops, and everyone disconnects when the watchdog restarts it

Symptoms: Freeze, Disconnect · Primary owner Game team (Server development)

Blocking calls on the game thread Synchronous DB / file I/O on the game loop

If the server waits for a DB response or a file write in the middle of a tick, all game progress on the server stops for that long.

Why: The tick waits on DB reads and writes, log writes, or external API calls → Effect: If the DB takes 100 ms, the tick stalls for 100 ms too → On screen: Every time the DB or disk slows down, the whole field hitches

Symptoms: Freeze, Stutter · Primary owner Game team (Server development)

Message queue backlog Mailbox / job queue backlog

When requests arrive faster than they’re processed and pile up in the queue, the ones at the back get processed only seconds later or are dropped.

Why: Requests arrive faster than they can be processed → Effect: The queue grows, and messages are dropped once it passes its limit → On screen: Skills and trades respond late or are dropped

Symptoms: Input lag, Dropped action / rollback · Primary owner Game team (Server development)

Timers firing all at once Synchronized timers

When every monster respawn, every buff expiry, and the on-the-hour reward all land on the same tick, that one tick becomes tens of times heavier.

Why: Respawn, expiry, reward, and autosave timers are all set to the same moment → Effect: That one tick has tens of times its usual work → On screen: A hitch at each of those scheduled times

Symptoms: Freeze, Stutter · Primary owner Game team (Server development)

Pathfinding storm Pathfinding storms

When hundreds of monsters chase players and compute paths at the same time, it takes a lot of CPU.

Why: Large mob pulls or mass spawns send many monsters chasing players at once → Effect: Each monster runs its own pathfinding → On screen: Slow motion in that hunting ground only

Symptoms: Slow motion · Primary owner Game team (Server development)

Serialization and compression cost Serialization / compression cost

Turning outgoing data into bytes and compressing it takes CPU too, and with many players this cost explodes.

Why: Structs are converted to bytes and compressed for every update → Effect: Cost grows with the square of the player count → On screen: Sends go out late: input lag

Symptoms: Input lag · Primary owner Game team (Server development)

Server crash Server process crash

When the server process dies from an unhandled error, everyone on that server disconnects at the same time.

Why: A fatal error such as a reference to something that doesn’t exist (null reference), bad data, or running out of memory → Effect: The server (or zone) process exits → On screen: Everyone disconnects at once, and progress since the last save may be rolled back

Symptoms: Disconnect, Dropped action / rollback · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Thread pool exhaustion Thread pool starvation

When every worker thread is tied up in slow work, new requests just wait with no end in sight.

Why: Worker threads are tied up waiting on external API or DB responses → Effect: No thread is free to take a new request → On screen: Infinite loading in specific features such as login or the shop

Symptoms: Can’t connect / infinite loading, Input lag, Freeze · Primary owner Game team (Server development)

Infinite loops and runaway logic Infinite loop / runaway logic

When a bug keeps a tick from ever finishing, the server stops, and the watchdog forces a restart.

Why: A loop that never ends because of a wrong condition, or runaway recursion → Effect: The tick never finishes, and the server stops → On screen: A freeze, then everyone disconnects

Symptoms: Freeze, Disconnect · Primary owner Game team (Server development)

Combat concentrated on one target (world boss) Hot entity / combat event fan-out

When hundreds of players hit one boss at the same time, the computation for that single boss piles up in one place, and hit information goes out to everyone watching.

Why: Hundreds of players use skills, buffs, and debuffs on one boss nonstop → Effect: The boss’s HP, aggro list, and debuff calculations pile up in one place, and every hit sends damage number and effect packets to everyone watching → On screen: Skills land late and damage numbers pop up in bursts; slow motion only around the boss

Symptoms: Input lag, Fast-forward, Slow motion · Primary owner Game team (Server development)

Spawn burst when entering a crowded area Spawn burst when entering a crowd

When you teleport into a town packed with players, the server has to send the appearance, gear, and status of hundreds of newly visible players all at once.

Why: You suddenly appear somewhere crowded by teleporting, logging in, or switching channels → Effect: Full data for hundreds of players is built and sent at once, and your PC also loads it all at once → On screen: A brief pause right after arrival, characters pop in late one by one, and input lags

Symptoms: Freeze, Input lag, Fast-forward · Primary owner Game team (Server development) · Also Game team (Client development)

Entity buildup (items and summons never cleaned up) Entity / timer buildup over uptime

When ground items that should have disappeared, summons, and finished timers pile up without being cleaned up, every tick has more work to do the longer the server stays up.

Why: Ground items, summons, expired timers, and empty party data aren’t removed on time → Effect: The lists walked every tick grow longer day by day → On screen: Fine right after maintenance, then after a few days only that server or area gets more and more sluggish

Symptoms: Slow motion, Stutter, Input lag · Primary owner Game team (Server development)

Patch changes the traffic pattern Patch changes traffic pattern

When new content, effects, or synced fields raise packet size and frequency, a server that ran fine starts hitting MTU, bandwidth, and packet-rate limits after the patch.

Why: The patch adds new skill effects, synced fields, or item data, making packets bigger or more frequent → Effect: Large packets exceed the MTU and get fragmented, and the extra volume runs into bandwidth limits, cloud PPS limits, and send buffers → On screen: Teleporting, skills not going off, and input lag in crowded places, starting right after the patch. Nothing changed in the infra, yet loss goes up

Symptoms: Teleporting, Dropped action / rollback, Input lag · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (Network infrastructure)

Memory

Everything the server remembers, from characters and monsters to items and maps, sits in memory. Memory itself is fast, but lag appears the moment the server stops for GC (reclaiming unused memory), memory slowly leaks away (a leak), or it runs short and swapping starts (moving part of memory to disk).

Languages that manage memory automatically, such as Java, C#, and Go, rely on a garbage collector (GC) to collect and reclaim memory that has been used and thrown away. Depending on the GC type, it may briefly stop every thread. A GC over the entire heap (the memory area a program allocates while running) takes longer the more live data there is, reaching hundreds of milliseconds to several seconds. Modern GCs like ZGC cut pauses to under 1 ms at the cost of more CPU and memory. C++ servers have no GC, but they suffer from leaks, where memory nobody freed piles up, and fragmentation, where free space gets chopped into small pieces so large blocks can’t be used. Even with a GC, if something keeps referencing objects that are no longer needed, you get a leak just the same.

Another reason memory gets slow is the memory hierarchy (how far the storage is from the CPU). The cache right next to the CPU takes 1 ns, RAM takes 100 ns, and reading back memory that was pushed out to disk (swap) takes more than 1,000 times as long as RAM. Try stretching these gaps out to human time scales in the “Latency numbers” table below.

Analogy

Memory is a cook’s workbench. Ingredients within reach (cache) are fast, a trip to the fridge (RAM) is a bit slower, and when the workbench is full and ingredients go to the storeroom (disk, swap), every trip takes ages. And while doing the dishes (GC), the cooking has to stop.

Causes of lag at this layer

Server GC stop-the-world pause Stop-the-world GC pause

While a Java or C# server halts every thread to collect garbage (stop-the-world), the whole server stalls.

Why: The heap fills up and GC starts → Effect: Every game thread stops while GC collects (longer the more live data there is) → On screen: Everyone on the server freezes at the same moment, then the game fast-forwards

Symptoms: Freeze, Fast-forward · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Script engine GC pause Scripting VM GC (Lua, etc.)

Even on a C++ server, if quests, AI, and skills run in a scripting language such as Lua, the zone stops while the script engine’s GC runs.

Why: Each zone’s script engine creates large numbers of temporary objects while running quests, AI, and events → Effect: When the script engine’s GC collects a lot at once, that zone’s tick stops → On screen: Periodic hitches only in certain zones or during certain events

Symptoms: Stutter, Freeze · Primary owner Game team (Server development)

Allocation surge Allocation storms

Creating large numbers of temporary objects during an event makes GC run far more often than usual.

Why: Item drops, combat logs, and event rewards create a flood of temporary objects → Effect: GC runs several times as often, and objects not yet discarded get promoted to the old generation, so full GCs come sooner too → On screen: Periodic hitches only during events

Symptoms: Stutter, Freeze · Primary owner Game team (Server development)

Memory leak Memory leak

Memory that is never freed piles up little by little and, days later, leads to GC storms, swapping, or the process getting killed.

Why: Data for logged-out characters and event handlers is never freed → Effect: Free memory shrinks over several days → On screen: Fine right after maintenance, laggier every day, and eventually the server goes down

Symptoms: Slow motion, Freeze, Disconnect · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

GC thrashing (too little heap headroom) GC thrashing (heap nearly full)

When live data gets close to the heap limit, each GC reclaims almost nothing, so GC runs over and over without a break.

Why: An event crowd or a leak pushes live data close to the heap limit → Effect: GC reclaims only a little, so another full GC follows right away and GC uses most of the CPU → On screen: The whole server alternates between slow motion and freezes for several minutes, then dies from running out of memory

Symptoms: Slow motion, Freeze, Disconnect · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Swap Swapping

When memory runs short and the OS moves part of it out to disk, every access to that memory waits on a disk more than 1,000 times slower.

Why: Memory in use exceeds physical RAM → Effect: The OS moves part of it to disk and reads it back when needed → On screen: Ticks balloon to hundreds of ms, and every player on the server sees slow motion and freezes

Symptoms: Slow motion, Freeze · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

Cache miss CPU cache misses

When data is scattered all over memory, the CPU has to go all the way out to slow RAM and wait every time.

Why: Objects scattered behind pointers and accessed in no particular order → Effect: Data isn’t in the CPU cache, so every read goes to RAM (roughly 100 times slower) → On screen: The same work costs several times more tick time; in bad cases, slow motion

Symptoms: Slow motion · Primary owner Game team (Server development)

Memory fragmentation Heap fragmentation

When repeated allocation and freeing chops free space into small pieces, the process holds far more memory than it actually uses.

Why: Many threads allocate and free blocks of varying sizes over a long time → Effect: Free space ends up scattered in small pieces that can’t be returned to the OS, so usage keeps growing like a leak → On screen: The longer it runs, the slower it gets from swapping and memory shortage, until it gets killed

Symptoms: Slow motion, Disconnect · Primary owner Game team (Server development)

Remote NUMA memory Remote NUMA access

On a server with two CPUs, using memory attached to the other CPU slows access down.

Why: Threads and their memory sit on different CPU sockets → Effect: Memory access slows down (1.5–2× depending on hardware) → On screen: Same specs, but performance differs from process to process

Symptoms: Slow motion · Primary owner Infra team (Server infrastructure)

Disk

Logs, character saves, map data, and DB files all live on disk. Disks are hundreds of times (SSD) to 100,000 times (HDD) slower than memory, so if the game server is built to wait on the disk, the game stops whenever the disk gets busy.

Disk performance is measured in “how many reads and writes per second” (IOPS). An old HDD manages about 150, an SSD tens of thousands to hundreds of thousands. Cloud disks get a limit set by what you pay (AWS’s default gp3 gives 3,000). Some cloud disks and small instance types hand out burst credits that allow briefly higher performance on top of the baseline, but when busy periods run long, the credits run out and speed drops suddenly. Reports like “it lags every evening after a few hours” have this shape.

The key is who waits. Ordinary file writes usually finish right away, because the OS takes them into memory first and flushes them to disk later. The trouble starts when you ask to wait “until it’s actually on disk” (fsync), or when the OS’s limit for holding writes in memory fills up. If the game thread waits directly (synchronous), a 100 ms disk backlog freezes the tick for 100 ms too. Handing writes to a separate thread (asynchronous) keeps the game from freezing, but if the server suddenly crashes, anything not yet written can be lost (dropped action / rollback).

Analogy

The disk is a warehouse, and IOPS is the number of warehouse doors. With few doors, the people carrying things in and out stand in line. Burst credits are stamina for a short sprint; once they’re used up, you’re back to walking speed.

Causes of lag at this layer

Synchronous log writes Synchronous logging

If the game thread waits for the disk to finish every log line, the game stalls too whenever the disk is busy.

Why: Combat and trade logs are written straight to a file from the game thread → Effect: When a durable write (fsync) is required or the OS write buffer (page cache) hits its limit, a single write takes tens of ms while the disk is busy → On screen: Hitches in log-heavy fights

Symptoms: Stutter, Freeze · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

fsync surge fsync storms

Asking for data to be written to disk “for sure” takes 0.1 ms to tens of ms per request depending on the disk, and when requests pile up, the queue grows.

Why: Scheduled saves and logout rushes send a flood of durable write requests → Effect: The disk queue grows → On screen: Lag at every save time, slow logouts and channel changes

Symptoms: Stutter, Input lag · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Infra team (DB infrastructure)

Cloud disk out of burst credits Burst credit depletion

Some cloud disks and small server sizes have burst credits that let them run faster than baseline for a while, so when a busy period drags on and the credits run out, speed drops suddenly.

Why: Sustained use above baseline performance → Effect: Burst credits run out and performance drops sharply to baseline → On screen: Lag starts a few hours into every evening

Symptoms: Stutter, Slow motion, Input lag · Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)

IOPS limit / queue saturation IOPS limit / queue saturation

When requests exceed what the disk can handle per second, the queue grows and latency explodes.

Why: Read and write requests approach the disk’s capacity → Effect: The queue grows (usually exploding above 90% utilization) → On screen: Slow saves and loading; freezes if the calls are blocking

Symptoms: Input lag, Freeze · Primary owner Infra team (Server infrastructure) · Also Game team (Server development), Infra team (DB infrastructure)

Disk full Disk full

When logs and dumps pile up and fill the disk, writes fail, and without safeguards the server crashes.

Why: Logs, dumps, and temp files pile up to 100% → Effect: Writes fail. Crash if there’s no error handling, failed saves if there is → On screen: Disconnects, rolled-back progress

Symptoms: Disconnect, Dropped action / rollback · Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure), Game team (Server development)

Backup / compression / scan jobs Backup / compression / scans

When early-morning backups, log compression, or security scans monopolize the disk, the game server’s reads and writes get held up.

Why: A scheduled backup or compression job starts → Effect: It takes most of the disk bandwidth and IOPS → On screen: Lag at the same time every day

Symptoms: Stutter, Input lag · Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure)

Server-side lazy loading Lazy loading on the server

If the server reads dungeon or map data from disk the first time it’s requested, everyone freezes for that tick.

Why: Someone enters a dungeon or area for the first time → Effect: The server reads the data from disk on the game thread → On screen: Everyone on that server freezes briefly

Symptoms: Freeze · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Writing a core dump Core dump writing

When the server crashes, writing several GB of memory to disk can delay the restart by several minutes.

Why: A server crash writes all of memory to a file → Effect: No restart until several GB have been written → On screen: After the server dies and players disconnect, they can’t connect again for a long time

Symptoms: Can’t connect / infinite loading · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

HDD seek latency HDD seek latency

An HDD has to move its head across the platter (a seek), so reading or writing scattered data takes close to 10 ms each time.

Why: HDDs in old servers or low-cost storage → Effect: About 10 ms for every scattered read or write → On screen: Slow saves and loading across the board

Symptoms: Input lag · Primary owner Infra team (Server infrastructure) · Also Infra team (DB infrastructure), Game team (Server development)

Database

The place that holds what must never be lost: characters, items, currency, and trade records. When the DB slows down, combat is fine but items arrive late, trades fail, and logins never finish. If the game server is built to wait on the DB, the whole field freezes.

The game server opens a few connections to the DB in advance (a connection pool) and takes turns using them. If one query (a request sent to the DB) takes a long time, that connection stays busy, and when every connection in the pool is busy, the remaining requests wait in a queue. Queries usually slow down for one of two reasons: there’s no index (like the index of a book), so the whole table gets read (a full table scan), or several requests try to modify the same row at once and wait for a lock.

Databases use several mechanisms to stay reliable and handle many requests: replicas that share the read load, a standby DB to fail over to, and checkpoints that periodically flush changes to disk in bulk. When checkpoints bunch up, things slow down briefly. When a replica falls behind, you get “I can’t see the item I just bought,” and when failover happens while replication is behind, you get “I logged in and it went back to where I was a little while ago,” the dropped action / rollback symptom. If the game server saves characters only once every few minutes, a server crash turns into “I got rolled back 10 minutes.”

Analogy

The DB is a row of bank teller windows. There’s a fixed number of windows (the connection pool), and if one request has a teller search the entire ledger (a full table scan), every request behind it waits. If everyone wants to open the same vault (a hot row), only one person can go in at a time.

Causes of lag at this layer

Queries with no index Missing index / full table scan

Without an index, finding the rows that match a condition means reading the entire table (a full table scan).

Why: A new feature ships with a search on a condition that has no index → Effect: Scanning millions of rows makes a single query take hundreds of ms to several seconds → On screen: Mailbox and trade history load slowly, and tied-up connections make other requests wait too

Symptoms: Input lag, Can’t connect / infinite loading · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

Hot row lock contention Hot row lock contention

When everyone tries to modify the same row (a guild vault, a popular auction house item, a server-wide counter), only one request at a time gets the lock.

Why: An event or a popular item concentrates updates on the same row → Effect: Requests wait until they get the lock → On screen: Failed trades, “Please try again later” messages, timeouts

Symptoms: Dropped action / rollback, Input lag · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

DB deadlock Database deadlock

When two transactions (groups of DB operations processed as one unit) each wait for a row the other has locked, the DB forcibly cancels one of them.

Why: Trade A locks in item→currency order, trade B in currency→item order → Effect: The DB detects the deadlock and rolls back one side → On screen: Trades and crafting fail now and then, items revert

Symptoms: Dropped action / rollback, Input lag · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

Connection pool exhaustion Connection pool exhaustion

The number of connections open to the DB is fixed, so when slow queries hold connections, every other request waits.

Why: Slow queries or a flood of requests put every connection in use → Effect: New requests wait until a connection frees up → On screen: Infinite loading at login, slow saves, timeouts

Symptoms: Can’t connect / infinite loading, Input lag · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

Replication lag Replication lag

Writes go to the primary and reads come from replicas, so when a replica falls behind, data that was just written isn’t visible yet.

Why: A burst of writes on the primary puts replicas several seconds behind → Effect: Reading just-saved data from a replica finds it missing → On screen: An item you just bought doesn’t show up, marketplace prices are stale, duplicate-reward bugs

Symptoms: Dropped action / rollback · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)

Checkpoint / log flush Checkpoint / log flush stalls

Queries slow down at the moments the DB periodically writes its accumulated in-memory changes to disk in bulk.

Why: Changes pile up and are periodically written to disk → Effect: The disk gets busy at that moment and queries slow down → On screen: Saves and loading slow down periodically

Symptoms: Input lag, Stutter · Primary owner Infra team (DB infrastructure)

Cold cache (right after a restart) Cold buffer pool after restart

After a DB restart, the memory cache is empty, so for a while every lookup reads from disk.

Why: The DB restarts for maintenance → Effect: Frequently used data isn’t in memory, so it’s read from disk → On screen: Logins and loading are slow for a while right after maintenance

Symptoms: Can’t connect / infinite loading, Input lag · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)

Login storm and N+1 queries Login storm, N+1 queries

If loading one character takes dozens of separate queries, tens of thousands of simultaneous logins turn into millions of queries.

Why: Character loading queries items, skills, and quests one by one → Effect: Simultaneous logins right after maintenance make the query count explode → On screen: Infinite loading at login, and even saves for players already in the game get held up

Symptoms: Can’t connect / infinite loading, Input lag · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

Bulk batch jobs Batch jobs during service

Running ranking aggregation, mass mail sends, or old-data cleanup during live service ties up locks and the disk.

Why: Bulk jobs run during service hours → Effect: Wide-range locks, disk and CPU tied up → On screen: Failed trades and saves, slow loading at certain times of day

Symptoms: Input lag, Dropped action / rollback · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

DB failover Database failover

When the primary DB dies, writes stop while it fails over to a standby, and the last data that hadn’t been replicated yet can be lost.

Why: The primary DB fails and a standby is promoted → Effect: No writes for seconds to minutes during the switch; with asynchronous replication, unreplicated data may be lost → On screen: Every save fails for a moment, items and XP roll back

Symptoms: Dropped action / rollback, Freeze, Disconnect, Can’t connect / infinite loading · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)

Lost progress from a long save interval Periodic save window

If the server saves only once every few minutes to reduce load, progress is lost when the server dies in between.

Why: Character state is saved once every few minutes → Effect: A server crash or outage hits in between → On screen: After reconnecting, the character is back to where it was minutes ago (rollback)

Symptoms: Dropped action / rollback · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

Cache stampede Cache stampede / thundering herd

When cache entries for popular data expire at the same time, thousands of requests hit the DB all at once.

Why: Popular data stored in Redis or a similar cache expires at the same time → Effect: Requests trying to rebuild the same data rush to the DB all at once → On screen: DB overload makes one feature after another slow down or freeze

Symptoms: Input lag, Freeze, Can’t connect / infinite loading · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

Transaction left open too long Long-running transaction / MVCC purge lag

When a transaction stays open for a long time, it keeps holding its locks and the DB can’t clean up (purge) old versions of data, so everything gradually slows down.

Why: A transaction stays open while waiting for another server’s response, or a long aggregate query runs on the primary during service → Effect: Its locks are never released, and old row versions awaiting cleanup keep piling up → On screen: Features that use those rows time out, and saves and lookups slow down across the board over several hours

Symptoms: Input lag, Dropped action / rollback · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

Slow Redis commands Redis blocking commands (single-threaded)

Redis processes commands one at a time, so a single slow command blocks every request behind it.

Why: A full key search with KEYS in production, or reading or deleting a ranking or list with millions of elements in one go → Effect: Every other request waits until that command finishes (tens of ms to several seconds) → On screen: Features that use sessions, rankings, or the cache all hitch at once, slow logins

Symptoms: Freeze, Input lag, Can’t connect / infinite loading · Primary owner Game team (Server development) · Also Infra team (DB infrastructure)

Query slowdown from a query plan change Query plan regression (stats, parameter sniffing)

Even with the code unchanged, if the DB changes how it executes a query (its query plan), a query that took 2 ms yesterday takes hundreds of ms today.

Why: Automatic statistics updates, a DB restart, or shifts in data distribution make the DB build a new query plan → Effect: A plan that skips the index gets picked, the same query becomes tens to hundreds of times slower, and connections get tied up → On screen: With no deploy at all, loading for a specific feature suddenly slows down and other requests wait too

Symptoms: Input lag, Can’t connect / infinite loading · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)

Schema change (DDL) lock during live service Schema change lock (DDL / metadata lock)

Adding a column or index to a table during live service can make every request that uses that table wait, all because of one lock that’s needed only briefly.

Why: A hotfix adds a column or index to a table in live use → Effect: The schema change waits for a long transaction opened earlier, and every request that comes after waits for the schema change → On screen: Features that use that table (inventory, mail, and so on) stop entirely and time out

Symptoms: Input lag, Dropped action / rollback, Can’t connect / infinite loading · Primary owner Infra team (DB infrastructure) · Also Game team (Server development)

Server architecture and operations

Most MMOs today run as a set of login, gateway, field, dungeon, chat, party, auction house, cache, and DB servers that call one another. A failure in one place spreads to whatever is connected to it, and operational work such as deploys, scaling, and maintenance creates lag too.

Splitting servers can keep a failure in one place from spreading to everything, but it creates call chains (servers calling other servers in sequence). The game server calls the auction house server, and the auction house server calls the cache and the DB, for example. When the server at the end of the chain slows down, the servers in front keep holding threads and connections while they wait for responses, and eventually even features that look unrelated stop. This is called a cascading failure, and timeouts and circuit breakers (which briefly block calls that keep failing) keep it from spreading.

Operational work causes lag too. Restarts during update deploys, the few minutes it takes to add servers automatically when crowds arrive, moving a character to another server during a zone transfer, and the invisible load from bots and macros all look like “lag” to players.

Analogy

Server architecture is a company where departments pass approvals back and forth. If one department at the end of the approval chain (the DB) slows down, the departments in front line up holding paperwork, and eventually the whole company’s work stops. A timeout is the rule “if there’s no answer within 10 minutes, send it back for now,” and a circuit breaker is the rule “if things keep getting sent back, stop sending paperwork to that department for a while and return it right away.”

Causes of lag at this layer

Routing through a gateway or proxy Gateway / proxy hop

Putting an intermediate server between the client and the game server adds processing time at every hop, and that server becomes a single point of failure.

Why: Client ↔ gateway ↔ game server architecture → Effect: The intermediate server adds processing and queueing time, and when it’s overloaded everyone is affected → On screen: Higher ping for everyone; if a gateway fails, every player routed through it disconnects

Symptoms: Input lag, Disconnect · Primary owner Game team (Server development) · Also Infra team (Server infrastructure), Game team (Client development)

Zone transfer (handoff between servers) Zone / server handoff

Entering another area or dungeon means handing the character’s data to another server, and that handoff can be slow or fail.

Why: Entering a dungeon or traveling to another continent changes which server is responsible → Effect: Save → transfer → load, with a wait if the target server is busy or has no free dungeon instance → On screen: Long loading screens, failed entry, disconnects mid-transfer

Symptoms: Can’t connect / infinite loading, Freeze, Disconnect, Rubber-banding · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Cascading failure Cascading failure

When one service slows down, the servers that call it get tied up waiting for responses, and even unrelated features stop.

Why: One service, such as the DB or authentication, slows down → Effect: Threads and connections on the calling servers are tied up waiting for responses, and retries of failed requests add more load → On screen: Everything slows down or stops, even features that look unrelated

Symptoms: Freeze, Input lag, Can’t connect / infinite loading · Primary owner Game team (Server development) · Also Infra team (Network infrastructure)

Auxiliary server outage Auxiliary service outage

When a server that runs separately from the game server, such as chat, party, or auction house, fails, only that feature stops working.

Why: A server dedicated to one feature slows down or dies → Effect: Only requests for that feature get no response → On screen: Chat doesn’t work, party invites do nothing, the marketplace loads forever (combat is fine)

Symptoms: Dropped action / rollback, Can’t connect / infinite loading · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Deploys and restarts Deploy / rolling restart

If you restart a server for an update without moving its connections, everyone on it disconnects, and the final saves before shutdown and the reconnects all hit at once.

Why: Servers restart one after another to roll out a hotfix → Effect: Each server shuts down without moving its connections, and saves for every player on it hit the DB at once → On screen: Disconnects without notice, a surge of reconnects

Symptoms: Disconnect, Can’t connect / infinite loading, Input lag · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Autoscaling delay Autoscaling lag

When players flood in, servers are added automatically, but getting them ready takes several minutes, and the existing servers are overloaded in the meantime.

Why: Connections spike when an event starts → Effect: Several minutes pass before new servers boot and are ready → On screen: Slow motion, and players who can’t connect, for the first few minutes after the event starts

Symptoms: Slow motion, Can’t connect / infinite loading · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

Logging and monitoring overload Logging / monitoring overhead

During an outage, log volume explodes, and servers that ship logs synchronously get even slower because of the logging.

Why: Errors make log and metric volume explode → Effect: The log collector falls behind, and servers that send synchronously wait on it → On screen: Stutter and freezes during an outage get worse because of logging

Symptoms: Stutter, Freeze · Primary owner Game team (Server development) · Also Infra team (Server infrastructure)

Clock skew between servers Clock skew between servers

When each server’s clock is slightly off from the others, cooldown, buff, and event start checks disagree from server to server.

Why: A server whose time sync stopped drifts hundreds of ms to several seconds away from the other servers → Effect: Passing absolute times, such as when a buff ends, between servers makes their checks disagree → On screen: A buff disappears or a cooldown starts over after moving to another server

Symptoms: Dropped action / rollback · Primary owner Infra team (Server infrastructure) · Also Game team (Server development)

Too many macros and bots Bots and macros

Bots send requests far more often than people do and eat into server capacity.

Why: Large numbers of bots connected, farming, moving, and trading nonstop → Effect: Server processing load and DB load go up → On screen: A specific farming spot or the whole server slows down (slow motion, input lag)

Symptoms: Slow motion, Input lag · Primary owner Game team (Server development) · Also Infra team (Network infrastructure)

External service dependency External dependencies (auth, billing, platform)

When an external service such as platform login, payments, or identity verification is slow or down, players get stuck at that step.

Why: An external authentication or payment service is down or slow → Effect: That step waits for a response → On screen: Can’t log in, payments fail. Players already in the game are fine

Symptoms: Can’t connect / infinite loading, Dropped action / rollback · Primary owner External (External) · Also Game team (Server development)

Matchmaking and region assignment errors Wrong region assignment (matchmaking / GeoDNS)

When a player lands on a server in a distant region while a closer region exists, that player’s ping stays high even though their connection is fine.

Why: Bad GeoIP data, a VPN, assigning a whole party by the members’ average ping, rules that widen the search to distant regions when there aren’t enough players, assignment by DNS resolver location → Effect: The player connects to a server across the ocean even though a nearby region exists → On screen: In a game with servers in several regions, only you (or only your party) always have high ping, with input lag, rubber-banding, and skills that don’t go off

Symptoms: Input lag, Rubber-banding, Dropped action / rollback · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Network infrastructure), External (External)

Expired or misconfigured TLS certificate TLS certificate expiry / misconfiguration

When the certificate on a login, API, or patch server expires or is missing its intermediate certificate, every client that connects from that moment on fails the TLS connection.

Why: The certificate is past its validity period, the server sends it without the intermediate certificate, or the date and time on the player’s device are wrong → Effect: The client fails certificate validation and drops the TLS connection → On screen: Can’t connect / infinite loading at the login or patch step, or only HTTPS features such as the store fail. Players already connected are usually fine

Symptoms: Can’t connect / infinite loading, Dropped action / rollback · Primary owner Infra team (Network infrastructure) · Also Infra team (Server infrastructure), Game team (Client development)

Login queue cap and insufficient reconnect grace Login queue cap / no reconnect grace

When players flood in right after launch or maintenance, the login queue hits its cap and turns away new arrivals, and players who were already waiting lose their place during a brief disconnect and go back to the end of the line.

Why: More people try to connect than the login server can take at once, so it keeps a queue, and when the queue gets too long it refuses new entries to protect the server → Effect: The longer the queue, the longer the wait, and a brief Wi-Fi or mobile network drop during that time costs the player their place → On screen: Can’t connect / infinite loading, the game quits with an error while waiting, the player starts over at the back of the line

Symptoms: Can’t connect / infinite loading, Disconnect · Primary owner Game team (Server development) · Also Game team (Client development), Infra team (Server infrastructure)

T1Tools

Triage helper

When a lag report comes in, pick just three things: “who, when, and what it looks like.” The helper shows the causes in this white paper that fit best, ranked by score. It isn’t a definitive diagnosis, but it’s enough to decide which team to ask first.

T2Tools

Diagnosing from monitoring data

When a report or alert comes in, narrow it down in the order scope → timing → layer. Where the anomaly is concentrated matters most in deciding the owner, what it coincides with narrows down the cause, and which layer’s metrics look abnormal confirms it. Combine this with the “On the graph” and “How to confirm” sections on every cause card, and you can pick candidates from the shape of a graph and go straight to the place to check.

Decision flow

1 Scope

Where is the anomaly concentrated?

  • Specific country or ISP (ASN) → Infra teamNetwork ExternalISP
  • Specific server, channel, or zone → Game teamServer if host metrics are normal, Infra teamServers/OS if not
  • Specific OS, device, or build → Game teamClient
  • One player or household → ExternalPlayer environment (Game teamClient if several players show the same pattern)
  • Everyone at once → shared resources (DB, load balancer, gateway) or a deploy that just went out
2 Timing

What does it coincide with?

3 Layer

Which layer’s metrics are abnormal?

  1. Network: RTT, loss, retransmission rate, interface errors and discards
  2. Host: per-core CPU, CPU steal, softirq, NIC discards, memory pressure
  3. Game server: tick time, socket receive queue (Recv-Q), per-thread CPU, GC logs
  4. DB: query latency, lock waits, replication lag
  5. Client: frame time, net graph, crash reports

Signal table

What to checkIf it looks like thisWho to call first
Server socket receive queue (Recv-Q)Builds up because the server process can’t read it in timeGame teamServer (tick stalls, GC, locks)
Per-connection retransmissions and RTTOnly some connections, clustered on specific ASNsInfra teamNetwork ExternalISPs/player connections
Every connection on one hostInfra teamServers/OS (NIC, kernel)
Server-wide retransmissions and bandwidth up right after a patchPacket size or frequency changedGame teamServer Infra teamNetwork (MTU, limits)
CPU steal, throttling, softirq, NIC discardsRisingInfra teamServers/OS
One thread at 100%, run queue latency, GC pausesRisingGame teamServer
DB latency up, query count unchangedIOPS, locks, other jobsInfra teamDB hosts
DB query count or shape changed after a patchN+1, new queriesGame teamServer
Synthetic monitoring from overseas locations (RTT, loss)BadInfra teamNetwork ExternalISP
Synthetic monitoring is normal, but players see bad numbersPlayer environment or clientExternalPlayer environment Game teamClient
Distribution of disconnect reasonsHeartbeat timeouts↑ / RSTs↑ / server-initiated disconnects↑NAT or path / equipment / server
Exact intervals (on the hour, every N minutes)Scheduled jobs, backups, GC, eventsWhoever owns that schedule

Find causes by graph shape

Just knowing the shape of a monitoring graph narrows the candidates a lot. For each of the 13 shapes below, we collected the causes that produce it. The small pictures on the cause cards use the same shapes. The solid line is the main metric to watch, the dashed line is a metric to watch alongside it (player count, waits, errors, and so on), and the faint dashed line is the normal level.

Random spikes

Spikes at irregular intervals, then quickly returns to normal.

Frame time spike, Excessive extrapolation (dead reckoning), Client-side prediction mismatch, Fixed-timestep catch-up spiral, Background processes taking up CPU, Client low on memory and swapping, NIC power saving and driver issues, Wi-Fi interference and weak signal, Frequent 5G↔LTE switching (at 5G coverage edges), Poor line quality, Ring buffer too small, Virtualization overhead and noisy neighbors, Kernel socket buffers too small, CPU steal (virtual machines), Stalls from memory reclaim and compaction, System clock jump (NTP step), Blocking sends caused by slow clients, Reliable UDP retransmission settings, Last data lost to an abortive close (RST), Blocking calls on the game thread, Synchronous log writes, Server-side lazy loading, DB deadlock, Slow Redis commands, Logging and monitoring overload, Lockstep waiting on the slowest player, Rollback netcode misprediction, Events played on arrival without timestamps, Overly strict server validation, Pathfinding mismatch in command sync, AOI registration race, Lost baseline snapshot, Missed despawn message (ghost entity), Entity ID reuse mix-up, Spurious retransmission from latency spikes

Step change

Steps up one level at a specific moment, such as a patch, config change, or route change, and stays there.

Submarine cable / international link outage, BGP route changes and convergence, DDoS protection detours and false positives, Performance changes after OS, kernel, driver, or firmware updates, Patch changes the traffic pattern, Queries with no index, Query slowdown from a query plan change, Schema change (DDL) lock during live service, External service dependency, Route change / bad ECMP path

Rises with load

Climbs as concurrent users or the crowd in one spot grows, and more steeply than the head count itself.

Rendering load from large crowds, Packet processing bottleneck on the main thread, Receive buffer overflow, Other apps on the same device using up bandwidth, Bufferbloat (router queue), Switch microbursts, Too many threads and context switching, Container CPU throttling (CFS quota), IP fragmentation of UDP packets, Blocking I/O design, Tick overrun, AOI calculation blowup (N²), Broadcast fan-out overload, Single-threaded zone overload (hotspot), Lock contention, Pathfinding storm, Serialization and compression cost, Combat concentrated on one target (world boss), Allocation surge, Hot row lock contention, Replication lag, Routing through a gateway or proxy, Zone transfer (handoff between servers), Per-connection send budget and priority, Send bursts overflow shallow buffers, Duplex mismatch

Hits a ceiling

Throughput or connection count reaches a value and can’t go higher; from then on, waits and errors pile up.

Out of graphics memory (VRAM), Underpowered or overheating router, ISP throttling and traffic management, Shared links saturated by DDoS, Firewall session table full, Cloud NAT gateway connection and port limits, Data center link saturation, NIC interrupts concentrated on one core, Cloud PPS limit exceeded, NIC bandwidth saturation, File descriptor limit, Server conntrack table full, Ephemeral port exhaustion on server-to-server connections, Message queue backlog, Thread pool exhaustion, GC thrashing (too little heap headroom), Cloud disk out of burst credits, IOPS limit / queue saturation, Connection pool exhaustion, Cascading failure, Login queue cap and insufficient reconnect grace, Streaming failure from memory or VRAM shortage, Policer drops excess traffic, Packet drops on the receiving host, Firewall and connection tracking drops, Middlebox over capacity (firewall, IPS, DDoS protection)

Always high

Stays high without spiking. This comes from structure: distance, routing, or design.

Missing or too-short interpolation buffer, V-Sync and the render queue, Timer resolution, Display, input device, and frame generation latency, Propagation delay (physical distance), Detour routing, Excessive interrupt coalescing, GRO/LRO batching delay, Latency spikes from server power management (C-states, frequency scaling), Nagle’s algorithm + delayed ACK, Cache miss, HDD seek latency, Feedback only after the server responds (request-response), Chatty protocol (many sequential round trips), No skill input buffering, Client authority, Double tick wait, Low snapshot send rate, Spurious fast retransmit from reordering, RTO settings that don’t fit the environment, Middlebox strips TCP options

Outliers only

Most are normal, while specific players, regions, ISPs, or devices are high on their own.

Slow storage delays asset streaming, Client crash, Packet inspection by security software, Overlay software interference, RRC state transition delay (mobile radio power saving), Weak mobile signal and dead zones, Public Wi-Fi and corporate network restrictions, Satellite internet (LEO/GEO), One faulty ECMP path, Country- or ISP-level UDP restrictions and packet inspection, DNS failures and delays, Routing through a VPN or game booster, Load balancer skew and misjudged health checks, Bad cables and port errors, MTU mismatch (only large packets vanish), Slow client (slow consumer) handling policy, Keepalive default of 2 hours, Slow start after idle, Uneven SO_REUSEPORT distribution, Remote NUMA memory, Too many macros and bots, Matchmaking and region assignment errors, Short timing windows eaten up by ping, Hit registration without lag compensation, Too much lag compensation, Player-hosted server (host), Server rejects after client-side feedback, A lagging player moves in bursts on others’ screens, Fast-forward on servers that process on arrival, Per-player input buffer size, One lagging party member and boss mechanics, A lagging client controls the monster, Bloated data on one character, Different channel, instance, or phase, Spawn messages dropped during loading, Fixed UDP port collision, Sessions keyed by IP or device (bug), Multi-client restriction, Simultaneous access to cache or asset files, Different display settings, Client version or data mismatch, Entities held back by clock estimate error, Wireless link loss, Physical errors (bad cable, optics, connectors), MTU black hole (only large packets keep getting lost), Late or lost ACKs (saturated upload)

Mass disconnect

The connection count plunges, or the disconnect count shoots up in an instant.

Mobile app sent to the background, Wi-Fi ↔ LTE/5G switching, NAT mapping expiry, ISP-shared IP addresses (CGNAT), Load balancer idle timeout, Cloud security group connection tracking expiry, Network equipment failover, OOM killer, WSAECONNRESET errors on Windows UDP sockets, Deadlock, Server crash, Infinite loops and runaway logic, Writing a core dump, DB failover, Lost progress from a long save interval, Auxiliary server outage, Deploys and restarts, Expired or misconfigured TLS certificate, NAT or load balancer mapping expires mid-connection

How much can you confirm without game code?

We counted the easiest way to confirm each cause. Infra tools means it can be confirmed with OS, network, cloud, and DB tools plus runtime startup options (GC logs and so on), with no changes to game code. Game logs/metrics means things you can only see if the game records them, like tick time and disconnect reasons. The more a layer depends on game logs and metrics, the stronger the case for asking the game team to add instrumentation.

Reading the numbers

Averages hide spikes. On a server running 20 ticks per second, if just 1% of ticks are slow, everyone hitches about once every 5 seconds, yet the average tick time barely moves. That’s why you look at percentiles too. p50 (the median) is the value half the samples come in under, and p99 is roughly the slowest 1 in 100. What players remember as “lag” is usually the p99 end.

The aggregation interval hides spikes too. On a 1-minute average graph, a 1-second freeze is diluted to 1/60. When you’re hunting for freezes, also look at the maximum or p99 of the same graph, or at a shorter interval.

Jitter is how much the gaps between arriving packets vary. Even with a low average ping, high jitter drains the interpolation buffer and causes stutter and teleporting.

MethodWhat it measuresWatch out for
ping (ICMP)Round-trip time to the deviceRouters and servers may process ICMP replies late or rate-limit them, so results can differ from game packets. If ICMP is blocked, there’s no reply at all
mtr·tracerouteLatency and loss per hopIf only one device in the middle shows high loss and the hops after it are normal, that device is most likely just limiting its ICMP replies. Only loss that carries through to the end is real loss
TCP RTT (rtt in ss -ti)Round-trip time the kernel measures for each connectionThe most trustworthy, since it comes from the actual game connection. Visible per player on the server side
In-game pingRound-trip time the game measures with its own messagesIf measured inside the game loop, frame and tick waits get mixed in. Rises when the server or PC is busy, even if the connection is fine

What you can do right now, and what to add to game code

Without game code
  • Add dimensions: tag the client IPs in connection and load balancer logs with country and ISP (ASN) so that “only overseas” or “only one ISP” becomes visible.
  • Connection quality: collect per-connection RTT and retransmissions on the server with ss -ti or eBPF tools and view them by ASN.
  • See inside the server from outside: socket queues, per-thread CPU (pidstat -t), run queue latency, and GC logs you can turn on with startup options alone.
  • Path measurement: synthetic monitoring from the target country and ISP (RIPE Atlas, probe servers in cloud regions) and mtr.
  • Change log: mark deploys, patches, config changes, and network work as vertical lines on every graph. It’s the starting point for judging whether something began “after the patch.”
Minimal game code
  • Client summary reports: every 30–60 seconds, RTT p50 and p95, jitter, loss, FPS, frame spike count, build, server and channel.
  • Server tick metrics: tick time p50 and p99, tick overrun count, player count per zone, per-connection send queue.
  • Disconnect reason codes: heartbeat timeout, RST, server-initiated disconnect, authentication failure, and maintenance, using the same codes on both sides.
  • Session IDs and timestamps: session, character, and server IDs plus synchronized UTC time in every log.
  • Report lag button: uploads the last 60 seconds of RTT, FPS, and tick gaps along with the session ID.
T3Tools

Incidents and playbooks

Step-by-step checks for two situations you’ll run into often, plus real outages that the original developers and operators published themselves. Every step and every incident links to the related cause cards.

Playbooks

Lag after a patch

When lag reports increase after a specific patch or deployment. Use it when reports like “it’s been weird since this update” pile up, or when a graph steps up at some point and stays there.

  1. Pin down the start time and gather every change around it: Find when reports first spiked and when the graph stepped up, and list every change that went out around that time. Cover client patches, server deployments, configuration changes, DB schema changes (DDL) and restarts, network and firewall work, and infrastructure swaps (instance type, kernel, drivers). If you use your monitoring tool’s annotation feature to mark every deployment with a vertical line on all graphs, this step goes quickly. If a game patch and infrastructure work went out in the same maintenance window, keep both as suspects. Who to call first: the game team and the infra team, for the changes each of them shipped.
  2. Break down the scope: build, device, server, region: Find the dimension where the problem clusters. Suspect the client first if only players on the new build are affected; client performance or drivers if only certain OSes, graphics cards, or devices are; the server if only certain servers, channels, or zones are; the network path if only certain countries or ISPs are; and shared resources (DB, load balancers, gateways) or the server deployment that just went out if everyone is affected at once. If client telemetry includes the build number, put ping, FPS, frame spikes, and disconnect counts for the old and new builds side by side. If ping is unchanged and only FPS got worse, client performance is a likelier suspect than the network. Who to call first: the game team (client) if it clusters by build or device; for servers or channels, the game team (server) if host metrics look normal, or the infra team (servers/OS) if they don’t; the infra team (network) if it clusters by country or ISP.
  3. Compare the new and old versions over the same time window: A plain before-and-after comparison mixes in changes from the day of the week, time of day, and events, which muddies the picture. If possible, roll the new version out to a few servers first (canary), and compare tick time p50 and p99, tick overrun count, CPU, memory, and error rate side by side with old-version servers (the control group) over the same time window. If it’s already deployed everywhere, compare against the same day and time last week. A server-wide average hides problems on individual servers or zones, so break it down by server and zone. Who to call first: the game team (server).
  4. Compare the traffic fingerprint before and after: Even without knowing the server code, you can tell from values visible on the network whether the patch changed the shape of the traffic. Compare before and after: packets per second (pps) and bytes per player, average and maximum packet size, connection count, and the size of the send burst that goes out all at once each tick. If UDP packets have started to exceed the path MTU (usually 1,500 bytes), IP fragmentation occurs. Losing a single fragment loses the whole packet, and some NATs and firewalls drop fragments outright. Players whose path crosses a segment with a smaller MTU (a tunnel or VPN) lose only the large packets. If pps went up, check whether you’re hitting the cloud instance’s PPS limit or the throughput limit of firewalls or DDoS protection appliances. Who to call first: if the fingerprint changed, the game team (server), with the evidence attached; if the fingerprint is the same and only loss and retransmissions went up, the infra team (network).
  5. Compare DB query types and counts before and after: If DB latency went up, start by checking whether query volume (QPS) went up with it. PostgreSQL’s pg_stat_statements and the digest summaries in MySQL Performance Schema group queries that differ only in their values into one entry and track execution count and total time. Comparing the top query lists from before and after the patch reveals new queries, queries whose count multiplied (N+1), and queries that read the whole table without an index (the SUM_NO_INDEX_USED column in MySQL). Who to call first: the game team (server) if QPS or the shape of the queries changed; the infra team (DB: query plans, IOPS, locks) if the queries are the same and only latency went up.
  6. Use host and server process metrics to find the layer: Use values visible from the OS, without needing the game code, to tell problems inside the server process from problems on the host. If the receive queue (Recv-Q) on the server socket is building up, the server process isn’t reading in time (tick stalls, GC, locks). If one thread alone is at 100%, it’s a single-thread bottleneck. If pause times in the GC log went up, the memory usage pattern changed. Also check whether the build shipped with a higher log level and log writes went up. Conversely, if CPU steal, throttling, or NIC drops went up, look at infrastructure that changed at the same time (instance type, kernel, container limits). Who to call first: the game team (server) for signals inside the process; the infra team (servers/OS) for host signals.
  7. Roll back to confirm, then record the result: Revert the most likely change for only some servers or some players (roll back, or turn off a feature flag), or set the configuration back to its previous value, and see whether the symptom goes away with it. If only the reverted side improves, the cause is confirmed. Reverting can itself cause a brief slowdown from restarts and cold caches, so if it isn’t urgent, do it during off-peak hours. Record the result in the incident log along with the cause ID, and add limits on packet size, query count, and tick time to the pre-deployment checklist for the next patch. Who to call first: the team that made the change.

Launching in a new country or region

When you launch the service in a new country or add a new region or data center. Use it both for pre-launch checks and for sorting out reports like “it’s fine back home, but players in the new country are lagging.”

  1. Measure path quality for each local ISP before launch: For each major ISP (ASN) in the target country, measure the round-trip time (RTT) distribution, jitter, and packet loss to each candidate game server location. A single average hides the differences between ISPs, so look at the median and 95th percentile per ISP, separately for the evening peak and the early morning hours. The public measurement network RIPE Atlas lets you pick countries and ASNs and send ping and traceroute from probes around the world, or you can spin up temporary VMs in the candidate regions and measure from there. Devices along the path sometimes rate-limit ICMP responses, so when possible, also measure with the same protocol and port the game uses. If one ISP’s traffic stands out by passing through distant cities, it’s a peering or routing problem. ISPs choose lower-cost routes even when lower-latency ones exist, so even nearby destinations can end up taking long detours. Who to call first: the infra team (network); external (ISP, IX) if the route problem is on the ISP side.
  2. Compare the measurements with the limits the game design can tolerate: Compare the measured RTT and jitter with the game’s timing windows (reaction times for dodges, parries, and so on), lag compensation limit, interpolation buffer length, and input buffer size. For example, if the parry window is 0.2 s, players on ISPs where round-trip delay plus the interpolation buffer adds up to more than that will be late even when they react in time. Widen lag compensation to make up for it, and now the players on the receiving end start reporting “I got hit behind a wall.” If many ISPs exceed the limits, the infra team should look at placing regions or edge PoPs closer, and the game team should review the timing window, interpolation, and lag compensation values. The white paper’s chapter “Same ping, different feel: netcode models” serves as the reference table. Who to call first: the game team (server and client: design limits) and the infra team (network: region and PoP locations).
  3. Check MTU and whether UDP gets through: Check that the game’s largest packets make it through local networks intact. Measure the path MTU by sending pings of various sizes with the Don’t Fragment (DF) bit set, and look for segments smaller than 1,500 bytes, such as PPPoE, tunnels, and mobile networks. The standard for datagram transports such as UDP (RFC 8899) recommends 1,200 bytes as the base size that can cross most paths on IPv4, so if the game’s largest packet is bigger than that, decide with the game team whether to shrink it or split it up. Also check whether UDP or the game’s ports are blocked or throttled on public Wi-Fi, corporate networks, or certain ISPs, and whether there’s a fallback path (TCP, port 443) for when they are. Who to call first: the infra team (network) and the game team (server: packet size).
  4. Measure NAT and CGNAT idle timeouts and set the heartbeat interval to match: Measure how long local home routers and mobile networks (CGNAT) keep the mapping for an idle UDP connection before deleting it. For each trial, send one packet from a test device to the server to create the mapping. The device then sends nothing, and the server sends a packet back to the device after a set wait (30 s, 60 s, 120 s …). The wait at which the device stops receiving that packet is the network’s idle timeout. The standard (RFC 4787) says UDP mappings must not expire in less than 2 minutes and recommends a default of 5 minutes or more, but values vary widely between devices, and some delete mappings sooner. Only outbound packets from the device reliably refresh the mapping, so have the client send the heartbeat, and check that its interval is at most half of the shortest of these: the measured value and the idle timeouts of the load balancer and cloud security groups. Who to call first: the game team (client: heartbeat interval; server: timeout values) and the infra team (load balancer and security group settings).
  5. Check the external services and security appliances on the local path: Check that local platform login, payment, and identity verification respond at normal speed, that local DNS resolves the login and patch server addresses correctly, and that the CDN serves patches from locations close to that country. Make sure the new country’s IP ranges aren’t caught by country-blocking rules or rate limits in DDoS protection and firewalls. In particular, make sure CGNAT ranges, where many subscribers share one IP, don’t get blocked wholesale. Who to call first: the infra team (security appliances, DNS, CDN) and external parties (platforms, payment providers, ISPs).
  6. After launch, break the data down by country and ASN: Tag client IPs in connection logs and load balancer logs with country and ASN, and look at RTT, retransmissions, and disconnect counts and reasons (heartbeat timeout, RST, server kick) by country and ISP. Free databases such as MaxMind GeoLite ASN map IPs to an ASN and organization name. To comply with local privacy rules, store IPs truncated to /24 or reduced to the ASN. Look first at that ISP’s route if problems concentrate in one ASN (infra team, external); at distance and design limits if the whole new country is bad (infra team, game team); and at peering congestion if it only gets worse in the evening. If some players always have high ping, check with the game team (server) whether they’re being assigned to a distant region because of GeoIP errors, VPNs, or region assignment based on the party leader. If synthetic monitoring looks normal and only players are having problems, it points to the player’s environment or the client.
  7. Check how distant players affect everyone else: When more players connect from far away, the damage doesn’t stop at their own screens. A slow player’s inputs arrive in bunches, so on other players’ screens that one character moves in fast-forward, and it trips the server’s speed and cooldown checks, causing rubber-banding or rejected skills. In party mechanics, one slow player’s late reaction can fail the whole party, and in lockstep, everyone waits for the slowest player. Check whether reports of “one character looks off” from existing players went up after the new country opened, and work out input buffers, validation tolerances, and separate matchmaking regions with the game team. Who to call first: the game team (server).

Real-world incidents

We picked only postmortems that game companies and infrastructure companies published themselves. The summaries stay within what the original posts disclose; see the original for the full account.

CCP Games 2014: Server overload in EVE Online’s massive HED-GP fleet battle

During the massive fleet battle in the HED-GP system, covered in a January 2014 retrospective, the server was badly overloaded. Even after Time Dilation (a feature that slows game time under overload) hit its 10% floor and the whole battlefield went into slow motion, load kept piling up. The backlog in processing module deactivations and repeat cycles (Dogma Lateness) peaked at 193 seconds of game time, about 32 minutes in real time. The July 2013 battle in 6VDT, of nearly the same size, peaked at 42 seconds (about 7 minutes in real time). CCP cautioned that it couldn’t be certain, because its profiling tools add load of their own and aren’t run in situations like this, and named two likely causes. The first was unprocessed load that kept building up as the battle dragged on. The second was heavier drone use: unique drones deployed during the battle went from 21,123 in 6VDT to 38,852 in HED-GP, 84% more. Telling everyone in view about each player’s actions takes traffic that grows with the square of the player count (O(n²)), and drones generate more messages per attack. The code drones use to pick targets also often scans every attackable target on the same battlefield, so its cost grows close to n².

When the processing load in one crowded area exceeds its limit, the whole area goes into slow motion, and the longer the battle runs, the more the backlog grows and the worse the input lag gets. Signals to check: tick time and backlog on the server (node) handling that area, plus player and entity counts. A telltale sign is that other areas stay fine. The primary owner is the game team (server), and the things to fix are how many recipients each action is sent to and the cost of AI target searches. Slowing game time can’t eliminate the overload, but it slows everyone down at the same rate, which keeps a subset of actions from falling behind indefinitely. Original post

Riot Games 2015: League of Legends traffic on roundabout routes, and Riot Direct

A technical post in which Riot Games explains why the internet is a poor fit for real-time games. Real traffic reported by a League of Legends player should have gone straight from San Francisco to Portland, but it went through Los Angeles, Denver, and Seattle, taking 70 ms for a trip that would take 14 ms on a direct path. Riot explained that when routers overflow and drop packets, other champions appear to jump around the screen and projectiles seem to teleport. Riot pointed to routes and routers. Backbone providers and ISPs send traffic along the cheapest path even when a lower-latency path exists, and when the route BGP settles on takes a long detour, traffic also passes through more routers. A router’s processing load depends on the number of packets, whatever their size. Game packets are around 55 bytes, so the same amount of data takes 27 times as many packets as it would in 1,500-byte packets, and fills router input buffers that much faster. According to Riot, many routers drop UDP packets first when they’re overloaded. As a fix, Riot built its own network, Riot Direct, with routers at 10 major internet hubs in the US and direct connections (peering) with as many ISPs as possible. According to Part II, the share of players with a ping under 80 ms rose from 31% to 50% in a little over 9 months, and hit 80% overnight after the game servers moved to Chicago.

If only customers of one ISP have unusually high ping, even within the same country, suspect the route. Signals to check: the RTT distribution per ISP (ASN) and the cities that show up as hops in traceroute. The primary owner is the infra team (network), and the fixes are direct peering with ISPs, connecting at IXs (internet exchanges), and choosing server locations. Routing policy on the ISP side has to be worked out with the external party (the ISP). The case also shows that simply moving servers closer to the center of the player base makes a big difference. Original post

Riot Games 2020: Edge host overload on League of Legends servers in Europe and Brazil

In late February 2020, the League of Legends EUW, EUNE, and BR servers had several outages, and the number of new games starting dropped sharply. Backend services such as matchmaking and game servers all reported healthy, yet almost no traffic was coming in. Riot pushed the tournament mode (Clash) back a week to avoid launching it on clusters that might be unstable. The postmortem doesn’t say how long each outage lasted. Three things came together. Requests to one service were malformed, so in certain cases they kept failing and being retried, and request volume exploded. A known compatibility problem between the container system and the OS version was leaking memory inside the OS. The OS upgrade was finished on only about 60% of Riot’s entire container environment and was still in progress on the Europe and Latin America clusters. Edge containers, which receive internet traffic, filter it, and pass it to the backend, were kept apart within a shard (server group), but nothing kept different shards apart, so in every outage edge containers from at least three shards were packed onto a single host. The retry surge landed on that host, and the memory leak brought it to a halt.

When every backend service reports “healthy, but no traffic is coming in,” look at what sits in front of them (edge, gateways, load balancers). Signals to check: inbound connection counts skewed toward particular hosts, and the failure and retry rate of specific requests. The primary owner is the game team (server: the malformed request and the retry behavior), and the infra team (servers/OS) handles container placement rules, OS upgrades, and skew alerts. Riot fixed the request code, changed retries so they wouldn’t spike, and put skew alerts in place until it could implement spreading across shards. Original post

Riot Games 2021: League of Legends EUW 5-hour outage: one auxiliary DB halted the whole server

On January 22, 2021, the League of Legends EUW server didn’t work properly for a little over 5 hours. The metrics for logged-in players and players in a game cut out at the same moment, and between the two restarts, logins went up but almost no games started. The primary server of a database behind a non-critical feature had a hardware failure, and that database had no automatic failover to a standby configured. Each database had its own connection pool, but all the pools shared one thread pool; work sent to the failed database never finished and held on to threads, until the whole system ran out of threads. Amid a flood of alerts, the team first suspected a recent malicious network attack and hardware work in another region, so the alert for the failed database wasn’t noticed until about 1 hour later. Because every system ran inside a single JVM, when GC paused the process for several seconds at a time under the reconnect load after the restart, metrics collection also developed large gaps. The login queue also didn’t hold to its configured limit, so players flowed in unevenly.

Even one auxiliary database that nobody considered critical can halt everything through a shared resource such as a thread pool. Signals to check: pending requests per database, thread pool utilization, and a ratio of game starts to logins that is far too low. The owners are the game team (server: thread pool isolation, timeouts) and the infra team (DB: automatic failover). When alerts flood in, it’s easy to suspect whatever hit you recently (an attack, for example) first, so rule things out one at a time in the decision order (scope → timing → layer). After a restart, also check that the login queue actually limits inflow as configured. Original post

Roblox 2021: Roblox 73-hour outage: contention in the service discovery (Consul) cluster

It began on the afternoon of October 28, 2021 (Pacific Time) with high CPU load on one Consul server. At 16:35 the number of players online fell to half of normal, and then the entire service went down. Not until 16:45 on October 31 could all players get back in, 73 hours after the outage began. Roblox said 50 million people use it every day. Roblox uses HashiCorp Consul for service discovery (how services find each other’s addresses), health checks, and a key-value (KV) store, and a single Consul cluster was handling several workloads at once. There were two root causes. First, the day before the outage, Consul’s new streaming feature, which had been rolled out gradually over several months, was turned on for the traffic routing service too, and that service’s node count was raised by 50%. Under very heavy read and write load, the feature caused contention on a single shared resource (a Go channel). The contention was even worse on the dual-socket (NUMA) servers with more cores that were swapped in during the outage. Second, free-page list (freelist) management in BoltDB, which Consul uses to store its Raft log, became pathologically slow and wrote 7.8 MB to disk for every append of 16 kB or less. Median KV write latency, normally under 300 ms, rose to 2 s, and zero windows (full TCP buffers) were seen on the slow leader server. Because telemetry depended on Consul, the metrics needed to find the cause disappeared along with it.

When a foundational system that many services rely on (service discovery, configuration store, authentication) slows down, every feature stops at once. Signals to check: that system’s write latency, leader changes, and CPU, plus any configuration change made just before the outage. Ownership is shared between the game team (server) and the infra team (servers/OS). Keep monitoring separate from the systems it watches, so you can still see metrics during an outage. During recovery, caches are empty, and letting everyone in at once can knock things over again, so Roblox used DNS to control the share of players let in and raised it about 10% at a time. Original post

Square Enix 2021: FINAL FANTASY XIV expansion launch congestion and login queue errors

From the start of early access for the Endwalker expansion in December 2021, every World was extremely congested. Login queues grew long, and Error 2002 appeared often when logging in from the character selection screen or while waiting in the queue. Some Worlds and zones also went down (Error 3001), and queues timed out (Error 4004). As of the December 11 notice, on day 8 of early access, the congestion was still going on. Error 2002 occurs in two cases. The first is when more than 17,000 players are waiting on a logical data center. This cap exists to keep the login server from going down under an overly long queue, and when it’s hit, the client shuts down completely. On December 7, spare development hardware was added to the lobby servers to raise the cap; this error became less common, but the queues actually got longer. The second case is when a waiting player’s connection is unstable. As waits got longer, brief disconnects caused by packet loss on the internet path or unstable Wi-Fi became more common. The lobby server waits somewhere from tens of seconds to about 1 minute for a reconnect. Players who reconnect in time keep their place in the queue, but anyone who takes longer goes to the back of the line. Square Enix said most reports fell into this case. The semiconductor shortage also meant new Worlds couldn’t be added right away.

The longer the queue, the more often a brief connection drop for a waiting player turns into a connection error. Under the same congestion, errors cluster among players on Wi-Fi or unstable connections, so it becomes a problem that hits “only some players.” Signals to check: queue length and wait time, and the share of disconnects that happen while players wait in the queue. The primary owner is the game team (server: the queue cap and the reconnect grace period), and the infra team joins in on adding lobby and World servers. A generous reconnect grace period keeps more of these brief drops from costing players their place in the queue. Original post

Cloudflare 2020: Traffic loss in some cities from a Cloudflare backbone configuration error

Many games rely on CDN providers for their websites, APIs, and DDoS protection, so this is the kind of infrastructure outage that affects games too. For 27 minutes on July 17, 2020, from 21:12 to 21:39 (UTC), traffic across Cloudflare’s entire network dropped by about 50%. The impact was limited to some city locations (PoPs) in the US, Europe, Russia, and Brazil that were connected to the backbone. Other locations were fine. An outage on the Newark–Chicago backbone link congested the Atlanta–Washington link, so an engineer changed a router configuration to take some backbone traffic off Atlanta. The change was supposed to disable an entire policy term, but it disabled only the condition inside it (the prefix-list), so the Atlanta router advertised all its BGP routes across the whole backbone with a higher preference (local-preference 200). Each location gave the routes to its own servers a preference of 100, so traffic from every backbone-connected location was pulled to Atlanta. Atlanta was overloaded, and the affected locations were left with almost no traffic to handle. Service recovered once the Atlanta router was removed from the backbone. Cloudflare stated that the outage was unrelated to any attack or breach.

If players in particular cities or regions all hit disconnects or can’t connect / infinite loading at once while everyone else is fine, first suspect a routing configuration change made just before. On the graph, CPU and traffic spike at one location only, while the affected locations actually drop to close to 0. The primary owner is the infra team (network), or external if the outage is on the provider’s side. Cloudflare decided to cap the number of routes each backbone BGP session can accept (maximum-prefix), and adjusted preferences so that one location can’t pull in traffic meant for other locations. Original post

Fastly 2021: Worldwide Fastly CDN errors

Many games deliver patch files, launchers, and web pages through a CDN, so this is the kind of infrastructure outage that affects games too. Starting at 09:47 (UTC) on June 8, 2021, 85% of Fastly’s network returned errors. Within 49 minutes, 95% of the network was back to normal, and the incident was resolved at 12:35. A software deployment that began on May 12 contained a bug that would trigger when a specific customer configuration met specific conditions. On June 8, a customer pushed a valid configuration change that met those conditions. Fastly detected the problem within 1 minute, and recovery began once it identified and disabled the customer configuration that triggered it. Deployment of the bug fix began at 17:25 the same day.

Code deployed weeks earlier can still cause a global outage in an instant when it meets a rare condition. On the game side, the signals are HTTP error rates for patch, launcher, and web requests rising in every region at the same time, and the CDN provider’s status page. A telltale sign is that game connections already in progress stay fine if they don’t go through the CDN, and only new connections, patch downloads, and web logins are blocked. The primary owner is external (the CDN provider). The game and infra teams should have a fallback ready, such as a second CDN or a path that fetches directly from the origin server. Original post

Meta 2021: Facebook outage: one backbone command took DNS down with it

An infrastructure outage whose lessons apply directly to a game company’s own network and DNS. On October 4, 2021, Facebook (now Meta) services were unreachable worldwide. The backbone linking its data centers went down completely, and Facebook’s DNS servers could no longer be found from the internet. The postmortem doesn’t say how long the outage lasted. During routine maintenance, a command issued to assess global backbone capacity unintentionally took down every connection in the backbone, and an audit tool meant to block commands like this failed to stop it because of a bug. DNS servers at smaller locations are designed to mark themselves unhealthy and withdraw their BGP advertisements when they can’t talk to the data centers, so the DNS servers became unreachable from the internet even though they were still running. The normal access paths and out-of-band access were both down, and internal tools had lost DNS as well, so engineers had to be sent to the data centers in person, and security procedures slowed that down further. By the time of recovery, power draw at each data center had dropped by tens of MW, and bringing everything back at once could put everything from electrical systems to caches at risk, so load was raised in stages.

If can’t connect / infinite loading hits every region and every ISP at the same time, check DNS and BGP routes before the game servers. You can confirm this from outside the company with external DNS lookups and public BGP route data. The primary owner is the infra team (network). Check in advance that your out-of-band access path and the internal tools you’d use in an outage don’t depend on the same DNS and network, and during recovery, raise load in stages so reconnects don’t all arrive at once. Original post

AWS 2021: AWS us-east-1 internal network congestion

Many games run their servers, login, and data on public clouds, so this is the kind of infrastructure outage that affects games too. At 7:30 AM PST on December 7, 2021, the internal network in the Northern Virginia region (us-east-1) became congested. From 7:33 AM, EC2 API errors and latency rose, making it hard to launch new instances (instance launches recovered at 2:40 PM), followed by console login failures, Route 53 configuration changes being blocked, and delayed and partly lost CloudWatch metrics. The network devices fully recovered at 2:22 PM. EC2 instances that were already running and existing DNS responses were not affected. An automated activity to scale capacity for a service in the main network triggered unexpected behavior from a large number of clients in the internal network, causing a surge in connection attempts. The devices linking the internal network to the main network were overwhelmed and communication slowed down. The delays in turn drove more connection attempts and retries, so the congestion persisted. The clients had backoff behavior that spaces out requests during congestion like this, but a latent defect kept it from working properly. Internal monitoring depended on the same network, so the operations team had to respond using logs, without real-time metrics.

When retries can’t back off, a brief bout of congestion turns into an outage that lasts hours. From the game’s point of view, game servers that are already running may be fine, but new server capacity (autoscaling), any login, matchmaking, or payment flow that calls cloud APIs, and monitoring can all be blocked at the same time. Signals to check: the cloud provider’s status page, cloud API error rates, and instance launch failures. The primary owner is external (the cloud provider). The game team should give every retry exponential backoff with randomized delays and a retry limit, and the infra team should keep enough spare capacity to ride out blocked scale-out, plus a fallback in another region. Original post

Cloudflare 2025: Cloudflare public DNS 1.1.1.1 outage

An outage of a public DNS resolver that players set up themselves on their devices or routers. This type of outage blocks every game and service at once, but only for players using that setting. For 62 minutes on July 14, 2025, from 21:52 to 22:54 (UTC), the 1.1.1.1 resolver stopped responding worldwide. Cloudflare said that for many users, this meant being unable to use essentially any internet service. Queries over UDP, TCP, and DNS over TLS were affected, while DNS over HTTPS, which connects by domain name, stayed relatively stable. On June 6, while a service topology (the configuration that decides which locations advertise which IP ranges) was being prepared for a different service not yet in use, the 1.1.1.1 resolver’s IP ranges were mistakenly included in it. When that service’s configuration was changed on July 14, the locations advertising the resolver ranges shrank from every location to a single offline one, and the BGP routes were withdrawn worldwide. The change skipped canary deployment and went straight to every data center. Reverting the configuration at 22:20 brought traffic back to about 77%, but in the meantime about 23% of edge servers had lost required IP configuration, and re-applying it meant service wasn’t back to normal until 22:54. Cloudflare said it was an internal configuration error unrelated to any attack or BGP hijack.

If the game servers and other players are fine but some players get can’t connect / infinite loading on the login or patch servers, suspect the DNS those players use. A telltale sign is that existing sessions stay connected and only new connections fail. Having those players change their DNS settings or look up the server address directly settles it right away. The primary owner is external (the DNS operator or ISP). If the game team (client) reports name resolution failures separately from other errors, customer support can make the call right away. Original post

AWS 2025: AWS us-east-1 DynamoDB DNS outage and long recovery

Many games run their servers, login, and data on public clouds, so this is the kind of infrastructure outage that affects games too. From 11:48 PM on October 19, 2025, to 2:20 PM on October 20 (PDT), the Northern Virginia region was affected in three stages. DynamoDB API errors were elevated until 2:40 AM on the 20th; new EC2 instance launches failed from 2:25 AM to 10:36 AM (connection problems on some new instances cleared up at 1:50 PM); and some Network Load Balancers (NLB) saw more connection errors from 5:30 AM to 2:09 PM. The automation that manages DynamoDB’s DNS had a latent race condition. Among the processes that apply DNS plans in different Availability Zones (DNS Enactors), one that was running unusually late overwrote a newer plan with an old one. Right after that, another Enactor’s cleanup job deleted that old plan, leaving the DNS record for the regional endpoint (dynamodb.us-east-1.amazonaws.com) empty. The automation couldn’t repair this, so it had to be fixed by hand. EC2’s system for managing physical servers depends on DynamoDB, so in the meantime the lease it kept for each physical server expired. After DynamoDB came back, there were so many physical servers that re-establishing the leases timed out before finishing, and retries piled up again, pushing the system into “congestive collapse.” Network configuration for newly launched instances propagated slowly, so NLB health checks flapped between passing and failing, and even healthy nodes were repeatedly removed from DNS and added back.

A DNS record error in one place spreads to the other services that depend on that service, and even after the cause is fixed, backlogged work and flapping health checks stretch recovery out by several more hours. From the game’s side, servers already running may hold up, but new servers can’t launch, so autoscaling stalls, and flapping health checks can make the load balancer pull healthy servers out of rotation. Signals to check: the cloud status page, managed service API error rates, instance launch failures, and the load balancer’s healthy target count. The primary owner is external (the cloud provider). The infra team should cap how many servers can drop out at once on failed health checks and have a fallback in another region ready. Original post

T4Tools

Lag reporting guide

What takes the game team and infra team longest when hunting for a cause is finding out “when, where, and who.” Fill in the items below, and they can jump straight to that moment in the logs and graphs.

T5Tools

Glossary

Words that come up often when talking with the game team and infra team. Type a term in English or Korean in the search box.

Ping Ping, RTT
The time it takes for a signal you send to reach the server and come back (round trip). The ping a game displays sometimes also includes time spent waiting for server processing.
Latency Latency
The time a packet takes to get from where it was sent to where it arrives. It often refers to one direction only, so it’s roughly half the ping.
Jitter Jitter
Variation in packet arrival intervals. Even at the same average ping, high jitter makes the screen stutter.
Packet Packet
A chunk of data sent over the network in one go. Usually 1,500 bytes at most; game updates are tens to hundreds of bytes.
Packet loss Packet loss
Packets that are sent but never arrive. In a game that uses TCP, even 1% loss causes a noticeable hitch every few seconds to ten-odd seconds, while a UDP game with interpolation and redundant input sending can sometimes hide a few percent.
Bandwidth Bandwidth
The maximum amount of data a connection can carry per second (Mbps). It’s a different concept from how quickly data arrives (latency).
Tick Tick
One step in which the server computes the game state. A 20-tick server computes 20 times per second, every 50 ms.
Tick rate Tick rate
How many ticks the server runs per second. Higher means faster response, but server cost and bandwidth go up. To save bandwidth, some games send packets less often than the tick rate.
Tick budget Tick budget
The time limit for finishing one tick. Go over it and the next tick starts late, stretching the tick interval.
FPS Frames per second
How many times per second the screen is drawn. At 60 FPS, each frame gets 16.7 ms.
Frame time Frame time
How long it took to draw one frame. Occasional spiky frames matter more to how a game feels than average FPS.
Snapshot Snapshot
A summary of “the game state right now” that the server sends every tick: position, health, status, and so on. It usually carries only what changed relative to what the receiver already has (delta compression).
Interpolation Interpolation
Drawing smooth motion between two received snapshots. The trade-off is that it shows a moment slightly in the past.
Interpolation buffer Interpolation buffer
Time the client deliberately waits before drawing, so it has something to interpolate. It’s the slack that absorbs jitter and one or two lost packets. It’s usually twice the packet interval (100 ms at 20 updates per second), and some games grow it automatically when jitter increases.
Extrapolation Extrapolation, Dead reckoning
When no new packet has arrived, guessing where an entity is headed from its last velocity and drawing it there. A wrong guess looks like teleporting, so many games extrapolate for only about 0.25 s and then stop (the Source engine default is 0.25 s).
Client-side prediction Client-side prediction
Moving your own character on screen right away, without waiting for server confirmation.
Server reconciliation Reconciliation
When the server’s result arrives, comparing it with the prediction and correcting your character’s position. The client starts from the position the server confirmed and replays the inputs the server hasn’t confirmed yet. A large correction looks like rubber-banding.
Lag compensation Lag compensation
When the server judges an attack, it rewinds to the moment the attacker was seeing and checks whether the attack hit. To keep things fair for the player being hit, the rewind is capped: 0.2–0.25 s is common in competitive shooters, and some games rewind as far as 1 s, like the Source engine default.
Authoritative server Authoritative server
A design in which only the server makes final decisions. It stops cheating, but every outcome needs a round trip to the server, so prediction and client-side feedback hide the wait.
Lockstep Deterministic lockstep
Everyone exchanges only inputs and runs the exact same computation on the same turn. Inputs get a fixed delay, and if anyone’s input is late, everyone waits.
Server-side input buffer Server-side input buffer
A per-player buffer in which the server collects a few inputs and consumes one per tick. Even a player with high jitter looks smooth to others, but that player’s actions are confirmed on the server correspondingly later.
Listen server Listen server
One player’s PC plays the game and acts as the server at the same time. The host has zero ping, but if the host’s connection or PC is slow, everyone lags.
Phasing Phasing
Showing different NPCs and terrain in the same place depending on quest progress. If two characters are at different points in a quest, it’s normal for an NPC to be missing for one of them.
Rollback netcode Rollback netcode (GGPO)
Predicting the opponent’s input to keep the game moving, then rewinding to a past frame and recomputing when the real input turns out different. Common in fighting games. Unrelated to a database rollback.
Input buffering Input buffer, spell queue
Accepting the next input pressed shortly before a cooldown or animation ends, and executing it the moment it ends. This keeps a round trip from sneaking in between combo steps.
Client-side feedback Client-side feedback
Playing animations, sounds, and effects right away, without waiting for server confirmation. Only results that need confirmation, like damage and rewards, wait for the server’s reply. If the server rejects the action, what was shown has to be reverted.
TCP Transmission Control Protocol
A protocol that delivers data in order, with nothing missing. It holds later packets back from the game until a lost packet has been received again.
UDP User Datagram Protocol
A protocol that delivers whatever is sent, with no guarantees. There’s no waiting, but the game has to handle loss and ordering itself.
Reliable UDP Reliable UDP (KCP, ENet…)
An approach that implements just as much retransmission and ordering as needed on top of UDP.
HOL blocking Head-of-line blocking
One stuck item at the front makes everything behind it wait. This is what causes fast-forward on TCP.
RTO Retransmission timeout
The retransmission timer: how long TCP waits before deciding a packet is lost and sending it again. On Linux it’s ping + 200 ms or more, doubling after each failure.
Nagle’s algorithm Nagle’s algorithm
A TCP feature that holds small pieces of data until the acknowledgment (ACK) for earlier data arrives, then sends them together to save packets. Games should usually turn it off.
TCP_NODELAY TCP_NODELAY
The socket option that turns off Nagle’s algorithm. Small messages go out immediately.
Delayed ACK Delayed ACK
Sending the acknowledgment of receipt a little later, bundled with other data. Linux typically waits 40 ms (up to 200 ms); older Windows versions wait 200 ms and current ones 40 ms.
Socket buffer SO_SNDBUF / SO_RCVBUF
The size of the send and receive queues the OS keeps for each socket. Too small and they overflow; too large and stale data piles up and waits.
keepalive SO_KEEPALIVE
A TCP feature that checks whether an idle connection is still alive. It’s off by default, and even when turned on, the default is to check only after 2 hours.
RST TCP reset
A TCP signal that kills a connection on the spot. Any data not yet sent is discarded.
Heartbeat Heartbeat
An “I’m alive” signal the game itself sends periodically. It’s used to detect dead connections and to keep connections open on devices along the path.
Timeout Timeout
How long to wait for a response before treating it as a failure. Too short causes false alarms; too long delays detection.
NAT Network Address Translation
A router feature that sends traffic from several devices at home out through one public IP, recording each connection in the NAT table.
CGNAT Carrier-grade NAT
Large-scale NAT in which an ISP shares one IP address among many subscribers.
MTU Maximum Transmission Unit
The largest packet that can be sent in one piece. Usually 1,500 bytes, and smaller on VPN and PPPoE segments.
Bufferbloat Bufferbloat
Network equipment letting its queues grow so large that latency climbs to hundreds of ms.
SQM Smart Queue Management (fq_codel, CAKE)
A router feature that keeps queues short and sends traffic out fairly per flow. The fix for bufferbloat.
QoS Quality of Service
A feature that gives important traffic priority so it goes out first.
Peering Peering
The points where ISPs connect their networks to each other. They tend to get congested in the evening.
BGP Border Gateway Protocol
The protocol ISPs use to tell each other which routes to take across the internet. When routes change, the path and ping change too.
DDoS Distributed Denial of Service
An attack that floods a service with traffic from many sources to knock it offline.
Scrubbing center DDoS scrubbing center
A DDoS mitigation provider’s site that receives traffic headed for your servers during an attack, filters out the attack, and forwards only legitimate traffic. A distant site makes the route longer.
Firewall Firewall
A device or program that lets through only permitted connections. It tracks connections in a session table.
Load balancer Load balancer
A device that spreads incoming connections across several servers.
Session table Session table, conntrack
The table a device or OS uses to track current connections. Its size is limited.
Microburst Microburst
Traffic piling up in very short windows of 1 ms or less, even though the average is low.
NIC Network Interface Card
A server’s network card.
Ring buffer Ring buffer
The buffer that holds packets the NIC has received until the CPU picks them up. It cycles through a fixed number of slots, and when every slot is full, new packets are dropped.
Interrupt Interrupt
A signal a device sends to tell the CPU that there’s work to handle.
RSS Receive Side Scaling
A NIC feature that spreads incoming packets across several receive queues so multiple CPU cores can process them.
PPS Packets per second
Packets per second. Game servers often hit this limit before they run out of bandwidth.
Kernel Kernel
The core of the operating system. It manages networking, memory, and how CPU time is shared out.
backlog Listen backlog
The queue where new connection requests wait until the server accepts them. When it’s full, Linux silently drops new requests and Windows sends a rejection.
TIME_WAIT TIME_WAIT
The state in which the side that closed a connection first keeps that port pair reserved for a while (60 seconds on Linux) in case late packets arrive.
CPU steal Steal time
Time a virtual machine spent waiting for CPU because the physical host was giving it to other virtual machines. Shown as the st value in top.
CPU throttling CFS throttling
Forcibly pausing a container until the next period once it has used up its CPU quota within a set period (CFS period, usually 100 ms).
File descriptor File descriptor
The number (fd) attached to each file or connection a process opens. There’s a limit on how many a process can have.
Thread Thread
A unit of work that runs independently inside a program. Several threads can run at the same time.
Context switching Context switch
The CPU swapping out the running thread for another one. It has a cost.
Lock Lock, Mutex
A guard that lets only one thread at a time use shared data.
Deadlock Deadlock
A state in which threads each wait for a lock the other holds and stay stuck forever.
Thread pool Thread pool
A set of worker threads created in advance. When all of them are busy, new work waits.
Asynchronous I/O epoll, IOCP, io_uring
A model in which a thread keeps doing other work while I/O is in progress and gets notified when it completes.
AOI Area of Interest
The range each player “can see.” Only changes inside it are sent, which cuts bandwidth. To reduce the cost of working out who’s in range, the map is usually divided into a grid and only nearby cells are checked.
Broadcast Broadcast, fan-out
Sending one change to everyone who can see it. If everyone in a crowd can see everyone else, the amount to send grows with the square of the player count.
GC Garbage collection
Automatically reclaiming memory that’s no longer in use. Programs can pause during GC.
Heap Heap
The memory area a program allocates from whenever it needs memory while running.
Memory leak Memory leak
A bug in which memory that’s no longer needed is never released, so usage keeps growing. It happens even with GC if something still holds a reference to objects the program is done with.
Swap Swap, paging
Moving part of memory to disk when RAM runs short. Using that memory again is more than 1,000 times slower than RAM.
OOM killer Out-of-memory killer
A Linux feature that, when memory runs out, picks the process using the most memory and forcibly kills it. In a container it kicks in as soon as the container’s memory limit is hit.
Cache miss Cache miss
The data isn’t in the cache close to the CPU, so it has to come from slower memory.
IOPS I/O operations per second
How many reads and writes a disk can handle per second. On cloud disks, the limit depends on what you pay for.
fsync fsync
A call that waits until data is safely written to disk. Normal writes land in OS memory first and reach the disk later, so they can be lost if the server loses power in between. fsync is safe but slow.
Burst credits Burst credits
A balance that cloud disks and servers build up so they can briefly run above their baseline performance. When it runs out, performance drops back to baseline.
Index Index
A lookup structure in a database. Without one, the database has to read the whole table.
Full table scan Full table scan
A query that checks every row of a table without using an index.
Query plan Query plan
The database’s plan for executing a query: in what order and with which indexes. Even with unchanged code, the same query can suddenly slow down if the database switches plans.
Transaction Transaction
A group of DB operations that either all succeed or all fail. Trades must always be handled in a transaction. Rows it modifies stay locked until it finishes, so shorter is better.
Connection pool Connection pool
A set of DB connections opened in advance. When all of them are in use, new requests wait.
Hot row Hot row
A single row that many requests try to modify at the same time. A cause of lock contention.
Replication lag Replication lag
How far a replica database has fallen behind the primary.
Rollback Rollback
A save is canceled and data returns to its previous state. Players experience it as “my item disappeared.”
Cache Cache (Redis etc.)
A copy of frequently used data kept somewhere fast. It reduces database load.
Checkpoint Checkpoint
The database periodically writing the changes it has collected in memory out to disk in one batch. Saves and queries can slow down briefly while it happens.
Failover Failover
Switching over to a standby when the primary server or DB goes down. Writes stop briefly during the switch, and if replication was lagging, the most recent data can be lost.
MVCC Multi-version concurrency control
A way for a database to keep old row versions for a while so readers and writers don’t block each other. A long-open transaction lets old versions pile up and slows things down.
Cache stampede Cache stampede
Cache entries empty out all at once and requests flood the origin (DB).
Gateway Gateway
An intermediate server that accepts client connections and forwards them to the game servers behind it.
Circuit breaker Circuit breaker
A mechanism that temporarily stops calling a service that keeps failing and returns a failure right away, which prevents cascading failures. After a while it tries a call or two, and if the service has recovered, it lets calls through again.
Cascading failure Cascading failure
A failure in one place spreading to other services along the call chain.
Autoscaling Autoscaling
Automatically adding and removing servers based on load. Adding servers takes time.
Watchdog Watchdog
A timer that watches whether the server has frozen. If the game loop stalls longer than a set time (a few seconds to tens of seconds), it writes a state dump and forcibly kills the server so it restarts.
Utilization Utilization
The fraction of time a worker (anything that processes requests, such as a CPU core, thread, or DB connection) is busy. Above 80–90%, waiting time climbs steeply.
p99 99th percentile
The value that 99 out of 100 measurements beat, with about 1 in 100 slower. It reflects the lag players feel better than the average does.
V-Sync Vertical sync
Sending frames in step with the display’s refresh cycle. It eliminates screen tearing but adds input lag, and when FPS drops below the refresh rate, it can bounce between 60 and 30 and stutter.
Variable refresh rate VRR, G-Sync, FreeSync
The monitor refreshes whenever a frame is ready. It reduces the stutter and input lag caused by V-Sync bouncing between 60 and 30.
Anti-cheat Anti-cheat
A security module that blocks game hacks. When its periodic scans or its heartbeats to the server fail, it can cause stutter or disconnects.
Overlay Overlay
A feature of chat, recording, or FPS-counter programs that draws on top of the game screen. It hooks into the game’s rendering and can cause stutter.
Shader compilation Shader compilation
Converting graphics effect programs into code for the GPU. If it isn’t done ahead of time, the screen hitches the first time an effect appears, and updating the graphics driver invalidates the saved results, so it happens all over again.
Main thread Main thread, Game thread
The game’s central thread, which runs game logic and prepares each frame in turn. If any one task on it takes long, the screen freezes for that long.
Timer resolution Timer resolution
The shortest interval at which the OS can wake a sleeping program. The Windows default is 15.6 ms, so unless a program changes it, even “wake me in 1 ms” wakes up late.
Thermal throttling Thermal throttling
A protective feature that lowers CPU and GPU speed when a device gets hot. On phones it’s common after a few to tens of minutes of play.
VRAM Video memory
Dedicated memory on the graphics card. Textures and models are loaded here for rendering. When it runs short, data shuttles to and from PC memory over a slower path and the game stutters.
Net graph Net graph
A dev and debug overlay that shows ping, packet loss, FPS, and tick as live graphs on the game screen. Having it in a lag report video makes finding the cause much easier.
Retransmission rate Retransmission rate
The share of sent TCP packets that had to be sent again. There’s no official threshold, but a server-wide average below 0.1% is generally healthy, and above 1% many players are likely to feel lag. Also check how many times higher it is than the usual level.
SACK Selective ACK
A TCP feature in which the receiver reports in detail: “I got this range; only this part is missing.” It can recover several losses at once.
RACK-TLP Recent ACK, Tail Loss Probe
A TCP feature that detects loss based on time and, when no ACK arrives for a while, resends the last packet to speed up recovery. It’s the default on current Linux and Android. On Windows, TLP and RACK are on by default from Windows 10 (1607) and Server 2016, and the newer RACK that also recovers lost retransmissions arrived with Server 2022. It works only on connections with SACK enabled.
Spurious retransmission Spurious retransmission
Resending a packet that wasn’t lost: it arrived late or out of order, and the sender mistook it for lost. It wastes bandwidth and needlessly cuts the sending rate.
Zero window Zero window
The receiver’s buffer is full, so it has told the sender “stop sending for now.” It looks like retransmission, yet the network is fine; the receiving program just didn’t read the data in time.
thin stream Thin stream
A connection that sends small packets sparsely, like a game. Fast retransmit signals rarely build up, so losses cause long stalls.
Policer Policer
A rate limiter that drops packets exceeding a set rate right away, without queuing them. A limiter that queues them and releases them slowly is called a shaper.
Pacing Pacing
Spreading outgoing packets evenly over time so they don’t all go out in one burst. It keeps small buffers from overflowing.
ECN Explicit Congestion Notification
Under congestion, marking packets with a “congested” flag, without dropping them, so the sender slows down. It signals congestion with no loss. It only works if both endpoints and the equipment on the congested segment all support it.
MSS Maximum Segment Size
The maximum amount of data TCP puts in one packet. Usually 1,460 bytes; lowering it to fit tunnel segments prevents MTU black holes.
Handover Handover
A moving phone switching the cell tower it’s connected to.
Percentile Percentile (p50, p95, p99)
The value at a given percentage position when all values are sorted from smallest to largest. p50 is the median; p99 is the value near the slowest 1 in 100. It reveals the spikes an average hides.
Tail latency Tail latency
Long delays that happen occasionally even though most requests are fast. They barely show up in the average, but they’re what players remember as lag.
Synthetic monitoring Synthetic monitoring
Measuring path quality with dedicated probes or servers that periodically run ping, traceroute, and similar tests from fixed locations, standing in for real players. RIPE Atlas is the best-known public tool.
Aggregation interval Aggregation interval
How many seconds or minutes of data one point on a graph combines. The longer the interval, the more short spikes get averaged away.
Postmortem Postmortem
A write-up after an incident covering what happened, why it happened, and what will change. It is written blamelessly, with the goal of preventing a repeat.
C-state CPU idle state
Power-saving states a CPU enters when idle. Deeper states save more power but take longer to wake from.
Live migration Live migration
A cloud provider moving a running virtual machine to another host, for example for host maintenance. The VM can pause briefly during the move.
SNAT Source NAT
NAT that rewrites the source address of outgoing packets to a public address. Each public address has a limited number of usable ports, and new connections fail once they run out.
NAT gateway NAT gateway
A cloud device that lets servers on a private network share one public address to reach the internet. It limits concurrent connections per destination.
LEO satellite internet LEO satellite internet
Internet service through satellites orbiting hundreds to thousands of km up. Latency is far lower than with geostationary satellites, but it can spike when the connection switches to another satellite.
GeoIP IP geolocation
A database that estimates country, city, and ISP from an IP address. Wrong or outdated entries can get players assigned to a distant server.
TLS certificate TLS certificate
A digital document proving a server is who it claims to be. It has an expiry date, and once it expires, encrypted connections fail and players can’t connect.
Frame generation Frame generation
A technology in which the graphics card inserts predicted frames between the frames it actually rendered to raise FPS. Motion looks smoother, but input-to-screen latency can increase.
T6Tools

References

The basis for the numbers, defaults, and descriptions of behavior in this white paper. We collected only authoritative sources: standards documents (RFCs), kernel and OS documentation, official cloud, engine, and DB documentation, and talks and papers. Cause cards and the “Sources” at the end of each chapter link to the same material. Defaults can change between versions, so check the documentation for the version you run before applying anything.

616 sources from 83 publishers. The full list is in the references of the text edition.