Game Lag White Paper › Browse by symptom
Input lag: 76 causes and who fixes them
Also called: delayed response, sluggish, mushy controls
Open in the illustrated symptom guide →
It takes a while from pressing a button to seeing the result. The screen itself can still be smooth.
A skill fires 0.2–0.5 s after you press it. Actions that wait for confirmation, like looting, talking to NPCs, and trading, feel slow.
Round-trip time (ping) is long, or a queue is building up somewhere. Check distance, router queues, Nagle (a TCP feature that collects small packets and sends them together), and server queues. If it always feels sluggish even with low ping, look at your PC (V-Sync, low FPS) or at a design that waits for server confirmation on every action (see the netcode models chapter).
Causes of this symptom
L1 Client game process
- Rendering load from large crowds: When hundreds of players fill one screen, as in a siege or a world boss fight, the cost of drawing them is more than the device can handle. (Game team (Client development))
- Packet processing bottleneck on the main thread: If the client processes only a fixed amount of received packets per frame, a flood of packets keeps getting pushed to the next frame. (Game team (Client development))
- V-Sync and the render queue: Input is delayed while several finished frames wait in a queue to be sent out in step with the monitor’s refresh. (Game team (Client development))
L2 Client OS and device
- Power saving and thermal throttling: Laptop battery mode, phone power-saving mode, and device heat slow down the CPU and GPU. With heat, the telltale sign is that the game runs fine at first and slows down only after a while. (External (External))
- Other apps on the same device using up bandwidth: When cloud sync, a large download, or a game patch runs on the same PC, game packets have to wait in a queue. (External (External))
- Display, input device, and frame generation latency: If ping is normal but controls feel heavy, a TV’s video processing, a wireless controller, or frame generation may be adding delay between your input and the screen. (External (External))
L3 Home network
- Congested Wi-Fi channel: Where there are dozens of routers, as in an apartment building, they share the same channel and have to wait for a chance to transmit. (External (External))
- Bufferbloat (router queue): When someone in the household uploads a video or downloads a large file, hundreds of ms worth of packets pile up in the router’s queue, and game packets wait behind them. (External (External))
- RRC state transition delay (mobile radio power saving): When a phone has no traffic for a while, it drops its radio connection to a low-power state, and the next packet is delayed while it powers back up. (Game team (Client development))
L4 Internet path
- Propagation delay (physical distance): Even light travels only about 200,000 km per second in optical fiber. A distant server is slow no matter how good it is. (Infra team (Server infrastructure))
- Satellite internet (LEO/GEO): Satellite signals have to travel to space and back. With geostationary satellites the round trip alone exceeds 0.5 seconds. Low Earth orbit satellites such as Starlink are usually fast, but latency fluctuates and the link can drop briefly at the moment routes are reassigned. (External (External))
- Detour routing: Because of interconnection agreements between ISPs, traffic to even a nearby server can take a long way around. (Infra team (Network infrastructure))
- Submarine cable / international link outage: When a submarine cable is cut, traffic takes long detours for weeks (sometimes months) until it is repaired, and the remaining links get congested. (External (External))
- ISP throttling and traffic management: When you go over your data allowance, or on plans that manage certain kinds of traffic, packets get delayed or dropped. (External (External))
- Routing through a VPN or game booster: With a VPN or game booster on, packets go through that company’s relay servers. If the relay is far away or busy, the connection can actually get slower. (External (External))
L5 Data center network equipment
- DDoS protection detours and false positives: Diverting traffic to a scrubbing center to stop attacks makes the route longer, and legitimate players are sometimes mistaken for attackers and blocked. (Infra team (Network infrastructure))
- Data center link saturation: When patch distribution, log shipping, or backups share a link with the game, the link fills up. (Infra team (Network infrastructure))
L6 Server network card
- NIC interrupts concentrated on one core: If the NIC sends every packet-arrival interrupt to a single CPU core, that core becomes the bottleneck. (Infra team (Server infrastructure))
- Excessive interrupt coalescing: When the NIC collects packets and notifies the CPU once per batch to reduce CPU load, packets arrive later by the time spent collecting. (Infra team (Server infrastructure))
- NIC bandwidth saturation: Running a 1 Gbps or 10 Gbps card at its limit makes the transmit queue grow until packets get dropped. (Game team (Server development))
- GRO/LRO batching delay: GRO and LRO bundle several packets into one to reduce CPU load. Depending on settings, a small game packet may wait briefly for the next packet to bundle with. (Infra team (Server infrastructure))
L7 Server OS (kernel)
- Latency spikes from server power management (C-states, frequency scaling): Idle CPU cores drop into deep power-saving states (C-states) and lower their frequency to save power. Waking up and raising the frequency when a packet or timer arrives takes time, which adds delay to handling small packets. (Infra team (Server infrastructure))
- Performance changes after OS, kernel, driver, or firmware updates: The game code hasn’t changed, but the server has been slower since an OS, kernel, driver, or firmware update. Updates can change defaults, the scheduler, CPU vulnerability mitigations, and driver behavior. (Infra team (Server infrastructure))
L8 Sockets and protocols
- Nagle’s algorithm + delayed ACK: Nagle’s algorithm, which batches small packets, and delayed ACK, which sends ACKs late, interact so that each message written in pieces is delayed by 40–200 ms. (Game team (Server development))
- Slow start after idle: When a connection has been idle for a while, TCP shrinks the congestion window (how much it can send at once) again, so a sudden large send goes out in several rounds. (Infra team (Server infrastructure))
- Sending rate plunges under congestion control: TCP treats loss as a sign of congestion and cuts its sending rate by 30–50%. It reacts the same way to Wi-Fi loss. (Infra team (Server infrastructure))
- Blocking I/O design: In a design where a thread can’t do anything else while it waits on one socket, everything slows down as the player count grows. (Game team (Server development))
L9 Server game process
- Tick overrun: When one tick has more work than its budget, the server’s tick interval stretches, and the whole area slows down or stutters. (Game team (Server development))
- Broadcast fan-out overload: Sending one player’s movement to everyone who can see them creates updates on the order of the square of the crowd size. (Game team (Server development))
- Single-threaded zone overload (hotspot): When each area runs on a single thread and players crowd into one place, only that one core hits 100%. (Game team (Server development))
- Lock contention: When several threads wait on one lock to write the same data, they run one at a time no matter how many threads you add. (Game team (Server development))
- Message queue backlog: When requests arrive faster than they’re processed and pile up in the queue, the ones at the back get processed only seconds later or are dropped. (Game team (Server development))
- Serialization and compression cost: Turning outgoing data into bytes and compressing it takes CPU too, and with many players this cost explodes. (Game team (Server development))
- Thread pool exhaustion: When every worker thread is tied up in slow work, new requests just wait with no end in sight. (Game team (Server development))
- Combat concentrated on one target (world boss): When hundreds of players hit one boss at the same time, the computation for that single boss piles up in one place, and hit information goes out to everyone watching. (Game team (Server development))
- Spawn burst when entering a crowded area: When you teleport into a town packed with players, the server has to send the appearance, gear, and status of hundreds of newly visible players all at once. (Game team (Server development))
- Entity buildup (items and summons never cleaned up): When ground items that should have disappeared, summons, and finished timers pile up without being cleaned up, every tick has more work to do the longer the server stays up. (Game team (Server development))
- Patch changes the traffic pattern: When new content, effects, or synced fields raise packet size and frequency, a server that ran fine starts hitting MTU, bandwidth, and packet-rate limits after the patch. (Game team (Server development))
L11 Disk
- fsync surge: Asking for data to be written to disk “for sure” takes 0.1 ms to tens of ms per request depending on the disk, and when requests pile up, the queue grows. (Game team (Server development))
- Cloud disk out of burst credits: Some cloud disks and small server sizes have burst credits that let them run faster than baseline for a while, so when a busy period drags on and the credits run out, speed drops suddenly. (Infra team (Server infrastructure))
- IOPS limit / queue saturation: When requests exceed what the disk can handle per second, the queue grows and latency explodes. (Infra team (Server infrastructure))
- Backup / compression / scan jobs: When early-morning backups, log compression, or security scans monopolize the disk, the game server’s reads and writes get held up. (Infra team (Server infrastructure))
- HDD seek latency: An HDD has to move its head across the platter (a seek), so reading or writing scattered data takes close to 10 ms each time. (Infra team (Server infrastructure))
L12 Database
- Queries with no index: Without an index, finding the rows that match a condition means reading the entire table (a full table scan). (Game team (Server development))
- Hot row lock contention: When everyone tries to modify the same row (a guild vault, a popular auction house item, a server-wide counter), only one request at a time gets the lock. (Game team (Server development))
- DB deadlock: When two transactions (groups of DB operations processed as one unit) each wait for a row the other has locked, the DB forcibly cancels one of them. (Game team (Server development))
- Connection pool exhaustion: The number of connections open to the DB is fixed, so when slow queries hold connections, every other request waits. (Game team (Server development))
- Checkpoint / log flush: Queries slow down at the moments the DB periodically writes its accumulated in-memory changes to disk in bulk. (Infra team (DB infrastructure))
- Cold cache (right after a restart): After a DB restart, the memory cache is empty, so for a while every lookup reads from disk. (Infra team (DB infrastructure))
- Login storm and N+1 queries: If loading one character takes dozens of separate queries, tens of thousands of simultaneous logins turn into millions of queries. (Game team (Server development))
- Bulk batch jobs: Running ranking aggregation, mass mail sends, or old-data cleanup during live service ties up locks and the disk. (Game team (Server development))
- Cache stampede: When cache entries for popular data expire at the same time, thousands of requests hit the DB all at once. (Game team (Server development))
- Transaction left open too long: When a transaction stays open for a long time, it keeps holding its locks and the DB can’t clean up (purge) old versions of data, so everything gradually slows down. (Game team (Server development))
- Slow Redis commands: Redis processes commands one at a time, so a single slow command blocks every request behind it. (Game team (Server development))
- Query slowdown from a query plan change: Even with the code unchanged, if the DB changes how it executes a query (its query plan), a query that took 2 ms yesterday takes hundreds of ms today. (Infra team (DB infrastructure))
- Schema change (DDL) lock during live service: Adding a column or index to a table during live service can make every request that uses that table wait, all because of one lock that’s needed only briefly. (Infra team (DB infrastructure))
L13 Server architecture and operations
- Routing through a gateway or proxy: Putting an intermediate server between the client and the game server adds processing time at every hop, and that server becomes a single point of failure. (Game team (Server development))
- Cascading failure: When one service slows down, the servers that call it get tied up waiting for responses, and even unrelated features stop. (Game team (Server development))
- Deploys and restarts: If you restart a server for an update without moving its connections, everyone on it disconnects, and the final saves before shutdown and the reconnects all hit at once. (Game team (Server development))
- Too many macros and bots: Bots send requests far more often than people do and eat into server capacity. (Game team (Server development))
- Matchmaking and region assignment errors: When a player lands on a server in a distant region while a closer region exists, that player’s ping stays high even though their connection is fine. (Game team (Server development))
Netcode design
- Feedback only after the server responds (request-response): Press a button and there’s no animation or sound until the server answers. Your ping becomes your response time. (Game team (Client development))
- Chatty protocol (many sequential round trips): If one action needs several server round trips one after another, your ping is multiplied by that many. (Game team (Server development))
- No skill input buffering: If you can’t press the next skill until the server confirms the previous one has finished, a round trip gets inserted between every skill in a rotation. (Game team (Client development))
- Short timing windows eaten up by ping: When the time you have to react is short, as with dodges, parries, and guards, ping eats up that time and some attacks become impossible to avoid. (Game team (Server development))
- Lockstep waiting on the slowest player: When everyone computes the same turn together, one player’s late input makes everyone wait. (Game team (Server development))
- Double tick wait: If requests wait for the next tick to be processed and the results wait for the tick after that to be sent, the tick interval is added twice. (Game team (Server development))
Problems only some players hit
- Per-player input buffer size: If the server holds a few inputs per player and takes out one per tick, others see smooth motion, but your own actions are confirmed on the server that much later. (Game team (Server development))
- One lagging party member and boss mechanics: In raid mechanics where everyone has to react together at a set moment, one lagging player’s late reaction fails the whole party. (Game team (Server development))
- Bloated data on one character: A character with thousands of items or mails piled up, or an unusually large friend list, block list, or set of buffs, has several times more to load, save, and announce to nearby players than others. It’s slow only on that character, regardless of connection. (Game team (Server development))
- Per-connection send budget and priority: If the server caps how much it sends per connection and sends the nearest things first, a connection with a low cap gets distant NPCs late or not at all. (Game team (Server development))
Root causes of TCP retransmission
- Packet drops on the receiving host: Packets reach the server but get dropped, because the NIC’s ring buffer (which briefly holds arriving packets) overflows or the kernel cores that handle receive processing are saturated. (Infra team (Server infrastructure))
- Spurious retransmission from latency spikes: A packet that isn’t lost, just very late for a moment, still gets retransmitted if the delay is longer than the RTO, because the sender treats it as lost. (External (External))
- Spurious fast retransmit from reordering: When packets get out of order crossing multiple paths or bundled links, the receiver signals “a packet is missing” with duplicate ACKs, and the sender resends a packet that arrived fine. (Infra team (Network infrastructure))
- Late or lost ACKs (saturated upload): Data arrives fine, but if the “got it” ACK is delayed or dropped in a full upload queue, the sender treats the data as lost and retransmits. (External (External))
- RTO settings that don’t fit the environment: Set the RTO minimum too low and even small delays cause spurious retransmissions; leave the default (200 ms) and it’s too long for games, so every loss means a long freeze. (Infra team (Server infrastructure))
View the illustrated symptom guide