Game Lag White Paper › Browse by symptom
Dropped action / rollback: 36 causes and who fixes them
Also called: skill didn’t go off, item reverted, failed trade
Open in the illustrated symptom guide →
Something you definitely did never happened, or its result gets reversed much later.
You press a skill and it doesn’t fire. An item you bought disappears, or you reconnect and find everything back the way it was minutes ago.
The request was lost (packet loss, queue overflow), the server ruled differently from what your screen showed (timing difference, rejection after client-side feedback), or saving failed partway (DB lock or outage, server crash).
Causes of this symptom
L1 Client game process
- Clock sync error: If the client’s estimate of the server time is wrong, interpolation timing and cooldown checks drift out of step with the server. (Game team (Client development))
L5 Data center network equipment
- Cloud NAT gateway connection and port limits: When servers in a private subnet connect out (platform authentication, payments, external APIs), a NAT gateway rewrites their address and port. If concurrent connections to the same destination exceed the gateway’s port limit, new connections fail. (Infra team (Network infrastructure))
- Switch microbursts: When several servers send packets to thousands of players at the same instant, the small buffer on the switch port where that traffic converges overflows in less than 1 ms. (Game team (Server development))
L6 Server network card
- Ring buffer too small: If the NIC’s ring buffer, which briefly holds incoming packets, is small, a sudden burst overflows it and packets get dropped. (Infra team (Server infrastructure))
- Cloud PPS limit exceeded: Each cloud instance type has limits on packets per second and bandwidth, and traffic over them is silently dropped. (Infra team (Server infrastructure))
L7 Server OS (kernel)
- OOM killer: When memory runs out, Linux picks the process using the most memory and kills it. Usually that’s the game server. (Game team (Server development))
- System clock jump (NTP step): When the server clock is moved forward or back by several seconds in one step, timers that depend on the system clock fire all at once or stop. (Game team (Server development))
- Ephemeral port exhaustion on server-to-server connections: When a game server opens and closes short connections to the DB or other servers very often, closed connections hold their ports for a while, and new connections can’t be opened. (Game team (Server development))
L8 Sockets and protocols
- Reliable UDP retransmission settings: When the retransmission rules you built on top of UDP are too conservative, recovery is slow; when they’re too aggressive, they clog the connection even more. (Game team (Server development))
L9 Server game process
- Message queue backlog: When requests arrive faster than they’re processed and pile up in the queue, the ones at the back get processed only seconds later or are dropped. (Game team (Server development))
- Server crash: When the server process dies from an unhandled error, everyone on that server disconnects at the same time. (Game team (Server development))
- Patch changes the traffic pattern: When new content, effects, or synced fields raise packet size and frequency, a server that ran fine starts hitting MTU, bandwidth, and packet-rate limits after the patch. (Game team (Server development))
L11 Disk
- Disk full: When logs and dumps pile up and fill the disk, writes fail, and without safeguards the server crashes. (Infra team (Server infrastructure))
L12 Database
- Hot row lock contention: When everyone tries to modify the same row (a guild vault, a popular auction house item, a server-wide counter), only one request at a time gets the lock. (Game team (Server development))
- DB deadlock: When two transactions (groups of DB operations processed as one unit) each wait for a row the other has locked, the DB forcibly cancels one of them. (Game team (Server development))
- Replication lag: Writes go to the primary and reads come from replicas, so when a replica falls behind, data that was just written isn’t visible yet. (Infra team (DB infrastructure))
- Bulk batch jobs: Running ranking aggregation, mass mail sends, or old-data cleanup during live service ties up locks and the disk. (Game team (Server development))
- DB failover: When the primary DB dies, writes stop while it fails over to a standby, and the last data that hadn’t been replicated yet can be lost. (Infra team (DB infrastructure))
- Lost progress from a long save interval: If the server saves only once every few minutes to reduce load, progress is lost when the server dies in between. (Game team (Server development))
- Transaction left open too long: When a transaction stays open for a long time, it keeps holding its locks and the DB can’t clean up (purge) old versions of data, so everything gradually slows down. (Game team (Server development))
- Schema change (DDL) lock during live service: Adding a column or index to a table during live service can make every request that uses that table wait, all because of one lock that’s needed only briefly. (Infra team (DB infrastructure))
L13 Server architecture and operations
- Auxiliary server outage: When a server that runs separately from the game server, such as chat, party, or auction house, fails, only that feature stops working. (Game team (Server development))
- Clock skew between servers: When each server’s clock is slightly off from the others, cooldown, buff, and event start checks disagree from server to server. (Infra team (Server infrastructure))
- External service dependency: When an external service such as platform login, payments, or identity verification is slow or down, players get stuck at that step. (External (External))
- Matchmaking and region assignment errors: When a player lands on a server in a distant region while a closer region exists, that player’s ping stays high even though their connection is fine. (Game team (Server development))
- Expired or misconfigured TLS certificate: When the certificate on a login, API, or patch server expires or is missing its intermediate certificate, every client that connects from that moment on fails the TLS connection. (Infra team (Network infrastructure))
Netcode design
- No skill input buffering: If you can’t press the next skill until the server confirms the previous one has finished, a round trip gets inserted between every skill in a rotation. (Game team (Client development))
- Short timing windows eaten up by ping: When the time you have to react is short, as with dodges, parries, and guards, ping eats up that time and some attacks become impossible to avoid. (Game team (Server development))
- Hit registration without lag compensation: If the server checks hits only against where targets are on the server right now, what you saw on your screen and the server’s call disagree. (Game team (Server development))
- Too much lag compensation: If the server rewinds too far in the attacker’s favor, the target gets hit even after they’ve already taken cover. (Game team (Server development))
- Client authority: When each client decides its own results, your own screen feels responsive, but results disagree with other players’ screens and the game is easy to hack. (Game team (Server development))
- Overly strict server validation: If the server checks movement speed, cooldowns, and range too strictly, it rejects even valid inputs that arrive bunched together because of jitter. (Game team (Server development))
- Server rejects after client-side feedback: When the server later refuses a hit or skill your screen already showed, the result you clearly saw never happened. (Game team (Client development))
- Low snapshot send rate: If the server sends position updates (snapshots) only a few times a second, the interpolation buffer has to be that much longer, and you see other characters further in the past. (Game team (Server development))
Problems only some players hit
View the illustrated symptom guide