Game Lag White Paper › Browse by symptom
Can’t connect / infinite loading: 45 causes and who fixes them
Also called: can’t log in, stuck on the loading screen
Open in the illustrated symptom guide →
You can’t get into the game, or you’re stuck on a loading or entry screen.
“Unable to connect to server” over and over, or the loading bar never finishes after character select.
Whatever accepts new connections (the server’s connection queue, firewall, login server, DB) is full. It’s especially common right after maintenance.
Causes of this symptom
L2 Client OS and device
- Packet inspection by security software: When antivirus software or a firewall inspects every packet, latency goes up, and in bad cases it mistakes the game for an attack and blocks it. (External (External))
L3 Home network
L4 Internet path
- Country- or ISP-level UDP restrictions and packet inspection: Some networks block specific UDP addresses and ports or throttle UDP, and packet inspection equipment filters out protocols it doesn’t recognize. Games that communicate over UDP can’t connect on those networks or disconnect often. (External (External))
- DNS failures and delays: If DNS, which turns server names into addresses, is slow or fails, the game can’t find its login or patch servers. (External (External))
- Shared links saturated by DDoS: Massive attacks aimed at the game company, or at someone else on the same network, fill up shared links. (Infra team (Network infrastructure))
- ISP-shared IP addresses (CGNAT): Mobile networks and some ISPs have many subscribers share one IP address, and they delete the mappings of idle connections after a short time. (Game team (Client development))
- Routing through a VPN or game booster: With a VPN or game booster on, packets go through that company’s relay servers. If the relay is far away or busy, the connection can actually get slower. (External (External))
L5 Data center network equipment
- Firewall session table full: A firewall tracks every connection it lets through by recording it in a session table. Once the table is full, it can’t accept new connections. (Infra team (Network infrastructure))
- DDoS protection detours and false positives: Diverting traffic to a scrubbing center to stop attacks makes the route longer, and legitimate players are sometimes mistaken for attackers and blocked. (Infra team (Network infrastructure))
- Cloud NAT gateway connection and port limits: When servers in a private subnet connect out (platform authentication, payments, external APIs), a NAT gateway rewrites their address and port. If concurrent connections to the same destination exceed the gateway’s port limit, new connections fail. (Infra team (Network infrastructure))
- Load balancer skew and misjudged health checks: Connections pile onto one server, or players keep getting sent to a server that’s already dead. (Infra team (Network infrastructure))
- MTU mismatch (only large packets vanish): If the MTU (the largest size that can be sent at once) shrinks somewhere along the path and the “packet too big” messages are blocked, only large packets keep vanishing. (Infra team (Network infrastructure))
L7 Server OS (kernel)
- Connection queue (listen backlog) overflow: When tens of thousands of players connect at once right after maintenance, the kernel’s connection queue (listen backlog) overflows and connection attempts are dropped. (Game team (Server development))
- File descriptor limit: Every connection needs a file descriptor (fd: the number the OS gives an open file or socket), and the number of fds one process can open is capped. (Infra team (Server infrastructure))
- Server conntrack table full: When the connection tracking (conntrack) table, where the Linux firewall records every connection, reaches its limit, new packets are dropped. (Infra team (Server infrastructure))
- Ephemeral port exhaustion on server-to-server connections: When a game server opens and closes short connections to the DB or other servers very often, closed connections hold their ports for a while, and new connections can’t be opened. (Game team (Server development))
L8 Sockets and protocols
- Keepalive default of 2 hours: When the other side vanishes without a close signal, TCP notices only much later. Keepalive (a TCP feature that checks whether an idle connection is still alive) is off by default, and even when it’s on, checks start only after 2 hours of idle time. (Game team (Server development))
- Uneven SO_REUSEPORT distribution: When several processes share one port, the kernel assigns each connection to a process by address hash and never reassigns it. If one of those processes stalls, only the players assigned to it wait. (Game team (Server development))
L9 Server game process
- Thread pool exhaustion: When every worker thread is tied up in slow work, new requests just wait with no end in sight. (Game team (Server development))
L11 Disk
- Writing a core dump: When the server crashes, writing several GB of memory to disk can delay the restart by several minutes. (Infra team (Server infrastructure))
L12 Database
- Queries with no index: Without an index, finding the rows that match a condition means reading the entire table (a full table scan). (Game team (Server development))
- Connection pool exhaustion: The number of connections open to the DB is fixed, so when slow queries hold connections, every other request waits. (Game team (Server development))
- Cold cache (right after a restart): After a DB restart, the memory cache is empty, so for a while every lookup reads from disk. (Infra team (DB infrastructure))
- Login storm and N+1 queries: If loading one character takes dozens of separate queries, tens of thousands of simultaneous logins turn into millions of queries. (Game team (Server development))
- DB failover: When the primary DB dies, writes stop while it fails over to a standby, and the last data that hadn’t been replicated yet can be lost. (Infra team (DB infrastructure))
- Cache stampede: When cache entries for popular data expire at the same time, thousands of requests hit the DB all at once. (Game team (Server development))
- Slow Redis commands: Redis processes commands one at a time, so a single slow command blocks every request behind it. (Game team (Server development))
- Query slowdown from a query plan change: Even with the code unchanged, if the DB changes how it executes a query (its query plan), a query that took 2 ms yesterday takes hundreds of ms today. (Infra team (DB infrastructure))
- Schema change (DDL) lock during live service: Adding a column or index to a table during live service can make every request that uses that table wait, all because of one lock that’s needed only briefly. (Infra team (DB infrastructure))
L13 Server architecture and operations
- Zone transfer (handoff between servers): Entering another area or dungeon means handing the character’s data to another server, and that handoff can be slow or fail. (Game team (Server development))
- Cascading failure: When one service slows down, the servers that call it get tied up waiting for responses, and even unrelated features stop. (Game team (Server development))
- Auxiliary server outage: When a server that runs separately from the game server, such as chat, party, or auction house, fails, only that feature stops working. (Game team (Server development))
- Deploys and restarts: If you restart a server for an update without moving its connections, everyone on it disconnects, and the final saves before shutdown and the reconnects all hit at once. (Game team (Server development))
- Autoscaling delay: When players flood in, servers are added automatically, but getting them ready takes several minutes, and the existing servers are overloaded in the meantime. (Infra team (Server infrastructure))
- External service dependency: When an external service such as platform login, payments, or identity verification is slow or down, players get stuck at that step. (External (External))
- Expired or misconfigured TLS certificate: When the certificate on a login, API, or patch server expires or is missing its intermediate certificate, every client that connects from that moment on fails the TLS connection. (Infra team (Network infrastructure))
- Login queue cap and insufficient reconnect grace: When players flood in right after launch or maintenance, the login queue hits its cap and turns away new arrivals, and players who were already waiting lose their place during a brief disconnect and go back to the end of the line. (Game team (Server development))
Netcode design
Problems only some players hit
- Bloated data on one character: A character with thousands of items or mails piled up, or an unusually large friend list, block list, or set of buffs, has several times more to load, save, and announce to nearby players than others. It’s slow only on that character, regardless of connection. (Game team (Server development))
- Fixed UDP port collision: If the client is built to use a fixed local port, a second client on the same PC either can’t get the port or ends up splitting packets with the first. (Game team (Client development))
- Multi-client restriction: If an anti-cheat module or server policy limits multiple clients on one PC, the second client is blocked from launching or connecting, or the first one gets disconnected. Some games only block features on the extra client. (Game team (Client development))
Root causes of TCP retransmission
- Firewall and connection tracking drops: Firewalls and Linux connection tracking (conntrack, which records passing connections in a table) drop packets when the table is full or when they decide a packet doesn’t match the connection’s state. (Infra team (Network infrastructure))
- MTU black hole (only large packets keep getting lost): If the largest packet size a link along the way can carry shrinks and the “too big” notice (ICMP) is blocked, large packets keep vanishing no matter how many times they’re resent. (Infra team (Network infrastructure))
- Connection request (SYN) retransmission: If a connection request is lost because the connection queue (backlog) overflows or a firewall blocks it, the client OS resends it starting 1 second later, at growing intervals. (Game team (Server development))
View the illustrated symptom guide