한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L13 Server architecture and operations

Cascading failure Cascading failure

Cause ID in-cascade · Primary owner Game team (Server development) · Also Infra team (Network infrastructure)

Open the interactive card with figures and simulations →

When one service slows down, the servers that call it get tied up waiting for responses, and even unrelated features stop.

Why One service, such as the DB or authentication, slows down → Effect Threads and connections on the calling servers are tied up waiting for responses, and retries of failed requests add more load → On screen Everything slows down or stops, even features that look unrelated

Symptoms
Freeze, Input lag, Can’t connect / infinite loading
Factors
Stall
Who’s affected
Whole server
When
When crowds gather, Randomly
Owner
Primary owner Game team (Server development) · Also Infra team (Network infrastructure)
Game team action items
Put a timeout on every call, add circuit breakers and per-feature isolation (bulkheads), retry with growing intervals and a capped count, keep health check responses separate from busy work.
Infra team action items
Give load balancer health checks slack in failure count and interval so a briefly slow server isn’t pulled right away, limit how many servers can be pulled at once.
On the graph
Hits a ceiling · Per-service response time and error rate, thread and connection usage
Where to look
Per-service response time, error rate, and retry count on one screen with aligned time axes, to find what slowed down first. Behind a load balancer: target response time (TargetResponseTime on AWS ALB), target 5xx count (HTTPCode_Target_5XX_Count), and number of targets pulled as unhealthy (UnHealthyHostCount)
Confirmed if
One service’s latency rises first, then thread and connection usage on its callers hits the limit, errors spread to other services, and retry count and pulled-target count rise together
Ruled out if
Several services slowed down at the same instant: check shared resources (DB, network, hosts) first
Check with
Infra tools (no game code needed)
Learn more
Health checks (probes that confirm a server is alive) also make cascades worse. When a busy server answers a check late, the load balancer pulls a server that is actually working, its traffic piles onto the remaining servers, and the next server falls behind too.
Real incidents
Riot Games 2020: Edge host overload on League of Legends servers in Europe and Brazil
Riot Games 2021: League of Legends EUW 5-hour outage: one auxiliary DB halted the whole server
Roblox 2021: Roblox 73-hour outage: contention in the service discovery (Consul) cluster
AWS 2021: AWS us-east-1 internal network congestion
AWS 2025: AWS us-east-1 DynamoDB DNS outage and long recovery

Sources

  1. Site Reliability Engineering, Chapter 22: Addressing Cascading Failures Google
    An overloaded server that fails health checks gets pulled, load piles onto the rest, and retries amplify it; recommends capped retries, randomized exponential backoff, and deadlines
  2. Circuit Breaker Pattern Microsoft Azure
    Requests blocked until their timeout hold threads and DB connections and make unrelated features fail; once failures pile up within a set time, calls are rejected immediately
  3. Timeouts, retries, and backoff with jitter AWS
    Amazon Builders’ Library. With 3 retries at each layer of a 5-deep call chain, DB load grows 243 times; retry at only one layer and cap retries with a token bucket
  4. CloudWatch metrics for your Application Load Balancer AWS
    TargetResponseTime (time from the request leaving the load balancer until the target starts responding), HTTPCode_Target_5XX_Count (5xx responses generated by targets), UnHealthyHostCount (number of unhealthy targets)

See also

Same layer: L13 Server architecture and operations

Same symptom (Freeze), other layers

View the interactive card with figures and simulations