When one service slows down, the servers that call it get tied up waiting for responses, and even unrelated features stop.
Why One service, such as the DB or authentication, slows down → Effect Threads and connections on the calling servers are tied up waiting for responses, and retries of failed requests add more load → On screen Everything slows down or stops, even features that look unrelated
Primary owner Game team (Server development) · Also Infra team (Network infrastructure)
Game team action items
Put a timeout on every call, add circuit breakers and per-feature isolation (bulkheads), retry with growing intervals and a capped count, keep health check responses separate from busy work.
Infra team action items
Give load balancer health checks slack in failure count and interval so a briefly slow server isn’t pulled right away, limit how many servers can be pulled at once.
On the graph
Hits a ceiling · Per-service response time and error rate, thread and connection usage
Where to look
Per-service response time, error rate, and retry count on one screen with aligned time axes, to find what slowed down first. Behind a load balancer: target response time (TargetResponseTime on AWS ALB), target 5xx count (HTTPCode_Target_5XX_Count), and number of targets pulled as unhealthy (UnHealthyHostCount)
Confirmed if
One service’s latency rises first, then thread and connection usage on its callers hits the limit, errors spread to other services, and retry count and pulled-target count rise together
Ruled out if
Several services slowed down at the same instant: check shared resources (DB, network, hosts) first
Check with
Infra tools (no game code needed)
Learn more
Health checks (probes that confirm a server is alive) also make cascades worse. When a busy server answers a check late, the load balancer pulls a server that is actually working, its traffic piles onto the remaining servers, and the next server falls behind too.
Circuit Breaker PatternMicrosoft Azure Requests blocked until their timeout hold threads and DB connections and make unrelated features fail; once failures pile up within a set time, calls are rejected immediately
Timeouts, retries, and backoff with jitterAWS Amazon Builders’ Library. With 3 retries at each layer of a 5-deep call chain, DB load grows 243 times; retry at only one layer and cap retries with a token bucket
CloudWatch metrics for your Application Load BalancerAWS TargetResponseTime (time from the request leaving the load balancer until the target starts responding), HTTPCode_Target_5XX_Count (5xx responses generated by targets), UnHealthyHostCount (number of unhealthy targets)
See also
Same layer: L13 Server architecture and operations