한국어English日本語简体中文繁體中文DeutschไทยTiếng ViệtРусскийPortuguês (Brasil)EspañolBahasa Indonesia

Game Lag White Paper › L9 Server game process

Thread pool exhaustion Thread pool starvation

Cause ID sp-threadpool · Primary owner Game team (Server development)

Open the interactive card with figures and simulations →

When every worker thread is tied up in slow work, new requests just wait with no end in sight.

Why Worker threads are tied up waiting on external API or DB responses → Effect No thread is free to take a new request → On screen Infinite loading in specific features such as login or the shop

Symptoms
Can’t connect / infinite loading, Input lag, Freeze
Factors
Stall
Who’s affected
One feature only, Whole server
When
When crowds gather, Right after login or maintenance
Owner
Primary owner Game team (Server development)
Game team action items
Put timeouts on slow calls, separate thread pools by feature, make calls asynchronous.
On the graph
Hits a ceiling · Thread pool thread count and queue length, request processing time
Where to look
For .NET, thread pool thread count and queue length in dotnet-counters monitor (dotnet.thread_pool.thread.count and dotnet.thread_pool.queue.length on .NET 9 and later, ThreadPool Thread Count and ThreadPool Queue Length on 8 and earlier), and where the worker threads are waiting, from dotnet-stack. For JVM and native servers, the same from a thread dump
Confirmed if
CPU utilization is well below 100%, yet the thread count keeps creeping up or sits at its cap, the queue builds up, and most workers are waiting on responses from the same external call (DB, HTTP)
Ruled out if
Queue empty yet still slow: the called service itself is slow, which points to a cascading failure or an external service dependency
Check with
Infra tools (no game code needed)
Learn more
If packet receiving and game logic share the same worker thread pool, the moment a few slow jobs occupy every worker, packet processing stops for the whole server.
Real incidents
Riot Games 2021: League of Legends EUW 5-hour outage: one auxiliary DB halted the whole server

Sources

  1. Debug ThreadPool Starvation Microsoft
    When the pool has no threads left and new work has to wait, responses slow down; blocking code that holds threads is the cause. In dotnet-counters, dotnet.thread_pool.thread.count creeping up while CPU is well below 100% signals exhaustion (dotnet.thread_pool.queue.length is often large too); dotnet-stack shows where threads are waiting
  2. Avoiding insurmountable queue backlogs AWS
    Concurrent requests = arrival rate × latency (Little’s law). At 100 requests per second, latency growing from 100 ms to 10 s turns 10 threads into 1,000 and the pool runs dry
  3. Bulkhead Pattern Microsoft Azure
    With separate connection and thread pools for each called service, a failure in one service blocks only its own pool
  4. .NET runtime metrics .NET
    dotnet.thread_pool.thread.count (thread pool thread count) and dotnet.thread_pool.queue.length (queued work items) since .NET 9
  5. Well-known EventCounters in .NET Microsoft
    ThreadPool Thread Count (threadpool-thread-count) and ThreadPool Queue Length (threadpool-queue-length) on .NET 8 and earlier

See also

Same layer: L9 Server game process

Same symptom (Can’t connect / infinite loading), other layers

View the interactive card with figures and simulations