Running far more threads than there are cores makes the OS spend CPU just switching between them.
Why Hundreds to thousands of threads, for example one thread per connection → Effect Higher context-switching cost (swapping out the running thread) and more cache misses → On screen CPU is busy but throughput is low and ticks are uneven: stutter, slow motion
Primary owner Game team (Server development) · Also Infra team (Server infrastructure)
Game team action items
Match thread count to core count, use asynchronous I/O (epoll, IOCP).
Infra team action items
Monitor context switches and runnable threads (cs and r in vmstat).
Ballpark numbers
One context switch costs a few µs, and more once you add the cache misses that follow.
On the graph
Rises with load · Context switches per second, runnable threads
Where to look
cs (context switches per second) and r (running or waiting for CPU) from vmstat 1 compared with the core count; voluntary (cswch/s) and involuntary (nvcswch/s) context switches per game server thread from pidstat -w -t
Confirmed if
As concurrent users grow, r climbs far above the core count, cs spikes with it, and hundreds of threads show many involuntary context switches
Ruled out if
r stays at or below the core count: not this cause. Mostly voluntary switches: threads are waiting on locks or I/O (“Lock contention,” “Blocking I/O design”)
vmstat(8) — Linux manual pageprocps-ng The cs (context switches per second) and r (processes running or waiting to run) fields
I/O Completion PortsMicrosoft Handle many asynchronous I/Os with a pre-created thread pool and IOCP, and match the number of concurrently running threads to CPU concurrency
pidstat(1) — Linux manual pagesysstat In -w, cswch/s counts voluntary context switches (the task stopped on its own to wait for a resource) and nvcswch/s counts involuntary ones (forced out after using up its time slice); -t shows them per thread