When several threads wait on one lock to write the same data, they run one at a time no matter how many threads you add.
Why Several threads use shared data at once, such as the auction house or guild storage → Effect The others wait until the thread holding the lock finishes → On screen Only certain features are slow; in bad cases, the whole tick is delayed
Split locks into finer-grained ones, do less work inside locks, move to a message-based design (give each piece of data an owning thread, and have other threads only send it requests as messages).
Ballpark numbers
If 20% of the work happens inside the lock, throughput tops out at 5 times that of one thread no matter how many threads you add; at 40%, it stops at 2.5 times.
On the graph
Rises with load · Request processing time, per-thread CPU and context switches
Where to look
Per-thread voluntary context switches (cswch/s, times a thread stopped to wait for a resource) from pidstat -w -t 1; where threads wait after leaving the CPU (wait time per call stack) from bcc offcputime -p. For .NET, the lock contention count in dotnet-counters (dotnet.monitor.lock_contentions on .NET 9 and later, Monitor Lock Contention Count on 8 and earlier)
Confirmed if
Processing time grows as load rises while CPU utilization stays low, most of the wait time is concentrated in call stacks trying to acquire a lock, and the lock contention count rises along with it
Ruled out if
CPU maxed out: a compute problem (tick overrun, single-threaded zone overload). Waiting on DB or file calls: points to blocking calls on the game thread
Check with
Infra tools (no game code needed)
Learn more
This happens in designs where several threads modify game data together. A design where one thread owns each area or feature and threads communicate only through messages has almost no locks, but you have to watch for work piling up on one thread (single-threaded zone overload). If the game thread waits for a lock held by a slow save operation, that whole tick stalls.
Amdahl's Law in the Multicore EraIEEE IEEE Computer 2008 paper (authors’ copy). If the fraction that can’t be parallelized is 1−f, the speedup can never exceed 1/(1−f) no matter how many cores you add (Amdahl’s law)
Request schedulingMicrosoft Orleans grains (actors) use a single-threaded execution model that runs each request to completion one at a time, so state is never modified concurrently; grains waiting on each other’s responses can deadlock
pidstat(1) — Linux manual pagesysstat cswch/s in -w is the number of voluntary context switches from stopping to wait for a resource; -t shows it per thread
Well-known EventCounters in .NETMicrosoft Monitor Lock Contention Count (monitor-lock-contention-count): number of times contention occurred when trying to acquire a monitor lock
.NET runtime metrics.NET dotnet.monitor.lock_contentions since .NET 9: number of times contention occurred when trying to acquire a monitor lock since the process started