The game code hasn’t changed, but the server has been slower since an OS, kernel, driver, or firmware update. Updates can change defaults, the scheduler, CPU vulnerability mitigations, and driver behavior.
Why A routine security patch or a new server image changes the kernel, drivers, or firmware → Effect Changed defaults or scheduler, or newly enabled vulnerability mitigations, make the same work take more CPU time and change the order in which threads get the CPU → On screen A server that ran fine is a little slower all the time from the day of the update: input lag, plus stutter and slow motion when crowds gather
Apply updates to a few servers first and compare tick time, latency, and CPU utilization against the previous version before rolling out wider, deploy on a different day from game patches, record kernel, driver, and firmware versions and key sysctl values before and after the update, boot into the previous kernel to confirm when problems appear, weigh the security risk before turning mitigations off (mitigations=off).
Ballpark numbers
A new kernel version brings new default behavior. For example, Linux began moving its scheduler from CFS to EEVDF in 6.6, and the default connection queue cap (somaxconn) changed from 128 to 4,096 in 5.4. CPU vulnerability mitigations add work, such as flushing internal CPU buffers when returning from the kernel to a program (at the end of every system call) and on context switches and VM transitions, so network servers that make a system call per packet are hit harder. Fully blocking some vulnerabilities requires turning off SMT (the feature that runs one core as two threads), and turning off SMT can cut performance sharply depending on the workload. The kernel parameter mitigations=off turns all these mitigations off and recovers the performance, but leaves the system exposed to the vulnerabilities.
On the graph
Step change · Server tick time, CPU utilization, latency under the same load
Where to look
Update history from the package manager and reboot times, the kernel version from uname -r, and NIC driver info from ethtool -i, lined up with when latency rose. Updated and non-updated servers compared under the same load with mpstat and pidstat, along with the mitigation status in /sys/devices/system/cpu/vulnerabilities/
Confirmed if
Latency and CPU utilization step up from the reboot after the update and stay there, and under the same load only the updated servers run high. Booting into the previous kernel or driver brings them back
Ruled out if
Updated and non-updated servers are equally slow under the same load: not this cause. A game patch went out the same day and packets per player or packet size changed: “Patch changes the traffic pattern”
Check with
Infra tools (no game code needed)
Learn more
Check mitigation status in the files under /sys/devices/system/cpu/vulnerabilities/. The default (mitigations=auto) mitigates with SMT left on, but auto,nosmt turns SMT off on vulnerable CPUs, so the logical core count can drop by half after a kernel upgrade. Updating the OS on the same day as a game patch makes it hard to tell which one caused a problem, so deploy them separately.
Sources
The kernel’s command-line parametersLinux kernel mitigations=: off disables all CPU vulnerability mitigations for more performance but leaves the system exposed, the default auto mitigates with SMT on, auto,nosmt turns SMT off when needed
MDS - Microarchitectural Data SamplingLinux kernel Mitigations flush CPU buffers when returning from the kernel to user space and when entering a VM; the files under /sys/devices/system/cpu/vulnerabilities/ show vulnerability and mitigation status; many CPUs need SMT off for full protection, and turning SMT off can have a large performance impact depending on the workload
Spectre Side ChannelsLinux kernel As a mitigation, branch prediction buffers are flushed on context switches and VM transitions, and stronger mitigations add overhead to every program
EEVDF SchedulerLinux kernel Linux began moving from CFS to the EEVDF scheduler in 6.6