If the NIC sends every packet-arrival interrupt to a single CPU core, that core becomes the bottleneck.
Why A single receive queue, or RSS (which spreads packets across cores) turned off → Effect One core hits 100% and can’t pull packets off in time → On screen Packet loss and latency across the whole server when players crowd in (teleporting, input lag)
Configure RSS (spreading by the NIC) and RPS (spreading by the kernel), spread interrupts across several cores, make UDP queue selection include ports (rx-flow-hash udp4 sdfn in ethtool -N), keep interrupt-handling cores separate from the game tick thread’s cores, watch %soft per core.
Ballpark numbers
One core can push roughly hundreds of thousands of packets per second through the kernel, depending on packet size and settings. If per-core utilization shows receive processing (%soft in mpstat) piled onto a single core, this is what’s happening.
On the graph
Hits a ceiling · %soft per core, packets received per second
Where to look
Check %soft (share of time spent on software interrupts) per core with mpstat -P ALL 1, which core each NIC queue’s interrupts go to in /proc/interrupts, the number of queues with ethtool -l, and packets per queue with ethtool -S (names vary by driver)
Confirmed if
One core’s %soft sits near 100% while the rest are idle, and interrupts and packets pile into one queue. From then on, packets received per second can’t climb any higher
Ruled out if
%soft spread evenly across cores: not this cause. CPU idle but loss present: “Cloud PPS limit exceeded” or “Ring buffer too small”
Check with
Infra tools (no game code needed)
Learn more
Even with multiple queues, if most traffic comes from a handful of addresses, such as gateways or proxies, it all lands in one queue. For UDP, some NICs pick the queue from addresses only by default, and traffic spreads evenly only after you change that to include ports.
Sources
Scaling in the Linux Networking StackLinux kernel RSS (the NIC spreads packets across multiple receive queues) and RPS (the kernel spreads them), giving each queue its own interrupt and spreading those across cores; RSS is recommended when receive interrupt handling is the bottleneck
How to receive a million packets per secondCloudflare Measurements where a receive queue served by a single core topped out at about 350,000–430,000 packets per second; a case where the NIC hashed UDP by IP address only and everything piled into one queue