When servers in a private subnet connect out (platform authentication, payments, external APIs), a NAT gateway rewrites their address and port. If concurrent connections to the same destination exceed the gateway’s port limit, new connections fail.
Why Servers open many short connections to the same external address, such as platform authentication or payments, or keep connections open for a long time → Effect The NAT gateway can’t allocate any more source ports for that destination, so new connections fail → On screen The game itself is fine, but only features that call external services, such as login, payments, and reward delivery, fail or slow down (can’t connect / infinite loading, dropped action / rollback)
Right after login or maintenance, Evening peak hours, When crowds gather
Owner
Primary owner Infra team (Network infrastructure) · Also Game team (Server development)
Game team action items
Reuse connections to external APIs (HTTP keep-alive, connection pools) and don’t open a new connection per request, send keepalives on idle pooled connections more often than the NAT idle timeout (350 seconds on AWS) or close them first, retry failures with growing, randomized intervals, record failure rates and latency per external call.
Infra team action items
Add IP addresses to the NAT gateway (an AWS public NAT gateway takes only 2 Elastic IPs by default, so request a quota increase for more), split gateways per availability zone and subnet, alert on port allocation failure metrics (AWS ErrorPortAllocation, Failed in Azure SNAT Connection Count, OUT_OF_RESOURCES in Google Cloud dropped_sent_packets_count), raise the minimum ports per VM or use dynamic port allocation on Google Cloud NAT.
Ballpark numbers
An AWS NAT gateway can open up to 55,000 concurrent connections to the same destination (IP, port, protocol) per IP address, and you can attach up to 8 IPs to raise that. It deletes connections that stay silent for 350 seconds and answers later packets on them with RST. Azure NAT Gateway has 64,512 SNAT ports per public IP (up to 16 IPs). Google Cloud NAT divides 64,512 ports per NAT IP among VMs, and the default minimum is 64 ports per VM (static allocation), so with default settings a single VM is usually limited to 64 concurrent connections to the same destination.
On the graph
Hits a ceiling · NAT gateway concurrent connections, port allocation failures
Where to look
Line up the NAT gateway metrics ErrorPortAllocation, ActiveConnectionCount, and PacketsDropCount in AWS CloudWatch (on Azure, SNAT Connection Count filtered by the Failed state and Dropped Packets; on Google Cloud, dropped_sent_packets_count with reason OUT_OF_RESOURCES) against the times the game server’s external calls failed
Confirmed if
ErrorPortAllocation (Failed SNAT Connection Count on Azure, OUT_OF_RESOURCES drops on Google Cloud) goes above 0 when external calls fail, and the failures concentrate on calls to one or two heavily used destinations such as authentication or payment servers
Ruled out if
Port allocation failures at 0, but the game server’s connect fails with EADDRNOTAVAIL and TIME_WAIT is close to the size of the ephemeral port range: “Ephemeral port exhaustion on server-to-server connections.” Connections succeed but responses are slow: “External service dependency”
Check with
Infra tools (no game code needed)
Learn more
“Ephemeral port exhaustion on server-to-server connections” is about one server running out of ephemeral ports. This limit sits on the NAT gateway and is shared by all the servers behind it (Google Cloud NAT divides it per VM). If only external calls fail while the servers still have plenty of room in TIME_WAIT and the ephemeral port range, this is the cause. Ports from closed connections also aren’t reused for the same destination right away (Azure applies a cooldown; Google Cloud blocks them during TIME_WAIT), so the more you repeat short connections, the sooner you hit the limit.
Sources
NAT gateway basicsAWS 55,000 concurrent connections per IPv4 address to the same destination (destination IP, port, protocol), expandable by attaching up to 8 IPs (public NAT gateways get 2 Elastic IPs by default, more through a quota increase request); bandwidth scales automatically from 5 to 100 Gbps and throughput from 1 million to 10 million packets per second, and packets beyond that limit are dropped
NAT gateway metrics and dimensionsAWS ErrorPortAllocation: number of times a source port couldn’t be allocated (above 0 means too many concurrent connections), ActiveConnectionCount, IdleTimeoutCount (connections cleaned up after 350 seconds idle), PacketsDropCount
Troubleshoot NAT gatewaysAWS Connections expire after 350 seconds idle and later sends get an RST; keepalives shorter than 350 seconds recommended; when hitting the connection limit, add gateways per availability zone, add IPs, or reduce connections
Source Network Address Translation (SNAT) with Azure NAT GatewayMicrosoft Azure 64,512 SNAT ports per public IP (up to 16 IPs); each connection to the same destination needs a different port; closed ports go through a cooldown before reuse for the same destination
Metrics and alerts for Azure NAT GatewayMicrosoft Azure SNAT Connection Count filtered by the Failed state above 0 suggests SNAT port exhaustion; Dropped Packets
IP addresses and portsGoogle Cloud 64,512 ports each for TCP and UDP per NAT IP; default minimum ports per VM is 64 (static allocation) or 32 (dynamic allocation); the number of ports reserved for a VM caps its concurrent connections to the same destination; ports of closed connections can’t be used during TIME_WAIT
Logs and metricsGoogle Cloud dropped_sent_packets_count with reason OUT_OF_RESOURCES: packets dropped for lack of NAT IPs or ports