When a cloud provider performs maintenance on a physical server (host), it moves VMs to another host (live migration) or pauses them briefly. The whole server freezes during that time, and if the pause is long, connections drop.
Why The provider moves the VM to another host, or pauses it briefly, for host maintenance or a predicted failure → Effect During the move, CPU, memory, and network slow down, and at the end the VM stops completely for a moment (from under 1 second to around 30 seconds, depending on the provider and method) → On screen Everyone on the server freezes at once and then sees fast-forward and teleporting; if the freeze outlasts the timeout, mass disconnects
Primary owner Infra team (Server infrastructure) · Also Game team (Server development), External (External)
Game team action items
Use timeouts that survive pauses of a few seconds, cap how many ticks the server catches up after a pause, compute elapsed time with a monotonic clock, have a procedure that saves progress and moves players to another server when a maintenance notice arrives.
Infra team action items
Subscribe to and alert on maintenance notices (Google Cloud maintenance-event, AWS scheduled events and AWS Health, Azure Scheduled Events), replace servers ahead of time during low-traffic hours when a notice arrives, reschedule maintenance where the provider allows it (Azure Maintenance Configuration, AWS scheduled events depending on type), compare maintenance records with incident records.
External action items
Ask the cloud provider about maintenance schedules and impact, report instances that keep pausing.
Ballpark numbers
Google Compute Engine says live migration pauses are usually much shorter than 1 second, and the system clock can jump forward by up to 5 seconds during the pause. The maintenance-event metadata value changes 60 seconds before the move (if you have queried it at least once beforehand). On Azure, maintenance that doesn’t need a reboot almost always pauses the VM for under 10 seconds, and rarely (no more than once every 18 months for general-purpose sizes) for about 30 seconds; live migration usually takes no more than 5 seconds. Azure Scheduled Events gives notice of these pauses (Freeze) at least 15 minutes ahead. If host hardware fails suddenly, though, recovery starts right away with no notice.
On the graph
Gap then burst · Server packets sent/received, tick interval
Where to look
Match the freeze time against the provider’s records. Google Cloud: compute.instances.migrateOnHostMaintenance in the audit logs; AWS: scheduled events in describe-instance-status and AWS Health; Azure: Microsoft.Compute/virtualMachines/liveMigration/action in the Activity Log and the time the VM availability metric (VmAvailabilityMetric) dropped to 0. Inside the server, check whether metrics and logs have a gap during the freeze and whether the clock jumped right after (time sync logs)
Confirmed if
The time the whole server froze overlaps with a maintenance or migration time in the provider’s records, and every metric and log inside the server is blank for those few seconds
Ruled out if
Not in the provider’s records and short freezes recur often: “CPU steal (virtual machines).” NIC reset entries in the kernel log: “NIC driver and firmware problems”
Check with
Infra tools (no game code needed)
Learn more
AWS gives notice through scheduled events. system-reboot means the instance will be rebooted and moved to a new host; system-maintenance means network or power maintenance may affect it briefly. Even if the pause lasts only a few seconds, clients that got no ACK for packets sent to the server during that time keep doubling their retransmission wait, so TCP connections can stay stalled for longer after the pause ends (“TCP RTO and exponential backoff”). When the VM wakes up, its clock can jump and lead to a “System clock jump (NTP step),” and failed load balancer health checks may take the server out of rotation for a while. Instances that can’t be moved (such as Google Cloud bare metal instances) are stopped or restarted during maintenance.
Sources
Live migration process during maintenance eventsGoogle Cloud Live migration pauses are usually much shorter than 1 second; the system clock jumps forward by up to 5 seconds during the pause; disk, CPU, memory, and network performance drop briefly during the move; VMs that don’t live-migrate are terminated for maintenance (bare metal instances don’t support live migration)
Query metadata server for maintenance event noticesGoogle Cloud The maintenance-event metadata value changes 60 seconds before live migration (when the VM is set to live-migrate and the value was queried at least once since the last maintenance)
Scheduled events for Amazon EC2 instancesAWS Scheduled event types (system-reboot reboots and moves to a new host; system-maintenance means brief impact from network or power maintenance), notified by email and AWS Health, checked with describe-instance-status, reschedulable for some types
Maintenance and updatesMicrosoft Azure Maintenance without a reboot almost always pauses for under 10 seconds, rarely (no more than once every 18 months for general-purpose sizes) for about 30 seconds, and live migration usually 5 seconds or less; the clock syncs automatically after the pause; long-lived TCP connections may drop, or recovery may take longer as peers retransmit data sent to the paused VM with exponential backoff; load balancer health checks mark the VM unhealthy within about 10 seconds; confirm with Microsoft.Compute/virtualMachines/liveMigration/action in the Activity Log and VmAvailabilityMetric dropping to 0 during the pause; pick when maintenance applies with Maintenance Configuration
Scheduled Events for Linux VMs in AzureMicrosoft Azure Freeze (a pause of a few seconds; CPU and network may stop) is announced at least 15 minutes ahead; for host hardware failures, recovery starts right away with no notice period