AWS directs engineers to reduce CPU waste amid EC2 capacity strain

1 hour ago 20

Amazon Web Services, the world’s largest cloud provider, has issued internal directives requiring its engineering teams to cut CPU waste and clean up underutilized EC2 instances. The reason is straightforward: demand for compute, largely driven by AI workloads, is outpacing supply, and engineers requesting new CPU capacity are now waiting days to get it.

What’s actually happening inside AWS

The directives trace back to internal meetings held in May, where leadership stressed the need to maximize existing infrastructure before expecting new hardware to arrive. Engineers have been given deadlines to identify and reduce idle EC2 instances, a category that reportedly accounts for roughly 65% of all running instances being underutilized.

The optimization playbook AWS is pushing internally includes better rightsizing (matching instance types to actual workload needs), automated shutdowns for inactive instances, and broader efficiency improvements across how teams provision and manage resources.

The multi-day delays for new CPU server capacity are the most telling symptom. AWS has historically operated with enough headroom that internal teams could spin up resources relatively quickly. When provisioning timelines stretch from hours to days, the infrastructure is running closer to the redline than anyone would prefer.

AI is eating the data center

AWS isn’t alone in feeling the squeeze. Microsoft Azure and Google Cloud have both acknowledged capacity constraints in recent earnings calls, with GPU availability in particular becoming a bottleneck across the industry. But CPU capacity getting tight at AWS is a newer development, and it suggests the strain is spreading beyond the GPU clusters that get most of the attention.

Amazon has been investing heavily in custom silicon, including its Graviton processors for general compute and Trainium chips for AI training, partly to reduce dependence on external suppliers and partly to improve performance per watt. But chip fabrication timelines mean those investments take quarters or years to translate into deployed capacity.

What this means for cloud customers and the broader market

For companies running workloads on AWS, internal capacity strain at the provider level can surface in several ways. Certain instance types may become harder to reserve in specific regions. Spot instance pricing, which fluctuates based on available surplus capacity, could rise as that surplus shrinks. AWS has already adjusted published pricing of reserved capacities, including a notable approximately 15% increase across GPU families in early 2026.

The enterprise customers most likely to feel the impact are those running resource-intensive applications without reserved capacity commitments. Companies that locked in capacity through savings plans or reserved instances are somewhat insulated. Those relying on on-demand provisioning may find themselves dealing with availability hiccups or paying premium pricing during peak demand periods.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article