AWS NVIDIA GPU Expansion: 7 Powerful Infrastructure Lessons

AWS and NVIDIA plan to deploy 2 million additional GPUs during 2027 and 2028. Here is what the expansion means for data center power, cooling, networking, and operations.

AWS NVIDIA GPU expansion infrastructure inside a high-density data center

The AWS NVIDIA GPU expansion will add 2 million NVIDIA GPUs across Amazon Web Services’ global infrastructure during 2027 and 2028, significantly increasing the companies’ joint AI computing capacity. The deployment will include Blackwell Ultra, Rubin, and Rubin Ultra systems while extending the partnership into CPUs, networking, custom memory, government AI infrastructure, and physical AI.

AWS and NVIDIA announced the expansion on August 26, saying customer demand had exceeded earlier expectations. AWS had previously announced plans to add more than 1 million NVIDIA GPUs beginning in 2026. The new commitment shows how quickly hyperscale AI infrastructure requirements are moving beyond individual server orders and becoming multiyear data center development programs.

1. What the AWS NVIDIA GPU Expansion Includes

The additional 2 million GPUs are planned for deployment across AWS’s global infrastructure, including purpose-built AI factories. The companies say the capacity will support workloads ranging from agentic AI and scientific computing to enterprise automation, simulation, and robotics.

The agreement goes considerably further than GPU procurement. AWS also plans to introduce infrastructure based on NVIDIA Vera CPUs. The companies are extending their work around NVIDIA Spectrum networking and integrating NVIDIA technologies with the AWS Nitro System and Elastic Fabric Adapter.

NVIDIA Nemotron open models will continue to be offered through Amazon Bedrock and Amazon SageMaker. Together, these commitments connect chips, servers, networks, models, cloud services, and facilities in one infrastructure roadmap.

2. Two Million GPUs Create a Data Center Challenge

The headline number is extraordinary, but the physical infrastructure required to operate those accelerators may be more consequential for data center leaders. Millions of high-performance GPUs require enormous quantities of electricity, cooling capacity, network bandwidth, floor space, storage, and supporting electrical equipment.

The deployment will be distributed across multiple facilities rather than concentrated in one location. Even so, its aggregate scale illustrates how cloud providers now plan AI capacity. AWS must ensure that data centers can energize, cool, connect, and operate successive GPU generations as quickly as the processors become available.

This makes the AWS NVIDIA GPU expansion dependent on the wider physical supply chain. Utility connections, transformers, switchgear, generators, cooling equipment, optical components, construction labor, and commissioning resources all need to arrive on compatible schedules.

3. Power and Cooling Become Strategic Constraints

Advanced AI systems can concentrate substantial electrical demand inside each rack. Moving from Blackwell Ultra to Rubin and Rubin Ultra may also change rack designs, power distribution, thermal loads, and cooling requirements.

Infrastructure teams therefore cannot assume that a data hall designed for one generation will accept the next without modification. They need flexible electrical distribution, liquid-cooling capability, accurate capacity models, and operating procedures that account for high-density equipment.

At this scale, efficiency improvements have material financial consequences. Small gains in power usage effectiveness, cooling efficiency, or GPU utilization can affect megawatts of demand across the fleet. Conversely, a bottleneck in facility infrastructure can strand expensive computing capacity.

4. AWS Is Combining NVIDIA With Custom Silicon

The partnership is notable because AWS is not abandoning its custom-silicon strategy. Amazon’s Annapurna Labs develops Trainium accelerators for AI workloads, and AWS previously announced support for NVIDIA NVLink Fusion in future Trainium systems. The expanded collaboration also adds work around NVIDIA’s custom high-bandwidth memory technology.

This points toward increasingly heterogeneous AI infrastructure rather than a winner-takes-all processor market. Cloud operators may combine custom accelerators with NVIDIA GPUs, CPUs, networking, and memory technologies according to workload economics, availability, and performance requirements.

For customers, heterogeneous infrastructure could offer more choice. It also makes benchmarking more important. Organizations should compare useful workload output, energy consumption, software compatibility, and total cost rather than relying only on peak processor specifications.

5. Government AI Capacity Adds Security Requirements

AWS and NVIDIA also plan AI factories for the U.S. government, including infrastructure containing 100,000 GPUs for federal and national-security workloads. The companies say the systems are intended for workloads classified at Impact Level 6 and above.

That adds another dimension to the buildout. Government AI capacity must satisfy performance and availability requirements while meeting strict controls for security, isolation, compliance, data handling, personnel access, and operational oversight.

These requirements can influence facility selection, network architecture, supply-chain assurance, monitoring, and incident response. A high-performance cluster is not operationally useful for sensitive workloads unless the complete environment meets the required security standard.

6. Networking Determines Useful GPU Capacity

A GPU fleet measured in millions places enormous pressure on network architecture. AWS and NVIDIA say they will continue integrating GPU-accelerated and Trainium-based EC2 systems with Elastic Fabric Adapter and AWS Nitro while collaborating further on NVIDIA Spectrum networking for large AI training clusters.

The reason is straightforward: adding accelerators without sufficient scale-out network capacity can leave expensive compute waiting for data. Congestion control, topology design, optics, telemetry, failure recovery, and storage access all influence real cluster performance.

For AI infrastructure, network performance increasingly determines how effectively installed GPU capacity becomes useful compute. Readers looking for additional context can review our guide to AI factories and accelerated-computing data centers.

7. Operations Must Scale With the Hardware

The AWS NVIDIA GPU expansion is also an operational challenge. Deploying equipment is only the beginning. AWS must maintain availability across many regions, hardware generations, network fabrics, cooling systems, and software layers.

Operators need repeatable commissioning, configuration management, capacity forecasting, spare-parts strategies, preventive maintenance, and incident-response procedures. Observability must connect facility conditions with IT performance so teams can identify whether a workload problem originates in software, networking, servers, cooling, or power.

Training also matters. As infrastructure becomes more heterogeneous, operations teams require skills that cross traditional boundaries between facilities, networking, hardware, and platform engineering.

What Data Center Leaders Should Watch Next

The additional GPUs are planned for 2027 and 2028, so the announcement represents future capacity rather than infrastructure available today. Operators should watch how quickly AWS brings the necessary power, cooling, networking, and building capacity online as NVIDIA progresses through Blackwell Ultra, Rubin, and Rubin Ultra.

Practical infrastructure readiness checklist

  • Validate power delivery: confirm utility milestones, substation capacity, backup-power design, and equipment lead times. Our guide to AI data center microgrids explains options for resilient onsite capacity.
  • Plan liquid cooling: model rack heat loads, coolant distribution, water chemistry, leak response, and maintenance access. Review these closed-loop liquid-cooling design considerations.
  • Test the network: benchmark topology, congestion control, optics, telemetry, and failure recovery before production deployment. See our analysis of GPU cluster architecture.
  • Prepare operations: define commissioning gates, staffing, spares, escalation paths, and cross-team drills. Use this AI-ready data center operations checklist as a planning reference.
  • Measure useful output: track application throughput, GPU utilization, energy per workload, job completion time, and service availability rather than installed accelerator count alone.

Important indicators include new region and availability-zone announcements, utility agreements, data center construction, liquid-cooling deployments, network upgrades, and the commercial availability of new EC2 instances. Official updates from Amazon Web Services and NVIDIA will help distinguish contracted future capacity from systems that customers can use.

The scale also raises a competitive question for enterprises: will scarce high-end AI compute become easier to consume through hyperscale clouds, or will more organizations pursue dedicated infrastructure to secure predictable capacity? The answer will depend on price, availability, data governance, workload scale, and internal operating expertise.

Frequently Asked Questions

When will the additional NVIDIA GPUs be deployed?

AWS and NVIDIA plan to deploy the 2 million additional GPUs during 2027 and 2028. The capacity is planned infrastructure and is not available in full today.

Which NVIDIA platforms are included?

The announcement includes Blackwell Ultra, Rubin, and Rubin Ultra systems, alongside broader collaboration involving Vera CPUs, Spectrum networking, custom memory, and NVIDIA software and models.

Why does the expansion matter to data center operators?

Operating millions of accelerators requires corresponding investment in power, liquid cooling, networking, storage, buildings, security, maintenance, and skilled personnel. The supporting infrastructure determines how much installed GPU capacity becomes productive compute.

Conclusion

The AWS NVIDIA GPU expansion is much more than an order for additional processors. It combines GPUs, CPUs, custom accelerators, high-bandwidth memory, networking, government infrastructure, models, security, and physical data center capacity into one cloud AI strategy.

For data center executives, the lesson is increasingly difficult to ignore: AI compute cannot scale independently of the infrastructure underneath it. Deploying 2 million additional GPUs means solving power, cooling, networking, security, and operational challenges at comparable scale.

THE INFRASTRUCTURE BRIEFING

Essential data center intelligence delivered to your inbox.


By: