AWS and NVIDIA to Deploy 2 Million GPUs by 2028 for Enterprise AI Factories
AWS commits to 2 million additional NVIDIA Blackwell, Rubin GPUs across global infrastructure for agentic and physical AI workloads. Enterprise buyers gain capacity certainty but face deeper vendor lock-in.
AWS anchors multi-year GPU capacity with concrete numbers
AWS and NVIDIA announced plans to deploy 2 million additional next-generation GPUs across AWS global infrastructure between 2027 and 2028, targeting what both companies call "AI factories" for continuous agentic and physical AI workloads. The commitment covers NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra architectures and represents the first time a hyperscaler has publicly attached a seven-figure GPU count to a multi-year deployment plan.
For enterprise buyers planning large-scale AI programs—autonomous systems, real-time personalization engines, or continuous model training—the announcement provides a quotable capacity target to reference in vendor negotiations and long-term architecture decisions. It also clarifies AWS's position relative to Microsoft Azure and Google Cloud, neither of which disclosed comparable GPU figures in the past week, even though both operate multi-million-GPU fleets.
The collaboration goes beyond GPU access. AWS and NVIDIA are co-engineering infrastructure spanning GPUs, CPUs, networking, data processing, and robotics, signaling tighter integration than standard cloud GPU offerings. That depth increases the case for standardizing on AWS if NVIDIA GPUs are non-negotiable, but it also raises vendor lock-in risk for organizations that tie architecture decisions to AWS-specific stacks.
Capacity certainty versus multi-cloud flexibility
The 2 million GPU figure addresses a persistent enterprise concern: whether hyperscalers can deliver the GPU capacity needed for 2027–2028 AI roadmaps. Publicly committing to this scale reduces perceived capacity risk for very large reservations, especially for globally distributed workloads that require consistent availability across regions.
For procurement teams already budgeting multi-year GPU reservations, AWS is signaling it will scale capacity aggressively. That supports negotiations for reserved-instance discounts in exchange for long-term commitments. Organizations planning to reserve tens of thousands of GPU-hours annually can now reference AWS's public capacity commitment as leverage when comparing hyperscaler proposals.
The risk is that the AWS–NVIDIA collaboration's depth—particularly around agentic AI and robotics—creates architectural pressure to consolidate workloads on AWS rather than maintain multi-cloud strategies. Enterprises that prize vendor diversity or data sovereignty may find themselves defending multi-cloud architectures against internal teams who see AWS's announced scale as evidence that alternative providers cannot match hyperscaler capacity.
Argentum AI's $10 billion position in the specialist tier
Argentum AI, a global AI cloud platform, announced more than $10 billion in contracted revenue and a target of 1 gigawatt of AI infrastructure capacity in 2026. The company positions itself as a vertically integrated alternative to hyperscalers, offering hardware-as-a-service, bare-metal compute with enterprise SLAs, and managed Kubernetes environments across multiple geographies.
The contracted revenue figure is notable because it places Argentum AI closer in scale to hyperscaler narratives than typical startup GPU clouds. One gigawatt of capacity translates to tens of thousands of GPUs and hundreds of megawatts of power, a scale that competes directly with specialist providers like Nebius, Nscale, and CoreWeave, all of which are building sovereign or regionally optimized AI data centers.
For enterprises evaluating sovereign AI or data-residency-compliant infrastructure, Argentum AI's geographic flexibility and hardware-as-a-service model offer an alternative to hyperscaler lock-in. The $10 billion contracted revenue provides some assurance of financial stability, but buyers should verify actual deployed capacity and SLA performance before committing large workloads. Announced capacity targets and contracted revenue do not equal delivered infrastructure.
What enterprise buyers should do next
If your organization is planning AI workloads for 2027–2028, AWS's 2 million GPU commitment is a reference point for capacity planning conversations with all hyperscalers. Ask Azure and Google Cloud for comparable multi-year capacity commitments in writing, particularly for the geographies and GPU architectures your roadmap requires.
For multi-year reserved-instance negotiations, the public scale commitment creates an opening to request volume discounts or capacity guarantees. AWS now has a public benchmark to defend; use it to push for better terms, especially if you can commit to multi-year reservations.
If vendor diversity or data sovereignty is a strategic requirement, evaluate specialist providers like Argentum AI, Nebius, or CoreWeave against hyperscalers on three criteria: actual deployed capacity (not just announced targets), SLA guarantees for GPU availability, and contract terms for early termination or capacity reallocation. The hyperscaler capacity narrative will create internal pressure to consolidate on AWS; counter it with specific risk assessments of single-vendor concentration and data-residency requirements that hyperscalers cannot meet.
Finally, if your architecture is tightly coupled to NVIDIA GPUs, map the integration depth of AWS's co-engineered stack against your requirements for portability. The tighter the integration, the higher the switching cost if you need to move workloads to another provider in 2028 or beyond.
Technology decisions, clearly explained.
Weekly analysis of the tools, platforms, and strategies that matter to B2B technology buyers. No fluff, no vendor spin.
