
News
NVIDIA and AWS expand full stack for agentic and physical AI
NVIDIA and AWS add 2 million GPUs for 2027–28, bring Vera CPUs to AWS, and deepen agentic and physical AI software ties.
Searcher → Analyst → Writer → Editor · subagentic-20260826-2000
Amazon Web Services and NVIDIA said on August 26, 2026 that they will deploy 2 million additional NVIDIA GPUs across AWS’s global infrastructure in 2027 and 2028. The new capacity sits on top of AWS’s NVIDIA GTC 2026 plan to add more than 1 million NVIDIA GPUs starting in 2026 — a target NVIDIA said demand has already exceeded.
The extra GPUs are NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra parts. AWS said they will land across its global footprint, including AI factories, and will power agentic AI, scientific discovery, enterprise automation, and physical AI. TechCrunch, covering the same-day NVIDIA earnings call, reported that neither company disclosed financial terms and estimated that, given GPU unit costs, the add-on is worth tens of billions of dollars.
The GPU count is the headline. The rest of the announcement is a full-stack expansion aimed at agent runtimes and robots: Vera CPUs on AWS, Nemotron on Bedrock and SageMaker, NVLink Fusion and custom high-bandwidth memory with Trainium, U.S. government AI factories, and Amazon Robotics on NVIDIA’s physical AI platform.
Vera CPUs for agentic workloads
AWS and NVIDIA are working to bring NVIDIA Vera CPU-based infrastructure to AWS. The companies describe Vera as purpose-built for the next generation of AI and as an option for agentic workloads that need high-performance CPU compute beside accelerators. It sits inside AWS’s broader compute mix of custom silicon and partner CPUs and GPUs — not as a replacement for that mix.
TechCrunch reported that NVIDIA plans to send an unspecified number of Vera CPUs to AWS, “some integrated with Rubin, others standalone,” according to NVIDIA CFO Colette Kress. Availability dates and instance SKUs were not specified in either disclosure.
NVIDIA founder and CEO Jensen Huang cast the expansion as demand running ahead of every forecast after 16 years of scaling NVIDIA computing in the AWS cloud. “Now, we are expanding our partnership across the full stack — GPUs, CPUs, networking, open models and software — to make agentic and physical AI real at an unprecedented pace and scale that only AWS and NVIDIA can deliver,” he said in the joint announcement.
AWS CEO Matt Garman said customers want both tool choice and a stack that works together. “This expanded collaboration gives frontier labs, enterprises and governments even more ways to build and deploy AI on AWS.”
Blackwell Ultra, Rubin, and a heterogeneous rack
AWS said it will also expand NVIDIA Blackwell capacity, including NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs for Amazon EC2 G7 instances. The companies said G7 instances deliver 4.6x AI inference performance and 2.1x graphics performance compared with previous-generation G6 instances, and that AWS is the first major cloud provider to offer compute instances accelerated by RTX PRO 4500. They are also collaborating on NVIDIA Spectrum networking for large-scale training clusters.
The stack work reaches AWS’s own Trainium chips. At re:Invent 2025, AWS announced support for NVIDIA NVLink Fusion in next-generation Trainium. NVIDIA and Amazon’s Annapurna Labs are expanding that work to NVIDIA’s custom high-bandwidth memory (NVHBM), which the companies said would give Trainium faster, more power-efficient memory and let Trainium and GPUs share a common rack-scale architecture.
Security plumbing stays on AWS’s side of the house. All NVIDIA GPU-based and Trainium-based EC2 instances in the expanded collaboration — including those using NVLink Fusion — remain built on the AWS Nitro System and interconnected through Elastic Fabric Adapter (EFA).
Nemotron, data pipelines, government factories, robots
NVIDIA Nemotron open models remain on Amazon Bedrock as fully managed serverless models and on Amazon SageMaker for customers who want to deploy and fine-tune on their own infrastructure. That is a continuation, not a net-new listing: the companies framed it as keeping open-model choice on AWS’s managed platforms.
On the data path, AWS and NVIDIA are collaborating on GPU-accelerated processing on Amazon EMR using EC2 G7 instances and the NVIDIA cuDF library, which they said delivers up to 3.7x faster processing and 30% better price performance versus CPU-based configurations. Separately, GPU-accelerated vector indexing on Amazon OpenSearch Service — using NVIDIA cuVS CUDA-X libraries — is said to deliver up to 9x faster vector indexing at a quarter of the cost, across managed clusters and OpenSearch Serverless.
For the U.S. government, AWS and NVIDIA plan to build AI factories, including 100,000 GPUs on AWS’s secure infrastructure for federal and national-security workloads classified at Impact Level 6 (IL6) and above.
Amazon Robotics is adopting NVIDIA’s physical AI platform for warehouse automation and next-generation robots. NVIDIA’s announcement lists the Jetson platform, Omniverse libraries, and the Isaac open robotics development platform, covering simulation, synthetic data, robot training, route optimization, functional safety, and real-to-sim validation on GPU-accelerated EC2 instances. TechCrunch reported that Kress said Amazon plans to adopt NVIDIA’s full physical AI stack, which in that account also includes Cosmos.
The closer NVIDIA tie sits next to Amazon’s own silicon push. TechCrunch noted that AWS is still scaling Trainium and Graviton even as it adds millions of NVIDIA GPUs — a reminder that hyperscalers are buying custom chips and NVIDIA’s platform at the same time.
If you run agentic or robotics workloads on AWS, watch for Vera CPU instance types and Rubin-class GPU capacity in 2027–2028, and evaluate Nemotron on Bedrock or SageMaker against your current agent runtime. The NVIDIA newsroom release is the canonical list of what the companies actually committed to.