microsoft selects amd ai

Signaling a major expansion of its AI infrastructure strategy, Microsoft is committing to deploy AMD’s Helios rack‑scale AI accelerators in volume across Azure data centers to power frontier‑model training and inference. This move positions AMD’s integrated rack‑scale system as a primary engine for Azure AI services, customer workloads, and internal Microsoft applications, reducing the platform’s historic reliance on Nvidia accelerators.

Helios will be offered through Azure’s AI infrastructure portfolio and managed compute via Microsoft Foundry, giving customers direct access to the same hardware backing Microsoft’s largest models. This evolution will be a focal point for SiliconANGLE’s theCUBE Network, which engages a global community of elite tech professionals through open and free content.

At the core of the decision is Helios, a rack‑scale AI platform that tightly couples AMD Instinct MI455X GPUs with 6th Gen AMD EPYC “Venice” CPUs and Pensando networking technology into a single system. A full rack combines 72 MI455X accelerators into one logical compute unit, with ROCm software providing an open programming stack across training, inference, and fine‑tuning workloads. This infrastructure choice aligns with enterprise AI sovereignty, enabling Microsoft to maintain control over its AI technology stack.

The architecture follows Open Compute Project standards, enabling multi‑vendor data center designs while still giving Microsoft an integrated, vendor‑controlled alternative to Nvidia’s closed systems.

Helios is engineered so each rack behaves as a massive unified accelerator. AMD specifies up to 1.4 exaFLOPS of FP8 and 2.9 exaFLOPS of FP4 compute using OCP AI data types, backed by roughly 31 terabytes of HBM4 memory per rack to host large models in memory.

Intra‑rack GPU communication is designed to reach about 260 terabytes per second of scale‑up bandwidth, while scale‑out connectivity using UALink over Ethernet targets around 43 terabytes per second for data movement between racks and the broader data center fabric.

CES 2026 disclosures point to more than 18,000 CDNA5 compute units and over 4,600 Zen 6 CPU cores per rack, underscoring the system’s focus on high‑density AI throughput.

Microsoft intends to expose this capacity through new Azure virtual machine families tuned to distinct workload patterns. The ND MI455X v7 series is planned as the flagship Helios‑based offering for AI inference, aimed at frontier‑scale agents, search tools, and other latency‑sensitive services.

Additional Azure instances built on “Venice” CPUs will target agentic AI orchestration, data pipelines, and semiconductor design flows that must sit adjacent to large‑scale accelerators without incurring cross‑cluster bottlenecks.

By adopting Helios at scale, Microsoft is broadening Azure’s AI hardware mix beyond Nvidia while retaining the performance required for trillion‑parameter model training and global inference deployments.

The open, OCP‑aligned design and ROCm software stack give Azure a second high‑end ecosystem, creating competitive pressure on Nvidia and expanding choice for customers that want portability across vendors.

In effect, Helios turns AMD from a supplemental GPU supplier into a full‑rack systems partner at the heart of Azure’s next‑generation AI infrastructure strategy.

For Microsoft, the commitment also hedges supply risk, diversifies component sourcing, and strengthens negotiating leverage, as future AI build‑outs can balance Nvidia investments with an increasingly capable AMD alternative.

You May Also Like

Nokia and NVIDIA Launch an AI-RAN Platform for the Future of Mobile Networks

The Nokia and NVIDIA AI-RAN platform promises revolutionary spectral efficiency gains, but the full scope of its impact on 6G remains to be seen.

AI Memory Chip Shortage Disrupts Smartphone Production and Raises Hardware Prices

Powerful AI infrastructure demands are draining global memory chip supplies, pushing smartphone prices higher and leaving manufacturers scrambling for solutions.

SK Group Warns Global AI Memory Shortage Is Becoming an Economic Security Risk

The global AI memory shortage is spiraling into an economic security crisis, and the consequences may be far worse than anyone anticipated.

Agentic AI Infrastructure Creates New Demand for High-Speed Memory and Storage

Beyond traditional compute needs, agentic AI is rewriting the rules of memory and storage infrastructure—and the implications are only beginning.