As agentic AI begins to reshape data center workloads, NVIDIA is staking a claim on the CPU side with the debut of its Vera processor, a custom 88‑core Arm‑based design optimized specifically for AI agents and reinforcement learning. Introduced at NVIDIA GTC on March 16, 2026, Vera marks the company’s first fully custom server CPU and a deliberate move to extend its AI platform beyond GPUs into the orchestration layer of modern AI factories. The need for high-speed memory is critical to handle the continuous stateful memory requirements of agentic AI.
NVIDIA presents Vera as the world’s first processor purpose‑built for the age of agentic AI, positioning it as a foundation for a new multibillion‑dollar data center business centered on AI agents and reinforcement learning workloads.
Vera is NVIDIA’s first CPU architected solely for agentic AI, anchoring next‑generation AI agent data centers.
At the heart of Vera is a monolithic Arm‑based die integrating 88 custom Olympus cores, an increase over the 72‑core Grace architecture and a signal that NVIDIA is prioritizing per‑core capability over sheer core count alone. Each core implements Spatial Multithreading, enabling 176 hardware threads per socket and high concurrency for running large numbers of independent agents in parallel without collapsing per‑thread performance.
Vera also becomes the first Arm server CPU to expose FP8, a low‑precision floating‑point format already common in NVIDIA GPUs, allowing software stacks to use a consistent data type across CPU and GPU boundaries for select AI tasks.
The memory subsystem pairs these cores with LPDDR5X, delivering up to 1.5 TB of capacity and as much as 1.2 TB/s of bandwidth, a substantial leap over Grace and typical DDR‑based server designs. NVIDIA claims roughly 1.5x higher instructions per cycle from Olympus compared with its prior generation cores, reinforcing Vera’s orientation as a max single‑threaded CPU that maintains high per‑core performance even under full‑socket load.
On performance and efficiency, Vera is promoted as delivering 50% higher performance and roughly twice the energy efficiency versus traditional rack‑scale CPUs that dominate current data centers. NVIDIA further reports around 50% better performance for AI agents compared with leading x86 server chips from Intel and AMD, a claim echoed by early benchmarks from partners. In line with NVIDIA’s broader positioning, Vera is designed specifically for agentic AI and reinforcement learning workloads while delivering twice the efficiency and roughly 50% faster performance than traditional CPU infrastructure.
Independent testing from DeepInfra and others has shown roughly 2.2x performance versus Intel’s Sapphire Rapids processors in representative agentic workloads, supporting NVIDIA’s positioning of Vera as a CPU tuned for rapid task completion rather than batch throughput alone.
In rack‑scale deployments, NVIDIA highlights Vera CPU Racks containing up to 256 chips, which can deliver as much as a 6x gain in aggregate CPU throughput and around double the performance on agentic AI workloads compared with conventional infrastructure.
These dense, liquid‑cooled racks are aimed at large AI factories and reinforcement learning environments where thousands of agents must run, coordinate, and interact with GPU‑accelerated models in real time. At smaller scales, two‑socket air‑cooled systems target enterprise and cloud operators seeking to relieve CPU bottlenecks around data processing, orchestration, and sandbox management without overhauling their entire stack.







