NVIDIA has begun delivering its first custom CPU, codenamed Vera, to top-tier AI research labs including Anthropic, OpenAI, and SpaceX's AI division, signaling a fundamental shift in the company's hardware strategy. The move represents NVIDIA's entry into the inference compute market—the phase where trained models run at scale to serve users—rather than relying solely on its dominant GPU franchise for both training and deployment. CEO Jensen Huang framed the opportunity at Dell Technologies World as demand that is "going parabolic," noting that agentic AI systems, which execute multi-step reasoning tasks autonomously, require different silicon characteristics than the matrix-multiplication-heavy training workloads that made H100s and H200 GPUs industry standards.

Vera is optimized for inference tasks where cost per token and latency matter more than raw throughput. According to NVIDIA's claims, the CPU delivers agentic AI inference at roughly one-tenth the cost per token compared to traditional processor-based inference, while sandboxed agent operations run 50 percent faster than legacy CPU alternatives and enterprise database queries execute three times faster. The distinction is critical: while GPUs excel at parallel computation during model training, inference often involves sparse, irregular workloads where CPU efficiency gains compound. Early adopters already include over 5,000 enterprises spanning pharma (Eli Lilly), manufacturing (Samsung), and industrial automation (Honeywell), suggesting substantial market validation before broader availability.

The Vera announcement arrives as competitors including AMD and Intel scramble to design inference-specific silicon. NVIDIA's vertical integration—controlling both GPU and CPU supply chains—gives it architectural flexibility that fabless rivals cannot match, though custom silicon startups remain threats in niche segments. The infrastructure implications are profound: cloud providers and enterprises that previously purchased GPUs exclusively may now adopt Vera for inference clusters, fragmenting the addressable market but potentially enlarging NVIDIA's total serviceable addressable market by capturing the cost-conscious deployment segment. This two-tier approach—premium GPUs for training, efficient CPUs for inference—signals NVIDIA's confidence that the AI hardware cycle has matured beyond pure scaling.