NVIDIA has shipped its first-generation Vera CPU to top-tier AI research facilities including Anthropic, OpenAI, and SpaceX AI, marking a significant expansion beyond the company's dominant GPU portfolio into custom processors. The Vera NVL72 is purpose-built for agentic AI inference workloads, addressing a critical gap in infrastructure as enterprises scale autonomous AI systems. This move reflects NVIDIA's recognition that the compute landscape is evolving beyond training-focused architectures toward specialized inference engines optimized for agent reasoning and decision-making at scale.

According to NVIDIA CEO Jensen Huang, Vera delivers agentic AI inference at one-tenth the cost per token compared to traditional solutions, a compelling efficiency metric that directly impacts enterprise economics at scale. Agent sandboxes run 50% faster on Vera than CPU alternatives, while enterprise data queries execute up to 3x faster. Early deployments at companies like Eli Lilly, Samsung, and Honeywell across 5,000 enterprise deployments demonstrate immediate production viability and market confidence in the architecture.

The Vera launch signals NVIDIA's strategic pivot toward comprehensive infrastructure solutions as AI workloads mature beyond foundational model training. By controlling both GPU and CPU tiers through custom silicon, NVIDIA strengthens its vertically integrated position in the AI stack while addressing the distinct computational requirements of inference-heavy agentic systems. This positions the company to capture value across the entire AI compute spectrum as enterprise adoption of autonomous agents accelerates.