NVIDIA has delivered its first Vera CPUs to leading AI labs including Anthropic, OpenAI, and SpaceX AI, signaling a major expansion beyond graphics processors into bespoke data center silicon. The Vera CPU is specifically architected for agentic AI inference—where autonomous systems execute reasoning and planning tasks—delivering up to 10 times lower cost per token compared to traditional solutions. At Dell Technologies World, CEO Jensen Huang highlighted the strategic shift, noting that "demand is going parabolic" and positioning Vera as critical infrastructure for the next phase of enterprise AI deployment.

The Vera architecture addresses a distinct workload gap in NVIDIA's portfolio. While the company's GPUs dominate training and traditional inference, agentic systems require different performance characteristics: lower latency for rapid decision-making, optimized memory bandwidth for context retrieval, and cost-efficient processing for long-running agent sandboxes. Testing shows agent sandboxes run 50% faster on Vera than traditional CPUs, while enterprise data queries execute up to 3x faster. Early adopters like Lilly, Samsung, and Honeywell—representing 5,000 enterprises total—are already running AI workloads on the platform, validating demand.

Vera's launch represents NVIDIA's strategic response to increasing CPU competition in data centers. By controlling both GPU and CPU design, NVIDIA can optimize the full compute stack for AI-specific workflows, deepening customer lock-in through the CUDA ecosystem. The move complements existing partnerships with cloud providers like Google Cloud, which continue accelerating developer adoption through curated learning paths and hands-on labs. As agentic AI infrastructure becomes increasingly specialized, NVIDIA's vertical integration positions the company to capture significant margin in this emerging market segment.