NVIDIA delivered its first Vera CPUs to leading AI research labs last week—Anthropic, OpenAI, and SpaceX's AI division received units first, followed by Oracle Cloud Infrastructure. The move marks NVIDIA's entry into CPU design, a significant pivot from its GPU-centric business model. According to NVIDIA CEO Jensen Huang, speaking at Dell Technologies World, demand for agent inference hardware is 'going parabolic.' Vera targets this emerging workload category, promising agentic AI inference at one-tenth the cost per token compared to traditional GPU-based approaches, with agent sandboxes running 50% faster than CPU alternatives while enterprise data queries deliver 3x performance gains.

The Vera NVL72 represents a calculated bet on where AI compute is heading. Unlike general-purpose processors, Vera is purpose-built for inference-heavy agentic workflows—scenarios where AI systems make repeated decisions with lower latency requirements than training demands. By optimizing for this specific task, NVIDIA claims efficiency advantages over both standard CPUs and GPUs in this domain. The company reports 5,000 enterprises including pharmaceutical giant Eli Lilly, Samsung, and Honeywell are already running AI workloads that Vera could optimize. However, the broader availability timeline remains unclear, and competitors like AMD and Intel have not yet publicly responded to the Vera announcement.

The strategic rationale is compelling: as agentic AI becomes mainstream, the economics of inference matter intensely. Cost per token directly impacts enterprise profitability at scale. Yet questions linger about switching costs and architectural lock-in—why would enterprises retool inference pipelines for a new CPU when GPUs already handle the job? NVIDIA's answer hinges on TCO: if Vera genuinely cuts costs tenfold while maintaining compatibility with CUDA-trained models, migration calculus shifts dramatically. The real test comes with broader hardware availability and independent benchmarking from non-NVIDIA sources.