NVIDIA has begun shipping its first custom-designed CPU, Vera, to leading AI research labs including Anthropic, OpenAI, and SpaceX AI, signaling a major architectural shift beyond its dominant GPU business. The Vera NVL72 is purpose-built for agentic AI inference, delivering a 50 percent performance improvement over traditional CPUs for agent sandboxes while reducing inference costs to one-tenth per token compared to GPU-based alternatives. This represents NVIDIA's strategic bet that the next wave of AI infrastructure will demand specialized processors tuned for autonomous agent workloads rather than generic compute.

The timing reflects market maturation in enterprise AI deployment. CEO Jensen Huang stated at Dell Technologies World that demand is "going parabolic," with 5,000 enterprises like Eli Lilly, Samsung, and Honeywell already running AI workloads on NVIDIA infrastructure. Vera's arrival at Oracle Cloud Infrastructure and deployment across hyperscaler platforms indicates that inference efficiency—not just raw training performance—has become a critical differentiator. The CPU delivers 3x faster enterprise data queries, addressing the bottleneck of real-time agent decision-making in production environments.

Vera's launch underscores NVIDIA's expansion into a full-stack AI platform beyond GPUs and CUDA. Parallel collaborations with Google Cloud to empower 100,000+ developers and announcements at GTC Taipei demonstrate NVIDIA's broader infrastructure vision encompassing CPUs, data center architecture, and developer ecosystems. This vertical integration strategy aims to lock in enterprise adoption by optimizing every layer of the AI infrastructure stack, making NVIDIA indispensable not just for training but for the cost-sensitive inference workloads that define profitability at scale.