OpenAI has introduced Ultrafast mode for GPT-5.6 Sol, a new API service tier powered by Cerebras infrastructure that delivers up to 14 times faster inference speeds compared to standard tiers. The service achieves throughput of up to 750 output tokens per second, dramatically reducing latency for real-time applications. This capability addresses a critical bottleneck for enterprises building production AI systems, where speed directly impacts user experience and operational efficiency. The move positions OpenAI's latest model as increasingly viable for latency-sensitive workloads including agentic AI applications and customer-facing systems.

Complementing this technical advancement, OpenAI appointed Dali Rajic as Chief Revenue Officer, tasking him with leading the company's global revenue organization and helping businesses extract maximum value from AI systems. This executive hire underscores OpenAI's strategic pivot toward enterprise monetization at scale. Rajic's appointment arrives as enterprise adoption accelerates—OpenAI research shows companies are moving from experimental chatbot implementations toward production agentic AI systems that execute business-critical tasks autonomously.

Together, these moves reflect OpenAI's dual-track strategy: delivering technical infrastructure that enables faster, more efficient AI deployment while simultaneously building sales and leadership capacity to capture revenue from widespread enterprise adoption. The Ultrafast tier combined with a Responses API redesign aims to reduce both latency and costs, creating competitive advantages for startups and enterprises building AI agents. With these product and personnel moves, OpenAI is structuring itself to maintain market leadership as AI transitions from emerging technology to mission-critical enterprise infrastructure.