OpenAI's latest frontier model, GPT-5.6, has become the default intelligence powering Microsoft 365 Copilot across Word, Excel, PowerPoint, Teams Chat, and Cowork—a critical integration that extends OpenAI's influence deeper into enterprise workflows. This represents a material escalation from previous arrangements, where Microsoft rotated between multiple models depending on workload. GPT-5.6 claims stronger performance per token spent, a metric that directly impacts Microsoft's infrastructure costs and customer billing at scale. The shift comes as Microsoft faces pressure to demonstrate ROI on its $13 billion OpenAI investment, with enterprise customers demanding tangible productivity gains rather than incremental improvements. For OpenAI, the deployment signals that its latest architecture is ready for production at the scale of Microsoft's 365 user base—potentially hundreds of millions of seats.
The integration includes ChatGPT Work, an agentic layer that can navigate multiple applications, maintain context across multi-hour projects, and execute tasks autonomously rather than simply generating suggestions. Microsoft positions this as moving beyond chat-based assistance toward active execution—agents that claim to turn high-level goals into finished deliverables. However, real-world performance data remains limited. Early deployments at enterprises like Deutsche Telekom show promise in customer service automation and employee workflow acceleration, but published KPIs focus on qualitative benefits rather than measurable productivity multipliers. Critically, ChatGPT Work still requires human oversight for complex decisions and cannot independently publish or submit work without approval—a meaningful limitation that complicates 'replacement for workers' narratives.
The GPT-5.6 strategy exposes a consolidation risk for Microsoft: deeper dependency on OpenAI for enterprise AI capabilities at precisely the moment Google (Gemini), Anthropic (Claude), and Meta (Llama) are closing the performance gap. Rivals can argue they offer lower switching costs or more transparent deployment. OpenAI counters with frontier performance per dollar—early benchmarks suggest 15-20% efficiency gains over GPT-4—but has not published independent third-party validation. For enterprises, the calculus involves both capability and vendor lock-in. Deutsche Telekom's integration across customer service, network operations, and employee workflows demonstrates comprehensive adoption, but also means switching costs escalate as integrations deepen. OpenAI's security record, including recent vulnerabilities in GPT-4 and the ongoing Bio Bounty program, remains a compliance concern for regulated industries. The real test arrives when competitive models reach parity; at that point, switching becomes purely economic.