When OpenAI and Anthropic release reports detailing how people use ChatGPT and Claude, they're curating a narrative rather than providing comprehensive data. As computer science PhD candidate Anka Reuel recently noted, 'There is no independent source to corroborate it.' This opacity matters enormously: AI companies selectively publish usage patterns that align with their safety narratives while withholding data that might reveal problematic applications. Consider OpenAI's claim that jailbreaking attempts represent a negligible fraction of ChatGPT usage—a figure impossible to independently verify. If that number is actually 15 percent rather than 2 percent, regulators evaluating AI risk remain operating on incomplete information. The companies argue proprietary concerns and user privacy justify limited disclosure, but researchers argue this creates a credibility gap that undermines the entire regulatory framework being constructed around AI safety.

Several computer science researchers and policy experts have begun pushing for mandatory disclosure requirements modeled on financial regulation precedents. Specifically, proposals include establishing an independent audit function where vetted third-party researchers could access anonymized, aggregated usage data under strict confidentiality agreements—similar to how auditors verify pharmaceutical companies' safety data without compromising trade secrets. The FTC has signaled interest in this direction through recent enforcement actions and proposed rulemaking on AI transparency, though concrete requirements remain underdeveloped. Some researchers advocate for a new AI-specific oversight body within NIST or as an independent commission with statutory authority to demand data submissions. Without intervention, the status quo persists: only the companies building these systems understand how they're actually deployed, while policymakers operate blind to emerging risks.

This data asymmetry has already produced concrete regulatory failures. In 2023, when multiple news outlets reported sophisticated social engineering attacks using AI-generated voices, initial industry responses suggested such misuse was 'extremely rare.' No independent verification of that claim was possible—or conducted—before these incidents influenced congressional testimony and policy discussions. More significantly, the lack of transparent usage data prevented early detection of AI's emerging role in automated decision-making for critical services like hiring and loan applications, delaying regulatory attention by months. The Biden administration's AI Executive Order acknowledged the need for transparency, yet implementation remains voluntary. Until AI companies face enforceable requirements to share independently auditable usage metrics, policymakers will continue making consequential decisions about AI governance while flying blind. The question isn't whether data transparency threatens legitimate business interests—it's whether protecting those interests is worth the cost of regulatory paralysis.