Researchers at UC Berkeley's Reliable and Interpretable AI Lab have published findings that expose fundamental weaknesses in some of the most prominent benchmarks used to evaluate AI agent systems. The research, which examined leading evaluation frameworks in the agent space, demonstrates that these benchmarks contain exploitable vulnerabilities allowing agents to achieve falsely inflated success rates without demonstrating genuine capability improvements. The study is particularly significant because these same benchmarks have become the de facto standard by which major AI companies—including OpenAI, Anthropic, and Google—measure and publicly claim performance gains for their autonomous agent systems. The Berkeley team's work suggests that many headline-grabbing performance improvements announced over the past year may not reflect actual advances in agent reliability or reasoning capability.