The AI research community is confronting a fundamental problem: conventional benchmarks no longer effectively measure how well large language models perform on genuinely complex tasks. A new framework called Xpertbench introduces rubrics-based evaluation specifically designed for expert-level challenges that require open-ended reasoning rather than multiple-choice answers. This addresses a critical blind spot in current AI assessment methodologies, as performance plateaus on standard tests have created uncertainty about whether improvements in model capability are actually occurring or whether evaluation metrics have simply become inadequate.