The race to deploy AI as an active participant in scientific discovery has accelerated dramatically, but measuring whether these systems actually advance research has remained surprisingly difficult. Researchers have now unveiled LABBench2, an enhanced benchmark designed to evaluate AI systems performing genuine biology research tasks. Unlike earlier evaluation frameworks focused on narrow question-answering, LABBench2 tests AI's capacity to formulate novel hypotheses, design experiments, and interpret results—simulating the core functions of a computational biologist. The benchmark addresses fundamental limitations in prior benchmarks, which often lacked the complexity and real-world constraints of actual laboratory work, making it difficult to assess whether AI agents could genuinely contribute to scientific discovery rather than simply answering predefined questions.