Graduate-level science question sets are among the more carefully constructed benchmarks in the field, and the care went into a place that rarely gets discussed: establishing what a determined non-expert scores. That control is what turns the number into a claim.
The design problem for a science benchmark is that most exam-style questions are solvable by retrieval. If a question can be answered by finding the right page, a high score measures search and memory, which is not what anybody is trying to measure.
The response, in benchmarks of the GPQA type, is a construction procedure with three steps. Domain experts — people holding or working towards doctorates in the relevant field — write questions in their own speciality. Other experts in the same field answer them, establishing that a knowledgeable person can get them right, which filters out questions that are merely ambiguous or wrong. Then skilled non-experts attempt them with unrestricted time and unrestricted web access, and questions they can solve are discarded.
That third step is the innovation and it is what makes the number interpretable. The benchmark ships with a measured human baseline for both groups, so a model score can be positioned against “determined person with a search engine” rather than against nothing. Most benchmarks do not have this, and a score without a human baseline is a number without a scale.
It is worth stating positively, because the criticisms are easier to make than the credit. A model scoring well above the non-expert baseline on questions of this kind is doing something real: it is selecting correct answers to hard, well-posed, closed-form problems across several sciences at a level that skilled people with search cannot match. Whatever else is true, the knowledge and the multi-step manipulation of it are present, and that was not the case a few model generations ago.
It is also a reasonable, if coarse, screening signal. If you are choosing a model for tasks that involve technical domain knowledge, the relative ordering on this kind of test carries some information about the ordering on your task — a weak inference, but not a worthless one, with the caveats in why benchmark results do not transfer.
There is also the standing issue that any public benchmark degrades as it circulates. Questions and discussions of them end up in training corpora, and a score on a set that has been public for years is a weaker signal than the same score on a freshly written set — the subject of benchmark contamination. This is why blind, prospective evaluations carry so much more weight, which is exactly the argument made for structure prediction in the CASP format.
Question answering is only one kind of scientific benchmark, and for machine learning applied within a science the more relevant families work differently.