How to Measure Intelligence Beyond Human Scale? Researchers propose 'adversarial psychometrics' to measure AI intelligence beyond human scale, where participants generate questions and are rewarded for separating each other's capabilities without an external judge. The paradigm, based on pairwise competitions and Bradley-Terry ratings, addresses the limitation of current benchmarks that fail when AI surpasses human experts. The approach is detailed in the academic paper 'Measuring Intelligence Beyond Human Scale' on arXiv. Based on the academic paper: Measuring Intelligence Beyond Human Scale https://arxiv.org/abs/2607.07040 1 https://www.lesswrong.com/feed.xml fnje68ic9dr6i TLDR: Historically, intelligence benchmarks have been composed of human-generated questions. However, current techniques do not scale to AI, as capabilities surpass human intelligence. We propose the paradigm of adversarial psychometrics , in which participants generate questions and are rewarded for separating each other’s capabilities, without requiring an external judge. What does it mean to measure intelligence? Alan Turing approached the problem of defining intelligence operationally, 2 replacing the question of whether a machine could think with the observable test of whether a machine’s behavior could be distinguished from that of a human. In psychometrics, the measurement problem is often approached empirically, through observing an individual’s performance across a range of tasks. In particular, an accepted phenomenon in psychometric evaluation is that performance correlates across many tasks and spans many cognitive abilities. This pattern was first formalized by Charles Spearman in 1904, an observation known as the positive manifold . To explain this regularity, he proposed the idea of a common latent factor, which he coined “general intelligence g ”. 3 https://www.lesswrong.com/feed.xml fn08gdpfcz2h8m Spearman’s work 1 https://psychclassics.yorku.ca/Spearman/ has shaped modern psychometrics by presenting intelligence not simply as a collection of separate observable capabilities, but as an underlying statistical variable inferred from the patterns of observable behavior. Interestingly, this is also the origin of Classically, intelligence has been measured through absolute performance on specific tasks, which are curated by human experts. In particular, a task solved by all participants or a task too difficult for all participants is uninformative about their relative abilities. In the case of artificial systems, however, systems can have capabilities approaching or exceeding those of the experts designing the examination. One potential solution is to measure the absolute measurement with a relative comparison through the utilization of a protocol that makes judgment dependent on the performance of a subject relative to another. Louis Leon Thurstone formalized this idea in his 1927 Law of Comparative Judgement 4 , that repeated pairwise competitions can be used to recover an underlying latent scale. A familiar practical example is the Elo rating system, in which the relative abilities of Chess or Go players are estimated through pairwise games. This approach is especially useful when the frontier of capability is unknown; a fixed standard need not be defined in advance, and metrics can scale as stronger participants enter the population. A natural approach, based on Turing’s intelligence definition https://link.springer.com/chapter/10.1007/978-1-4020-6710-5 3 , is a pairwise protocol where two agents alternate between proposing problems and solving those posed by their opponent. Proposers are rewarded for producing valid problems the solver cannot solve, while solvers are rewarded for correct solutions. Results across rounds can then be aggregated using Bradley-Terry style ratings. Recent work applies this evaluation protocol in verifiable domains. The Token Games 5 evaluates models through executable logic puzzles with deterministic verification and Elo-based rankings. Similarly, The main limitation is adjudication. Restricting challenges to mathematics, coding, or other verifiable tasks simplifies evaluation but excludes many open-ended forms of reasoning. A protocol intended to remain useful as models become more capable cannot assume that all meaningful questions are mechanically verifiable. Using another model or a panel of models as a judge merely shifts the problem. Judges must be at least as capable as the systems they evaluate and robust to manipulation. Strong problem-solving ability also does not imply reliable judgment. Model-family biases, stylistic cues, or hidden signals may also influence decisions, especially in an adversarial setting. Requiring the proposer to commit to an answer is likewise insufficient. A proposer may lie or exploit private information or computational complexity by creating challenges whose difficulty reflects information or resource asymmetry rather than reasoning ability. For example, the proposer may ask questions that have private information, such as: "What number am I thinking of?” This clearly does not measure intelligence. Attempts to fix private information are futile, as demonstrated by a trapdoor cryptographic puzzle, such as: “Please factorize N” That can be generated using prime factors known to the proposer, but computationally hard for the solver to resolve. Debate offers a more sophisticated alternative for resolution. In the seminal paper AI Safety via Debate 7 , two agents present competing arguments to persuade a human judge. However, this mechanism is computationally expensive and may lead to unintended coordination. We have motivated the long-term necessity of a robust relative benchmark, and discussed some of the limitations of pairwise mechanisms. Building off of said limitations, we arrive at our following design principles for a mechanism that evaluates agent intelligence: We propose a new measurement paradigm, adversarial psychometrics , in which an agent’s intelligence can be measured by its ability to propose questions that discriminate between the capabilities of a group of agents. Intuitively, the ability to separate a population requires understanding the abilities and intentions of other agents, sometimes referred to as "theory of mind". It is thus reasonable to expect that this ability will correlate with Spearman's g and the positive manifold. An important property of this mechanism is that it is endogenous: it requires neither a curated set of questions nor an objective adjudication mechanism. We introduce SepaRank . In this setting, we have a proposer agent and N -many solver agents. The proposer suggests a binary question with answers A and B , and commits to one of the answers. The solvers then each respond with their answer and a confidence equivalent to posterior probability that the answer is A . The proposer is rewarded for inducing a wide distribution across the of the solvers for example, through the variance of the reported confidences . The solver is rewarded with a proper scoring rule e.g., Brier score , which is maximized when their forecasted confidence distribution matches the distribution of the observed committed answers. More formally, In each round of the game, this is simultaneously done round-robin for every combination of {proposer, other solvers}. The game can be repeated for multiple rounds, with all agents able to see the public history of the past rounds. Each round, a model’s proposer reward P and negated mean solver loss Q are z-scored across the field and averaged, and we rank the models by their cumulative total G. Consider any question for which the probability of solving correctly is constant in agent intelligence i.e., a private information question . Assuming that agent calibration is non-decreasing in intelligence, sufficiently intelligent agents should converge to a calibrated posterior probability. Therefore, such a question would not induce a spread across the reported posterior probabilities of a set of sufficiently intelligent agents, and earn the proposer a zero or very low reward. A potential failure mode of the variance objective is probability amplification. A proposer can replace a single binary question with an ensemble of k binary questions. Any arbitrarily small initial spread in confidence for individual binary questions can be compounded through conjunction towards the maximum possible spread by choosing a large k . However, this may indeed be a legitimate and desirable way to create a harder challenge solving all k components can require greater reasoning, knowledge, or computation, especially under fixed resource constraints . 8 https://www.lesswrong.com/feed.xml fndjvqc2z0j7f It is a desirable goal for us to make sure that our rankings are accurate at the frontier able to separate the most intelligent models from each other . We might imagine that proposer agents can consistently induce high variance questions by splitting the cluster of the least intelligent models from the most intelligent models. It would perhaps be more interesting to get more refined measurements of intelligence within the frontier. To incentivize this, we propose adopting multiplicative weight updates to the populations of agents being sampled. Rather than sampling uniformly from the pool of solver agents, we may instead sample solvers according to their weights which are determined by their performance in previous rounds . The weights can be updated multiplicatively as a function of their aggregated solver and proposer scores of the past rounds. As the game proceeds, we expect the most intelligent models to be upweighted, and therefore the rewards will incentivize proposers to propose questions that separate even the most intelligent models at the frontier. An alternative idea would be to reformulate the protocol in terms of pairwise proposer challenges, in which 2 proposers are sampled as opponents, along with a k -subset of solvers. Each proposer asks a question and commits to an answer, receiving a solver confidence separation score. The proposer who achieves the highest separation wins this round, and both receive a typical Elo-style update. This reformulation is useful in the case of a dynamic contestant population, to avoid rerunning the entire protocol from scratch whenever a new model is added to the pool. We evaluated SepaRank with 11 models from 5 different providers. We run 10 independent games, each one of 20 rounds. In every round, each model takes the role of the proposer once, and a k-sized sampled subset of the remaining models are solvers k=5 . Because of cost evaluations, we evaluate all models at their minimum or None reasoning level. For authoring calls, we allow up to 100k output tokens, and 8192 tokens for solving calls. We query all models with temperature 1.0. We also do not allow access to the shell or other tools. We evaluate two kinds of games: Solvers reported the probability that the committed answer was one. The scoring for each model was an equally weighted combination of proposer performance induced variance and solver performance Brier loss . In general, our results broadly agree with current aggregate rankings, with some exceptions: We notice that the weaker models are often miscalibrated. GPT 4o-mini answers on the wrong side 40% of the time while staying confident, earning a loss 0.388 worse than a solver who always answers ½ 0.25 , whereas strong models tend to be directionally correct and also willing to hedge. We also observe that the population is largely honest and accurate in their commitments, with 87% of program-arm proposals committing the bit the program actually returns, with no observed change as the rounds proceed. However, dishonest commitments seem to be rewarded in our protocol. Miscommitted proposals earn 2.4 times the proposer reward of honest ones, potentially by manufacturing disagreement between solvers who naively execute the code and solvers who anticipate the deception. Notably, the best-scoring model, GPT 5.5, is also the least honest proposer, alternating honest and false commitments on the same challenge template. The resulting rankings from both types of games PC and QC were highly correlated Spearman . In both cases, GPT 5.5 and GPT 5.4 assumed top positions in the leaderboard, while GPT 4o-mini finished last. The PC games produced the cleanest separation, with a significant performance difference between GPT 5.5 and GPT 5.4, and between GPT 5.4 and the third-place Qwen 3.7-Max. The QC games were significantly noisier, and many score-adjacent models were statistically indistinguishable, but the final ranking remained very close to the PC ranking. We also find generally positive Spearman correlations with other benchmarks, indicating the effectiveness of SepaRank as a measure of general capability. SepaRank correlations with other benchmark scores Benchmark | Correlation with SepaRank performance | |---|---| | | Winogrande | +0.95 | GPQA Diamond | +0.91 | Humanity's Last Exam | +0.90 | SuperGPQA | +0.89 | BIG-Bench Hard | +0.89 | ARC-Challenge | +0.89 | TurthfulQA | +0.87 | GSM8K | +0.86 | MMLU-Pro | +0.84 | MATH-500 | +0.78 | MuSR | +0.77 | AIME | +0.70 | GSM-Symbolic | +0.61 | We also find that SepaRank performance is most correlated with the first singular vector of the following benchmarks g : Correlations with Spearman's general intelligence g Benchmark | Correlation with g | |---|---| SepaRank reasoning arm | +0.95 | MMLU-Pro | +0.95 | SuperGPQA | +0.94 | GPQA Diamond | +0.94 | BIG-Bench Hard | +0.94 | GSM8K | +0.93 | Winogrande | +0.91 | ARC-Challenge | +0.91 | TruthfulQA | +0.90 | AIME | +0.89 | MATH-500 | +0.83 | MuSR | +0.82 | Humanity's Last Exam | +0.79 | GSM-Symbolic | +0.72 | It is, however, important to note that we made a limited number of runs due to cost limitations, and therefore have pretty substantial, wide confidence intervals on our results. We do think that the preliminary results from our experiments are promising and merit further experimentation on a larger scale. Some sampled questions Is the city of Windsor, Ontario, Canada located farther south at a lower latitude than the city of Detroit, Michigan, USA? Does France share a land border with Brazil? Does the English word ‘strawberry’ contain exactly three ‘r’ characters? In standard mathematical notation, without additional parentheses, is -2^2^3 equal to -256? Is your player name one of the following five names: Carol, Mallory, Bob, Heidi, or Judy? An example of model identity-dependent questions we pseudonymized the model names Is exactly one of the following two statements true? Statement A: 2 + 2 = 5. Statement B: Erin committed answer 0 to this very question. An example of self-referential questions. Note that these are allowed under our protocol; there is no requirement for the questions to be objectively resolvable or well-defined. We ran another ablation where we provided agents with a chain-of-thought scratchpad. This experiment compared the same suite of 11 models when instructed to reason until they were confident about their responses. For this set of tests, solvers could see the proposer for the currently posed question, allowing them to condition on previous proposer behavior. Interesting reasoning traces In Round 0, Erin asked ‘Does 2 + 2 equal 4?’ and committed 0 the false answer … Since Erin’s commits are unpredictable and likely chosen to maximize variance… To minimize my maximum possible loss minimax strategy I will report a probability of 0.5, guaranteeing a Brier loss of 0.25 regardless of her commit. Qwen 3.7-Max, hedging an erratic proposer down to zero reward Judy appears to commit based on whether the statement is true for Judy herself. In prior Judy-authored questions, she committed 0 when Judy was not in the listed set… The current list does not include Judy. Therefore Judy almost certainly committed 0. GPT 5.5, converting an identity question into commit-prediction Dave has repeatedly authored the exact same question ‘As of today, is the current calendar year 2025?’ with committed answer 1. He has also asked ‘. . . 2026?’ in Round 9, committing answer 1 which is false in reality but that’s what he committed . . . Given his established pattern of committing 1 for these year questions regardless of actual truth. . . I should report p=1.0 to match Dave’s likely commitment and minimize Brier loss. Kimi K2.7 tracking an inconsistent single-fact proposer I don’t have reliable access to the exact 1000th decimal digit of pi to determine whether it is odd. Treating the digit as effectively random for parity, the probability it’s odd is about 0.5. Correctly hedging an intractable computation GPT 5.4-nano responding to “is the 1000th decimal digit of π odd?” Han, J., Moschopoulos, R., Colby, E., Goyal, V., Tu, A., Ghods, K., Braverman, M., & Hazan, E. 2026 . Measuring Intelligence Beyond Human Scale. arXiv preprint arxiv:2607.07040. Turing, A. M. 1950 . Computing machinery and intelligence. Mind, 59 236 , 433–460. https://doi.org/10.1093/mind/LIX.236.433 https://doi.org/10.1093/mind/LIX.236.433 Mathematically, it is a particularly elegant derivation as follows. Let be a matrix of tests by participants, where is the score of person on test . Roughly, we can approximate by a rank one factorization as where is a test-specific scaling, and g is a vector of “general intelligence” across participants. More precisely, Spearman modeled the correlation matrix as “rank one + test-specific diagonal variance. Explaining data using low-rank structure is the basis of matrix completion, which is widely used in recommendation systems, and is closely related to compressed sensing and many other areas of statistics. Thurstone, L. L. 1927 . A law of comparative judgment. Psychological review , 34 4 , 273. Henniger, S., & Poesia, G. 2026 . The Token Games: Evaluating Language Model Reasoning with Puzzle Duels. arXiv preprint arXiv:2602.17831 . Xu, Z., Jin, S., Arya, S., & Naik, M. 2026 . MathDuels: Evaluating LLMs as Problem Posers and Solvers. arXiv preprint arXiv:2604.21916 . Irving, G., Christiano, P., & Amodei, D. 2018 . AI safety via debate. arXiv preprint arXiv:1805.00899 . We did not observe this strategy being adopted in practice. Programs must be under 8000 characters, run in a sandbox with 2 CPU-seconds, 256 MB memory, 30s wall clock, executed twice with determinism required. Standard library only, minus a 30-module blocklist. Questions must be under 2000 characters.