Essay · Part 1 of 2

Are your eval scores measuring the model or the raters?

A psychometric teardown of the most-copied judging pattern in AI.

Part 1 of a two-part series. Part 2 runs the experiments.

Shawn Fraine

· 18 min read

I was trained to read validation studies for psychological tests, so that is how I read evaluation papers now. I want to know what construct the authors are claiming to measure, what evidence supports the scores, and where the error is coming from. When I read the most influential evaluation result in modern AI that way, it looks different from the way it usually gets cited. I don't think this is a niche complaint, either. A 2025 systematic review of 445 LLM benchmarks found construct-validity failures across the field, including benchmarks that never defined the phenomenon they claimed to measure (Bean et al., 2025).

Zheng et al. (2023) reported that GPT-4, used as a judge, agreed with human preferences about 85% of the time. That one number did more than any other single result to convince the industry that a language model could grade language models. The paper's judging setup, where a strong model weighs candidate answers against a short list of natural-language criteria, became the standard pattern in modern eval pipelines.

Before going further, I want to separate two things. The paper set out to build a cheap approximation of human preference for benchmarking models, and it succeeded. The authors were also unusually honest about the limitations. What happened afterward is a different story. The industry picked that preference heuristic up as a measurement instrument, used it to make fine-grained claims about model quality, and never added the calibration, anchors, or reliability analysis that measurement requires. This teardown is about the gap between what the paper validated and what the industry assumed.

What the instrument is

The canonical judge is a prompt template called pair-v2, which lives in the FastChat repository. It instructs the model to act as an impartial judge, weigh six qualities (helpfulness, relevance, accuracy, depth, creativity, level of detail), write a short explanation, and output [[A]], [[B]], or [[C]] for a tie. It also tells the judge to disregard response order, response length, and assistant names.

There is a second template, and the difference between the two matters more than it seems. The agreement study validated both as preference approximators. Table 5 shows single-answer grading matching pairwise GPT-4 and human preferences well, which is why the leaderboard ships the more scalable 1-to-10 version. It is important to note the difference between validating a procedure as a preference approximator and validating a scale. The 1-to-10 template has ten points and no anchors, and the paper offers no test-retest, anchor, or scale-invariance evidence that a 7 means the same thing across items, categories, or judge versions. The authors themselves warn that absolute scores are likely to fluctuate more than relative pairwise results if the judge model changes.

Back to pair-v2. The whole instrument is six nouns, three possible verdicts, and a few lines telling the judge what to ignore. None of the six qualities is defined. Nothing says whether accuracy should outrank creativity when the two conflict, or whether "level of detail" means the right amount of detail or simply more of it. When the judge picks A over B, the record does not show whether A won on accuracy, on formatting, or on something the rubric never named. The explanation reads like a breakdown by criterion, but nothing in the procedure validates it as one. There is no evidence that it decomposes the verdict faithfully rather than reconstructing a plausible story around it.

Unpacking the 85%

The headline comes from Table 5, and to the authors' credit they reported it two ways. S2 computes agreement using only non-tie votes. S1 includes ties and recodes position-inconsistent verdicts as ties. The famous 85% is the S2 figure for GPT-4 pairwise against humans on first-turn questions. Under S1 it is 66%. The same pattern holds for humans against each other: 81 to 82% on decisive votes, 63 to 67% with ties included. On the larger, noisier Chatbot Arena sample (Table 6), GPT-4 pairwise reached 87% on non-tie votes and 64% with ties. Two measurement details are worth stating plainly. First, "agreement with humans" means agreement with a randomly selected human vote on a randomly selected question. There is no adjudicated gold label behind it. Second, S1 and S2 have different chance baselines (33% versus 50%), so the drop from 85 to 66 mixes a change in scoring rule with a change in what counts as chance.

Those baselines make a rough correction possible. The paper never needed it, but anyone buying judge scores should ask for it. Percent agreement counts the agreement you would get by chance along with the real thing, so the standard move is to correct for the chance portion. Using the uniform-random baselines the paper reports (50% for two outcomes, 33% for three), GPT-4's binary 85% comes out to about 0.70 on a chance-corrected scale, human-versus-human agreement to about 0.62, and the famous three-way 66% to about 0.49. I want to be clear that these are rough numbers. The paper does not report the voting patterns a precise correction would need, so I have stated the assumption up front, and the true values will move with the real patterns.

The exact decimals matter less than the pattern. Chance inflates every one of the headline figures, and nobody in the citation chain applied the correction. Human agreement corrects down as well, which I think is the honest footnote here. It is also the reason the checklist at the end gates the judge against adjudicated exemplars instead of single raters. Agreement with a resolved standard typically runs higher than agreement between two individuals, and that is what keeps a real bar within reach.

The authors hid nothing. They published both numbers. My criticism is about what happened to the 85% after publication. It traveled as a general reliability claim, stripped of the qualifier that it describes agreement computed on decisive votes only. Excluding ties removes many of the cases where a judge declines to state a strict preference, which includes pairs that are genuinely equal and pairs that are genuinely uncertain. Those are the close calls, and close calls are where you most want to know whether your judge holds up.

The strongest defense of the number goes like this. GPT-4's agreement with humans is numerically close to human agreement with other humans under both setups, the authors describe the two as the same level, and so the judge is "as reliable as a second human." That is true on the paper's terms, and I think it misses the point. Matching crowd taste at human rates validates the judge as a preference approximator. It does not validate the six named criteria as measurable constructs, and it does not show that the scores carry the meaning people attach to them downstream. A 2026 paper formalizes exactly this gap. Reliability does not establish construct validity, and a judge can hold its verdicts steady under perturbation while responding only weakly to changes in the construct itself. In one audit, predictors that read only surface form reproduced 67.4% of MT-Bench's human votes (Chen et al., 2026). In other words, human disagreement sets a baseline for a taste-matching task. It says nothing about whether "depth: 7" means anything.

Table 7 adds a second qualifier that the citations usually drop. GPT-4's agreement with humans is highest when the two models being compared differ a lot in quality and lowest when they are close. That is intuitive, and it is awkward for how the number gets used, because leaderboards and procurement decisions live in the narrow margins between strong systems, which is exactly where agreement is thinnest.

The paper stress-tested its own judge

This is the part of the paper that deserves the most attention. The authors ran the attacks themselves rather than stopping at the headline.

Measurement science has names for the two kinds of damage on display here, and I think the names are worth learning because each one tells you what the problem costs. The first is construct-irrelevant variance, which means score movement driven by features outside the thing you meant to measure. The prompt itself says order should not matter, so when swapping A and B flips the verdict, that is the textbook case. The same goes for preferring a padded answer over an informative one once you hold the content constant, for self-preference that survives controls for response quality, and for verdicts that change when someone rephrases the instructions without changing their meaning. Each of these can move a score independently of the qualities the evaluation set out to measure. In business terms, you are paying for contamination. Some of it is noise. Some of it is systematic bias, which is worse, because bias does not average out.

The second problem is a construct specification failure. Six named considerations with no definitions, no anchors, and no weights means the procedure never states what "quality" covers. The technical name for the risk this creates is construct underrepresentation: the score ends up representing a narrower construct than the one claimed, whole dimensions go unmeasured on a given call, and nobody notices. In business terms, you are measuring a narrower thing than you think you are. The first three findings below are three demonstrations of the first problem. Prompt wording belongs with them, and it shows up in the follow-up literature two sections down. The fourth finding, on scaffolding, is a separate question about whether the judge is competent at all without help. The second problem gets its own section after that.

Position. The authors built a set of comparisons designed to be hard, with answers that were very similar and occasionally indistinguishable even to humans, and then swapped the answer order. The verdict changed often. GPT-4 stayed consistent 65% of the time, GPT-3.5 46%, and Claude-v1 24%. The prompt explicitly forbids letting order matter, and the instruction did not eliminate the bias. A contemporaneous study showed how far this reaches: reordering the candidates let Vicuna-13B "beat" ChatGPT on 66 of 80 queries (Wang et al., 2023). The standard mitigation is to run both orders and score the inconsistent cases as ties, which is what S1 effectively does. That is honest bookkeeping, and it has a cost. It doubles your judging budget and converts instability into ties, and the headline statistic then excludes those ties. In other words, the protocol contains the instability. It does not demonstrate that the judge is order-invariant.

Verbosity. The authors took 23 answers containing numbered lists, padded them with redundant rephrasings that added no information, and checked whether the judge preferred the longer version. Claude-v1 and GPT-3.5 fell for it 21 times out of 23. GPT-4 failed twice. So the best judge resisted and the weaker ones did not, which tells you something useful: the rubric leaves the door open even if it does not push anyone through it. "Level of detail" is one of the six criteria, and the prompt never distinguishes useful detail from padding. I think that is the sharper psychometric criticism. Appropriate detail is a legitimate quality dimension, but the rubric never says how to tell it apart from verbosity, so the judge decides what the term means on every call.

Wu and Aji (2023) made the same point from a different angle at about the same time. In their main experiment, a self-evaluation setup with GPT-4 judging GPT-4 outputs, the judge assigned an Elo of 1206 to responses constructed to contain several minor factual errors and only 1096 to responses that were correct but short. A longer answer with errors outscored a shorter answer without them.

Self-preference. GPT-4 awarded itself a win rate roughly 10 points higher than humans awarded it, and Claude-v1's gap was about 25 points. The authors called this inconclusive on their data, which was the right call. Later work tested it under tighter conditions. Pombal et al. (2026) found judges more than 50% more likely to wrongly mark a criterion as satisfied for their own outputs (among rubrics their own outputs fail), with score distortions of up to 10 points on HealthBench, and the effect persisted on verifiable instruction-following criteria. I read that as evidence that self-preference survives more controlled, objectively checkable settings, and I would stop short of treating it as a retroactive explanation of the original gaps. Wataoka et al. (2024) found evidence consistent with a familiarity account, in which judges score low-perplexity outputs higher regardless of who generated them.

Scaffolding. On math questions, GPT-4's judging failures went from 14 out of 20 with the default prompt to 6 out of 20 with chain-of-thought prompting and 3 out of 20 with a reference answer. The honest reading cuts both ways. Judge competence depends heavily on the setup, and much of the gap closes with prompt scaffolding before you ever supply a reference answer, at least in this 10-question diagnostic. Either way, reading the default prompt as a domain-general rater does not survive contact with these numbers. The authors knew this. They used reference-guided judging for math themselves.

One more result deserves a fair hearing. When humans disagreed with GPT-4 and then saw its reasoning, they called the reasoning reasonable 75% of the time and changed their own verdict 34% of the time. That is real evidence that the judge's rationales are plausible to people. Plausibility and calibration are different properties, though. A fluent explanation can convince a reader without being a valid decomposition of the verdict.

What a conventional rating instrument would specify

A fair objection at this point: MT-Bench was built for cheaper benchmarking rather than for high-stakes psychological measurement, and not every benchmark needs a full psychometric dossier. That is fair as far as it goes. The premise underneath it is wrong, though. In measurement, validity belongs to the interpretation of scores for a specific use, and that is the position the Standards for Educational and Psychological Testing take. What Zheng et al. validated is a narrow family of claims: these judging procedures approximate expert and crowdsourced human preferences well enough to support scalable chatbot benchmarking in the settings studied. A use close to that inherits some of the evidence. A use farther out does not. "Model X reasons better" needs reasoning-specific evidence, and a ship-or-no-ship decision needs threshold and decision-consistency evidence. Existing evidence can carry over when the new use is close enough, but it never transfers automatically.

"It was just a benchmark" is a statement about where the scores came from. Whether your decision is a valid use of them is a separate question, and your decision is the thing your budget actually depends on. The requirements below bite whenever an organization treats small score differences as stable measurements and makes deployment or procurement decisions from them.

Instrument specifications still matter, because they are the raw material any wider validity argument is built from. The question is whether this instrument can supply them. To the paper's credit, it provides more measurement evidence than most ML benchmarks: stated criteria, output rules, human-machine agreement, human baselines, position-swap testing, and bias analyses. The gap is between that solid start and what a conventional rating instrument would still require.

One more thing before the list. Nothing below depends on the rater being a model. A human annotation team working from the same six nouns has the same problem. A contractor with no anchors and no calibration improvises a private scale the same way GPT-4 does, agreement between two such contractors sits under the same chance ceiling, and drift over a long labeling session is the human version of the judge checkpoint changing underneath you. If your pipeline uses an LLM judge for scale and a human team for the gold set, which is a common arrangement, you are running two uncalibrated instruments and treating one as the standard for the other. I wrote each requirement below for both.

Behavioral definitions. "Depth" could mean conceptual sophistication, causal explanation, or completeness. "Creativity" is a virtue in fiction and a liability in factual QA. Six nouns do not establish six measurable constructs. The authors concede the point in their own limitations section, noting that, within helpfulness, accuracy, relevance, and creativity "are all combined into a single metric in this study" (Zheng et al., 2023). That is a construct specification failure in plain sight. The prompt names six considerations and collapses them into one holistic verdict, with no evidence showing how each one is represented. This is the gap that makes underrepresentation likely, because nothing in the design guarantees that the six labels cover the quality domain or that the judge honors all six.

Anchors. Nothing shows what poor, adequate, or excellent looks like on any dimension, and the 1-to-10 scale has no score-level descriptors. That leaves the scale's meaning free to drift across prompts, categories, and model checkpoints, with no evidence that equal numerical differences mean comparable things. Practitioners run into the same wall in the field. Uncalibrated 1-to-5 scales across multiple dimensions leave "what makes something a 3 versus a 4" unanswered, and different evaluators end up interpreting the scale differently (Husain, 2024).

Weights. Accurate-but-terse versus detailed-with-one-factual-error: the rubric never says which wins. The judge invents a tradeoff policy on the spot, invisibly, on every call.

Calibration. A serious rating operation trains raters on exemplars and gates them against adjudicated standards before they score anything real. The counterargument is that a pretrained model brings evaluative knowledge with it, so rater training is unnecessary. I agree that capability was never the problem. What is missing is a documented, stable interpretation of the six criteria that you could audit, plus a check that this particular judge instance meets it before you trust its scores.

A reliability model. Percent agreement and order consistency are legitimate reliability evidence, and the paper deserves credit for reporting them. What they do not give you is a decomposition of the variance: how much of a score comes from answer quality and how much from the question, the order, the judge checkpoint, the prompt wording, or sampling noise. Each of those is a potential source of unwanted score variance, and each has to be modeled, controlled, or justified before you can read the score as answer quality.

Tie semantics. [[C]] is one label covering several distinct states the procedure cannot tell apart: the two answers are genuinely equal, the judge is uncertain, or the dimensions genuinely conflict. The headline analysis then excludes that bucket entirely.

Validity scope. Agreement with crowd preference on the paper's tasks does not transfer to medical QA, safety evals, or RAG faithfulness. The authors' own switch to reference-guided judging for math is the paper telling you that the default instrument has boundaries.

The follow-up literature kept finding the same cracks

The three years since read like a slow inventory of the problems above.

Length-controlled AlpacaEval stopped asking the judge to ignore length. It controlled for length statistically instead, estimating what the preference would have been at equal length, and correlation with human Arena rankings rose from 0.94 to 0.98. The lesson generalizes. If a nuisance variable matters, measure it and model it. A sentence in the prompt telling the judge to ignore it has already been shown to fail.

Arena-Hard's tooling added an explicit Style Control mode (length and markdown features) in late 2024, and its confidence-aware analysis found MT-Bench separating only about 23% of model pairs among the top 20, counting a pair as separated when its 95% confidence intervals do not overlap. The conventional ranking-agreement formulation had put the figure at 91%. It is important to note that these are two different statistics. The 23% measures separability and the 91% measures agreement, so the drop is a change in the question being asked rather than an error bar added to one number. The substantive point survives that caveat: once you require statistical confidence, MT-Bench cannot separate strong models from each other.

Position effects turned up everywhere anyone looked. Shi et al. (2025) found order bias varying systematically by judge, task, and quality gap across more than 150,000 evaluations. A 2026 benchmark gave 880 items to 25 judges under two meaning-preserving wordings of the same instructions and found that the rewording alone cost the judges agreement with themselves on all four tasks tested (Bellibatlu et al., 2026). If rewording the instructions changes the verdict, then the wording is part of the instrument, and nobody ever specified it.

How far the pattern spread

One correction before I close. The sources do not support the strong claim that tools like RAGAS and DeepEval copied the MT-Bench prompt text. What spread was the pattern: candidate answers, natural-language criteria, a strong model as judge, a written rationale, and a verdict. That architecture is now widely adopted, and to me the interesting part is what the adopters did next. Current versions of RAGAS ship anchored 1-to-5 scales with behavioral descriptions. DeepEval supports explicit rubric objects that tie score ranges to described performance levels. So the paradigm spread, the same classes of confound spread with it, and the repairs are arriving piecemeal, one vendor at a time, with no shared standard for what "calibrated" even means.

I think that is the more damning version of the story, and also the more useful one. The prompt itself matters less than the pattern of adoption. An industry took an uncalibrated measurement procedure as its standard way to generate numbers, and it is still discovering, paper by paper, what proper instrumentation would have required up front.

What to do Monday morning

Start with one question, and answer it honestly: how do you currently check whether your raters agree, and who owns that? If the answer is "we don't" or "the vendor does," you do not have a reliability problem yet. You have an unmeasured one, which is worse, because you cannot size it. If there is a named owner and a number they can quote, you are ahead of most teams, and the three checks below are a way to pressure-test that number.

If you run evals with an LLM judge or buy results from someone who does, three checks will tell you how much of your score is signal:

  1. Run every pairwise judgment in both orders. Report the inconsistency rate alongside the win rate. If swapping A and B flips more than a trivial share of verdicts, presentation order is partly driving your leaderboard.
  2. Report agreement with ties included, using a chance-corrected statistic suited to your design. The decisive-vote figure on its own is the 85% everyone has seen, which is the S2 number. Ask your vendor for the S1 equivalent.
  3. Gate the judge before you trust it: a set of adjudicated exemplars with known-correct verdicts, a pre-specified agreement threshold appropriate to the decision you are making (set it before you see the scores, and report the uncertainty around it), and re-gating whenever the judge checkpoint changes.

And the question behind all three: are your eval scores measuring the model, or the raters?

In Part 2, I stop arguing from the literature and start measuring. I will use the same rubric and run controlled experiments (position swaps, length padding, and deliberately conflicting quality dimensions) across multiple judge models. The KPI to watch is the flip rate: how often a judge's verdict changes when the only thing you perturb is order, length, or style. We will see what breaks, by how much, and what a sound version of this instrument would actually require.

References

Bellibatlu, R. R., Raff, E., & Zhang, W. (2026). JudgeSense: A benchmark for prompt sensitivity in LLM-as-a-judge systems. arXiv. https://arxiv.org/abs/2604.23478

Bean, A. M., Rocher, L., Kearns, R. O., Romanou, A., et al. (2025). Measuring what matters: Construct validity in large language model benchmarks. arXiv. https://arxiv.org/abs/2511.04703

Chen, J., Chen, W., Lin, Z., & Vong, C. M. (2026). A judge should know what changed: Construct validity for LLM-as-a-judge evaluation. arXiv. https://arxiv.org/abs/2608.24419

Dubois, Y., Galambosi, B., Liang, P., & Hashimoto, T. B. (2024). Length-controlled AlpacaEval: A simple way to debias automatic evaluators. arXiv. https://arxiv.org/abs/2404.04475

Husain, H. (2024). Creating a LLM-as-a-judge that drives business results. https://hamel.dev/blog/posts/llm-judge/

Li, T., Chiang, W.-L., Frick, E., Dunlap, L., Wu, T., Zhu, B., Gonzalez, J. E., & Stoica, I. (2025). From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. In Proceedings of the 42nd International Conference on Machine Learning (PMLR 267, pp. 34209–34231). https://proceedings.mlr.press/v267/li25h.html

lmarena. (2024). Arena-Hard-Auto [Computer software]. https://github.com/lmarena/arena-hard-auto

LMSYS Org. (2023). FastChat [Computer software]. https://github.com/lm-sys/FastChat

Pombal, J., Rei, R., & Martins, A. F. T. (2026). Self-preference bias in rubric-based evaluation of large language models. arXiv. https://arxiv.org/abs/2604.06996

Shi, L., Ma, C., Liang, W., Diao, X., Ma, W., & Vosoughi, S. (2025). Judging the judges: A systematic study of position bias in LLM-as-a-judge. In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (pp. 292–314). https://aclanthology.org/2025.ijcnlp-long.18/

Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., & Sui, Z. (2023). Large language models are not fair evaluators. arXiv. https://arxiv.org/abs/2305.17926

Wataoka, K., Takahashi, T., & Ri, R. (2024). Self-preference bias in LLM-as-a-judge. arXiv. https://arxiv.org/abs/2410.21819

Wu, M., & Aji, A. F. (2023). Style over substance: Evaluation biases for large language models. arXiv. https://arxiv.org/abs/2307.03025

Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems (Vol. 36). https://arxiv.org/abs/2306.05685