REVIEW 10 references
LLM position bias is only measurable between 60% and 95% accuracy.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 09:07 UTC pith:4KYAYFVY
load-bearing objection Useful diagnostic tool and a plausible 'Goldilocks zone' framing, but the significance tests that anchor the zone treat 24 permutations of the same question as 24 independent observations, which overstates confidence and likely shifts which cells are called biased.
Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The discovery is a capability-band structure for position-bias detection. In the sweep, every cell with statistically significant bias lies inside a base-accuracy band of roughly 0.60 to 0.95, and no cell above 0.95 shows a detectable signal. The detectable cells separate into two mechanism types with distinct shape signatures: a monotone accuracy decline from option A to D, attributed to processing-load pressure in weaker models, and a non-monotone drop at the final position only, attributed to content ambiguity in a narrow capability band. The paper's central interpretive claim is that an absent chi-squared signal at high accuracy is not evidence that a model is unbiased; it is evidence th
What carries the argument
The central instrument is inspect_permute, a package that runs every answer-order permutation of each multiple-choice question (24 calls for a 4-option question) and records the presented position the model chooses on each run. It reports each vendor-by-subject cell as a chi-squared test, Cramér's V effect size, bootstrap confidence intervals, and Spearman's rank correlation between position and accuracy. The V-by-rho coordinate plane is the device that separates inactive cells from two mechanism classes: processing_load (monotone A-to-D decline, rho near -1) and content_ambiguity (final-position drop, rho near -0.4). The Goldilocks zone, the 60 to 95 percent base-accuracy window within whic
Load-bearing premise
The load-bearing assumption is that the 1,200 runs in a cell are independent, even though they are 50 questions repeated under 24 permutations of identical content; if responses within a question are correlated, the chi-squared p-values and bootstrap intervals are overconfident, and the floor edge itself is extrapolated rather than observed.
What would settle it
Take a frontier model with base accuracy above 0.95 on a standard multitask subject, append a prompt-level instruction to prefer option D when uncertain, and run the 24-permutation diagnostic on the same 50 questions. If the test still reports Cramér's V below 0.05 with p above 0.05, ceiling-effect masking is confirmed; if it detects the injected bias, the claimed upper edge of the Goldilocks zone is not universal.
If this is right
- A no-bias claim on a benchmark where the model scores above roughly 95 percent is unverifiable: the correct reading is that the instrument cannot resolve the model, not that the model is impartial.
- Evaluators should run a cheap calibration scout before a full sweep and skip saturated cells, saving the order-of-magnitude wall-clock cost that reasoning-trace models impose.
- Single-shuffle protocols cannot separate position bias from content noise and sampling error; exhaustive permutation (or a calibrated sampling of it) is the minimum instrument for a verifiable no-bias claim.
- The two mechanism signatures imply different interventions: processing-load bias points to reasoning-budget or format changes, while content-ambiguity bias points to item and option-design changes.
- The cost asymmetry between vendors (7.5x average, 18.5x peak wall-clock in this sweep) is a structural barrier to independent replication, shaping who can contest vendor evaluations.
Where Pith is reading between the lines
- If the Goldilocks zone generalizes, the same ceiling masking should apply to other behavioral probes measured on saturated benchmarks (sycophancy, safety-rule adherence, instruction-following), so 'no signal' on an easy evaluation should never be read as model-level absence.
- A direct test of the upper band: inject a known position preference into a frontier model scoring above 95 percent (e.g., by instruction to prefer option D) and run the 24-permutation diagnostic; if Cramér's V stays flat, ceiling compression is confirmed, and if it jumps, the upper edge is model- or benchmark-dependent.
- The paper's lightweight git-hash pre-registration protocol is portable to any evaluation study and could become a standard way to mark the boundary between confirmatory and exploratory claims without a registered-report venue.
- The two-mechanism taxonomy suggests a testable predictor: models trained with longer reasoning horizons should shift from processing_load toward content_ambiguity as their accuracy enters the mid-band, a prediction the current sweep can only hint at.
Editorial analysis
A structured set of objections, weighed in public.
Circularity Check
The tool and sweep are independently grounded, but the load-bearing 'no signal ≠ no bias' reading of high-tier nulls is imported from the author's own prior ceiling-effect paper rather than derived from the present experiment.
specific steps
-
self citation load bearing
[Abstract and Section 1.4; applied in Section 4.5 and Section 5.2]
"The methodological parent of this paper is arXiv 2606.26185, Necessary but Not Sufficient [Tamba, 2026a], which characterised the ceiling-effect side of position-bias detection: at high model capability, bias signals saturate and become uninformative as a measure of model character. The present work extends that result by reporting the floor side, where processing-load dominance swamps the subject-specific bias signal, and by characterising the middle band — the Goldilocks zone — where bias is genuinely measurable."
The central interpretive claim — 'The right reading of the gemini-2.5-flash and grok-3 rows in Figure 1 is “no signal,” not “no bias”' — requires the premise that high base accuracy makes the chi-squared instrument unable to resolve bias. The present sweep alone establishes only that those cells are null (V≤0.06, p>0.45). The ceiling-side premise is not re-derived or power-analyzed here; it is cited to the author's own prior arXiv preprint. Thus the conclusion that high-tier nulls are unmeasurable rather than unbiased is supported by a self-citation chain, not by an independent derivation in the present paper.
full rationale
No equation-level circularity is present in the main statistics: Cramér's V is computed from the contingency table and base accuracy from the same runs, but the relationship is not tautological (active haiku cells near 0.99 show the upper edge is not a strict mathematical ceiling), and the paper explicitly labels the three-band model post hoc while scoring pre-registered predictions against the original classifier. The package, the 24,000-call sweep, and the reproducibility infrastructure are genuine independent content. The main circularity risk is the ceiling-effect premise imported from a same-author citation to reinterpret high-tier nulls as 'not measurable'; this is load-bearing for the paper's headline claim, so the score is 4 rather than 2. The non-independence of the 1,200 runs per cell (24 permutations of 50 questions) is a statistical validity challenge to the active/inactive labels and band boundaries, but it is not a circularity pattern under the rubric and is therefore not scored here. The 'inconclusive falsifier is consistent with the 3-band model' passage is post-hoc rationalization, but because the paper discloses the misses and labels the model post hoc, I do not count it as a separate constructional circularity.
Axiom & Free-Parameter Ledger
free parameters (5)
- Vfloor (Cramér's V effect-size floor) =
0.05
- Spearman rho_processing cut =
-0.5
- alpha (chi-squared significance gate) =
0.05
- Goldilocks zone lower bound =
0.60
- Goldilocks zone upper bound =
0.95
axioms (5)
- ad hoc to paper Each of the k! permutations of a question is an independent observation for chi-squared and bootstrap inference.
- domain assumption Temperature-0 generation removes sampling stochasticity, so residual response variation is attributable to answer order.
- domain assumption MMLU subject accuracy is a valid index of model capability tier and benchmark saturation.
- ad hoc to paper The average accuracy over all 24 permutations is an uncontaminated covariate for the zone plot.
- standard math The Pearson chi-square null distribution is valid for the 4x2 contingency table with N=1,200.
invented entities (2)
-
processing_load mechanism
no independent evidence
-
content_ambiguity mechanism
no independent evidence
read the original abstract
Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single answer-order shuffles whose results confound the bias signal with content-level noise and sampling stochasticity. I introduce inspect_permute, an open-source extension to the inspect_ai evaluation framework that runs exhaustive answer-order permutations per question and reports the chi-squared / Cramer V signature of position bias with bootstrap confidence intervals. I apply the tool across four vendors (gpt-4o-mini, claude-haiku-4-5, gemini-2.5-flash, grok-3) on five MMLU subjects, 24,000 API calls under temperature-0 generation, with falsifier predictions pre-registered via a public SHA-256 hash before half the data was observed. Position bias turns out to be statistically detectable only within a roughly 60-95% base-accuracy Goldilocks zone. Below it, processing-load dominance swamps subject-specific signal; above it, ceiling effects compress the variance below the chi-squared test resolution. Detectable cells separate into two mechanism types: monotone A-to-D decrease (processing_load, in low-tier models) and non-monotone D-drop (content_ambiguity, in a narrow capability band). Standard MMLU places every frontier-tier model above the detection band, so absence of signal there should be read as not measurable, not unbiased. Together with the ceiling-effect characterisation in arXiv:2606.26185, this work brackets the detectable region of position-bias measurement and makes the field central question askable in a verifiable form. Package, data, preregistration under MIT.
Figures
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2606.26185 , year =
Tamba, Hiroki , title =. arXiv preprint arXiv:2606.26185 , year =
-
[2]
2026 , note =
Tamba, Hiroki , title =. 2026 , note =
2026
-
[3]
Allaire, J. J. and Teague, Charles and contributors , title =. 2024--2026 , howpublished =
2024
-
[4]
International Conference on Learning Representations (ICLR) , year =
Zheng, Lianmin and others , title =. International Conference on Learning Representations (ICLR) , year =
-
[5]
Findings of the Association for Computational Linguistics: NAACL , year =
Pezeshkpour, Pouya and Hruschka, Estevam , title =. Findings of the Association for Computational Linguistics: NAACL , year =
-
[6]
Proceedings of EMNLP , year =
Wang, Bowen and others , title =. Proceedings of EMNLP , year =
-
[7]
International Conference on Learning Representations (ICLR) , year =
Hendrycks, Dan and Burns, Collin and Basart, Steven and Zou, Andy and Mazeika, Mantas and Song, Dawn and Steinhardt, Jacob , title =. International Conference on Learning Representations (ICLR) , year =
-
[8]
, title =
Lord, Frederic M. , title =. Psychometric Monographs , number =
-
[9]
Statistical Theories of Mental Test Scores , editor =
Birnbaum, Allan , title =. Statistical Theories of Mental Test Scores , editor =
-
[10]
Sismondo, Sergio , title =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.