Pith. sign in

REVIEW 10 references

LLM position bias is only measurable between 60% and 95% accuracy.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 09:07 UTC pith:4KYAYFVY

load-bearing objection Useful diagnostic tool and a plausible 'Goldilocks zone' framing, but the significance tests that anchor the zone treat 24 permutations of the same question as 24 independent observations, which overstates confidence and likely shifts which cells are called biased.

arxiv 2607.20864 v1 pith:4KYAYFVY submitted 2026-07-23 cs.LG cs.CL

Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks

classification cs.LG cs.CL
keywords position biasLLM evaluationmultiple-choice benchmarksceiling effectspermutation testingCramér's Vbenchmark calibrationpre-registration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that position bias (a multiple-choice language model's tendency to answer according to the screen position of options rather than their content) can be statistically detected only when the model's base accuracy on a benchmark falls in a roughly 60 to 95 percent band. Below that band, processing noise swamps the bias signal; above it, ceiling effects compress response variance below the resolution of the chi-squared test. Because current frontier-tier models score above 95 percent on standard multitask benchmark subjects, the paper concludes that their flat permutation results should be read as 'not measurable on this instrument,' not as evidence of impartiality. This matters because it implies that many published no-bias claims on saturated benchmarks are unverifiable, and it reframes the field's central question: benchmark calibration must come before bias measurement. The claim is supported by a 24,000-call sweep across four vendor models and five subjects using exhaustive answer-order permutations, with pre-registered falsifier predictions.

Core claim

The discovery is a capability-band structure for position-bias detection. In the sweep, every cell with statistically significant bias lies inside a base-accuracy band of roughly 0.60 to 0.95, and no cell above 0.95 shows a detectable signal. The detectable cells separate into two mechanism types with distinct shape signatures: a monotone accuracy decline from option A to D, attributed to processing-load pressure in weaker models, and a non-monotone drop at the final position only, attributed to content ambiguity in a narrow capability band. The paper's central interpretive claim is that an absent chi-squared signal at high accuracy is not evidence that a model is unbiased; it is evidence th

What carries the argument

The central instrument is inspect_permute, a package that runs every answer-order permutation of each multiple-choice question (24 calls for a 4-option question) and records the presented position the model chooses on each run. It reports each vendor-by-subject cell as a chi-squared test, Cramér's V effect size, bootstrap confidence intervals, and Spearman's rank correlation between position and accuracy. The V-by-rho coordinate plane is the device that separates inactive cells from two mechanism classes: processing_load (monotone A-to-D decline, rho near -1) and content_ambiguity (final-position drop, rho near -0.4). The Goldilocks zone, the 60 to 95 percent base-accuracy window within whic

Load-bearing premise

The load-bearing assumption is that the 1,200 runs in a cell are independent, even though they are 50 questions repeated under 24 permutations of identical content; if responses within a question are correlated, the chi-squared p-values and bootstrap intervals are overconfident, and the floor edge itself is extrapolated rather than observed.

What would settle it

Take a frontier model with base accuracy above 0.95 on a standard multitask subject, append a prompt-level instruction to prefer option D when uncertain, and run the 24-permutation diagnostic on the same 50 questions. If the test still reports Cramér's V below 0.05 with p above 0.05, ceiling-effect masking is confirmed; if it detects the injected bias, the claimed upper edge of the Goldilocks zone is not universal.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A no-bias claim on a benchmark where the model scores above roughly 95 percent is unverifiable: the correct reading is that the instrument cannot resolve the model, not that the model is impartial.
  • Evaluators should run a cheap calibration scout before a full sweep and skip saturated cells, saving the order-of-magnitude wall-clock cost that reasoning-trace models impose.
  • Single-shuffle protocols cannot separate position bias from content noise and sampling error; exhaustive permutation (or a calibrated sampling of it) is the minimum instrument for a verifiable no-bias claim.
  • The two mechanism signatures imply different interventions: processing-load bias points to reasoning-budget or format changes, while content-ambiguity bias points to item and option-design changes.
  • The cost asymmetry between vendors (7.5x average, 18.5x peak wall-clock in this sweep) is a structural barrier to independent replication, shaping who can contest vendor evaluations.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the Goldilocks zone generalizes, the same ceiling masking should apply to other behavioral probes measured on saturated benchmarks (sycophancy, safety-rule adherence, instruction-following), so 'no signal' on an easy evaluation should never be read as model-level absence.
  • A direct test of the upper band: inject a known position preference into a frontier model scoring above 95 percent (e.g., by instruction to prefer option D) and run the 24-permutation diagnostic; if Cramér's V stays flat, ceiling compression is confirmed, and if it jumps, the upper edge is model- or benchmark-dependent.
  • The paper's lightweight git-hash pre-registration protocol is portable to any evaluation study and could become a standard way to mark the boundary between confirmatory and exploratory claims without a registered-report venue.
  • The two-mechanism taxonomy suggests a testable predictor: models trained with longer reasoning horizons should shift from processing_load toward content_ambiguity as their accuracy enters the mid-band, a prediction the current sweep can only hint at.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Circularity Check

1 steps flagged

The tool and sweep are independently grounded, but the load-bearing 'no signal ≠ no bias' reading of high-tier nulls is imported from the author's own prior ceiling-effect paper rather than derived from the present experiment.

specific steps
  1. self citation load bearing [Abstract and Section 1.4; applied in Section 4.5 and Section 5.2]
    "The methodological parent of this paper is arXiv 2606.26185, Necessary but Not Sufficient [Tamba, 2026a], which characterised the ceiling-effect side of position-bias detection: at high model capability, bias signals saturate and become uninformative as a measure of model character. The present work extends that result by reporting the floor side, where processing-load dominance swamps the subject-specific bias signal, and by characterising the middle band — the Goldilocks zone — where bias is genuinely measurable."

    The central interpretive claim — 'The right reading of the gemini-2.5-flash and grok-3 rows in Figure 1 is “no signal,” not “no bias”' — requires the premise that high base accuracy makes the chi-squared instrument unable to resolve bias. The present sweep alone establishes only that those cells are null (V≤0.06, p>0.45). The ceiling-side premise is not re-derived or power-analyzed here; it is cited to the author's own prior arXiv preprint. Thus the conclusion that high-tier nulls are unmeasurable rather than unbiased is supported by a self-citation chain, not by an independent derivation in the present paper.

full rationale

No equation-level circularity is present in the main statistics: Cramér's V is computed from the contingency table and base accuracy from the same runs, but the relationship is not tautological (active haiku cells near 0.99 show the upper edge is not a strict mathematical ceiling), and the paper explicitly labels the three-band model post hoc while scoring pre-registered predictions against the original classifier. The package, the 24,000-call sweep, and the reproducibility infrastructure are genuine independent content. The main circularity risk is the ceiling-effect premise imported from a same-author citation to reinterpret high-tier nulls as 'not measurable'; this is load-bearing for the paper's headline claim, so the score is 4 rather than 2. The non-independence of the 1,200 runs per cell (24 permutations of 50 questions) is a statistical validity challenge to the active/inactive labels and band boundaries, but it is not a circularity pattern under the rubric and is therefore not scored here. The 'inconclusive falsifier is consistent with the 3-band model' passage is post-hoc rationalization, but because the paper discloses the misses and labels the model post hoc, I do not count it as a separate constructional circularity.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

The central statistical analysis depends on run-level independence of repeated permutations, a standard chi-squared approximation, and a post-hoc band definition on the same 20 cells. The paper honestly discloses the post-hoc nature of the 3-band model and ships data, but the statistical assumptions are not independently validated and the lower edge of the Goldilocks zone is asserted from extrapolation.

free parameters (5)
  • Vfloor (Cramér's V effect-size floor) = 0.05
    Chosen as a conventional rule-of-thumb for the active/inactive classifier; sensitivity swept over 0.04-0.06 in Appendix B. It directly determines which cells count as showing position bias.
  • Spearman rho_processing cut = -0.5
    Separates monotone A-to-D processing_load from non-monotone content_ambiguity; sensitivity swept over -0.4 and -0.6 in Appendix B. Only the boundary grok cell changes label under perturbation.
  • alpha (chi-squared significance gate) = 0.05
    Conventional significance threshold; Appendix B shows alpha=0.10 would move the gpt-4o-mini x college_math borderline cell into the active processing_load set.
  • Goldilocks zone lower bound = 0.60
    Lower edge of the claimed detection band. No observed cell sits below 0.61, and Sections 4.5/5.5 explicitly state the floor edge is predicted by extrapolation, not measured.
  • Goldilocks zone upper bound = 0.95
    Upper edge of the claimed detection band, chosen from visual post-hoc inspection of Figure 4. The paper concedes the edge is subject-dependent: haiku is active at accuracy 0.99 on elementary_mathematics.
axioms (5)
  • ad hoc to paper Each of the k! permutations of a question is an independent observation for chi-squared and bootstrap inference.
    Sections 2.3-3.2 and 4.1 compute chi-squared statistics and run-level bootstraps over 1,200 runs per cell, treating the 24 runs of the same question as independent. No clustering or repeated-measures correction is provided; this is load-bearing for the p-values that define the Goldilocks zone.
  • domain assumption Temperature-0 generation removes sampling stochasticity, so residual response variation is attributable to answer order.
    Section 4.1 sets temperature 0 'wherever the vendor API supported it'; for vendors that do not support temperature-0, the residual stochasticity may not be controlled. The retry/backoff behavior during the Anthropic credit interruption is also assumed not to affect measurements.
  • domain assumption MMLU subject accuracy is a valid index of model capability tier and benchmark saturation.
    Section 4.5 uses base accuracy as the independent axis of the Goldilocks zone and classifies gemini/grok as 'above the detection band' based on their accuracy. If MMLU accuracy is not a good capability proxy, the claimed capability-band interpretation is unsupported.
  • ad hoc to paper The average accuracy over all 24 permutations is an uncontaminated covariate for the zone plot.
    Section 4.5 plots Cramér's V against base accuracy computed from the same permuted runs. Averaging over all permutations makes the covariate unbiased for the model's true accuracy, but its sampling error is ignored and the plot includes only 20 cells.
  • standard math The Pearson chi-square null distribution is valid for the 4x2 contingency table with N=1,200.
    This is a standard result, but it requires the run-level independence listed above. Under the actual repeated-measures design, the effective sample size for per-position accuracy differences is closer to 50 questions than 1,200 runs.
invented entities (2)
  • processing_load mechanism no independent evidence
    purpose: Label for a monotone A-to-D accuracy decline, interpreted as the model committing to the first plausible option when processing budget is exhausted.
    Observed in the four gpt-4o-mini cells from the same sweep; no independent replication or falsifiable handle outside the present dataset is provided. The shape signature overlaps with previously described 'first-option preference' in the cited literature.
  • content_ambiguity mechanism no independent evidence
    purpose: Label for a non-monotone accuracy drop specifically at option D, interpreted as anchoring produced by genuinely ambiguous subjects.
    Inferred from two claude-haiku-4-5 cells in the same sweep; no independent replication or external falsifiable handle is given, and the small D-drop may not survive a proper repeated-measures reanalysis.

pith-pipeline@v1.3.0-alltime-deepseek · 16840 in / 19670 out tokens · 185389 ms · 2026-08-01T09:07:30.404079+00:00 · methodology

0 comments
read the original abstract

Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single answer-order shuffles whose results confound the bias signal with content-level noise and sampling stochasticity. I introduce inspect_permute, an open-source extension to the inspect_ai evaluation framework that runs exhaustive answer-order permutations per question and reports the chi-squared / Cramer V signature of position bias with bootstrap confidence intervals. I apply the tool across four vendors (gpt-4o-mini, claude-haiku-4-5, gemini-2.5-flash, grok-3) on five MMLU subjects, 24,000 API calls under temperature-0 generation, with falsifier predictions pre-registered via a public SHA-256 hash before half the data was observed. Position bias turns out to be statistically detectable only within a roughly 60-95% base-accuracy Goldilocks zone. Below it, processing-load dominance swamps subject-specific signal; above it, ceiling effects compress the variance below the chi-squared test resolution. Detectable cells separate into two mechanism types: monotone A-to-D decrease (processing_load, in low-tier models) and non-monotone D-drop (content_ambiguity, in a narrow capability band). Standard MMLU places every frontier-tier model above the detection band, so absence of signal there should be read as not measurable, not unbiased. Together with the ceiling-effect characterisation in arXiv:2606.26185, this work brackets the detectable region of position-bias measurement and makes the field central question askable in a verifiable form. Package, data, preregistration under MIT.

Figures

Figures reproduced from arXiv: 2607.20864 by Hiroki Tamba.

Figure 1
Figure 1. Figure 1: Position-bias Cramér’s V across vendor × subject. Outline colors use the refined classifier of Section 3.3: red for processing_load, blue for content_ambiguity, no outline for inactive. • gpt-4o-mini registers significant position bias on four of five subjects, V from 0.10 (pro￾fessional_medicine) to 0.18 (elementary_mathematics, formal_logic). The exception is college_mathematics, where V = 0.08 falls jus… view at source ↗
Figure 2
Figure 2. Figure 2: Accuracy by correct-answer presented position. Line color from the refined classifier of Section 3.3. Red lines indicate processing_load (monotone A→D decrease). Blue lines indicate content_ambiguity (non-monotone D-drop). Grey lines indicate inactive. In every gpt-4o-mini cell with significant bias, accuracy decreases monotonically from position A to position D. The clearest example is formal_logic, where… view at source ↗
Figure 3
Figure 3. Figure 3: plots each cell in the V × ρ plane to bring the two-mechanism distinction into a single view. Vendor identity is encoded by marker shape; subject by colour [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Goldilocks zone — capability vs. detectability. V vs. base accuracy under the refined classifier. Shaded band marks the 60–95% accuracy range within which position bias is statistically detectable. Every cell with detected bias lies inside the shaded band. Outside the band, on the high-accuracy side, twelve cells cluster below the V = 0.05 floor: nothing detectable. Inside the band, ten cells span the full… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

10 extracted references · 1 linked inside Pith

  1. [1]

    arXiv preprint arXiv:2606.26185 , year =

    Tamba, Hiroki , title =. arXiv preprint arXiv:2606.26185 , year =

  2. [2]

    2026 , note =

    Tamba, Hiroki , title =. 2026 , note =

  3. [3]

    Allaire, J. J. and Teague, Charles and contributors , title =. 2024--2026 , howpublished =

  4. [4]

    International Conference on Learning Representations (ICLR) , year =

    Zheng, Lianmin and others , title =. International Conference on Learning Representations (ICLR) , year =

  5. [5]

    Findings of the Association for Computational Linguistics: NAACL , year =

    Pezeshkpour, Pouya and Hruschka, Estevam , title =. Findings of the Association for Computational Linguistics: NAACL , year =

  6. [6]

    Proceedings of EMNLP , year =

    Wang, Bowen and others , title =. Proceedings of EMNLP , year =

  7. [7]

    International Conference on Learning Representations (ICLR) , year =

    Hendrycks, Dan and Burns, Collin and Basart, Steven and Zou, Andy and Mazeika, Mantas and Song, Dawn and Steinhardt, Jacob , title =. International Conference on Learning Representations (ICLR) , year =

  8. [8]

    , title =

    Lord, Frederic M. , title =. Psychometric Monographs , number =

  9. [9]

    Statistical Theories of Mental Test Scores , editor =

    Birnbaum, Allan , title =. Statistical Theories of Mental Test Scores , editor =

  10. [10]

    Sismondo, Sergio , title =