REVIEW 3 major objections 4 minor 38 references
All ten evaluated frontier LLMs deviate substantially from expert-validated reference answers, and eight of them form a statistically indistinguishable fidelity ceiling.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 13:55 UTC pith:SU7WKYNU
load-bearing objection A rigorous, fully crossed human evaluation with a carefully analyzed but unvalidated single-reference fidelity metric; the bimodal ceiling is interesting but may partly be an artifact of reference-style matching. the 3 major comments →
Response drift across frontier large language models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that response drift is a universal property of frontier LLMs and is experimentally accessible only through reference-anchored human evaluation. In a fully crossed, blinded study, every model deviates substantially from expert-validated references, but the magnitude is bimodal: Claude (47.0%) and Gemini (49.4%) form a high-fidelity tier, while eight models (Llama, Mistral, Grok, Perplexity, Copilot, DeepSeek, Qwen, ChatGPT) are statistically indistinguishable within a 77.6–80.5% deviation ceiling (confirmed by equivalence testing at a 5 pp bound). Drift profiles are structured: ceiling models share nearly identical per-question patterns (mean pairwise r = 0.90), w
What carries the argument
The core object is the fidelity deviation metric, defined as (5 − rating)/4 on a five-point Likert scale, averaged per model–question cell against a single expert-validated reference answer. This metric is embedded in a fully crossed repeated-measures design (every evaluator rates every model on every question), which enables variance decomposition—model identity accounts for 52% of variance (η² = 0.524) with question and participant effects separated—and per-question correlation analysis across models. Equivalence testing (TOST) is used to positively confirm the ceiling models' statistical indistinguishability, rather than merely failing to reject a null difference.
Load-bearing premise
The evaluation assumes that a single expert-validated reference answer per open-ended question is the right gold standard; if a question admits many valid responses that differ in style or structure from that reference, then the measured 47–81% 'drift' could reflect divergence from one reference rather than actual content error.
What would settle it
A decisive test would use the released 29,140-evaluation dataset: train an LLM-as-a-judge or a modern embedding-based regressor on a training split and check whether it can predict held-out human fidelity ratings with positive cross-validated R²; the paper predicts it cannot (all models currently yield negative R²). Alternatively, build multi-reference evaluations where each question has several expert-validated references and a model is judged faithful if it matches any one; if the 30+ percentage-point gap between tiers collapses to near zero, the single-reference assumption would be the caus
If this is right
- Model selection is the primary determinant of response fidelity: model identity explains over half the variance, outweighing question difficulty and evaluator variability.
- The two high-fidelity models show complementary strengths (near-zero correlation), suggesting that domain-specific routing between them could reduce overall deviation—though the paper notes no ensemble was tested.
- The eight ceiling models share systematic limitation patterns (pairwise correlations >0.85 even after controlling for question difficulty), implying convergent failure modes across independently developed systems.
- Automated similarity metrics and modern machine-learning predictors (including neural networks) cannot predict human fidelity judgments, so tracking response drift requires human evaluation.
- Reference-based fidelity evaluation and preference-based ranking measure different constructs; rankings of models can change substantially depending on which paradigm is used.
- The fidelity gap between tiers is robust to refusal handling: excluding refusals widens the gap by about 3.6 percentage points, because the high-fidelity models refused most often.
- The ceiling is not absolute: question-level tier gaps range from 0 to 73 percentage points, so individual questions can sharply separate models even within the ceiling.
Where Pith is reading between the lines
- If the ceiling's cross-model correlation reflects shared training data and RLHF procedures, then raising human-perceived fidelity may require changing training objectives rather than merely scaling compute or data.
- A testable extension: construct a question set designed to maximally separate ceiling models (e.g., the five most discriminating questions identified here) and rerun the evaluation; this could reveal hidden sub-tiers or stable ordering within the band.
- The near-zero correlation between the high-fidelity pair suggests a concrete ensemble test: compare each model alone against a per-question oracle that picks the better of Claude and Gemini, using the released dataset as a simulated oracle.
- Multi-reference scoring—rating each response against several expert-validated references—would directly test whether the measured drift is content error or stylistic divergence from a single reference; if the 30+ point tier gap collapses, the single-reference assumption is the main driver.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a fully crossed human evaluation of ten frontier LLMs on 62 open-ended questions, with 47 raters producing 29,140 ratings. Fidelity is defined as normalised deviation from an expert-validated reference answer (Eq. 1–3). The main findings are: (i) all models deviate substantially, with Claude and Gemini at 47.0% and 49.4% deviation and eight models in a statistically indistinguishable 77.6–80.5% band (TOST at Δ=5 pp); (ii) drift profiles vary by domain and question, with ceiling models highly correlated (r>0.85) and the two high-fidelity models uncorrelated; (iii) automated NLP similarity metrics explain <2% of variance in human judgments, and cross-validated ML models yield negative R². The paper interprets these results as evidence of a universal, human-accessible 'response drift' and a 'fidelity gap' between automated and human evaluation.
Significance. The study's design is a strength: the fully crossed 47×10×62 matrix, zero attrition, bootstrap CIs, mixed-effects modelling, TOST equivalence testing, and multiple sensitivity analyses are reported in unusual detail, and the analysis code is deposited. If the reference-based fidelity construct is valid, the two-tier structure with a convergent ceiling is an important empirical finding with implications for benchmark design and model selection. The negative results for automated metrics are also a useful caution, although they concern the specific feature set tested rather than a proof of fundamental inaccessibility. The main weakness is that the construct validity of the central metric rests on a single expert reference per open-ended question and on raters who may not be effectively blinded; these issues are acknowledged in the Discussion but are load-bearing for the headline claims.
major comments (3)
- [Methods, 'Task design and reference-answer development'; Eq. (1)–(3)] The primary endpoint is deviation from a single expert-validated reference per open-ended question. The Discussion concedes that 'open-ended questions admit multiple valid responses' and that the metric 'may partly capture stylistic conformity rather than content quality.' Because the construct-validity analyses (Results, 'Construct validity'; Methods, 'Construct validity analyses') are all computed on the same human ratings generated against those single references, they establish internal consistency of the ratings but do not independently establish that the ratings measure content fidelity rather than agreement with one canonical answer's content selection, emphasis, or structure. The 0–73 pp question-level gap variation with identically styled references is equally compatible with reference-specific content selection. This is load-bearing: the headline two-tier structure (47–49% vs.
- [Methods, 'Evaluation procedure and blinding', Stage 1 vs. Stage 3] The same participants who collected the 620 responses in Stage 1 rated them in Stage 3. Although model-identifying metadata was stripped, this is not effective blinding: participants may remember distinctive responses or recognize model-specific formatting, refusal patterns, or RAG citations (notably Perplexity). Because the raters were recruited for LLM experience and are the same individuals who interacted with each model, expectation effects could inflate the between-tier contrast. The manuscript asserts 'blinded conditions' without reporting any check for recognition (e.g., asking raters whether they recognized models). Independent raters, or at minimum a post-rating recognition probe, are needed.
- [Results, 'Construct validity: the human–automated gap is fundamental'] The claim that human fidelity judgements capture a dimension that is 'fundamentally inaccessible' to automated NLP metrics overstates what the evidence shows. The negative cross-validated R², mediation, and clustering results demonstrate that the seven chosen NLP features cannot predict these ratings; they do not demonstrate fundamental inaccessibility, especially since the features are simple surface/semantic similarities. Also, the 'evaluator consensus analysis' in that section refers to 'standard deviation of semantic similarity ratings across 47 evaluators,' but semantic similarity is computed automatically per model–question cell (Methods, 'Construct validity analyses'), so it has no per-evaluator distribution; this sentence appears to conflate the automated metric with human ratings. This analysis is used to support the construct-validity argument, so it needs correction or removal
minor comments (4)
- [Abstract and Results, 'Variance decomposition'] The abstract says 'Model identity accounts for over half of total variance in fidelity scores.' The two-way ANOVA on the aggregated 620-cell matrix gives η²=0.524 (Table 2A), but the participant-level linear mixed-effects model (Table 2B) gives a fixed-effect variance proportion of 0.41. Please qualify the claim by aggregation level.
- [Results, 'Automated metrics fail to capture human-perceived fidelity differences'] The text says automated metrics 'explained less than 2% of fidelity variance,' citing r²=0.018. However, Supplementary Table S11 reports r²=0.049 for Average Similarity and r²=0.016–0.018 for individual metrics. Please clarify which metric is used for the <2% claim and report the relevant value consistently.
- [Methods, 'Rating rubric'] The rubric allows 'minor stylistic differences' at rating 5 but does not operationalize how raters distinguish style from content. Given that the single-reference construct is central, the rubric would benefit from anchor examples illustrating acceptable versus unacceptable structural divergence.
- [Discussion, limitations] The limitation that 'fidelity is measured against a single reference' is candidly stated, but it appears only near the end of the paper. Consider foregrounding this caveat in the abstract or at the end of the introduction, since it conditions the interpretation of every headline number.
Circularity Check
No significant circularity: response drift is a direct measurement, not a derivation that reduces to its inputs.
full rationale
The paper's primary endpoint, fidelity deviation, is defined directly from human Likert ratings (Eqs. 1–3) and is not fitted to a subset of data and then 'predicted' on a closely related quantity. The ceiling/equivalence analysis, variance decomposition, correlations, and the human–automated gap are descriptive statistics or independent comparisons from the collected ratings and the separate NLP pipeline; none of these results is equivalent by construction to an input. The paper does not invoke a uniqueness theorem from the authors' prior work, nor does it smuggle in an ansatz via self-citation; the cited prior work is standard benchmark/method literature or an illustrative parallel ('densing law') that is non-load-bearing. The main weakness—that construct validity is assessed using the same single-reference human ratings that produced the headline numbers—is an acknowledged limitation rather than a circular derivation: the Discussion explicitly concedes that 'the metric may partly capture stylistic conformity rather than content quality' and that 'open-ended questions admit multiple valid responses,' and the paper refrains from equating reference deviation with absolute correctness. The construct-validity ML analyses use independent automated features; obtaining negative cross-validated R2 is an empirical result, not a tautology. No load-bearing step reduces, by the paper's own equations or by self-citation, to its own inputs.
Axiom & Free-Parameter Ledger
free parameters (2)
- Equivalence bound Δ (SESOI) =
5 pp (3 pp in sensitivity)
- Performance tier thresholds =
60% and 80% deviation
axioms (5)
- domain assumption A single expert-validated reference answer is a valid gold standard for each open-ended question.
- domain assumption 47 self-selected STEM professionals, calibrated to ICC=0.84, provide representative human judgments.
- ad hoc to paper Raters were effectively blinded to model identity during rating.
- domain assumption The 62 questions and six domains are representative of frontier-LLM use.
- standard math Standard linear-model assumptions (normality, random intercepts, full data) apply.
read the original abstract
All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet the magnitude and structure of this drift remain uncharacterised by systematic human evaluation. Here we report a fully crossed evaluation in which 47 geographically diverse participants each assessed all 62 multidomain questions across ten frontier LLMs under blinded conditions, yielding 29,140 independent assessments. Every model drifts, but drift magnitude varies substantially: eight models converge on a statistically indistinguishable ceiling (78-81% deviation), while two achieve lower deviation (47-49%). Drift profiles differ across six domains and 62 questions, with pairwise correlations among ceiling models exceeding r = 0.85. Automated similarity metrics explain less than 2% of variance in human judgements. These findings reveal that response drift is universal across frontier LLMs, domain- and question-dependent in structure, and accessible only through human-centred evaluation.
Figures
Reference graph
Works this paper leans on
-
[1]
B.et al.Language models are few-shot learners
Brown, T. B.et al.Language models are few-shot learners. InProceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20 (Curran Associates Inc., Red Hook, NY, USA, 2020)
2020
-
[2]
Chen, Z.et al.Revisiting scaling laws for language models: The role of data quality and training strategies, DOI: 10.18653/v1/2025.acl-long.1163 (2025)
-
[3]
InProceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22 (Curran Associates Inc., Red Hook, NY, USA, 2022)
Hoffmann, J.et al.Training compute-optimal large language models. InProceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22 (Curran Associates Inc., Red Hook, NY, USA, 2022)
2022
-
[4]
InProceedings of the International Conference on Learning Representations(2021)
Hendrycks, D.et al.Measuring massive multitask language understanding. InProceedings of the International Conference on Learning Representations(2021)
2021
-
[5]
InFirst Conference on Language Modeling (2024)
Rein, D.et al.GPQA: A graduate-level google-proof q&a benchmark. InFirst Conference on Language Modeling (2024)
2024
-
[6]
Du, X.et al.Evaluating large language models in class-level code generation. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering, ICSE ’24, DOI: 10.1145/3597503.3639219 (Association for Computing Machinery, New York, NY, USA, 2024)
arXiv 2024
-
[7]
InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23 (Curran Associates Inc., Red Hook, NY, USA, 2023)
Zheng, L.et al.Judging llm-as-a-judge with mt-bench and chatbot arena. InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23 (Curran Associates Inc., Red Hook, NY, USA, 2023)
2023
-
[8]
In Che, W., Nabende, J., Shutova, E
Zhao, Q.et al.MMLU-CF: A contamination-free multi-task language understanding benchmark. In Che, W., Nabende, J., Shutova, E. & Pilehvar, M. T. (eds.)Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13371–13391, DOI: 10.18653/v1/2025.acl-long.656 (Association for Computational Linguistics, Vi...
-
[9]
Phan, L.et al.A benchmark of expert-level academic questions to assess ai capabilities.Nature649, 1139–1146, DOI: 10.1038/s41586-025-09962-4 (2026). 15/33
-
[10]
Agrawal, A., Suzgun, M., Mackey, L. & Kalai, A. Do language models know when they’re hallucinating references? In Graham, Y. & Purver, M. (eds.)Findings of the Association for Computational Linguistics: EACL 2024, 912–928, DOI: 10.18653/v1/2024.findings-eacl.62 (Association for Computational Linguistics, St. Julian’s, Malta, 2024)
-
[11]
Bommasani, R., Liang, P. & Lee, T. Holistic evaluation of language models.Annals New York Acad. Sci.1525, 140–146, DOI: https://doi.org/10.1111/nyas.15007 (2023). https://nyaspubs.onlinelibrary.wiley.com/doi/pdf/10. 1111/nyas.15007
-
[12]
InProceedings of the 41st International Conference on Machine Learning, ICML’24 (JMLR.org, 2024)
Chiang, W.-L.et al.Chatbot arena: an open platform for evaluating llms by human preference. InProceedings of the 41st International Conference on Machine Learning, ICML’24 (JMLR.org, 2024)
2024
-
[13]
& Hashimoto, T
Dubois, Y., Liang, P. & Hashimoto, T. Length-controlled alpacaeval: A simple debiasing of automatic evaluators. InFirst Conference on Language Modeling(2024)
2024
-
[14]
In Proceedings of the 42nd International Conference on Machine Learning, ICML’25 (JMLR.org, 2025)
Li, T.et al.From crowdsourced data to high-quality benchmarks: arena-hard and benchbuilder pipeline. In Proceedings of the 42nd International Conference on Machine Learning, ICML’25 (JMLR.org, 2025)
2025
-
[15]
Y.et al.Wildbench: Benchmarking llms with challenging tasks from real users in the wild (2024)
Lin, B. Y.et al.Wildbench: Benchmarking llms with challenging tasks from real users in the wild (2024). 2406.04770
Pith/arXiv arXiv 2024
-
[16]
Min, S.et al.FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Bouamor, H., Pino, J. & Bali, K. (eds.)Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12076–12100, DOI: 10.18653/v1/2023.emnlp-main.741 (Association for Computational Linguistics, Singapore, 2023)
-
[17]
Sci.28, 517–540, DOI: https://doi.org/10.1016/j.tics.2024.01.011 (2024)
Mahowald, K.et al.Dissociating language and thought in large language models.Trends Cogn. Sci.28, 517–540, DOI: https://doi.org/10.1016/j.tics.2024.01.011 (2024)
-
[18]
Steyvers, M.et al.What large language models know and what people think they know.Nat. Mach. Intell.7, 221–231, DOI: 10.1038/s42256-024-00976-7 (2025). 19.Cohen, J.Statistical Power Analysis for the Behavioral Sciences(Routledge, New York, 1988), 2 edn
-
[20]
Equivalence tests: A practical primer for t tests, correlations, and meta-analyses.Soc
Lakens, D. Equivalence tests: A practical primer for t tests, correlations, and meta-analyses.Soc. Psychol. Pers. Sci.8, 355–362, DOI: 10.1177/1948550617697177 (2017). PMID: 28736600, https://doi.org/10.1177/ 1948550617697177
-
[21]
Bates, D., Mächler, M., Bolker, B. & Walker, S. Fitting linear mixed-effects models using lme4.J. Stat. Softw. 67, 1–48, DOI: 10.18637/jss.v067.i01 (2015)
-
[22]
Nakagawa, S. & Schielzeth, H. A general and simple method for obtaining r2 from generalized linear mixed- effects models.Methods Ecol. Evol.4, 133–142, DOI: https://doi.org/10.1111/j.2041-210x.2012.00261.x (2013). https://besjournals.onlinelibrary.wiley.com/doi/pdf/10.1111/j.2041-210x.2012.00261.x
Pith/arXiv arXiv 2041
-
[23]
Shrout, P. E. & Fleiss, J. L. Intraclass correlations: Uses in assessing rater reliability.Psychol. Bull.86, 420–428, DOI: 10.1037/0033-2909.86.2.420 (1979)
-
[24]
InProceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22 (Curran Associates Inc., Red Hook, NY, USA, 2022)
Ouyang, L.et al.Training language models to follow instructions with human feedback. InProceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22 (Curran Associates Inc., Red Hook, NY, USA, 2022)
2022
-
[25]
Preprint at arXiv https://arxiv.org/ abs/2212.08073
Bai, Y.et al.Constitutional AI: Harmlessness from AI feedback (2022). Preprint at arXiv https://arxiv.org/ abs/2212.08073
Pith/arXiv arXiv 2022
-
[26]
Surv.55, DOI:10.1145/3571730 (2023)
Ji, Z.et al.Surveyofhallucinationinnaturallanguagegeneration.ACM Comput. Surv.55, DOI:10.1145/3571730 (2023)
doi:10.1145/3571730 2023
-
[27]
Preprint at arXiv https://arxiv.org/abs/2210.01790
Shah, R.et al.Goal misgeneralization: Why correct specifications aren’t enough for correct goals (2022). Preprint at arXiv https://arxiv.org/abs/2210.01790
Pith/arXiv arXiv 2022
-
[28]
Xiao, C.et al.Densing law of llms.Nat. Mach. Intell.7, 1823–1833, DOI: 10.1038/s42256-025-01137-0 (2025)
-
[29]
& Hilton, J
Gao, L., Schulman, J. & Hilton, J. Scaling laws for reward model overoptimization. InProceedings of the 40th International Conference on Machine Learning, ICML’23 (JMLR.org, 2023)
2023
-
[30]
Koo, T. K. & Li, M. Y. A guideline of selecting and reporting intraclass correlation coefficients for reliability research.J. Chiropr. Medicine15, 155–163, DOI: https://doi.org/10.1016/j.jcm.2016.02.012 (2016). 16/33
-
[31]
Multi-agent ai systems need transparency.Nat
Nature Machine Intelligence. Multi-agent ai systems need transparency.Nat. Mach. Intell.8, 1–1, DOI: 10.1038/s42256-026-01183-2 (2026). 32.Bommasani, R.et al.On the opportunities and risks of foundation models.ArXiv(2021). 33.Likert, R. A technique for the measurement of attitudes.Arch. Psychol.22, 1–55 (1932)
-
[34]
& Shazeer, N
Fedus, W., Zoph, B. & Shazeer, N. Switch transformers: scaling to trillion parameter models with simple and efficient sparsity.J. Mach. Learn. Res.23(2022)
2022
-
[35]
Touvron, H.et al.Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288 (2023)
Pith/arXiv arXiv 2023
-
[36]
InProceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20 (Curran Associates Inc., Red Hook, NY, USA, 2020)
Lewis, P.et al.Retrieval-augmented generation for knowledge-intensive nlp tasks. InProceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20 (Curran Associates Inc., Red Hook, NY, USA, 2020)
2020
-
[37]
In First Conference on Language Modeling(2024)
DURMUS, E.et al.Towards measuring the representation of subjective global opinions in language models. In First Conference on Language Modeling(2024)
2024
-
[38]
Reliability in content analysis: Some common misconceptions and recommendations.Hum
Krippendorff, K. Reliability in content analysis: Some common misconceptions and recommendations.Hum. Commun. Res.30, 411–433, DOI: 10.1111/j.1468-2958.2004.tb00738.x (2006). https://academic.oup.com/hcr/ article-pdf/30/3/411/22338169/jhumcom0411.pdf
arXiv 2004
-
[39]
Hothorn, T., Bretz, F. & Westfall, P. Simultaneous inference in general parametric models.Biom. J.50, 346–363, DOI: https://doi.org/10.1002/bimj.200810425 (2008). 40.Holm, S. A simple sequentially rejective multiple test procedure.Scand. J. Stat.6, 65–70 (1979)
-
[40]
— High-resolution versions of all main and supplementary figures in PNG format. 33/33
-
[41]
Response drift across frontier large language models
Schuirmann, D. J. A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability.J. Pharmacokinet. Biopharm.15, 657–680, DOI: 10.1007/BF01068419 (1987). 17/33 Supplementary Information This Supplementary Information accompanies the main manuscript “Response drift across frontier large lang...
arXiv 1987
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.