{"id":"0437056b-1260-4689-9d20-e90c8fe02741","arxiv_id":"2607.20454","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"All ten tested frontier LLMs deviate substantially from expert-validated references, with eight models forming an indistinguishable ~78–81% fidelity ceiling while Claude and Gemini reach ~47–49%.","lead":"In a blinded human evaluation of 29,140 responses, all ten frontier LLMs deviated substantially from expert-written reference answers; eight models clustered near 78–81% deviation while Claude and Gemini sat lower near 47–49%. The paper argues this 'response drift' is invisible to automated text-similarity metrics, which explained under 2% of human judgments.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-expert reference as gold standard could drive the measured deviation ceiling; without multi-reference scoring, the 47–81% drift numbers may reflect reference-style mismatch rather than content infidelity.","rationale":"The reader's weakest_assumption identifies the same core issue: fidelity is measured against a single expert-validated reference, and the construct-validity checks rely on the same human ratings they are meant to validate. This is the single most load-bearing concern because every headline quantity—the 47–49% high-fidelity tier, the 78–81% ceiling, the two-tier separation, and the claim that drift is 'substantial'—depends on interpreting deviation from one reference as content infidelity. The paper itself acknowledges the limitation and proposes multi-reference scoring as future work, which is exactly the missing control. The proposed test would settle whether the effect is real or an artifact. Since the reader already marked the verdict CONDITIONAL, and the concern supports that conditionality without disproving the central phenomenon, no change to the reader's verdict is needed.","tokens_in":31361,"tokens_out":3114,"duration_ms":31928,"concrete_test":"Select 10–12 questions from the existing set, balanced across domains. For each, have independent experts draft two additional reference answers that are factually equivalent but differ in structure, emphasis, and formatting; retain the original reference as the third. Recruit a fresh set of calibrated raters (or reuse the same 47 after a washout) to rate a fixed sample of model responses (e.g., 6 models × 10 questions) against each of the three references in a counterbalanced, fully blinded design. Compare model-level deviation distributions across reference versions. If the Claude/Gemini versus ceiling gap shrinks by more than 10 percentage points, or if ceiling models drop below 70% deviation under an alternative valid reference, then single-reference anchoring drives the headline. If deviations stay within roughly 5 pp across references, the ceiling is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that deviation from a single expert-validated reference is a valid measure of response fidelity for open-ended questions (Methods: 'Task design and reference-answer development'; Eq. 1–3). The Discussion concedes the metric 'may partly capture stylistic conformity rather than content quality' and that 'open-ended questions admit multiple valid responses', yet the headline numbers (47.0–80.5%) and the two-tier structure are reported as properties of the models. The construct-validity checks do not resolve this: predictive models, mediation, clustering, and evaluator-consensus analyses all use the same human ratings generated against the same single references, so they can validate internal consistency of those ratings but not whether the ratings measure content fidelity rather than agreement with one reference's structure or emphasis. The question-level gap variation of 0–73 pp with 'identically styled references' is also compatible with reference-specific content selection—a valid alternative answer choosing different key facts can be marked down regardless of style. Because the rubric permits 'minor stylistic differences' but does not operationalize how raters distinguish style from content, the measured ceiling might be an artifact of comparing every model to one canonical answer.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a fully crossed human evaluation of ten frontier LLMs on 62 open-ended questions, with 47 raters producing 29,140 ratings. Fidelity is defined as normalised deviation from an expert-validated reference answer (Eq. 1–3). The main findings are: (i) all models deviate substantially, with Claude and Gemini at 47.0% and 49.4% deviation and eight models in a statistically indistinguishable 77.6–80.5% band (TOST at Δ=5 pp); (ii) drift profiles vary by domain and question, with ceiling models highly correlated (r>0.85) and the two high-fidelity models uncorrelated; (iii) automated NLP similarity metrics explain <2% of variance in human judgments, and cross-validated ML models yield negative R². The paper interprets these results as evidence of a universal, human-accessible 'response drift' and a 'fidelity gap' between automated and human evaluation.","tokens_in":31640,"tokens_out":6005,"duration_ms":53143,"significance":"The study's design is a strength: the fully crossed 47×10×62 matrix, zero attrition, bootstrap CIs, mixed-effects modelling, TOST equivalence testing, and multiple sensitivity analyses are reported in unusual detail, and the analysis code is deposited. If the reference-based fidelity construct is valid, the two-tier structure with a convergent ceiling is an important empirical finding with implications for benchmark design and model selection. The negative results for automated metrics are also a useful caution, although they concern the specific feature set tested rather than a proof of fundamental inaccessibility. The main weakness is that the construct validity of the central metric rests on a single expert reference per open-ended question and on raters who may not be effectively blinded; these issues are acknowledged in the Discussion but are load-bearing for the headline claims.","major_comments":[{"comment":"The primary endpoint is deviation from a single expert-validated reference per open-ended question. The Discussion concedes that 'open-ended questions admit multiple valid responses' and that the metric 'may partly capture stylistic conformity rather than content quality.' Because the construct-validity analyses (Results, 'Construct validity'; Methods, 'Construct validity analyses') are all computed on the same human ratings generated against those single references, they establish internal consistency of the ratings but do not independently establish that the ratings measure content fidelity rather than agreement with one canonical answer's content selection, emphasis, or structure. The 0–73 pp question-level gap variation with identically styled references is equally compatible with reference-specific content selection. This is load-bearing: the headline two-tier structure (47–49% vs.","section":"Methods, 'Task design and reference-answer development'; Eq. (1)–(3)"},{"comment":"The same participants who collected the 620 responses in Stage 1 rated them in Stage 3. Although model-identifying metadata was stripped, this is not effective blinding: participants may remember distinctive responses or recognize model-specific formatting, refusal patterns, or RAG citations (notably Perplexity). Because the raters were recruited for LLM experience and are the same individuals who interacted with each model, expectation effects could inflate the between-tier contrast. The manuscript asserts 'blinded conditions' without reporting any check for recognition (e.g., asking raters whether they recognized models). Independent raters, or at minimum a post-rating recognition probe, are needed.","section":"Methods, 'Evaluation procedure and blinding', Stage 1 vs. Stage 3"},{"comment":"The claim that human fidelity judgements capture a dimension that is 'fundamentally inaccessible' to automated NLP metrics overstates what the evidence shows. The negative cross-validated R², mediation, and clustering results demonstrate that the seven chosen NLP features cannot predict these ratings; they do not demonstrate fundamental inaccessibility, especially since the features are simple surface/semantic similarities. Also, the 'evaluator consensus analysis' in that section refers to 'standard deviation of semantic similarity ratings across 47 evaluators,' but semantic similarity is computed automatically per model–question cell (Methods, 'Construct validity analyses'), so it has no per-evaluator distribution; this sentence appears to conflate the automated metric with human ratings. This analysis is used to support the construct-validity argument, so it needs correction or removal","section":"Results, 'Construct validity: the human–automated gap is fundamental'"}],"minor_comments":[{"comment":"The abstract says 'Model identity accounts for over half of total variance in fidelity scores.' The two-way ANOVA on the aggregated 620-cell matrix gives η²=0.524 (Table 2A), but the participant-level linear mixed-effects model (Table 2B) gives a fixed-effect variance proportion of 0.41. Please qualify the claim by aggregation level.","section":"Abstract and Results, 'Variance decomposition'"},{"comment":"The text says automated metrics 'explained less than 2% of fidelity variance,' citing r²=0.018. However, Supplementary Table S11 reports r²=0.049 for Average Similarity and r²=0.016–0.018 for individual metrics. Please clarify which metric is used for the <2% claim and report the relevant value consistently.","section":"Results, 'Automated metrics fail to capture human-perceived fidelity differences'"},{"comment":"The rubric allows 'minor stylistic differences' at rating 5 but does not operationalize how raters distinguish style from content. Given that the single-reference construct is central, the rubric would benefit from anchor examples illustrating acceptable versus unacceptable structural divergence.","section":"Methods, 'Rating rubric'"},{"comment":"The limitation that 'fidelity is measured against a single reference' is candidly stated, but it appears only near the end of the paper. Consider foregrounding this caveat in the abstract or at the end of the introduction, since it conditions the interpretation of every headline number.","section":"Discussion, limitations"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern lands: the single-expert-reference issue is real and load-bearing. The paper is methodologically careful in many respects, and the authors acknowledge the limitation, but the construct-validity checks do not break the circularity. The blinding issue is also more serious than the paper suggests: the same participants collected and rated responses, and no recognition check is reported. I recommend a major revision asking for a multi-reference validation subset and a genuine blinding or recognition probe. The reviewer should also ask the authors to temper the 'fundamental inaccessibility' language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper before trusting any of its headline numbers: it's a large, carefully executed human evaluation (47 raters x 10 models x 62 questions = 29,140 assessments) with a fully crossed design, and it finds a striking bimodal split—Claude and Gemini around 47-49% deviation, eight other models clustered at 78-81%. The statistical work is genuinely careful: bootstrap CIs, mixed-effects variance decomposition, TOST equivalence tests, robustness checks, and a refusal-exclusion analysis all hang together. The per-question correlation structure (ceiling models >0.85, high-fidelity pair ~0.12) is a nice contribution. If the metric were clearly valid, this would be an important paper.\n\nThe problem is the metric. Deviation is measured against a single expert-validated reference per open-ended question. The paper itself concedes open-ended questions admit multiple valid answers, and that the metric 'may partly capture stylistic conformity rather than content quality.' That's not a minor caveat—it's the load-bearing assumption. The construct-validity battery (ML prediction, mediation, clustering) uses the same human ratings generated against the same references, so it can show the ratings are internally consistent but cannot show they measure content fidelity rather than agreement with one reference's structure or emphasis. The stress-test note is right: the 0-73 pp question-level gap is compatible with reference-specific content selection, not just style. The blinding is also weaker than claimed: participants collected the responses themselves in Stage 1, then rated them in Stage 3, so they could recognize models by formatting, memory, or RAG citations. The 'fundamental' inaccessibility to automated metrics outruns the evidence—seven NLP features and a few off-the-shelf models getting negative R2 does not prove a fundamental perceptual gap.\n\nThat said, the paper is honest about its limitations and doesn't hide the single-reference issue. It explicitly calls for multi-reference scoring as future work. The data and code aren't deposited yet, only 'available on request,' which is a real barrier for independent checking. The ceiling finding could survive a multi-reference design, but it could also dissolve if Claude and Gemini happen to match reference style better. So I'd treat the specific percentages as provisional. Who's this for? Anyone working on LLM evaluation methodology or deployment decisions; it's a well-powered study that raises the right questions, even if it doesn't yet answer the validity challenge. It deserves a serious referee, but the referee should insist on either multi-reference scoring or release of the full dataset so others can test the alternative explanation.","headline":"A rigorous, fully crossed human evaluation with a carefully analyzed but unvalidated single-reference fidelity metric; the bimodal ceiling is interesting but may partly be an artifact of reference-style matching.","tokens_in":32104,"tokens_out":1948,"would_cite":false,"duration_ms":21747,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"All ten evaluated frontier LLMs deviate substantially from expert-validated reference answers, and eight of them form a statistically indistinguishable fidelity ceiling.","keywords":["response drift","LLM evaluation","human evaluation","fully crossed design","fidelity deviation","automated metrics","reference-based assessment","model comparison"],"falsifier":"A decisive test would use the released 29,140-evaluation dataset: train an LLM-as-a-judge or a modern embedding-based regressor on a training split and check whether it can predict held-out human fidelity ratings with positive cross-validated R²; the paper predicts it cannot (all models currently yield negative R²). Alternatively, build multi-reference evaluations where each question has several expert-validated references and a model is judged faithful if it matches any one; if the 30+ percentage-point gap between tiers collapses to near zero, the single-reference assumption would be the caus","tokens_in":31271,"feed_emoji":"📉","tokens_out":3933,"duration_ms":36436,"temperature":0.7,"pith_summary":"This paper tries to establish that response drift—the systematic deviation of model outputs from expert-validated reference answers—is universal across frontier large language models, and that its magnitude and structure can only be measured through human evaluation. Using a fully crossed design in which 47 evaluators rated all 62 questions across all 10 models (29,140 assessments), the authors find a bimodal distribution: two models (Claude and Gemini) deviate 47–49%, while the other eight cluster in a statistically equivalent 78–81% ceiling. Drift is domain- and question-dependent, and model identity accounts for about half of all variance in fidelity scores. Crucially, automated NLP similarity metrics and machine-learning predictors explain less than 2% of the variance in human fidelity judgments, indicating that the phenomenon is invisible to current automated evaluation pipelines. If correct, these results imply that leaderboard rankings and preference-based evaluation miss a large, human-perceived quality dimension.","feed_headline":"Eight of ten frontier LLMs share a 78–81% drift ceiling","feed_subtitle":"A 29,140-rating human study finds every model drifts from expert answers; two lead by ~30 points, and automated metrics miss the gap.","key_machinery":"The core object is the fidelity deviation metric, defined as (5 − rating)/4 on a five-point Likert scale, averaged per model–question cell against a single expert-validated reference answer. This metric is embedded in a fully crossed repeated-measures design (every evaluator rates every model on every question), which enables variance decomposition—model identity accounts for 52% of variance (η² = 0.524) with question and participant effects separated—and per-question correlation analysis across models. Equivalence testing (TOST) is used to positively confirm the ceiling models' statistical indistinguishability, rather than merely failing to reject a null difference.","core_discovery":"The paper's central claim is that response drift is a universal property of frontier LLMs and is experimentally accessible only through reference-anchored human evaluation. In a fully crossed, blinded study, every model deviates substantially from expert-validated references, but the magnitude is bimodal: Claude (47.0%) and Gemini (49.4%) form a high-fidelity tier, while eight models (Llama, Mistral, Grok, Perplexity, Copilot, DeepSeek, Qwen, ChatGPT) are statistically indistinguishable within a 77.6–80.5% deviation ceiling (confirmed by equivalence testing at a 5 pp bound). Drift profiles are structured: ceiling models share nearly identical per-question patterns (mean pairwise r = 0.90), w","pith_inferences":["If the ceiling's cross-model correlation reflects shared training data and RLHF procedures, then raising human-perceived fidelity may require changing training objectives rather than merely scaling compute or data.","A testable extension: construct a question set designed to maximally separate ceiling models (e.g., the five most discriminating questions identified here) and rerun the evaluation; this could reveal hidden sub-tiers or stable ordering within the band.","The near-zero correlation between the high-fidelity pair suggests a concrete ensemble test: compare each model alone against a per-question oracle that picks the better of Claude and Gemini, using the released dataset as a simulated oracle.","Multi-reference scoring—rating each response against several expert-validated references—would directly test whether the measured drift is content error or stylistic divergence from a single reference; if the 30+ point tier gap collapses, the single-reference assumption is the main driver."],"forward_implications":["Model selection is the primary determinant of response fidelity: model identity explains over half the variance, outweighing question difficulty and evaluator variability.","The two high-fidelity models show complementary strengths (near-zero correlation), suggesting that domain-specific routing between them could reduce overall deviation—though the paper notes no ensemble was tested.","The eight ceiling models share systematic limitation patterns (pairwise correlations >0.85 even after controlling for question difficulty), implying convergent failure modes across independently developed systems.","Automated similarity metrics and modern machine-learning predictors (including neural networks) cannot predict human fidelity judgments, so tracking response drift requires human evaluation.","Reference-based fidelity evaluation and preference-based ranking measure different constructs; rankings of models can change substantially depending on which paradigm is used.","The fidelity gap between tiers is robust to refusal handling: excluding refusals widens the gap by about 3.6 percentage points, because the high-fidelity models refused most often.","The ceiling is not absolute: question-level tier gaps range from 0 to 73 percentage points, so individual questions can sharply separate models even within the ceiling."],"fun_headline_variants":["All 10 frontier LLMs drift, but two drift far less","Claude and Gemini break drift ceiling; eight others share it","29,140 human ratings show LLM drift is universal","Drift in every frontier LLM: automated checks miss it","Two LLMs cut expert-miss rate to 48%; eight stay at 80%"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The evaluation assumes that a single expert-validated reference answer per open-ended question is the right gold standard; if a question admits many valid responses that differ in style or structure from that reference, then the measured 47–81% 'drift' could reflect divergence from one reference rather than actual content error.","fun_headline_variants_meta":{"raw":{"variants":["All 10 frontier LLMs drift, but two drift far less","Claude and Gemini break drift ceiling; eight others share it","29,140 human ratings show LLM drift is universal","Drift in every frontier LLM: automated checks miss it","Two LLMs cut expert-miss rate to 48%; eight stay at 80%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00037,"raw_usage":{"total_tokens":1803,"prompt_tokens":711,"completion_tokens":1092,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":455,"completion_tokens_details":{"reasoning_tokens":1000}},"tokens_in":455,"tokens_out":1092,"duration_ms":9647,"temperature":1.0,"reasoning_tokens":1000,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T13:55:38.480922+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would use the released 29,140-evaluation dataset: train an LLM-as-a-judge or a modern embedding-based regressor on a training split and check whether it can predict held-out human fidelity ratings with positive cross-validated R²; the paper predicts it cannot (all models currently yield negative R²). Alternatively, build multi-reference evaluations where each question has several expert-validated references and a model is judged faithful if it matches any one; if the 30+ percentage-point gap between tiers collapses to near zero, the single-reference assumption would be the caus","supporting_citations":[],"review_version":1}