{"id":"6a4e1e22-a8fe-431c-ab9a-66db848413e1","arxiv_id":"2604.06820","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The body finds LLM judges weakly track human credibility/sharing judgments on deceptive articles, while the abstract claims a direct-prediction failure that the body never tests.","lead":"This preprint's abstract announces a result about direct versus indirect prediction of human sharing from LLMs, but the body reports a different study: eight LLM judges are harsher than human readers and align more with each other than with people. The mismatch between the abstract and the manuscript means the paper's central claim is unclear as submitted.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The advertised central claim — direct sharing judgments add nothing beyond credibility for predicting human sharing — is never tested; §3.3 reports only rank correlations, not a predictive comparison.","rationale":"I read the full package in good faith and tried to identify the single most load-bearing condition for the paper's advertised central claim. That claim is the direct-prediction result in the title and abstract: credibility scores predict human sharing at least as well as sharing scores, and sharing scores provide no incremental benefit for unseen scenarios. The body's actual design is a proxy-validity audit: it correlates judge scores with human item means and compares judge–judge agreement. These correlations are not a test of the direct-prediction claim. The missing element is a within-judge, out-of-sample comparison of credibility and sharing scores as predictors of human sharing. No amount of reanalysis of §3.3's Spearman matrix can yield that comparison. This is not a disagreement with a consensus result; it is an internal mismatch between the stated contribution and the reported experiment. The reader's weakest-assumption point about noisy human item means (median six ratings per text, no reliability coefficient) is real and would matter for the body's proxy-validity claim: λ ≈ 0.45 may be attenuated, and the judge–human gap could be overstated. But even a perfect human reference would not rescue the abstract's direct-prediction claim, because the required predictive model is absent. I therefore agree with the reader's rejection verdict, but for a slightly different primary reason. The paper could be repaired either by removing the direct-prediction claims from the abstract and title and framing the body as a proxy-validity audit (with reliability analysis of the human reference), or by actually running the missing predictive comparison. Until one of those happens, the package is not acceptable as a coherent scientific claim.","tokens_in":19624,"tokens_out":4122,"duration_ms":45191,"concrete_test":"Run the missing direct-prediction analysis on the 290-text paired dataset. For each of the eight judges, predict the human item-mean willingness-to-share using (i) the judge's credibility score, (ii) the judge's willingness-to-share score, and (iii) both scores together, with cross-validation across texts (e.g., leave-one-scenario-out or 5-fold). Report out-of-sample Spearman and incremental R² for adding sharing scores to credibility, plus the sign and significance of the sharing coefficient. The abstract's claim requires that credibility-only ≥ sharing-only and that sharing adds no significant incremental validity for every judge. If this analysis cannot be performed on the released data, the abstract's central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The submitted abstract/title claims a direct-prediction result: \"For every evaluator, credibility scores track human sharing at least as closely as sharing scores, while sharing scores offer no detectable benefit beyond credibility when predicting responses to unseen scenarios.\" The full body, however, is a different study: \"Beyond Surface Judgments\" audits LLM judges as proxies for human readers. Section 3.3 reports item-level Spearman correlations between each judge and human item means (human–judge ρ≈0.45 for credibility, 0.24 for sharing; judge–judge ρ≈0.81 and 0.69). Section 3.4 reports signal-dependence deltas, and Section 3.5 reports a prompt ablation. None of these analyses compares, for any judge, the predictive validity of credibility scores versus willingness-to-share scores for human sharing. There is no regression, no nested-model comparison, no cross-validation, and no out-of-sample evaluation. The claimed pattern — \"the assumption holds for credibility but fails for sharing\" — requires exactly that missing comparison. The full-text abstract, by contrast, claims only that judges are weakly aligned with human readers and that internal agreement is not validity; that is a different claim. The mismatch is internal to the package: the title and abstract of the submission describe one experiment, while the body reports another. A reader cannot verify the headline finding from the reported methods, and no amount of reanalysis of the provided correlations would produce the missing predictive comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission is internally inconsistent. The arXiv title and abstract announce a direct-prediction result: for each LLM evaluator, asking for credibility is at least as good as asking for willingness-to-share for predicting human sharing, and sharing scores add nothing beyond credibility for unseen scenarios. The full text, however, reports a different study ('Beyond Surface Judgments'): eight LLM judges are audited as proxies for human readers on 290 deceptive articles, with findings that judges are harsher, recover human item rankings only weakly (Spearman ρ ≈ 0.45 for credibility, 0.24 for sharing), and rely on different textual signals despite high judge–judge agreement (ρ ≈ 0.81 and 0.69). The body never tests the direct-prediction claim. The human reference is the item mean of a median of six ratings per text, with no reliability analysis reported.","tokens_in":19943,"tokens_out":10305,"duration_ms":101406,"significance":"If the body's proxy-validity finding survives proper handling of measurement error, it is a useful contribution: it is one of the few human-grounded audits of LLM judges for disinformation, with matched ratings, multiple frontier judges, topic-level breakdowns, and a prompt ablation. The finding that judges agree more with each other than with humans is a valuable caution against using internal consistency as evidence of reader-response validity. However, the advertised direct-prediction claim—which is more novel and practically actionable—is entirely absent from the reported analyses, and the human-reference reliability issue directly affects the central negative result. As submitted, the paper's significance is compromised by the mismatch between its title/abstract and its actual content.","major_comments":[{"comment":"The abstract claims: 'For every evaluator, credibility scores track human sharing at least as closely as sharing scores, while sharing scores offer no detectable benefit beyond credibility when predicting responses to unseen scenarios.' This is a predictive claim comparing two score types as predictors of human sharing. Section 3.3 reports only rank correlations of judge credibility with human credibility and judge sharing with human sharing, plus judge–judge correlations. There is no regression, no nested-model comparison, no cross-validation, and no out-of-sample evaluation. The phrase 'unseen scenarios' is never operationalized. The headline result of the title/abstract cannot be verified from the reported methods.","section":"Abstract vs §3.3"},{"comment":"The human reference is the item mean of a median of six ratings per text. The paper reports no reliability coefficient (e.g., ICC, Cronbach's alpha), no confidence intervals for the Spearman correlations, and no multilevel model that treats human ratings as noisy draws. Item-mean measurement error attenuates human–judge correlations, so the reported gap (human–judge ρ ≈ 0.45/0.24 vs judge–judge ρ ≈ 0.81/0.69) may substantially overstate the true misalignment. Table 9's restriction to texts with at least two ratings is vacuous because all 290 texts meet that threshold. The authors should report inter-rater reliability and provide either disattenuated correlations or a model that accounts for rater noise.","section":"§2.2, §3.3"},{"comment":"The signal-dependence conclusion rests on textual-signal annotations produced by three LLM annotators, not by human annotators. The 'human' signal reliance is therefore a correlation between human outcomes and LLM-judged signal scores. If those annotations carry systematic bias, the judge–human deltas in Figure 5 could be partly an artifact of the annotation procedure. This concern is secondary to the main ordering result, but it should be acknowledged and ideally validated on a human-annotated subsample before the paper claims that judges 'rely on different textual signals.'","section":"§3.4"}],"minor_comments":[{"comment":"The abstract says 317 participants while the full text (Table 1, §2.2, Appendix A) reports 392 participants. This numerical inconsistency must be resolved.","section":"Abstract vs §2.2"},{"comment":"The arXiv title ('When Direct Prediction Fails...') and the full-text header ('Beyond Surface Judgments...') describe different papers. The authors must align the submission's framing with the actual study.","section":"Title"},{"comment":"Figure 5 reports N=286 while the main aligned set is N=290. The reason for the difference is not explained.","section":"Figure 5"},{"comment":"The analytical-role prompt includes a soft_refusal mechanism and can return the string 'soft_refusal' in place of integer scores, but the paper never reports how often this occurred. This should be stated for transparency.","section":"§D.4"},{"comment":"The generation prompt optimizes for high judge scores and explicitly simulates a sharing competition, so the texts may be atypical of real-world disinformation. This limits external validity and should be discussed more explicitly.","section":"§2.1"}],"recommendation":"major_revision","confidential_remarks":"The mismatch between the arXiv abstract and the full-text abstract is severe: the advertised direct-prediction experiment is never run. The editor should ask the authors to clarify which contribution is intended. If the direct-prediction claim is the main contribution, the missing analysis is essential and the current body is insufficient. If the proxy-validity audit is the intended contribution, the title/abstract must be rewritten and the reliability of the human reference must be addressed. Given the severity of the mismatch and the load-bearing nature of the human-reference reliability issue, I do not see how the paper can proceed without major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead this one if you're deciding whether to trust LLM judges for disinformation evaluation, but read the body, not the abstract. The advertised result—that asking for sharing directly is no better than asking for credibility—is never tested. The body reports a different study: a human-grounded audit of eight LLM judges on 290 deceptive articles, and that audit is a reasonable piece of empirical work.\n\nWhat's actually new: a reader-aligned benchmark with generated goal-directed deceptive texts, paired human ratings on credibility and sharing, and a three-part comparison (score distributions, item ranking, signal reliance). The finding that judges agree with each other at rho~0.8 while only weakly tracking human orderings (rho~0.45 credibility, 0.24 sharing) is a useful addition to the existing LLM-as-judge audit literature. The prompt ablation and the worked examples help make the failure concrete. Prompt templates and instruments are in the appendix, which is good practice.\n\nSoft spots, in order:\n\n1. The title/abstract don't match the body. The full text never compares, for any judge, the predictive validity of credibility vs sharing scores for human sharing. There's no regression, no nested model, no cross-validation. That missing analysis is not recoverable from the correlations reported. This is load-bearing because the submission's headline claim depends on it.\n\n2. Human reference reliability. The paper uses item means from a median of six ratings per text as the gold standard, with no reliability coefficient, no confidence intervals, no multilevel model. Noisy means will attenuate human–judge rank correlations. The gap is large, so I suspect it survives, but the paper doesn't show that.\n\n3. Participant counts. The submitted abstract says 317; the body says 392; Appendix A says 423 after resolving conflicts. That may be different cleaning stages, but as written it reads as an inconsistency and undercuts confidence in the data handling.\n\n4. Signal annotations come from three LLM annotators, which shares the exact bias family being audited. The signal-dependence analysis is suggestive, not probative.\n\nThe body's main negative result is plausible and consistent with prior audits. If the authors align the framing to what they actually did, or run the predictive test they claim, this is a solid contribution. As submitted, it needs major revision. I'd send it to peer review rather than desk reject; the benchmark and the question are worth referee time, but the advertised claim needs to be fixed or removed.","headline":"A useful proxy-validity audit is packaged with a headline predictive claim the body never tests; worth serious revision, not desk rejection.","tokens_in":20398,"tokens_out":3166,"would_cite":false,"duration_ms":31854,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper tries to show that LLM judges of deceptive articles are not valid proxies for human readers: eight frontier models agree strongly with one another but recover human credibility and sharing rankings only weakly, and asking a judge","keywords":["LLM-as-a-judge","proxy validity","misinformation risk","human–LLM alignment","perceived credibility","willingness to share","deceptive content","rank correlation"],"falsifier":"Take a random subset of the 290 articles and collect 50 or more human ratings per article; if the corrected human–judge Spearman correlations rise to near judge–judge levels, the proxy-validity gap would be largely a measurement artifact. Separately, run the abstract's stated unseen-scenario test — train a predictor on human sharing using judge credibility scores versus judge sharing scores and compare out-of-sample predictions — to check whether direct sharing scores really add nothing.","tokens_in":19528,"feed_emoji":"🤖","tokens_out":10468,"duration_ms":97318,"temperature":0.7,"pith_summary":"The paper tries to establish that LLM judges are not valid proxies for human readers when evaluating deceptive articles, and that the common shortcut of asking a model to predict a reader response directly can fail. Using 290 deceptive articles, 2,043 paired human ratings of credibility and willingness to share, and eight frontier judges, it reports that judges are harsher, compress the score scale, and recover human item-level rankings only weakly — average human–judge Spearman correlation is 0.45 for credibility and 0.24 for sharing, versus 0.81 and 0.69 among judges. The abstract adds a sharper claim: credibility scores track human sharing at least as closely as sharing scores do, so a direct question about sharing offers no detectable benefit for predicting sharing. If these claims hold, internal agreement among models cannot be treated as evidence that an evaluation measures reader-facing risk, and choosing what to ask a judge becomes as important as how to ask it.","feed_headline":"Eight leading LLM judges barely track human readers","feed_subtitle":"Eight frontier models agree with each other (0.81/0.69) but recover human credibility and sharing rankings at only 0.45/0.24.","key_machinery":"The carrying mechanism is the proxy-validity audit: for each of 290 deceptive articles, the mean of about six human ratings serves as the reader-response gold standard, and each judge's scores are compared along three axes — mean bias, item-level Spearman rank alignment, and signal dependence measured through judge-minus-human correlation deltas on four annotated textual cues (emotional intensity, logical rigour, authority reliance, data intensity). The decisive comparison is human–judge rank alignment versus judge–judge rank alignment on the same texts, which separates internal coherence from fidelity to readers.","core_discovery":"The paper's central discovery, stated on its own terms, is that LLM judges form a coherent evaluative group that is far more aligned with itself than with human readers. Across all eight judges and 290 texts, human–judge rank alignment averaged 0.45 for credibility and 0.24 for willingness to share, while judge–judge alignment averaged 0.81 and 0.69; judges were also uniformly harsher, compressed the human score scale (regression slopes of 0.42 and 0.29), and leaned more on logical rigour and against emotional intensity than human readers did. The paper pairs this audit with the claim that direct prediction fails: asking a judge for the target response — a sharing score — does not predict hu","pith_inferences":["The paper's abstract promises a predictive test — unseen-scenario prediction with direct versus indirect scores — but the body reports only correlation and calibration analyses; a reader should treat the 'direct prediction fails' conclusion as the authors' interpretation of correlation results, not as a demonstrated out-of-sample result.","If the direct-vs-indirect pattern is real, it suggests that perceived credibility functions as a broader latent judgment that partly drives sharing, so indirect questions may be cheaper or more robust; a testable extension is to compare intermediate judgments such as 'how credible would most readers find this?' against direct sharing scores.","Because each text has only about six human ratings, the reported human–judge correlations are probably attenuated; collecting more ratings per text on a subset would give a truer estimate of the gap, though the judge–judge vs human–judge contrast is large enough that it would likely persist.","The signal-level finding — judges overweight logical rigour and penalize emotional intensity — suggests an actionable prompt experiment: instructing judges to emulate a distracted first-impression reader may or may not close the gap; the paper's analytical-role ablation suggests prompt changes shift behavior without improving human alignment."],"forward_implications":["Model-based safety or benchmark scores for disinformation can look stable and self-consistent while still misrepresenting real-world reader risk.","Optimizing systems against judge scores can produce a reward-hacking failure: content may be tuned to what judges reward (structured, low-emotion, internally coherent) without reducing — possibly while increasing — what humans believe and share.","Judge-in-the-loop or self-improving evaluation pipelines can amplify the judge–human mismatch instead of correcting it.","Evaluation design should compare direct questions with indirect routes through related judgments; the target question is not automatically the best predictor of the target outcome.","LLM judges may remain useful for monitoring human-facing content, but only after validation against human outcomes, not on the strength of judge–judge agreement."],"fun_headline_variants":["LLM judges align with each other, not humans","Why LLM misinfo checks miss human sharing","Direct LLM scores fail to predict sharing","LLM judges: 0.81 self-agreement, 0.24 human match","Asking LLMs to predict sharing backfires"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The gold standard for reader response is the mean of about six human ratings per article, and the paper reports no reliability coefficients, confidence intervals, or multilevel model for those means; if the means are noisy, the human–judge correlations are attenuated and the central gap is overstated.","fun_headline_variants_meta":{"raw":{"variants":["LLM judges align with each other, not humans","Why LLM misinfo checks miss human sharing","Direct LLM scores fail to predict sharing","LLM judges: 0.81 self-agreement, 0.24 human match","Asking LLMs to predict sharing backfires"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000423,"raw_usage":{"total_tokens":2013,"prompt_tokens":754,"completion_tokens":1259,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1179}},"tokens_in":498,"tokens_out":1259,"duration_ms":7758,"temperature":1.0,"reasoning_tokens":1179,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T16:35:54.013990+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random subset of the 290 articles and collect 50 or more human ratings per article; if the corrected human–judge Spearman correlations rise to near judge–judge levels, the proxy-validity gap would be largely a measurement artifact. Separately, run the abstract's stated unseen-scenario test — train a predictor on human sharing using judge credibility scores versus judge sharing scores and compare out-of-sample predictions — to check whether direct sharing scores really add nothing.","supporting_citations":[],"review_version":2}