{"id":"14fad711-08ba-436d-9211-fffd1fd162a8","arxiv_id":"2607.09783","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Judge-level effects and judge-specific slopes explain 61% of variance in H1 Connect research quality ratings versus 7% for papers and journals combined.","lead":"In a large post-publication peer-review database, differences between evaluators explain far more variance in quality ratings than differences between papers. This shows research assessment outcomes can be driven more by who judges than by what is judged, motivating routine noise audits.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The 61% judge-dominated partition rests on treating same-judge classification tags as independent paper attributes, inflating pattern-noise slopes.","rationale":"The Reader correctly isolates the co-production of tags and ratings as the weakest assumption supporting the strongest claim. The progressive models, singleton, multi-tag, ordinal, and permutation checks demonstrate that the variance partition is stable inside the H1 Connect data-generating process, yet they do not break the dependence between the predictors used for pattern noise and the judges whose slopes are estimated. Because the 49.5% slope term is what converts a modest judge-intercept advantage into the dramatic 61%/7% headline, any inflation of that term directly threatens the central quantitative claim. The proposed leave-one-judge-out factor averaging is a concrete, feasible check that would quantify how much of the slope variance survives once the factors are made exogenous to the rating judge. Until that (or an equivalent external paper-attribute measure) is shown, the verdict remains CONDITIONAL for the same reasons the Reader gave: the internal robustness suite is strong, but the interpretation of pattern noise as pure evaluator weighting of independent scientific attributes is not yet secured, and data/code release would allow independent verification. No stronger load-bearing flaw (e.g., model misspecification of the ordinal scale or journal confounding) appears once this issue is isolated.","tokens_in":11015,"tokens_out":649,"duration_ms":9099,"concrete_test":"Re-estimate the full model after replacing each paper’s latent-factor scores with the mean of the factor scores assigned by all other judges who rated that paper (or, for singletons, leave the paper out). If the sum of judge-slope variances falls substantially below ~49% (e.g., below 20%) while paper+journal variance remains ~7%, the 61% figure is an artefact of same-judge co-production and the headline ratio no longer holds.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that judges explain 61% vs papers+journals 7% is driven almost entirely by the 49.5% judge-slope term in the full model (Results: “Unmasked pattern noise”; Method: latent-factor construction and the formula with (0 + Latent_Factor1+…||judge)). Those six latent factors are derived from the twelve binary classification tags that the same judges assign in the same evaluation act that produces the rating. Consequently the random slopes largely capture within-judge consistency between a judge’s own tags and own rating, not differential weighting of independently measured paper attributes. The paper’s own permutation test (shuffling factors across cases) collapses slope variance to 0.6%, confirming that the association is real but not that the factors are exogenous to the judge. Baseline models without slopes already show judge intercepts > paper intercepts (26% vs 18%, or 25% vs 15% with journals), yet the headline 61%/7% ratio and the “more than seven times” language rely on the co-produced slopes. This is the softest link in the strongest claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper analyses 239,521 H1 Connect post-publication ratings by 12,649 judges of 193,128 papers and partitions rating variance with multilevel models. Baseline random-intercept models attribute more variance to judges (severity) than to papers; adding journals absorbs paper variance but not judge variance. A fuller model adds six latent factors from judge-assigned classification tags as fixed effects and judge-specific random slopes, yielding the headline partition: judge intercepts plus slopes ≈61% of total variance versus paper+journal ≈7%, with demographic/affiliation bias covariates explaining <1%. The authors interpret this as level and pattern noise dominating paper quality differences and argue for routine noise audits in high-stakes research assessment.","tokens_in":11251,"tokens_out":1337,"duration_ms":19414,"significance":"If the core claim holds—that evaluator-related variation substantially exceeds between-paper variation in a real large-scale assessment system—the paper supplies a concrete empirical basis for treating noise audits as standard practice alongside bias checks. Strengths include the unusually large N, progressive model building, and a battery of robustness checks (no-singletons subset, multi-tag subset, permutation collapse of slope variance from ~49% to 0.6%, Bayesian ordinal replication of the intercept partition, multiple optimisers). The bias analyses are a useful negative result. Even a carefully qualified version of the result would be of clear interest to research evaluation, scientometrics, and science policy.","major_comments":[{"comment":"Results “Unmasked pattern noise” and Method (latent-factor construction; formula with (0 + Latent_Factor1+…||judge)): the 49.5% judge-slope term that drives the headline 61%/7% partition is estimated from six latent factors derived from the twelve binary classification tags assigned by the same judges in the same evaluation that produces the rating. Those factors are therefore co-produced with the outcome, not independent measures of paper attributes. The random slopes largely capture within-judge consistency between a judge’s own tags and own rating. The permutation test shows the association is real rather than pure model flexibility, but does not establish that the factors are exogenous paper features. The abstract, Results, and Discussion should not present the 61% figure or the “more than seven times” language as pure pattern noise in the Kahneman sense without a clear qualification","section":"Results: Unmasked pattern noise; Method: Multilevel models"},{"comment":"Baseline models without slopes already show judge intercepts exceeding paper intercepts (26% vs 18%; 25% vs ~15% with journals). That qualitative ordering is the load-bearing empirical result and is robust across the no-singletons and Bayesian ordinal checks. The manuscript should lead with, and rest the central claim on, that intercept partition, and treat the slope model as an exploratory decomposition of residual judge structure whose interpretation is limited by co-production. Reframing the abstract and Discussion around the intercept result would make the central claim defensible without overstating the 61%/7% ratio.","section":"Abstract; Results: Judge-level variation exceeded paper-level variation; Discussion"},{"comment":"Discussion acknowledges that H1 Connect is a signed recommendation system with expert self-selection of papers, not gatekeeping peer review, and notes the compressed quality range. That selection and range restriction is still under-analysed as a threat to the paper-level variance component: if experts only rate papers they already view as above a high threshold, between-paper variance is mechanically reduced and the judge/paper ratio is inflated. A quantitative sensitivity discussion (or bounds) on how much paper variance could be missing under plausible selection would strengthen the claim that judges dominate papers rather than that the platform samples a narrow quality band.","section":"Discussion"}],"minor_comments":[{"comment":"Figure 1 caption and text: report exact variance percentages for each bar segment (or a companion table) so readers can verify the 61%/7% and robustness comparisons without estimating from the stacked bars.","section":"Figure 1"},{"comment":"Method: “no singletons” case counts are given as both 76,698 and 76,398 in adjacent paragraphs; reconcile the figure.","section":"Method: Robustness analyses"},{"comment":"Typos: “paritioning” (Discussion), “difference in the evaluated research” (Abstract, should be “differences”), and occasional “rankings” where “ratings” is meant (Results, journal model).","section":"Abstract; Results; Discussion"},{"comment":"Clarify whether latent factor scores are paper-level aggregates or evaluation-level (judge×paper) scores; the text says “for each evaluated paper” but tags are assigned per evaluation. This affects interpretation of the random slopes.","section":"Method: The H1 Connect dataset"},{"comment":"Appendix bias models: state sample sizes after gender_guesser exclusions more prominently so the <1% variance claim is easy to locate next to the main partition.","section":"Appendix"}],"recommendation":"major_revision","confidential_remarks":"The SKEPTIC concern about co-produced tags and slopes is real and load-bearing for the headline 61%/7% claim; it is not a reason to reject the paper, because the intercept-only result already supports a weaker but still important “judges > papers” conclusion. I would accept after a revision that demotes the slope-driven ratio and qualifies pattern-noise language. Scope fit for a cs.DL / research-evaluation venue is good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new result is a clean variance partition on 239k H1 Connect ratings: even the simplest models put judge intercepts above paper intercepts (26% vs 18%, or 25% vs ~15% once journals are in), and bias covariates (gender, North/South, affiliation match) explain almost nothing. That is useful and, at this scale with the progressive models and robustness suite, new relative to the IRR literature they cite.\n\nWhat they do well: the data are large and signed, the multilevel build is transparent, the no-singletons / multi-tag / multi-optimiser / Bayesian ordinal checks all hold, and the permutation collapse of the slope term shows the association is real rather than pure flexibility. The bias appendix is careful and correctly downplays the tiny effects. The REF parallel is fair: both systems rate already-filtered good work on a short scale.\n\nThe soft spot is real but limited. The 49.5% judge-slope term that produces the 61%/7% and “seven times” language comes from six latent factors built from the same judges’ own classification tags. Those slopes largely capture within-judge consistency between tags and rating, not independent paper attributes. The authors’ own permutation confirms the link is structural, not that the factors are exogenous. Drop the slopes and the qualitative claim still stands—judges still exceed papers—but the dramatic ratio does not. That is the main place the strongest claim over-reaches.\n\nGeneralizability is acknowledged: this is post-publication recommendation of already-published biomedical work, not gatekeeping peer review. Data and code are not released, which is a practical limit for re-analysis.\n\nThis is for people who care about research assessment design, REF-style exercises, or noise audits. The math and citation pattern look solid; the circularity burden is low once you treat the slopes carefully. I would send it to referees. They should force a clearer separation of the intercept-only result from the co-produced slope result and push for data release, but the paper is worth the referee time.","headline":"Solid large-scale variance partition showing judges dominate papers in H1 Connect ratings; the 61%/7% headline is inflated by same-judge tags, but the core result survives without them.","tokens_in":11865,"tokens_out":514,"would_cite":true,"duration_ms":5821,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"In post-publication research ratings, differences between judges explain far more variance than differences between papers.","keywords":["Research assessment","Peer review","Evaluator noise","Variance partitioning","Multilevel modelling","Systematic bias","Post-publication peer review"],"falsifier":"A replication in which independent, judge-blind characterizations of the same papers (or an evaluation system with pre-specified fixed criteria) reverse the variance partition so that paper-level effects exceed judge-level effects.","tokens_in":11875,"feed_emoji":"⚖️","tokens_out":553,"duration_ms":4608,"temperature":0.7,"pith_summary":"Expert research assessment is meant to track the quality of the work under review, yet human judgment is noisy. This paper asks which source of variation is larger: genuine differences among papers, or systematic differences among the people who rate them. Using multilevel models on hundreds of thousands of signed quality ratings from a large biomedical post-publication review platform, the authors partition the variance into paper effects, journal effects, judge severity, and judge-specific ways of weighting scientific attributes. Across models, judge-related components dominate: in the fullest specification they account for roughly 61 percent of total variance, while paper and journal effects together account for only about 7 percent. Demographic and affiliation-related biases explain less than 1 percent. The authors conclude that assessment outcomes are shaped more by who is judging than by what is being judged, and that routine noise audits that separate judge variance from paper variance should become standard in high-stakes evaluation.","feed_headline":"Judges drive research ratings more than the papers do","feed_subtitle":"In a large review database, evaluator differences explain 61% of rating variance; papers and journals only 7%.","key_machinery":"Multilevel variance partitioning that decomposes ratings into paper intercepts, journal intercepts, judge intercepts (level noise), and judge-specific random slopes on latent factors derived from classification tags (pattern noise).","core_discovery":"In a large post-publication peer-review dataset of quality ratings, judge-related variation (overall severity plus judge-specific weighting of scientific attributes) accounts for substantially more of the variance in ratings than variation attributable to the papers and journals being rated; directional biases such as gender and global affiliation explain almost none of it.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Judges explain 61% of rating variance, papers and journals just 7%","Evaluator differences outweigh paper quality in research ratings","Judge-level noise dominates post-publication quality scores","Ratings shaped more by judges than by the papers assessed","Severity and attribute weights of judges dwarf paper effects"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The classification tags that judges themselves assign can be treated as measures of paper attributes when estimating how differently judges weight those attributes.","fun_headline_variants_meta":{"raw":{"variants":["Judges explain 61% of rating variance, papers and journals just 7%","Evaluator differences outweigh paper quality in research ratings","Judge-level noise dominates post-publication quality scores","Ratings shaped more by judges than by the papers assessed","Severity and attribute weights of judges dwarf paper effects"]},"model":"grok-4.5","effort":"low","cost_usd":0.002556,"raw_usage":{"total_tokens":1014,"prompt_tokens":786,"num_sources_used":0,"completion_tokens":81,"cost_in_usd_ticks":25560000,"prompt_tokens_details":{"text_tokens":786,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":147,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":786,"tokens_out":81,"duration_ms":2192,"temperature":1.0,"reasoning_tokens":147,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T15:47:28.237784+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A replication in which independent, judge-blind characterizations of the same papers (or an evaluation system with pre-specified fixed criteria) reverse the variance partition so that paper-level effects exceed judge-level effects.","supporting_citations":[],"review_version":1}