{"id":"aa4fdbd6-fa50-41e4-90a0-5f30fb839167","arxiv_id":"2608.04463","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Open-ended LLM conformity is not captured by answer flips: wrong peers degrade revisions, and judges' ratings shift when peer context is visible, so evaluation must be modeled explicitly.","lead":"This paper tests how language models revise answers after reading peer comments, and how judges score those revisions. It finds wrong peers lower revision quality and that the same answer is judged differently when the peer context is visible to the judge.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The evaluator-side peer-context sensitivity (Conclusion 3) rests on a blind-then-informed comparison with no reported counterbalancing; pass-order drift is confounded with gamma_j and must be ruled out.","rationale":"The reader's weakest-assumption analysis identifies the pass-order confound in the evaluator-side blind-informed contrast, and my reading of the full paper confirms it is the single most load-bearing concern. The paper's central new claim is that evaluators themselves exhibit heterogeneous sensitivity to visible peer context; this claim is entirely supported by the paired blind/informed ratings, and the appendix gives no evidence that the two passes were counterbalanced or otherwise controlled for order. The concern is concrete: gamma_j is estimated from a difference between two passes, so anything that changes judge behavior between passes (fatigue, scale drift, learning, conditioning on the first rating) is attributed to peer-context sensitivity. The fact that the signed context d_i is a fixed property of each item makes the problem worse, because second-pass effects can correlate with condition rather than being a pure overall shift. I agree that this does not undermine the generator-side findings, which are robustly supported by blind ratings, consistency across all 12 cells, self-rating exclusion, and fixed-classifier validation. The verdict CONDITIONAL is appropriate: the paper should provide evidence about pass ordering or run a counterbalanced subset before Conclusion 3 is accepted as stated. The proposed randomized-order test would decisively resolve the concern.","tokens_in":22052,"tokens_out":2506,"duration_ms":25796,"concrete_test":"Run a counterbalanced evaluation subset on (say) 500 items per judge: randomize per item whether the blind or informed rating is collected first, then estimate gamma_j separately for the two order groups and test for a group difference. If the order groups produce materially different gamma_j (beyond overlapping credible intervals), the current estimates are confounded and Conclusion 3 must be re-qualified. As a cheaper diagnostic on existing data, refit Eq. (7) with an additional judge-specific second-pass effect (e.g., a judge-by-pass cutpoint shift or a pass-index term); if gamma_j moves by more than the width of its credible interval, order effects are present.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Conclusion 3 is identified from Eq. (7): same answer, blind rating versus informed rating, with gamma_j capturing directional judge sensitivity to the signed peer context d_i. The Appendix B.1-B.2 shows the blind and informed prompts as fixed templates in a fixed order, and no randomization or counterbalancing of the two rating passes is reported anywhere. If informed ratings always follow blind ratings, any systematic within-judge change across the passes (fatigue, scale drift, learning, or condition-correlated carryover) is absorbed into gamma_j. The paper itself concedes in Section 3.3 that a condition-independent leniency shift is not separately identified, but the problem is more general: because d_i is a fixed property of the item, second-pass effects that correlate with condition would also be attributed to peer-context sensitivity. The striking heterogeneity (Mistral +0.64 toward peers; Gemma -0.33 and Llama -0.24 away; Qwen ~0) is exactly the pattern that could arise from judge-specific order effects. The generator-side claims (Conclusions 1, 2, 4) are not threatened: they rely on blind ratings only, are consistent across all 12 generator-dataset cells, survive self-rating exclusion, and are corroborated by a fixed RoBERTa-MNLI classifier. But Conclusion 3, the paper's most novel and distinctive claim, is only as strong as the assumption that pass order is inert.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an experimental protocol for measuring open-ended LLM conformity that separates generator-side revision effects from evaluator-side reactions to visible peer context. Four open-weight generators are run on three benchmarks with a branched design that includes no-peer re-answering, content-only prompts, and peer-presented prompts; all generated answers receive blind ratings from a four-model judge panel, and peer-presented answers are additionally rated with the peer block visible. The authors fit a hierarchical ordinal model to estimate condition shifts and judge-specific peer-context sensitivity, and they corroborate the generator-side ordering with a fixed RoBERTa-MNLI classifier and other automated metrics. They report four conclusions: flip rates are insufficient for open-ended conformity; all-wrong peers produce the lowest-quality revisions in all 12 generator-dataset cells; evaluators exhibit heterogeneous, judge-specific sensitivity to visible peer context; and anchor calibration must be audited explicitly.","tokens_in":22288,"tokens_out":7868,"duration_ms":94981,"significance":"If the claims hold, the paper makes a useful methodological contribution: it demonstrates that graded, evaluator-sensitive measurement matters for open-ended conformity and that same-answer blind-versus-informed contrasts can, in principle, isolate evaluator-side effects. The generator-side results are strong: the all-wrong ordering is consistent across raw means, a hierarchical model, self-rating exclusions, and a non-generative fixed classifier. The paper also ships with substantial validation: reproducibility seeds, anchor audits, scale-sensitivity refits, and a detailed computational appendix. The most novel and distinctive claim — that evaluators shift identically-worded ratings in response to visible peer context, with direction varying by judge — is plausible but currently rests on an identification assumption that is not demonstrated.","major_comments":[{"comment":"The evaluator-side loading γ_j is identified from blind and informed ratings of the same answer, but the paper reports no randomization or counterbalancing of the two rating passes; the appendix shows the blind prompt first and the informed prompt second. If informed ratings always follow blind ratings, any systematic within-judge change across the passes (fatigue, scale drift, learning, or condition-correlated carryover) is absorbed into γ_j. The paper's concession in Section 3.3 that a generic informed-pass shift is not separately identified does not resolve the issue, because the mixed condition (d_i=0) could in principle estimate a condition-independent shift. I request either evidence that pass order was randomized or counterbalanced, a robustness model that includes an informed-pass intercept estimated from the mixed cell, or an analysis demonstrating that the γ_j estimates in Table 4 are invariant to pass order.","section":"Section 2.3, Eq. (7); Appendix B.1-B.2"},{"comment":"The text states that a 'reference-free intermediate fit also changed the direction of the estimated asymmetry despite satisfactory sampling diagnostics,' but no such fit is reported in the paper or appendix. This claim is load-bearing for Conclusion 4 because it is the only direct evidence that a poorly calibrated anchor scale can reverse a substantive conclusion. Please either document the fit with its point estimates, credible intervals, and convergence diagnostics, or remove the claim and rely solely on the anchor-recognition audit in Table 5.","section":"Section 3.4"}],"minor_comments":[{"comment":"The 9.0% statistic for all-wrong revisions that a binary flip indicator would score as unchanged is not operationalized; please specify how the flip indicator is defined for free-text answers and how the 9.0% value is computed.","section":"Section 3.1"},{"comment":"Krippendorff's α = 0.235 is reported as the raw blind agreement, but the paper does not report pairwise judge agreement or a justification for why the hierarchical model adequately accounts for such low raw agreement; a short clarification would help.","section":"Section 2.2"},{"comment":"The notation '⊮' for the indicator of the informed pass is nonstandard and could be confused with a negation symbol; please use a standard indicator notation such as '\\mathbb{1}'.","section":"Section 2.3, Eq. (7)"},{"comment":"The GPT-4o and GPT-5.4-mini rows come from separate five-judge refits, not from the same posterior as the four open-weight rows; the footnote says this, but the presentation would be clearer if those rows were in a separate table or clearly labeled as non-comparable posterior quantities.","section":"Table 4"},{"comment":"The count of 147,000 generated answers includes reproducibility executions (seeds 43 and 44) and anchor responses, while the primary hierarchical analysis uses only seed 42; please make this distinction explicit in the main text when referencing corpus sizes.","section":"Appendix A.2"},{"comment":"The illustrative scores 4/5, 2/5, and 3/5 in Figure 1 are not labeled as hypothetical ratings; adding a caption note that these are schematic values would avoid confusion with actual data.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The generator-side claims are well supported and should be publishable after a revision addressing the pass-order confound for Conclusion 3 and the undocumented reference-free fit in Section 3.4. The pass-order issue is the key risk: the paper's most distinctive contribution is exactly the claim that evaluators react to visible peer context, and that claim is currently identified by an assumption that blind and informed passes are exchangeable. If the authors can provide a counterbalanced evaluation or a model with an informed-pass intercept, I would view the paper favorably. The reliance on the authors' own prior work for the prompt templates is acceptable but should be scoped explicitly in the related-work discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a serious piece of work on a real measurement problem, and the generator-side findings are robust. The most novel claim—that judges shift ratings when peer context is visible—is real but rests on a blind-then-informed order that the paper doesn't counterbalance, so treat the specific gamma_j values with caution.\n\nWhat's new: most conformity work counts answer flips under verifiable labels. This paper sets up a branched generation design with no-peer, content-only, and peer-presented arms from the same Round 1 answer, and pairs blind and informed ratings of the identical answer. That decomposition—re-answering, content exposure, peer-presentation residual, evaluator-side sensitivity—is a genuine contribution. The all-wrong-peer result is consistent across all 12 generator-dataset cells, survives self-rating exclusion, and is corroborated by a fixed RoBERTa-MNLI classifier. That's the right way to validate an LLM-judge result. The anchor audit is also well done: it shows that reference-free scoring misreads terse anchors often enough to flip a sign, and they refit with reference-augmented anchors and report the scale sensitivity. Credit where due.\n\nThe soft spot is Conclusion 3. Eq. (7) identifies gamma_j from blind and informed ratings of the same answer, but Appendix B.1-B.2 shows the blind prompt first and the informed prompt second, with no randomization or counterbalancing reported anywhere. If judges change between passes—fatigue, scale drift, learning—any second-pass effect that correlates with the signed peer context d_i is absorbed into gamma_j. The heterogeneity they find (Mistral +0.64, Gemma -0.33, Llama -0.24, Qwen ~0) is exactly the pattern a judge-specific order effect could produce. The paper explicitly concedes a condition-independent leniency shift isn't identified, but the problem is broader than that. This doesn't touch Conclusions 1, 2, and 4, which rely on blind ratings only and are corroborated independently. But the paper's most distinctive claim is only as strong as the assumption that pass order is inert. That needs to be fixed—either by counterbalancing or by demonstrating order effects are absent—before I'd rely on the evaluator-side numbers.\n\nThe citation pattern looks fine. The self-cited template reuse is transparent and doesn't drive the estimates. Data/code release is claimed; I'd check the repo but no reason to doubt it.\n\nWho this is for: anyone building multi-agent systems or using LLM judges. It deserves a serious referee. The pass-order issue is a revise-and-resubmit scale problem, not a desk-reject.","headline":"Solid generator-side result on open-ended conformity; the evaluator-sensitivity claim is well-designed but confounded by an uncounterbalanced blind-then-informed order.","tokens_in":22829,"tokens_out":3427,"would_cite":true,"duration_ms":35777,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In open-ended tasks, showing peer discussion to an LLM judge can change the rating of an identical answer, and wrong peers consistently produce the lowest-quality revisions.","keywords":["open-ended conformity","LLM-as-judge","peer context sensitivity","latent quality measurement","anchor calibration","multi-agent revision","hierarchical ordinal model","blind-informed evaluation"],"falsifier":"Randomize the order of the blind and informed rating passes, or use separate judges for each pass, and re-estimate $\\gamma_j$; if the estimated evaluator-side sensitivity changes materially with order, the claim that judges react to visible peer context is confounded by pass order, while stability would support the claim.","tokens_in":21765,"feed_emoji":"⚖️","tokens_out":5138,"duration_ms":51865,"temperature":0.7,"pith_summary":"Open-ended LLM conformity is usually measured by whether a model flips its answer after seeing peers. This paper argues that for free-form answers that is the wrong yardstick, because quality is graded and the evaluator can be influenced by the same peer context that shaped the answer. The authors build a branched experiment in which each answer is rated blind and then rated again with the peer block visible, and they separate ordinary re-answering, content exposure, and peer presentation. Across four open-weight generators and three benchmarks, all-wrong peer input yields the lowest-quality revisions in every generator-dataset cell, and the same generated answer receives different ratings from several judges when the peer block is shown. The paper concludes that open-ended conformity is a joint generation-and-measurement problem, not a discrete flip phenomenon.","feed_headline":"Same answer, different score when peers are visible","feed_subtitle":"A blind-versus-informed rating design shows evaluators react to peer context, not just generators.","key_machinery":"The load-bearing object is the paired blind-informed evaluation contrast, formalized as the judge-specific loading $\\gamma_j$ in a hierarchical ordinal model. Each generated answer is rated once with the peer block hidden and once with it visible; because the answer is identical, the difference isolates the evaluator side of peer-context sensitivity. On the generator side, the branched Round 2 design defines three estimands — ordinary re-answering $\\Delta_{sp}$, content exposure $\\Delta_{\\mathrm{content},k}$, and the peer-presentation residual $\\Delta_{pp,c}$ — whose sum gives the total peer-arm shift, separating what the candidate text does from what the attributed social packaging does.","core_discovery":"On the paper's own terms, the central discovery is that evaluators are part of the conformity experiment: peer context can change the rating of an identical answer, and this evaluator-side sensitivity is heterogeneous across judges. One judge shifts ratings toward the peer-endorsed position, two shift away, one is approximately neutral, and two frontier API judges also show credibly negative directional sensitivity. On the generator side, the total quality shift decomposes into a small ordinary re-answering component, a content-exposure component, and a bundled peer-presentation residual; the residual is negative in every tested cell, and all-wrong peers produce the lowest latent-quality revisions in all twelve generator-dataset cells. Because blind and informed ratings hold the answer text fixed, the paper attributes the rating contrast to the evaluator's reaction to visible peer context rather than to changed generator output. The paper further shows that the anchor items fixing the latent scale can be misordered by judges unless scored with the gold reference visible.","pith_inferences":["The same paired blind-informed protocol transfers directly to other treatment-correlated judge context, such as retrieved passages in retrieval-augmented generation or rubric exemplars, where context is assumed informative but may bias the rating.","The reported $\\gamma_j$ is conditional on an antisymmetric coding in which mixed context is zero; a condition-independent shift in leniency between the blind and informed passes is not separately identified, so future designs should add an informed-pass intercept to capture it.","The 54% Round 1 reproduction rate under greedy decoding means the decomposition corpora had to be self-contained; this suggests that non-deterministic batching, not sampling temperature, can break reproducibility, which is worth checking in other greedy pipelines."],"forward_implications":["Flip-based conformity measures will understate harm in open-ended settings because a revision can keep its nominal answer while losing quality.","Systems that forward peer answers in multi-agent loops should expect the social packaging itself to lower rated quality, independent of content.","Evaluator panels should report blind-versus-context differences per judge rather than assuming judges add only information about quality.","Latent-scale conclusions need an explicit anchor-recognition audit, since reference-free anchors are misordered on roughly a quarter of questions."],"supporting_citations":[{"why":"Documents discrete sycophantic answer shifts that the paper argues are insufficient for open-ended tasks.","marker":"Ranaldi and Pucci, 2023"},{"why":"Provides the flip-based conformity operationalization the paper contrasts with graded quality measurement.","marker":"Weng et al., 2025"},{"why":"Supplies the harmful-versus-beneficial revision baseline and the 500-question benchmark samples.","marker":"Qu et al., 2026"},{"why":"Provides the graded-response measurement logic underlying the hierarchical ordinal latent-quality model.","marker":"Samejima, 1969"},{"why":"TruthfulQA is one of the three benchmark datasets used for generation and evaluation.","marker":"Lin et al., 2022"},{"why":"MMLU-Pro is one of the three benchmark datasets used for generation and evaluation.","marker":"Wang et al., 2024"},{"why":"ARC-Challenge is one of the three benchmark datasets used for generation and evaluation.","marker":"Clark et al., 2018"},{"why":"Documents judgment biases in LLM judges that motivate the evaluator-side controls.","marker":"Chen et al., 2024"},{"why":"Documents position bias in LLM-as-a-judge, supporting the claim that judges are part of the measurement pipeline.","marker":"Shi et al., 2025"}],"fun_headline_variants":["Evaluators are in the experiment: peer context shifts ratings","Peer context moves evaluator scores, not just generator outputs","Wrong peers lower revision quality; judges aren't neutral","Blind vs informed ratings diverge: judges react to peers","Conformity isn't just answer flips: evaluators matter too"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The peer-context sensitivity of judges rests on comparing blind and informed ratings of the same answers, but the paper does not report randomizing or counterbalancing which rating pass comes first; if judge behavior drifts from the first to the second pass, that drift is absorbed into the estimated sensitivity.","fun_headline_variants_meta":{"raw":{"variants":["Evaluators are in the experiment: peer context shifts ratings","Peer context moves evaluator scores, not just generator outputs","Wrong peers lower revision quality; judges aren't neutral","Blind vs informed ratings diverge: judges react to peers","Conformity isn't just answer flips: evaluators matter too"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1335,"prompt_tokens":943,"completion_tokens":392,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":308}},"tokens_in":559,"tokens_out":392,"duration_ms":4240,"temperature":1.0,"reasoning_tokens":308,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:37:28.958627+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Randomize the order of the blind and informed rating passes, or use separate judges for each pass, and re-estimate $\\gamma_j$; if the estimated evaluator-side sensitivity changes materially with order, the claim that judges react to visible peer context is confounded by pass order, while stability would support the claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the flip-based conformity operationalization the paper contrasts with graded quality measurement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the graded-response measurement logic underlying the hierarchical ordinal latent-quality model."}],"review_version":1}