{"id":"8be43be0-fc45-456c-a006-8028576b9ba5","arxiv_id":"2608.07762","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"AI judges score some models higher once the model's identity is revealed, especially on political questions, and a blockchain commit-reveal protocol is proposed to enforce blind evaluation.","lead":"This paper tests whether AI judges score answers differently when they know which company made the answering AI, and finds the biggest effects on politically sensitive questions. It also proposes a blockchain-based system to make blind AI evaluation auditable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post-hoc subcategory selection and unadjusted multiple comparisons leave the identity-bias effects indistinguishable from noise.","rationale":"The reader's weakest assumption is stylometric leakage in the anonymous condition. That is a real concern, but it is not the most load-bearing issue: if stylometry leaks identity, the paired difference between anonymous and transparent conditions is still causally attributable to explicit identity disclosure, because the response content is held constant. The more decisive vulnerability is that the reported significant gaps may be false positives from an unstated number of comparisons combined with post-hoc cell selection. Section 6.1 presents the GLM, Qwen, and Llama effects as strong evidence of content-dependent bias, yet Section 4 shows the subcategories were chosen after seeing significance. The paper's own Discussion limits the claim and calls for larger per-cell samples to support multiple-comparison connections. Without access to the full score matrix, the existence of any identity-aware bias—let alone its content-dependence—is not verifiable. The blockchain protocol is a reasonable architectural idea, but it does not rescue the empirical premise. Consequently, the reader's REJECT verdict is appropriate; this analysis does not change it.","tokens_in":8939,"tokens_out":6403,"duration_ms":59502,"concrete_test":"Require the authors to release the full paired score matrix (every judge, author, subcategory, merged category, n, mean gap, and t-statistic) and to recompute all p-values with a Benjamini-Hochberg FDR correction or a permutation-based family-wise error control across the complete set of comparisons. Check specifically whether GLM 5.1 +7.00, Qwen3 32B +5.20, and Llama 3.3 70B +1.56 remain significant after correction; if they do not, the central claim fails. A stronger version is to pre-register a replication with a fixed primary comparison and confirm the effect only in that pre-specified cell.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 6.1 rests on a handful of nominally significant paired t-tests (GLM 5.1 +7.00, p=0.0249; Qwen3 32B +5.20, p=0.0426; Llama 3.3 70B +1.56, p=0.0033). The design implicitly runs a full matrix of 7 verifiers × 3 primary models × subcategories/merged categories, but the total number of comparisons is never stated and no multiple-comparison correction is applied. Section 4 confirms post-hoc selection: 'these politically sensitive questions were subcategorised and analysed separately. We report findings from both the subcategory level, where we saw statistical significance.' Selecting cells because they reach significance and then reporting those p-values is not a confirmatory test. With dozens of comparisons, p-values of 0.03–0.05 are expected under the null. The claimed replication across GLM and Qwen may just be the two cells that passed the unadjusted filter. Section 8 itself concedes the need to 'increase per-cell sample sizes to support multiple-comparison connections across full judge × author matrix.' Because the full score matrix and code are not released, one cannot verify how many tests were run or whether any reported effect survives correction. The alternative of selective reporting or noise is not ruled out, so the empirical foundation for content-dependent identity bias is not established.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies identity-aware bias in LLM-as-a-judge evaluation. Seven verifier models (GPT-OSS 120B, Llama 3.3 70B, GLM 5.1, Qwen3 32B, DeepSeek V4 Pro, Mistral Large 3, Sarvam M) score responses from three primary models (DeepSeek V4 Pro, GPT-OSS 120B, Sarvam M) on 58 questions spanning factual, reasoning, political, and preference-based categories, under both anonymous and identity-disclosed conditions. The authors report that identity disclosure has negligible effects on factual questions, marginal effects on stress-reasoning tasks, and significant effects on geopolitically sensitive content, citing specific paired t-test results such as GLM 5.1 awarding Sarvam M +7.00 points (p=0.0249). They then propose a blockchain-based commit-reveal protocol using autonomous economic agents on an Ethereum-compatible ledger, where judges first commit a hash of their score and a secret salt, then reveal the raw score and salt after identity disclosure, creating a tamper-evident audit trail for blind evaluation.","tokens_in":9170,"tokens_out":5004,"duration_ms":47272,"significance":"If the empirical claim were well supported, this would be an important contribution to the growing literature on LLM-as-a-judge bias and to the practical integrity of AI benchmarks, with direct economic and policy implications. The paper also provides a concrete protocol design for verifiable blind evaluation, which is a useful idea even independent of the bias results. However, the statistical analysis is not convincing: the paper does not account for the large number of comparisons, selects subcategories post hoc, and concedes that the anonymous condition may be compromised by stylometric leakage. The protocol, while interesting, is not evaluated beyond a description. The strengths are the breadth of the model set and the direct measurement of identity gaps rather than a fitted derivation, but these do not outweigh the methodological weaknesses.","major_comments":[{"comment":"The central empirical claim rests on selected pairwise t-tests without any multiple-comparison correction. The design implicitly runs a matrix of 7 verifiers × 3 primary models × multiple categories, but the total number of comparisons is never reported. Section 4 states that politically sensitive questions were 'subcategorised and analysed separately' and that findings are reported 'where we saw statistical significance,' which is post-hoc selection. Under the null, dozens of tests would produce p-values around 0.03–0.05, so the reported effects (GLM 5.1 +7.00, p=0.0249; Qwen3 32B +5.20, p=0.0426) are not distinguishable from noise. Section 8 concedes the need to 'increase per-cell sample sizes to support multiple-comparison connections across full judge × author matrix,' but the full score matrix is not released, so the reader cannot verify how many tests were run or whether any effect survives correction.","section":"§6.1 / §4"},{"comment":"The sign convention in Eq. (1) is inconsistent with the reported results. Eq. (1) defines the Identity Gap as IG = S^A_n − S^T_n, where S^A is the anonymous score and S^T is the transparent score. However, every reported positive gap is described as higher scores 'upon identity revelation' (i.e., transparent score higher). For example, GLM 5.1 +7.00 is said to be awarded to Sarvam M when identity is revealed. This is a direct contradiction. Either the equation should be IG = S^T_n − S^A_n, or the text misdescribes the signs. This error affects the interpretation of every reported effect and must be corrected.","section":"§5 / §6.1"},{"comment":"The validity of the anonymous condition is not established. The paper concedes in the Discussion that 'A judge's underlying model could in principle infer authorship from response stylometry even without explicit metadata.' If stylometric leakage occurs, the anonymous condition is not actually blind, and the observed differences between anonymous and transparent scores cannot be attributed to the explicit identity disclosure. The paper needs a concrete test of leakage, such as asking each judge to guess the author of the anonymous responses, or paraphrasing/perturbing responses to remove stylistic signatures before scoring. Without such a control, the central measurement may be confounded.","section":"§8 / §3.1"},{"comment":"The sample sizes per cell are very small and no power analysis is provided. For instance, the merged political category has n=17, and several reported comparisons use n=9, n=10, or n=14. Paired t-tests on such small samples are extremely low-powered and yield fragile p-values. The paper does not justify the choice of 58 questions or the distribution across subcategories, and it is unclear whether these are independent observations or repeated measures from a few prompts. At minimum, the authors should report effect sizes, confidence intervals, and a formal justification for the per-comparison sample sizes.","section":"§4 / §6.1"}],"minor_comments":[{"comment":"The abstract contains incomplete sentences and typos, such as 'p = 0.00' (likely a truncated value) and a missing closing parenthesis after 'p = 0.00'.","section":"Abstract"},{"comment":"The phrase 'n different runs of various modes they tend to change sides' is grammatically unclear and should be rephrased to clarify how the 'conflicting' political answers were selected.","section":"§4"},{"comment":"The term 'control group' for factual questions is misleading; these questions are not a control condition for blinding but a separate treatment category. A control for identity disclosure would need to hold content fixed while varying the disclosed identity of the same underlying response.","section":"§5"},{"comment":"There are several reference formatting issues, including missing URL spaces and inconsistent capitalization, and the 'summary of cited literature' table appears after the bibliography without a clear caption or alignment.","section":"References"}],"recommendation":"reject","confidential_remarks":"The paper's central empirical claim is not supported by the statistical evidence as presented, and the blinding validity issue is a fundamental threat. The proposed protocol is a plausible idea but is not evaluated, and the lack of released code/data prevents any verification. I believe the manuscript would require a substantially redesigned study to meet the standards of a peer-reviewed venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The new bit here is measuring identity-aware bias on politically sensitive content across seven verifier models, and pairing that with a commit-reveal scheme for blind evaluation on a blockchain. Credit where due: they spot a real problem, set up a sensible pilot, and are unusually candid about limitations in Section 8, including stylometry leakage and small per-cell samples.\n\nThe soft spots are where the stress-test lands. The analysis is cherry-picked. They ran a large matrix of paired tests and report only the cells that passed p<0.05, with no multiple-comparison correction. The claimed replication across GLM 5.1 and Qwen3 32B is just the two cells that happened to reach significance; with dozens of comparisons that's expected under the null. Section 8 itself admits per-cell samples are too small for multiple-comparison corrections. The design also admits selecting 'conflicting' political answers, which biases the political category. No code or data are released, so you can't check the full test matrix. The stylometry concern is real, and they acknowledge it, but it undermines the central comparison between anonymous and transparent conditions.\n\nThe blockchain protocol is a reasonable application of a standard commit-reveal primitive. It's not deeply novel, and it's presented as mathematical enforcement without a security analysis or shared implementation. Still, it's a concrete idea that might be useful if developed further.\n\nThe paper is a pilot. It doesn't establish content-dependent identity bias, but it doesn't pretend to be definitive either. The topic matters—benchmark integrity has market consequences. A serious referee could push the authors to do a proper confirmatory analysis with pre-registered hypotheses or at least report all comparisons and correct for them, and to release the data and code.\n\nVerdict: reject for now, but not a waste of time. I'd say yes to peer review: the question is important, and the flaws are fixable in principle.","headline":"Pilot study with an interesting protocol, but the identity-bias claim is statistically under-supported; worth a serious referee, not publication as is.","tokens_in":9692,"tokens_out":2175,"would_cite":false,"duration_ms":21793,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Revealing a model's identity changes how LLM judges score answers, and the paper argues a blockchain commit-reveal protocol can make blind evaluation provable.","keywords":["identity-aware bias","LLM-as-a-judge","benchmark integrity","commit-reveal protocol","autonomous economic agents","blockchain verification","politically sensitive content","paired t-test"],"falsifier":"Run the same paired protocol with strict API-level blinding and, after each anonymous score, ask the judge to name the likely author model; if identification accuracy is above chance in the anonymous condition, the identity gap reflects authorship detection rather than identity disclosure, and the central comparison collapses. Alternatively, a pre-registered replication with larger per-cell samples and multiple-comparison correction would settle whether the headline gaps (for example, the 7.00-point GLM 5.1 shift) survive.","tokens_in":8707,"feed_emoji":"🤖","tokens_out":9640,"duration_ms":74949,"temperature":0.7,"pith_summary":"This paper tries to establish that identity-aware bias in LLM-as-a-judge evaluation is real, measurable, and content-dependent: revealing which model wrote an answer moves scores mainly when the content is culturally or politically loaded or when the author model carries authority, not across the board. It supports this with paired experiments where seven verifier models scored answers from three primary models anonymously and then with identity disclosed, using 58 factual, reasoning, political, and preference questions. The measured gaps are small or absent on factual items but large on sensitive ones, for example a 7.00-point bump from one verifier for a disclosed identity on geopolitical content. If the claim holds, non-blind benchmark evaluations on sensitive topics cannot be treated as quality measures, which is why the paper also proposes a blockchain commit-reveal protocol that binds a judge's score before identity is revealed.","feed_headline":"Revealing a model's identity bends LLM judge scores","feed_subtitle":"Identity disclosure shifted judge scores up to 7 points on sensitive topics; blockchain makes blind scoring provable.","key_machinery":"The carrying object is the Identity Gap, defined as $S^A_n - S^T_n$, the difference between the score a judge gives an answer in anonymous and transparent conditions for the same stored response, with a paired t-test used to test whether the gap is noise. Because each underlying response is fixed in a database, any score movement is attributed to identity revelation rather than response quality. The second mechanism is a two-phase cryptographic commit-reveal cycle: Phase 1 commits a one-way hash of score and salt to an Ethereum-compatible ledger before identity is revealed; Phase 2 reveals identity and the raw score, and the contract verifies the hash. This makes blind evaluation auditable so post hoc claims of blindness can be checked against the record.","core_discovery":"On the paper's own terms, the discovery is that identity-aware bias is not a universal property of LLM peer evaluation but is 'activated by the cultural and political relevance of the content being scored and the authority of the model.' Using paired anonymous versus transparent scoring, the authors report negligible identity gaps on objective factual questions, marginal gaps on stress-reasoning tasks, and statistically significant gaps on politically sensitive and preference-based items: GLM 5.1 awarded Sarvam M responses 7.00 points more when identity was revealed (p = 0.0249), Qwen3 32B awarded 5.20 points more (p = 0.0426), and Llama 3.3 70B awarded GPT OSS 120B responses 1.56 points more on merged political content (p = 0.0033). The paper interprets this as evidence that blindness is a necessity for evaluation integrity, not a procedural nicety, and introduces a two-phase commit-reveal protocol in which each judge lodges a one-way hash of its score and a secret salt before candidate identity is disclosed, then reveals the raw score and salt for on-chain verification.","pith_inferences":["Beyond the paper, the content-dependence pattern suggests the same judges may shift scores on other high-stakes culturally charged domains, such as legal rulings or public-health guidance; the paper's own question set only samples politics and preference.","By extension, the protocol proves when a score was committed, not whether the judge was genuinely blind; pairing it with stylometry checks or API-level isolation would address the paper's own caveat that a judge might infer authorship from prose style.","The same commit-reveal pattern could be applied outside LLM benchmarking, for example to human peer review or content moderation, wherever the evaluator's verdict should be fixed before the author's identity is known."],"forward_implications":["Blind evaluation is a necessity: on geopolitically sensitive content, simply disclosing the author model's identity shifts scores by several points, so non-blind benchmark results in such domains conflate source with quality.","Content sensitivity, not the judge model alone, determines bias: factual items show negligible identity gaps while political and preference items show significant ones, so bias checks need to be domain-specific.","Even objectively verifiable reasoning questions are not immune: a verifier awarded 2.40 extra points on stress-reasoning questions after identity disclosure, showing the effect can touch right-or-wrong content.","The commit-reveal protocol turns 'we ran blind' from an unverifiable assertion into a cryptographic fact: judge commitments exist on-chain before identities are known, so later score changes are detectable.","If deployed with participating providers, the protocol would cut the verification burden on independent researchers and third-party leaderboards by letting anyone audit the blinding sequence."],"supporting_citations":[{"why":"Establishes the LLM-as-a-judge framework and documents judge biases, the baseline phenomenon this paper measures under identity disclosure.","marker":"[1]"},{"why":"Provides a statistical method for measuring self-bias and family bias in LLM judges, the prior result the paper extends to content-dependent identity-aware bias.","marker":"[2]"},{"why":"Gives a parliamentary-voting method for measuring political bias in LLMs, which motivates the politically sensitive question category.","marker":"[10]"},{"why":"Shows LLMs reflect the ideology of their creators, supporting the expectation that identity and cultural relevance affect evaluation.","marker":"[12]"},{"why":"Introduces a blockchain-based decentralized framework for collaborative LLM evaluation, the closest prior infrastructure to the proposed commit-reveal protocol.","marker":"[5]"},{"why":"Supplies zero-knowledge proof machinery for verifying LLM output authenticity, a complementary verifiability tool the paper positions alongside its commit-reveal design.","marker":"[23]"},{"why":"Documents the January 2025 market reaction to unverified DeepSeek benchmark claims, the economic stake motivating verifiable benchmarking.","marker":"[14]"}],"fun_headline_variants":["LLM judges score known models higher on political topics","Blockchain commit-reveal protocol counters LLM judge bias","Identity disclosure shifts LLM judge scores by up to 7 points","Blind LLM evaluation with blockchain: tamper-evident audit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the anonymous condition is actually blind: in the control setting the judge model cannot tell which primary model wrote the answer, so the measured score gap is caused by explicit identity disclosure rather than by stylometric inference or leaked context.","fun_headline_variants_meta":{"raw":{"variants":["LLM judges score known models higher on political topics","Blockchain commit-reveal protocol counters LLM judge bias","Identity disclosure shifts LLM judge scores by up to 7 points","Blind LLM evaluation with blockchain: tamper-evident audit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000386,"raw_usage":{"total_tokens":2150,"prompt_tokens":1163,"completion_tokens":987,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":779,"completion_tokens_details":{"reasoning_tokens":916}},"tokens_in":779,"tokens_out":987,"duration_ms":9079,"temperature":1.0,"reasoning_tokens":916,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:19:03.398759+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same paired protocol with strict API-level blinding and, after each anonymous score, ask the judge to name the likely author model; if identification accuracy is above chance in the anonymous condition, the identity gap reflects authorship detection rather than identity disclosure, and the central comparison collapses. Alternatively, a pre-registered replication with larger per-cell samples and multiple-comparison correction would settle whether the headline gaps (for example, the 7.00-point GLM 5.1 shift) survive.","supporting_citations":[{"cited_title":"In: Advances in Neural Information Processing Systems (NeurIPS 2023), vol","cited_arxiv_id":null,"evidence_quote":"Establishes the LLM-as-a-judge framework and documents judge biases, the baseline phenomenon this paper measures under identity disclosure."},{"cited_title":"npj Artificial Intelligence 2, 7 (2026)","cited_arxiv_id":null,"evidence_quote":"Shows LLMs reflect the ideology of their creators, supporting the expectation that identity and cultural relevance affect evaluation."},{"cited_title":"The Guardian, Jan- uary 27 (2025)","cited_arxiv_id":null,"evidence_quote":"Documents the January 2025 market reaction to unverified DeepSeek benchmark claims, the economic stake motivating verifiable benchmarking."}],"review_version":1}