{"id":"501185e0-beb8-4362-9903-5210a47f293a","arxiv_id":"2503.04756","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A simulation suggests LLM judges can favor models fine-tuned on their own outputs, implying that private data-curator evaluations carry conflict-of-interest and annotator-bias risks.","lead":"This paper argues that private AI evaluation companies, which both supply training data and rank models, face conflicts of interest and annotator bias that can skew leaderboard results. It supports the argument with a small experiment in which two AI judges each showed some preference for a model fine-tuned on their own outputs, and it calls for transparency and independent evaluation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.2's 'Self Bias' formula measures cross-judge disagreement, not self-preference; GPT-4o's 50.68% direct preference is a statistical tie, so the claim that both evaluators show self-bias is unsupported.","rationale":"The reader's verdict identifies the LLM-to-human transfer and the untested near-tie as the weakest assumptions. I agree partially, but the most load-bearing issue is more internal: the metric in Eq. (1) does not measure self-bias. The central claim in the abstract is that subjective preferences of private expert annotators will lead to inherent evaluation bias toward models trained with the curators' data. The paper's empirical evidence for this is Section 3.2, which claims both models exhibit self-bias. However, Table 1 shows GPT-4o prefers its own fine-tuned model only 50.68% of the time, a statistical tie. The reported 12.9% 'self-bias' for GPT-4o is an artifact of combining GPT-4o's and Claude's preferences for Model A; it measures inter-judge disagreement, not self-preference. This is an internal inconsistency, not merely a missing significance test. The ELO claim inherits this: the 7-point gap for GPT-4o is negligible and no uncertainty is reported. If the metric is corrected, only Claude exhibits self-preference, so the 'inherent' universality claim fails. The paper could be revised to claim a risk rather than a demonstrated inherent bias, and to make the Claude-only result explicit. That would still be a CONDITIONAL acceptance, but with substantially weakened empirical support. I do not think the qualitative risk argument should be rejected outright, which is why I keep the verdict unchanged rather than moving to REJECT. The concrete test above would settle whether the concern lands: if GPT-4o's direct self-preference is significant, the original conclusion would be partially rescued; if not, the paper must be reframed.","tokens_in":6516,"tokens_out":9618,"duration_ms":75571,"concrete_test":"Recompute the within-evaluator self-preference rates directly from Table 1: GPT-4o prefers its own fine-tuned model in 407/803 = 50.68% of comparisons, Claude in 489/803 = 60.90%. Run a two-sided binomial test for each evaluator (n=803). If GPT-4o's self-preference is not significantly different from 50% (p > 0.05), and if a bootstrap 95% confidence interval for the GPT-4o ELO difference (1003 vs 996) includes zero, then the paper's claim that both evaluators exhibit self-bias and 'significant' ELO differences is unsupported for GPT-4o. Also report the discrepancy between the stated 805 queries and the 803 per-evaluator preferences listed in Table 1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is in the quantification of self-bias in Section 3.2. Equation (1) defines 'Self Bias_A' as (Judge A Prefers M_A − Judge B Prefers M_A) divided by the sum of both judges' preferences for M_A. This is not a measure of a judge's tendency to prefer its own model; it is a measure of cross-judge disagreement about M_A. For GPT-4o, Table 1 shows it prefers Model A (its own fine-tuned model) 407 times and Model B 396 times — a 50.68% rate, essentially a coin flip. The paper reports a 12.9% 'Self Bias (GPT-4)' from Eq. (3), but that number comes from comparing GPT-4o's 407 preferences for Model A against Claude's 314 preferences for Model A; it does not reflect GPT-4o's own self-preference. Consequently, the Section 3.2 statement that 'both models exhibit self-bias' is contradicted by the raw preferences for GPT-4o. Only Claude shows a clear within-judge self-preference (60.90% for its own fine-tuned model). This undermines the abstract's claim that the subjective preferences of private expert annotators will lead to 'inherent' evaluation bias: one of the two simulated evaluators shows no such bias. The ELO analysis inherits the problem: the 1003 vs 996 (7-point) gap under GPT-4o is not reported with any uncertainty and is indistinguishable from noise. Additionally, the paper's own Figure 1 note admits 'We don't see this effect in the current SEAL rankings,' further undercutting the claim of an inherent, general phenomenon. The paper's central empirical support therefore rests on a misdefined metric and a single evaluator's behavior.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that private LLM evaluation by data curation companies poses financial conflict-of-interest and evaluation-bias risks, focusing on the overlap between annotators who create training data and those who evaluate models. It reports a simulation in which GPT-4o and Claude-Sonnet-3.5 act as stand-ins for two companies' expert annotators, two Mistral models are fine-tuned on each evaluator's outputs, and each evaluator then compares the two models on 805 AlpacaEval queries. The authors report self-bias for both evaluators, quantify it with a formula, and simulate ELO ratings that they claim show 'significant differences' based on evaluator preferences. The paper also discusses financial-incentive disclosures and policy remedies such as 'Chinese walls.'","tokens_in":6831,"tokens_out":6149,"duration_ms":52467,"significance":"The topic is timely, and the experimental design directly targets a plausible mechanism for conflict of interest in private evaluation: fine-tuning on a judge's data and then measuring that judge's preference. This directness is a strength, as is the paper's transparency in noting in the Figure 1 caption that the effect is not observed in the current SEAL rankings. However, the quantitative support is undermined by a mis-specified 'self-bias' metric, the absence of any statistical inference, and the lack of validation that LLM self-preference transfers to human annotators. If the analysis were corrected and the claims suitably narrowed, the paper could serve as a useful cautionary note; in its present form, the abstract's and Section 3.2's strong claims are not supported by the data as reported.","major_comments":[{"comment":"Equation (1) is labeled 'Self Bias_A' but it actually measures the cross-judge disagreement about model A, namely the difference between Judge A's and Judge B's preference counts for M_A normalized by their sum. It is not a within-judge measure of a judge's tendency to prefer its own fine-tuned model. For Evaluator Alpha (GPT-4o), Table 1 shows that GPT-4o itself preferred Model A only 407/803 times (50.68%), which is statistically indistinguishable from chance (two-sided binomial test p≈0.72). The reported 12.90% in Eq. (3) arises from comparing GPT-4o's 407 preferences for Model A with Claude's 314 preferences for Model A, not from GPT-4o's own self-preference. Consequently, the Section 3.2 statement that 'The results highlight a clear bias aligned with each evaluator's preferences' is not supported for GPT-4o, and the claim that 'both models exhibit self-bias' is contradicted by the raw within-judge preference rates.","section":"Section 3.2, Eq. (1)"},{"comment":"The ELO simulation reports a 1003 versus 996 rating gap for Evaluator Alpha, a difference of only 7 points, yet the text states that these ELO ratings 'show significant differences based on evaluator preferences' and that the differences 'further highlight the significance of the bias.' No uncertainty, confidence interval, or significance test is provided for either ELO estimate or for the 1003–996 gap. Given that the underlying preference counts are 407 versus 396, this gap is plausibly pure noise. The paper should report bootstrap or Bradley–Terry-model uncertainties and state whether the gap exceeds the simulation's noise floor before claiming significance.","section":"Section 3.3, Table 2"},{"comment":"The experiment uses GPT-4o and Claude-Sonnet-3.5 as simulators for expert human annotators at private data-curation companies, but no human-subject data, prior human-annotation validation, or argument for why LLM self-preference transfers to human annotators is provided. The abstract and Section 2.2 assert that 'the subjective preferences of private expert annotators will lead to inherent evaluation bias,' which is a claim about human behavior. As the current design only demonstrates (at best) a property of LLM evaluators, the external-validity gap is load-bearing. The paper should either include human-subject evidence or explicitly restrict its conclusions to LLM-based evaluators and discuss the conditions under which the mechanism might extend to humans.","section":"Section 3.1, Experimental Protocol and Evaluators"},{"comment":"The caption of Figure 1 contains a direct self-limitation: 'We don't see this effect in the current SEAL rankings.' This admission is in tension with the abstract's claim that private expert annotators' subjective preferences lead to 'inherent evaluation bias.' If the proposed mechanism is 'inherent,' then one would expect some trace of it in the real private leaderboard that motivated the paper; the authors should explain why the SEAL non-effect does not contradict their central claim, or they should weaken the claim from 'inherent' to 'a demonstrated risk under specific conditions.'","section":"Figure 1 and Section 1.2.2"}],"minor_comments":[{"comment":"The label 'GPT-4' in Eq. (3) should be 'GPT-4o' for consistency with the rest of the paper, including Table 1.","section":"Section 3.2, Eqs. (2)-(3)"},{"comment":"The protocol states that 805 queries were used, but Table 1 sums to 803 preference judgments; the discrepancy should be explained (for example, two queries may have produced ties or invalid outputs).","section":"Section 3.1, step 3"},{"comment":"The ELO simulation is described only as using 'the same method as employed by the LMSys leaderboard'; the initialization, update rule, handling of ties, and number of iterations should be specified so that the results are reproducible.","section":"Section 3.3"},{"comment":"The sentence 'specific models look better than some models might appear to perform better than they actually do' is garbled and should be rewritten for clarity.","section":"Section 1.1"},{"comment":"Adding exact binomial confidence intervals for the preference proportions (e.g., 95% CIs for 50.68% and 60.90%) would help readers see which differences are meaningfully distinguishable from chance.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be cited for its claims about private-leaderboard bias, so the statistical flaws in the current version are consequential. The 'self-bias' metric is fundamentally mislabeled, and the near-tie GPT-4o result is treated as a clear effect. These issues are fixable with a reanalysis and more cautious wording, but the revision is substantial enough to warrant another round of review rather than acceptance now."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper makes a valid point worth public discussion: private data curators who both supply training data and evaluate models create a conflict of interest that can skew leaderboards. The qualitative case is well argued and the call for transparency, Chinese walls, and independent evaluation is sensible. The specific setup—two simulated curation companies fine-tuning the same base model on their own data, then having their evaluator LLMs judge both outputs—is a reasonable first attempt to quantify the effect.\n\nBut the empirical core is shaky in exactly the way the stress-test note says. The 'Self Bias' formula in Section 3.2 does not measure a judge's tendency to prefer its own model; it measures cross-judge disagreement about that model. For GPT-4o, the direct within-judge preference is 407 vs 396, a coin flip. Only Claude shows a clear self-preference (489 vs 314). So the claim that 'both models exhibit self-bias' is unsupported. The ELO simulation inherits this: the 7-point gap under GPT-4o is noise; the 81-point gap under Claude is real but driven by one evaluator. There are no confidence intervals, no significance tests, and no released code or data. The LLM-as-human-annotator assumption is also unvalidated, and the paper's own note that the effect is absent from current SEAL rankings undercuts the word 'inherent' in the abstract.\n\nThat said, I don't think this should be dismissed. The qualitative risk is legitimate, and the 81-point ELO gap under Claude is a striking illustration of how much an evaluator's own data can matter. The paper is transparent about its limitations and cites prior work (Panickssery et al.) rather than pretending the phenomenon is new. The writing is clear and the simulation is easy to follow.\n\nWho is this for? Anyone working on evaluation methodology, AI governance, or leaderboard design. It is more useful as a position piece than as a measurement paper. A serious editor could send it to review, but only with the expectation of major revision: fix the metric, add significance testing and error bars, release the artifacts, and soften the broad claims. Right now the central argument is defensible but the evidence is overstated.","headline":"A timely conflict-of-interest argument undermined by a misdefined self-bias metric and a statistical coin flip.","tokens_in":7405,"tokens_out":2452,"would_cite":false,"duration_ms":23378,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Private AI evaluators favor models trained on their own data.","keywords":["private evaluation","LLM evaluation bias","data curator","conflict of interest","ELO rating","annotator bias","self-preference","leaderboard"],"falsifier":"A direct falsifying experiment would recruit two teams of human annotators at a real data-curation firm, have each team write training answers used to fine-tune the same base model, then have each team evaluate both models on fresh prompts while blinded to provenance. If teams do not consistently prefer the model fine-tuned on their own answers—or if GPT-4o's 50.68% versus 49.32% split turns out to be statistically indistinguishable from chance—the paper's central claim loses its empirical support.","tokens_in":6290,"feed_emoji":"⚖️","tokens_out":9847,"duration_ms":80833,"temperature":0.7,"pith_summary":"This paper argues that private data-curation companies that sell training data and also grade models create an inherent evaluation bias toward the models trained on their own data. To demonstrate the mechanism, the authors used GPT-4o and Claude-Sonnet-3.5 as stand-ins for the expert annotators of two rival companies, fine-tuned an open-weight base model on each company's answers to 10,000 prompts, and then had each evaluator judge the two fine-tuned models on 805 fresh queries. Each evaluator preferred the model trained on its own data, with self-bias scores of 12.90% for the GPT-4o side and 10.51% for the Claude side. Translated into the ELO ratings used by public leaderboards, the preferences produce an 81-point gap when Claude evaluates both models. The paper concludes that even a good-faith private evaluator can mislead clients and investors, and that separating evaluation from training-data provision is necessary.","feed_headline":"Private AI evaluators favor models trained on their own data","feed_subtitle":"A simulated showdown between two data curators finds 10-13% self-bias and an 81-point ELO gap.","key_machinery":"The load-bearing mechanism is the dual-role pipeline: the same private company both supplies instruction-following training data and evaluates the resulting models. The paper operationalizes this by fine-tuning the same base model on answers generated by two different language models, using those same two models as judges over a fresh set of queries, and then converting pairwise preference counts into a self-bias score and into ELO ratings with the same method used by the public leaderboard it criticizes. The self-bias score measures how much more often one evaluator prefers a model than the other evaluator does, normalized by total preferences, which gives the structural conflict a concrete numerical size.","core_discovery":"The paper's central claim is that a private data curator's dual role—providing fine-tuning data to model developers and then evaluating those same developers' models—produces a systematic evaluator bias toward models trained on the curator's own data, even when the curator acts in good faith and does not leak test prompts. The experiments support this by showing that Evaluator Alpha (GPT-4o) preferred its own fine-tuned model 407 times versus 396 for the rival model, while Evaluator Beta (Claude) preferred its own model 489 times versus 314 for the rival. The largest effect appears in ELO simulation: Evaluator Beta assigns its own model a 1040 ELO versus 959 for the rival, an 81-point gap that the paper argues is large enough to shift public perception and investment decisions. The authors present this as the mildest form of bias, since there is no intentional tampering, only the subjective preferences of annotators who wear both hats.","pith_inferences":["If LLM self-preference is a reliable model for human annotator bias, the same distortion should appear in any preference-based leaderboard where the human voters overlap with the people who created a model's training data, including open platforms whose chat logs are released.","The paper only tests one base model and two GPT-class evaluators; a natural extension would vary the base model and evaluator families to see whether self-bias scales with model capability, dataset size, or task difficulty.","A concrete, policy-relevant test would be to have a private data curator run its leaderboard alongside an independent blinded evaluation on the same models; if the rankings diverge, disclosure alone cannot fix the bias and governance changes would be needed."],"forward_implications":["A private leaderboard's ranking reflects the preferences of its own annotators as much as model quality, so two private leaderboards can rank the same pair of models in opposite order.","Commercial evaluations should be treated as conflicted unless the curator discloses its training-data relationships, analogous to the Chinese walls used in finance.","An ELO gap of 81 points from self-bias is large enough to shift public perception and investment decisions even when the underlying models are otherwise comparable.","Separating training-data provision from evaluation, or using multiple independent evaluator pools, would reduce but not eliminate this bias.","The near-tie GPT-4o result (50.68% vs 49.32%) shows that the bias magnitude depends on which evaluator model is used, so a single private leaderboard may not even be internally stable across evaluator teams."],"supporting_citations":[{"why":"Supplies the 10,000 instruction prompts used to collect annotator answers from both simulated curators.","marker":"Taori et al., 2023"},{"why":"Supplies the 805 queries used to generate model outputs for pairwise evaluation.","marker":"Li et al., 2023"},{"why":"Supplies the ELO simulation method used to convert preference counts into leaderboard ratings.","marker":"LMSYS-Org, 2024"},{"why":"Provides independent evidence that LLM evaluators recognize and favor their own generations.","marker":"Panickssery et al. (2024)"},{"why":"Documents the real-world private leaderboard and business relationships that motivate the conflict-of-interest concern.","marker":"ScaleAI, 2024"},{"why":"Supplies the base model that is fine-tuned on each curator's data to create the two models being evaluated.","marker":"Mistral-Team, 2024"}],"fun_headline_variants":["Private data curators tilt AI evals toward their own models","Dual-role evaluators skew LLM rankings by 81 ELO","When AI evaluators train the models they judge","Curator self-bias skews AI model evals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the unverified assumption that the self-preference bias shown by GPT-4o and Claude-Sonnet-3.5 is a reliable stand-in for the preferences of the human expert annotators who create training data and grade models, which the paper does not test with human subjects.","fun_headline_variants_meta":{"raw":{"variants":["Private data curators tilt AI evals toward their own models","Dual-role evaluators skew LLM rankings by 81 ELO","When AI evaluators train the models they judge","Curator self-bias skews AI model evals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001403,"raw_usage":{"total_tokens":5664,"prompt_tokens":931,"completion_tokens":4733,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":4663}},"tokens_in":547,"tokens_out":4733,"duration_ms":29111,"temperature":1.0,"reasoning_tokens":4663,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T16:49:53.569469+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct falsifying experiment would recruit two teams of human annotators at a real data-curation firm, have each team write training answers used to fine-tune the same base model, then have each team evaluate both models on fresh prompts while blinded to provenance. If teams do not consistently prefer the model fine-tuned on their own answers—or if GPT-4o's 50.68% versus 49.32% split turns out to be statistically indistinguishable from chance—the paper's central claim loses its empirical support.","supporting_citations":[{"cited_title":"Hashimoto","cited_arxiv_id":null,"evidence_quote":"Supplies the 10,000 instruction prompts used to collect annotator answers from both simulated curators."},{"cited_title":"Lmsys org","cited_arxiv_id":null,"evidence_quote":"Supplies the ELO simulation method used to convert preference counts into leaderboard ratings."},{"cited_title":"Language model leaderboard","cited_arxiv_id":null,"evidence_quote":"Documents the real-world private leaderboard and business relationships that motivate the conflict-of-interest concern."},{"cited_title":"Mistral-7b-v0.3, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the base model that is fine-tuned on each curator's data to create the two models being evaluated."}],"review_version":1}