{"id":"d3dbe297-5c75-4890-a3a0-500f2f27361f","arxiv_id":"2602.23293","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A winrate-only benchmark pushes AI model makers to homogenize and lowers consumer welfare; weighting winrate by answer value incentivizes specialization and improves welfare in the model's regime.","lead":"This paper shows that when users route between AI models with noisy choices, the standard \"winrate\" metric pushes model makers to make similar models and can reduce consumer welfare. It proposes \"weighted winrate,\" which also rewards answer quality, and proves that it incentivizes specialization and improves welfare in the modeled regime.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Weighted winrate is defined on latent per-task values; if those values are noisy/unobservable in real LLM evaluation, the mechanism cannot be implemented as specified and Theorem 4's welfare guarantee need not transfer.","rationale":"The reader's weakest assumption—that weighted winrate relies on a known, comparable value v_{m,t}—is precisely the most load-bearing threat to the central claim. The paper's own conclusion acknowledges the issue but does not model it, and the abstract's statement that weighted winrate 'provably improves... consumer welfare' in LLM evaluation goes beyond what is proven once values are noisy or costly to obtain. I do not find a separate internal flaw that overturns the main theorems under their stated assumptions; Theorem 4's proof is defensible in the exact-value setting, and the empirical section is appropriately framed as illustrative. Since the reader already assigned CONDITIONAL with this concern, no verdict change is warranted, but the condition should be made explicit in the final version: either add a noise-robustness analysis or restrict the abstract's claims to settings where exact per-task values are available.","tokens_in":37298,"tokens_out":9895,"duration_ms":96410,"concrete_test":"On the MT-Bench-101 setup used for Table 2 (V=120, beta=1, 3-model market), let true values v_{m,t} be the benchmark scores and define noisy signals v_hat_{m,t}=v_{m,t}+epsilon_{m,t} with epsilon ~ N(0,sigma^2). Compute the producer's best response to weighted winrate using the noisy signals, then evaluate realized consumer welfare under true BTL probabilities. Sweep sigma from 0 to 5. If the weighted-winrate allocation ceases to strictly dominate the winrate allocation for any sigma below the typical cross-task value gap, the practical force of Theorem 4 depends critically on perfect value knowledge.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that weighted winrate provably improves producer incentives and consumer welfare—depends on Definition 5, which requires a known scalar value v_{m,t} for every model-task response. This value enters twice: once in the choice probability p_m(v_t) and once as the multiplicative reward. The theorems then compare producer best responses under exact knowledge of these values. In the LLM-benchmark setting the paper targets, such values are not directly observable; they would have to be elicited from users or inferred, and the paper itself concedes in Section 8 that 'If user-provided measures of value are sufficiently expensive or noisy... other mechanisms could improve upon weighted winrate.' The concern is therefore not an internal inconsistency but a load-bearing applicability gap: if the values used to compute weighted winrate are noisy, producers best-respond to a different objective, and the strict welfare improvement in Theorem 4 can fail even when all other assumptions hold. The theoretical results are sound conditional on perfect value information, but the abstract's claim about real 'LLM evaluation' overreaches without an explicit noise model.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a marketplace of AI models with discrete tasks, scalar per-task values v_{m,t}, BTL choice-based aggregation, consumer welfare equal to expected selected value, and producers maximizing average winrate. It first shows that BTL and other separable choice functions are non-monotone in value, and that monotone aggregation and guaranteed benefits from entry are mutually exclusive under substitutability/anonymity. It then analyzes two producer actions: model creation and model replacement. For model creation, Theorem 3 states that a sufficiently high-value entrant maximizing winrate equalizes winrate across tasks (homogenization), while Lemma 9 states that the proposed weighted-winrate mechanism induces full specialization on one task, and Theorem 4 claims strictly higher consumer welfare under the Theorem 3 regime. For model replacement with N=2, Lemmas 13–15 and Theorem 5 develop analogous local-derivative results. The empirical section uses MT-Bench-101 to illustrate the predicted specialization patterns. Proofs are in the appendix and code is provided.","tokens_in":37574,"tokens_out":15909,"duration_ms":158344,"significance":"If established, the paper's central mechanism-design point is valuable: standard winrate leaderboards can push producers toward homogenization under noisy BTL routing, while a value-weighted selection metric can improve specialization incentives. The theoretical calculus is mostly explicit, the proofs are organized in appendices, and the release of code is a strength. However, the headline claim is conditional on exact, known per-task response values; on large total value V for entry and N=2/local-derivative regimes for replacement; and on the BTL model. The abstract states the results more universally than the theorems support. The experiments are illustrative rather than a quantitative test. Conditional on these caveats, the contribution is a useful addition to the literature on benchmark incentives and model diversity.","major_comments":[{"comment":"Weighted winrate is defined as p_m(v_t)·v_{m,t}, and all welfare theorems assume the planner and producers know the exact scalar value v_{m,t} for every model-task response. In the LLM-evaluation setting the abstract targets, such values are not directly observable and would need to be elicited or inferred; the paper's own conclusion concedes that 'if user-provided measures of value are sufficiently expensive or noisy... other mechanisms could improve upon weighted winrate,' but no noise or elicitation model is provided. Without such a model, Theorem 4's strict welfare improvement is not robust to perturbing the inputs to Definition 5: producers would best-respond to a different objective, and the comparison can fail even when all other assumptions hold. This is a load-bearing applicability gap for the central claim, not an internal inconsistency. The abstract and the theorem statements","section":"§2.3, Definition 5; §8"},{"comment":"The proof of Theorem 3 in Appendix C establishes only a pairwise improvement. For a task k with v_k < β log E_k, a pigeonhole argument produces a task j whose combined budget with k exceeds β(log E_j + log E_k), and then Lemma 21 shows that the pair's total winrate could be increased by equalizing winrate across tasks j and k. This demonstrates that a non-equalizing allocation is not pairwise optimal, but it does not prove that the global maximizer equalizes winrate across all T tasks, which is exactly what the theorem states and what Theorem 4 invokes. A rigorous proof would need a global optimality condition (e.g., Lagrange equations showing p_t(1−p_t) is constant across tasks) or an explicit convergence argument for repeated pairwise equalization. As written, the load-bearing theorem is not fully established.","section":"§5.2, Theorem 3 / Appendix C proof"},{"comment":"The statement of Lemma 10 says that the welfare suboptimality of weighted winrate is bounded by u(v_{i*})−u(v_{j*}), where i* maximizes p_{N+1}(v_i⊕V) and j* is the task with smallest current utility u(v_{j*}). The proof instead bounds the gap by u(v_2)−u(v_1), where task 2 is the weighted-winrate pick and task 1 is the consumer-welfare-optimal task. There is no step showing that the consumer-optimal task has the smallest current utility, and no use is made of j* in the proof. Under the stated definitions the claimed bound does not follow; the lemma should either be corrected to use the consumer-optimal task in the bound, or removed/relegated to a less prominent claim.","section":"§5.3, Lemma 10"},{"comment":"The abstract and introduction state the results as universal: winrate 'can incentivize model creators to homogenize for both types of model changes, reducing consumer welfare,' and weighted winrate 'provably improves incentives for producers to specialize and increases consumer welfare.' The actual theorems are substantially more conditional: Theorem 3 requires V ≥ 2Tβ·max_j log(Σ_i exp(v_{j,i}/β)); the model-replacement results are for N=2 and for instantaneous derivative-based changes; and Theorem 5 explicitly excludes the non-monotone regime where one player's improvement reduces welfare on both tasks (Observation 5, Table 1). The abstract should carry the main qualifications; otherwise readers will infer a stronger mechanism-design guarantee than is proven.","section":"Abstract / Theorems 3–5"}],"minor_comments":[{"comment":"The experiments choose V=120 and β=1 precisely to 'demonstrate patterns from our theoretical results,' and the text says the values are 'picked so as to demonstrate.' This is illustrative rather than a test of the theory. Please label it as such and, ideally, add robustness plots over V and β or a random subset of tasks.","section":"§7, Table 2 / Figure 5"},{"comment":"In the final chain of inequalities, a weak inequality is written as '>2·T·β·max...' where '≥' is intended; the subsequent conclusion that V ≥ 2T·max_j u(v_j) then carries a minor strength-of-inequality mismatch. This is easy to fix but should be corrected for precision.","section":"Theorem 4 proof"},{"comment":"There is a typo in the heading: 'F uture directions' should be 'Future directions.'","section":"§8"},{"comment":"In Definition 7, the condition '∂f/∂v > 0 in at least one place' is used in the proof of Theorem 1 at a constructed v_1. The phrase 'at least one place' should be clarified to mean 'at the value v_1 used in the construction,' since the proof does not show monotonicity fails at every point.","section":"Definition 7 / Theorem 1"},{"comment":"The paper alternates between 'winrate' and 'total winrate' (Assumption 2 vs. Theorem 3) and between 'averaged across tasks' and 'summed across tasks.' Since additive constants do not change the argmax but do affect statements like Lemma 8, a single convention and a sentence clarifying that the argmax is unchanged would help.","section":"Notation"}],"recommendation":"major_revision","confidential_remarks":"The main reason I did not recommend minor revision is the combination of the unmodeled value-elicitation gap and the gap in the proof of Theorem 3. Both are fixable within scope: the first by explicitly qualifying the claims to the exact-value case or adding a robustness model, the second by supplying a global optimality argument. If the authors address Lemma 10 and the abstract overclaims as well, the paper would be a solid contribution to the benchmark-incentives literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core result is real: under BTL choice, a producer maximizing winrate will spread value evenly when the entrant is strong enough, while weighted winrate pulls effort to a single task, and in the large-capital regime that strictly improves consumer welfare. That's a genuine addition to the benchmark-incentive literature, which mostly works with Tullock functions and convex win probability. The proofs are mostly careful: Theorem 3's pigeonhole argument is sound, and Theorem 4's proof is compressed but the direction is correct. I don't think Lemma 10 is actually broken—the inequality in the proof forces u(task 2) ≥ u(task 1) when task 1 is welfare-optimal, so the stated bound goes through, though a referee will want the statement rewritten for clarity.\n\nSoft spots: the abstract's 'LLM evaluation' phrase oversells. The homogenization result holds only for sufficiently large V, and the replacement section is local derivatives with N=2. The experiments choose V=120 and β=1 to display the predicted pattern rather than test it; that's fine as illustration, not validation. The bigger caveat is one the authors concede: weighted winrate needs a known scalar value per model-task response. With noisy user ratings, the producer's best response changes and Theorem 4's guarantee is not robust. They mention it in Section 8 but don't model it; that should be called out in revision. Code and data are available, which helps reproducibility.\n\nWho should read: anyone designing leaderboards or studying AI market structure gets value. I'd send it to a serious referee. It's not acceptance-ready, but it deserves engagement; the model and results are workable.","headline":"The homogenization result and the weighted-winrate fix are real and new; the paper needs a scope check and a model of noisy values before I'd buy the 'LLM evaluation' framing.","tokens_in":38046,"tokens_out":7383,"would_cite":true,"duration_ms":66500,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A10","91B03"],"pacs":[],"model":"deepseek-v4-flash","headline":"Winrate, the standard LLM benchmark, incentivizes model homogenization and lower consumer welfare; a value-weighted winrate provably reverses this.","keywords":["model marketplace","winrate","weighted winrate","Bradley-Terry-Luce","consumer welfare","model diversity","benchmark incentives","LLM evaluation"],"falsifier":"Construct or observe a setting with two tasks where the entrant's value allocation under weighted winrate is not fully concentrated: for instance, introduce zero-mean noise on v_{m,t} and show that the best response spreads value across tasks, or empirically perturb MT-Bench scores and check whether weighted winrate's welfare advantage over winrate disappears under modest noise.","tokens_in":37195,"feed_emoji":"🤖","tokens_out":4852,"duration_ms":49202,"temperature":0.7,"pith_summary":"The paper asks what a healthy marketplace of AI models looks like when users and routers pick models noisily, and how to get producers to create such a marketplace. It models aggregation with the Bradley-Terry-Luce choice rule and shows that winrate—the probability that a model is selected—gives producers an incentive to spread their strengths evenly across tasks, a homogenization that hurts consumers. It then proposes weighted winrate, which rewards a model by its response value times selection probability, and proves that under the same conditions this metric encourages full specialization and strictly raises consumer welfare. Along the way it shows that even without strategic producers, noisy aggregation makes consumer welfare non-monotone: improving a model or adding a better one can reduce total utility. The results are illustrated on a 13-task LLM benchmark.","feed_headline":"Winrate pushes AI makers to homogenize, weighted winrate to specialize","feed_subtitle":"Rewarding selection probability alone hurts consumers under noisy choice; weighting by value restores specialization.","key_machinery":"The argument is carried by the Bradley-Terry-Luce (BTL) choice model, the workhorse probabilistic choice rule in which selection probability is exp(v/β) over the sum of exp(v/β). Winrate is simply a model's selection probability; weighted winrate is that probability multiplied by the response's value v. The key structural results are Theorem 3, which uses a log-sum-exp threshold to show that winrate best responses homogenize, and Lemma 9, which shows the weighted-winrate best response concentrates all value on a single task. The proof machinery also uses the observation (Lemma 1) that the change in welfare from adding a model is a 'swish'-shaped function of value differences.","core_discovery":"The central claim is that the choice of evaluation metric shapes the entire model ecosystem through producer incentives. With Bradley-Terry-Luce aggregation, a producer who maximizes average winrate in response to a large enough value budget equalizes winrate across tasks, producing homogeneous models; yet consumer welfare is maximized by concentrating value on a single task. Weighted winrate—defined as the probability a model is selected multiplied by the value of its response—aligns the two: the producer's best response is to specialize fully, and the resulting consumer welfare strictly beats the winrate equilibrium whenever the winrate-equalizing condition of Theorem 3 holds. For model re","pith_inferences":["Editorial inference: The same homogenizing pressure should apply to router-based leaderboards like the Max system that aggregate many models: any ranking metric based on selection frequency will push entrants toward the average, not toward complementary strengths.","Editorial inference: A natural testable extension is to deploy weighted winrate on a public leaderboard using coarse human value ratings (1-5 stars) and measure whether the distribution of model strengths across tasks becomes more heterogeneous over time.","Editorial inference: The theory suggests a calibration condition for the mechanism: the value signal must be accurate enough that the log-sum-exp ordering of tasks is preserved; when value noise is comparable to the gap between tasks, the welfare gain may vanish or reverse.","Editorial inference: The non-monotonicity results connect to AI oversight concerns: model homogenization may be driven not only by training data but by the evaluation metric itself, implying that oversight interventions could operate at benchmark design."],"forward_implications":["If leaderboards keep using winrate, stronger entrants will keep homogenizing; the paper shows this is not an accident of any one benchmark but a property of any separable choice function.","Adopting weighted winrate in public leaderboards would give producers a direct incentive to specialize, increasing the diversity of available models and the utility of users who aggregate over them.","The non-monotonicity results imply that 'improvements' to individual models can backfire under noisy routing, so evaluation of model updates must account for the selection context, not just the model in isolation.","Weighted winrate has a known gap: it leaves producers indifferent between returning value-0 and abstaining, and in the worst case this can still hurt consumers; this is the paper's stated open direction.","Under bounded task maxima (as on real benchmarks), the specialization results generalize to a greedy allocation rule (Lemma 11)."],"fun_headline_variants":["Winrate homogenizes AI models, new metric spurs specialization","Why winrate kills model diversity—and the fix that works","Weighted winrate: an incentive fix for AI model markets","Metric choice shapes AI ecosystems: winrate vs weighted","To boost consumer utility, reward value not just selection"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The welfare-improving claims for weighted winrate require a trusted scalar value v_{m,t} for each model-task pair; the paper assumes such values are available and correct, and the conclusion acknowledges that if user-provided measures are too expensive or noisy, weighted winrate may not deliver its guarantees.","fun_headline_variants_meta":{"raw":{"variants":["Winrate homogenizes AI models, new metric spurs specialization","Why winrate kills model diversity—and the fix that works","Weighted winrate: an incentive fix for AI model markets","Metric choice shapes AI ecosystems: winrate vs weighted","To boost consumer utility, reward value not just selection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1130,"prompt_tokens":687,"completion_tokens":443,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":431,"completion_tokens_details":{"reasoning_tokens":375}},"tokens_in":431,"tokens_out":443,"duration_ms":4646,"temperature":1.0,"reasoning_tokens":375,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T20:24:54.908172+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct or observe a setting with two tasks where the entrant's value allocation under weighted winrate is not fully concentrated: for instance, introduce zero-mean noise on v_{m,t} and show that the best response spreads value across tasks, or empirically perturb MT-Bench scores and check whether weighted winrate's welfare advantage over winrate disappears under modest noise.","supporting_citations":[],"review_version":1}