{"id":"836c0b9c-dd0d-456e-9561-abd37a9da0bf","arxiv_id":"2507.13302","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"In a public LLM comparison arena, showing users that the larger model consumes more energy caused about 46% of users who preferred it to say they would switch to the smaller model.","lead":"This paper built an online arena where people compare answers from two versions of the same AI model, one larger and one smaller, and then see which one uses more energy. When told the larger model uses more energy, around 46% of people said they would switch to the smaller model, suggesting energy information can change AI preferences.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equation (1) assigns all ties to the smaller model without measurement; removing that tie transfer likely drops the pooled small-model win rate below the reported >75% threshold.","rationale":"The reader's REJECT verdict is supported. I focused on Eq. (1) because the headline 'more than 75%' is a direct output of that formula, and the formula contains an assumption not supported by the protocol: all ties are assigned to the smaller model. This is not a minor rounding issue; it is the difference between the central quantitative claim and a weaker statement. I also note the definitional tension around Ec: if Ec is the overall fraction of votes changed, Eq. (2) is arithmetically inconsistent; if it is a conditional back-down rate, the text should say so and the tie transfer remains unjustified. The one-sided protocol and the hypothetical switch question are additional validity threats, but the tie transfer is the most precise, checkable defect. The paper's own limitations section acknowledges the small sample and limited model families but does not flag the tie-transfer assumption, which is the key issue. A raw-data reanalysis would settle the matter. No change to the reader's verdict is needed.","tokens_in":6359,"tokens_out":7112,"duration_ms":82291,"concrete_test":"Recompute the pooled and per-family win rates from the archived GEA database without moving ties: set WS(E) = WS + WL·Ec and WL(E) = WL·(1 − Ec), with Ec as the measured conditional switch rate among users who initially chose the large model, and keep ties as a separate category. Compare the resulting small-model share with the reported >75% in Section 4. If the share falls below 75% for the pooled data, the headline claim is an artifact of the unsupported T transfer in Eq. (1).","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.3 defines Ec as the fraction of responses in which users changed their vote, then gives Eq. (1): WS(E) = WS + T + WL·Ec, and Eq. (2): WL(E) = WL·(1 − Ec). The printed definition is ambiguous: the equations only make sense if Ec is the conditional switch rate among users who initially chose the large model, since the switch count is WL·Ec. If Ec is instead interpreted literally as an overall switch fraction, Eq. (2) is not the correct transformation (it would be WL − Ec), so the equation is either mis-specified or relies on the conditional reading. Under the conditional reading, Eq. (1) still transfers the entire tie rate T to the small model. The protocol in Section 2.2 asks the energy question only to users who selected the larger model; tied users are never asked how energy information would affect them, so T is not measured. Adding all ties to S is an unverified preference assumption, and it is exactly the step that produces the Section 4 headline that users choose the smaller models more than 75% of the time. Recomputing without the T transfer lowers the small model's share by T, which can move the pooled result below 75% for realistic tie rates. The one-sided protocol and the hypothetical 'assuming a loss in quality' question further push the direction, but the tie transfer is the specific arithmetic step on which the quantitative headline rests.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GEA, a two-step LLM comparison arena in which a user first selects the better response between two models of the same family and, if the user chose the larger model, is then asked whether they would switch to the smaller model knowing that it consumes less energy, assuming a loss in quality. The paper defines a switch fraction Ec and derives post-energy win rates via Eqs. (1)–(2), reporting that about 41–52% of responses changed and that smaller models are chosen more than 75% of the time once energy information is accounted for. It concludes that energy awareness should be included in human LLM evaluations and that larger models are not worth their extra cost for most user queries.","tokens_in":6641,"tokens_out":4558,"duration_ms":53144,"significance":"If the headline result were robust, the paper would make a valuable contribution to LLM evaluation by showing that energy information materially changes human preferences and by providing a public arena artifact (code and deployment on Hugging Face) for further study. The underlying question—whether energy-aware users still prefer larger models—is timely and practically important, and the raw switch fractions (41–52%) do indicate that a sizable share of users state willingness to switch. However, the quantitative headline (>75% smaller-model win rate) is not directly measured; it is produced by a deterministic formula that transfers all tie votes to the smaller model even though ties were never queried about energy, and it relies on a one-sided, hypothetical switch question asked only to large-model voters. These issues currently undermine the central claim as stated, though the data gathering approach and raw measurements are a reasonable starting point that could support a more carefully qualified conclusion.","major_comments":[{"comment":"The definition of Ec is ambiguous and the equations are only consistent under one reading. The text says Ec is 'the fraction of the responses for which the users changed their vote', but Eq. (2), WL(E) = WL · (1 − Ec), is correct only if Ec is the conditional fraction of initial large-model voters who switched, because the number of switches is WL · N · Ec. If Ec were instead the unconditional fraction of all responses that changed, the correct transformation would be WL(E) = WL − Ec, not WL · (1 − Ec). Please clarify which quantity Ec denotes and correct the text or the equations accordingly. Additionally, regardless of the intended reading, Eq. (1) transfers the entire tie rate T to the smaller model even though Section 2.2 shows that the energy question is posed only to users who selected the larger model; tied users are never asked how energy information would affect their choice, so T is not measured. This unverified assumption is precisely what pushes the pooled WS(E) above the 75% threshold reported in Section 4, so the paper must either measure tie responses, remove the T transfer, or explicitly flag it as a modeling assumption with sensitivity analysis.","section":"Section 2.3, Eqs. (1)–(2)"},{"comment":"The protocol is one-sided: only users who initially chose the larger model are offered the switch to the smaller model, while users who chose the smaller model are never asked whether energy information could make them switch to the larger one. Combined with the leading wording of the switch question ('would you change your choice assuming a loss in quality?'), the measured responses cannot support the summary statement that 'for most user interactions, the extra cost and energy incurred by the more complex and top-performing models do not provide an increase in the perceived quality of the responses that justifies their use.' The >75% figure is not an observed preference share; it is a deterministic transformation of a one-sided hypothetical response through Eq. (1). Please report the directly measured conditional switch rates and the win rates computed without the unmeasured tie transfer, and rephrase the conclusions to match what the data actually show.","section":"Section 4, Figures 3–4 and the >75% claim"},{"comment":"No confidence intervals, significance tests, or user-level clustering are reported. With 694 questions, with each MOOC student answering ten questions, and with initial family-level preferences that differ by up to roughly 20 percentage points (e.g., GPT families initially favoring the small model), the claims that initial preferences are 'comparable' (within 2%) and that smaller models win 'more than 75% of the time' need bootstrap or hierarchical estimates to establish that these differences are not sampling noise. The paper should also report the number of unique users, not just the number of questions, since votes from the same user are not independent.","section":"Section 4, statistics and sampling"}],"minor_comments":[{"comment":"There is a typo in the first paragraph: 'becnhmarks' should be 'benchmarks'.","section":"Section 1"},{"comment":"The estimate that 'at least 83% of the questions were made by the MOOC students' is not derived anywhere. If 295 fixed questions account for 42.5% of 694 total questions, then the remaining 399 questions were user-created, which would imply the MOOC students authored all of the remaining questions (assuming only students created questions), not 83%. Please provide the computation or correct the estimate.","section":"Section 4"},{"comment":"The paper states that the arena code is publicly available, but no repository link or URL is given. Please provide the exact location of the code and data to support reproducibility.","section":"Section 3"},{"comment":"The paper assumes that larger models in each family consume more energy than smaller ones, but this assumption is not validated with any measured energy data. A sentence acknowledging that the relative ordering is based on the model-size heuristic would make the presentation more precise.","section":"Section 2.1"},{"comment":"Some entries are formatted inconsistently (e.g., Chung et al. 2025 with many authors is abbreviated while others are not). Please standardize the reference list to the journal's style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is better described as a preliminary study or technical report than a full archival paper. The main quantitative claim is not supported by the protocol, but the raw switch fractions and the public arena infrastructure are worth preserving. A revised version that removes the tie-transfer artifact, adds proper uncertainty quantification, and hedges the conclusions could be publishable; I suggest the editor make the need for these changes explicit in the revision letter."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The raw per-family switch rates—41–52% of users saying they'd move to the smaller model—are new, and they're worth having. But the paper's headline claim that users pick the smaller model 'more than 75% of the time' comes from an arithmetic step that is not supported by the data: Equation (1) simply adds the entire tie rate T to the smaller model's win rate, even though the protocol never asks tied voters how energy information would change their choice. Remove that transfer and the pooled win rate drops below the 75% threshold for any realistic tie rate.\n\nWhat's good: GEA is a sensible extension of the Chatbot Arena / ML.ENERGY Colosseum idea, and the two-step design—quality vote first, energy reminder only for large-model voters—is a clean way to elicit a stated switch. The authors are honest about the small sample, the Spanish-only MOOC population, and the limited model families. The measured 41–52% switch fraction is genuinely interesting as a first estimate of stated willingness to trade quality for energy.\n\nThe problems are concentrated in the analysis. Ec is defined as 'the fraction of responses for which users changed their vote' but the equations only make sense if Ec is the conditional switch rate among voters who initially chose the large model—the text should say that. More importantly, adding all ties to the small model is an unverified preference assumption, and it's exactly the step that produces the >75% headline. The switch question is hypothetical and leading ('assuming a loss in quality'), so the 41–52% is stated intent, not observed behavior. Also, because only large-model voters are asked, the small model's computed win rate can only go up; the design can't detect any 'switch to large' effect, which would matter if users value energy but also care about quality.\n\nThe raw data could support a modest claim—'a substantial fraction of users say they'd switch to the smaller model when told it uses less energy'—but the paper currently overclaims. The limitation section doesn't mention the tie-transfer assumption or the one-sided protocol, which should be flagged.\n\nWho it's for: researchers working on green AI, LLM evaluation methodology, and human preference elicitation. With the win-rate equations fixed and the claims scaled back to the raw switch fractions, this would be a useful preliminary study. As it stands, it needs major revision before I'd trust the quantitative conclusions.","headline":"Useful raw data on energy-aware switching, but the 75% headline rests on an unsupported tie-transfer assumption.","tokens_in":7145,"tokens_out":3315,"would_cite":false,"duration_ms":35057,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Displaying relative energy consumption to LLM evaluators shifts their votes substantially, giving smaller models energy-aware win rates above 75 percent.","keywords":["LLM evaluation","energy awareness","human evaluation","arena-based evaluation","model efficiency","sustainability","user preferences","Generative Energy Arena"],"falsifier":"Run a randomized arena experiment in which the energy prompt is neutral rather than leading, such as 'This answer used less energy; you may change your vote,' and keep ties as ties; if the switch rate drops to near zero or the small-model win rate falls below 50 percent, the reported 75 percent result is an artifact.","tokens_in":6199,"feed_emoji":"⚡","tokens_out":6934,"duration_ms":74183,"temperature":0.7,"pith_summary":"Public arenas ask users to compare two anonymous model responses, but they ignore how much energy each answer cost. This paper builds GEA, an arena that first collects a blind quality vote and then tells users which response consumed less energy and asks whether they would switch if it meant a possible quality loss. Across 694 questions, 41–52 percent of users changed their vote after seeing the relative energy information, and the authors' energy-adjusted win-rate calculation puts the smaller models ahead more than 75 percent of the time. The paper argues that for most everyday queries the larger, more energy-hungry models do not add perceived value that justifies their cost, and that energy information belongs inside LLM human evaluation.","feed_headline":"Energy-aware users pick smaller LLMs over 75% of the time","feed_subtitle":"A new arena shows 41–52% of voters switch to the efficient model once relative energy use is revealed.","key_machinery":"The carrying mechanism is the two-stage GEA protocol plus the reweighting identity. Stage one is a blind quality vote; stage two asks only users who preferred the more energy-hungry model whether they would back down on seeing that the other answer consumed less energy. The fraction who switch, called $E_c$, is inserted into $W_S(E) = W_S + T + W_L E_c$ and $W_L(E) = W_L(1 - E_c)$, which moves all ties and every back-down from the large model to the small one. This makes relative energy information the tested intervention and gives the paper a single number to measure its effect on rankings.","core_discovery":"GEA compares pairs of models from the same family that differ mainly in scale, so relative energy use is simple and trustworthy. Users vote on which answer is better before any energy information appears; only when they prefer the larger model are they asked whether they would change their choice knowing the other response consumes less energy, assuming a loss in quality. With a switch fraction of roughly 46 percent, the energy-aware win rates computed from $W_S(E) = W_S + T + W_L E_c$ and $W_L(E) = W_L(1 - E_c)$ reverse the initial ranking: smaller models win more than 75 percent of the time, compared with a difference of under 2 percent before energy was shown. The paper concludes that energy awareness is a critical factor in human evaluation and that larger models are only worth their extra cost for specific questions.","pith_inferences":["The switch question's phrasing, 'assuming a loss in quality,' may overstate real willingness to change because it invites agreement; actual adoption of energy-aware routing could be lower than 75 percent.","The MOOC student sample and Spanish-language questions mean the effect may not transfer to expert users, high-stakes tasks, or other languages; a replication with diverse users would test that.","If the result holds at scale, it points to a practical routing policy: automatically send routine queries to small models and reserve large models for requests where users demonstrably value the extra quality.","Counting every tie as a small-model win in Eq. (1) is a modeling choice; a conservative analysis that keeps ties unresolved would show a smaller but still non-negligible shift."],"forward_implications":["Arena-style human evaluations that omit energy information can systematically overstate the effective quality of large models relative to user preferences.","For common conversational and generative tasks, same-family smaller models are sufficient, so providers can serve many queries with lower energy and cost without losing perceived quality.","Energy-aware win rates, not raw quality votes, should be the reported metric when the goal is to reflect what users would actually choose.","LLM evaluation and development should treat energy as a design dimension: improving efficiency is a way to improve the user-facing product, not just an environmental side goal."],"supporting_citations":[{"why":"Introduces the open pairwise human-preference arena method whose ranking design GEA adapts to include energy information.","marker":"Chiang et al. (2024)"},{"why":"Provides the earlier energy-aware model comparison arena and the inference-energy benchmarking context that GEA builds on.","marker":"Chung et al. (2025)"},{"why":"Shows that LLM speed and energy vary with hardware and configuration, supporting GEA's choice of same-family model pairs for relative comparisons.","marker":"Conde et al. (2024)"},{"why":"Surveys LLM evaluation approaches and their limitations, motivating the human-arena evaluation method that GEA extends.","marker":"Chang et al. (2024)"},{"why":"Documents the large environmental impact of creating and running language models, justifying energy as an evaluation dimension.","marker":"Morrison et al. (2025)"},{"why":"Shows that LLM judges favor their own outputs and can diverge from human preference, motivating arena-based human evaluation.","marker":"Zheng et al. (2024)"}],"fun_headline_variants":["When energy costs are shown, users flip to smaller LLMs","Revealing energy use makes 75% of votes favor small models","Energy-aware LLM arena: users switch to efficient models","User polls: small LLMs win 3 out of 4 once energy is shown","GEA shows energy info flips LLM rankings to small models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a user's 'yes' to the energy question reflects a genuine preference change, and that all ties would shift to the smaller model in the energy-aware tally.","fun_headline_variants_meta":{"raw":{"variants":["When energy costs are shown, users flip to smaller LLMs","Revealing energy use makes 75% of votes favor small models","Energy-aware LLM arena: users switch to efficient models","User polls: small LLMs win 3 out of 4 once energy is shown","GEA shows energy info flips LLM rankings to small models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000241,"raw_usage":{"total_tokens":1550,"prompt_tokens":1000,"completion_tokens":550,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":467}},"tokens_in":616,"tokens_out":550,"duration_ms":6087,"temperature":1.0,"reasoning_tokens":467,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:26:26.117956+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a randomized arena experiment in which the energy prompt is neutral rather than leading, such as 'This answer used less energy; you may change your vote,' and keep ties as ties; if the switch rate drops to near zero or the small-model win rate falls below 50 percent, the reported 75 percent result is an artifact.","supporting_citations":[],"review_version":1}