{"id":"583e4e8c-b9f9-40d2-892d-8d0ce07b31a2","arxiv_id":"2410.02736","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM-as-a-Judge systems exhibit significant biases in specific tasks despite strong overall performance, as measured by the new CALM quantification framework.","lead":"This paper identifies 12 potential biases in using LLMs as judges for evaluating AI outputs and introduces the CALM automated framework to quantify them via principle-guided input modifications. Smart generalists should read it because LLM judges are now common in benchmarks and training, so hidden biases could distort which models get selected or improved.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Automated principle-guided modifications may confound bias isolation without external validation against human judgments","rationale":"The reader's weakest assumption directly identifies the methodological risk. Full-text access does not remove it; the claim's strength is still conditional on unverified isolation. No other internal inconsistency was evident from the abstract and described method.","tokens_in":1644,"tokens_out":294,"duration_ms":17622,"concrete_test":"Take 50 examples from one task; manually apply each of the 12 principle-guided modifications exactly as described in the methods; collect human ratings of bias presence on a 5-point scale for both original and modified versions; compute correlation between human delta scores and CALM automated deltas. If mean correlation < 0.6 or if >30% of modifications affect non-target biases, the isolation assumption fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim rests on CALM's ability to quantify 12 distinct biases via automated modifications. For the empirical results to support 'room for improvement,' each modification must affect only its target bias dimension while leaving others unchanged. The paper does not appear to include controls (e.g., ablation on modification prompts, cross-bias correlation checks, or human validation of isolated effects) that would rule out interactions or model-induced artifacts. If modifications inadvertently alter multiple biases simultaneously, the reported per-bias scores become uninterpretable and the headline conclusion weakens.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper identifies 12 key biases in LLM-as-a-Judge, proposes the CALM automated bias quantification framework that uses principle-guided modifications to measure each bias, reports experiments across multiple popular LLMs showing that advanced models exhibit commendable overall performance yet retain significant biases on specific tasks, and concludes there remains room for improvement in reliability while offering suggestions for cautious application.","tokens_in":1727,"tokens_out":324,"duration_ms":17106,"significance":"If the modifications in CALM can be shown to isolate individual biases without confounding, the work would be significant for the many benchmarks and training pipelines that rely on LLM judges, by supplying a systematic diagnostic that could guide mitigation and increase trust in automated evaluation.","major_comments":[{"comment":"The central claim that CALM quantifies 12 distinct biases rests on the assumption that each automated principle-guided modification affects only its target bias dimension. No ablation on modification prompts, cross-bias correlation analysis, or human validation of isolated effects is described, so interactions or model-induced artifacts cannot be ruled out; this directly undermines interpretability of the per-bias scores and the headline conclusion that biases persist in specific tasks.","section":"CALM framework and experimental results"}],"minor_comments":[{"comment":"The abstract states results cover 'multiple popular language models' but provides no model names, sizes, or prompting details; these should be listed explicitly in the experimental setup section.","section":"Abstract and Experiments"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comment on the isolation of biases within the CALM framework. We respond to the major comment below.","responses":[{"response":"We agree that empirical confirmation of isolated effects is essential for the interpretability of the per-bias scores. Each modification in CALM is constructed from explicit, bias-specific principles that alter only the targeted dimension (e.g., swapping option order for positional bias while holding content fixed). This principle-guided design aims to minimize confounding by construction. Nevertheless, the current manuscript does not include ablations on the modification prompts, cross-bias correlation matrices, or human validation of the isolated effects. To address this directly, we will add (i) an ablation study varying prompt phrasing for a subset of biases, (ii) pairwise correlation analysis across all 12 bias scores, and (iii) a small-scale human study verifying that the modifications produce the intended isolated changes. These additions will appear in the revised manuscript and appendix.","revision_made":"yes","referee_comment":"The central claim that CALM quantifies 12 distinct biases rests on the assumption that each automated principle-guided modification affects only its target bias dimension. No ablation on modification prompts, cross-bias correlation analysis, or human validation of isolated effects is described, so interactions or model-induced artifacts cannot be ruled out; this directly undermines interpretability of the per-bias scores and the headline conclusion that biases persist in specific tasks."}],"tokens_in":1220,"tokens_out":317,"duration_ms":52666,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core contribution is a concrete list of 12 biases that can appear when LLMs judge other model outputs, paired with CALM, a framework that tries to measure them by making automated, principle-guided changes to the evaluation prompts. They apply this across several popular models and report that even strong models still show measurable biases on certain tasks, which leads to their claim that reliability still needs work. They close with notes on how the biases show up and some usage guidelines.","headline":"The paper lists 12 biases in LLM judges and offers CALM as an automated quantifier, but the method's ability to isolate each bias cleanly is not yet demonstrated.","tokens_in":2213,"tokens_out":173,"would_cite":false,"duration_ms":16042,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"Cost.FunctionalEquation","rs_theorem":null,"paper_passage":"We identify 12 key potential biases and propose a new automated bias quantification framework—CALM—which systematically quantifies and analyzes each type of bias in LLM-as-a-Judge by using automated and principle-guided modification."},{"relation":"unclear","rs_module":"PhiForcing","rs_theorem":null,"paper_passage":"Empirical results suggest that there remains room for improvement in the reliability of LLM-as-a-Judge."}],"headline":"Paper on LLM-as-a-Judge biases uses empirical perturbation metrics with no relation to RS cost J, φ, or recognition ledger","alignment":"orthogonal","rationale":"The paper's core machinery (CALM framework with principle-guided answer modifications, robustness/consistency rates, and 12 bias types) operates entirely in empirical ML evaluation space. It never references J-cost, golden-ratio fixed points, 8-tick periodicity, or ledger conservation. No theorems or structures from RS modules (Cost.FunctionalEquation, PhiForcing, LedgerCanonicality, etc.) are invoked or paralleled.","tokens_in":289490,"confidence":"high","tokens_out":280,"duration_ms":28885,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"lean_confirmation":{"model":"grok-4.3","status":"out_of_scope","citations":[],"rationale":"Paper is empirical (LLM bias experiments); falls under out_of_scope as it cannot be Lean-proved. No theorem in shape-of-logic applies.","tokens_in":289266,"confidence":"moderate","tokens_out":121,"duration_ms":22034,"inferential_bridge":"The paper's claims are empirical (bias measurements, robustness rates) and cannot be Lean-proved; no load-bearing math premise exists.","load_bearing_premise":"The paper's central result rests on empirical quantification of 12 biases in LLM-as-a-Judge via automated modifications, not a mathematical/structural claim.","cache_read_input_tokens":245824,"cache_creation_input_tokens":0},"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLM-as-a-Judge systems carry 12 measurable biases that automated tests can isolate and that persist in specific tasks.","keywords":["LLM-as-a-Judge","bias quantification","evaluation reliability","automated framework","language models","prejudice detection","CALM"],"falsifier":"Repeating the CALM measurements on the same model and inputs but obtaining substantially different bias scores when a different set of guiding principles is used would indicate that the isolation procedure does not reliably separate the biases.","tokens_in":2551,"feed_emoji":"⚖️","tokens_out":414,"duration_ms":20661,"temperature":0.7,"pith_summary":"The paper sets out to establish that LLM-as-a-Judge, already used for benchmarks and training rewards, is undermined by 12 distinct biases that reduce its reliability. It introduces the CALM framework, which applies automated principle-guided modifications to inputs in order to quantify each bias separately across popular language models. Experiments show that while overall performance is strong, certain tasks still display significant biases, implying that the method requires further refinement before it can be trusted without reservation. A sympathetic reader would care because biased judges can distort evaluation scores and training signals throughout AI development pipelines.","feed_headline":"Tests reveal 12 biases in LLM judges","feed_subtitle":"Automated measurements show that even advanced models retain significant biases in specific evaluation tasks.","key_machinery":"The CALM framework, which isolates and measures each of the 12 biases by applying automated principle-guided modifications to evaluation inputs.","core_discovery":"The paper claims that its CALM framework systematically quantifies 12 potential biases in LLM-as-a-Judge through automated and principle-guided input modifications, with empirical results across multiple models indicating that significant biases persist in certain specific tasks even when overall performance remains commendable.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLM judges retain 12 biases","CALM measures 12 biases in LLM judges","Biases persist in specific LLM tasks","Framework quantifies LLM judge biases","Advanced models show LLM evaluation biases"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That automated principle-guided modifications can cleanly isolate each bias without introducing new confounding effects or missing interactions between biases.","fun_headline_variants_meta":{"raw":{"variants":["LLM judges retain 12 biases","CALM measures 12 biases in LLM judges","Biases persist in specific LLM tasks","Framework quantifies LLM judge biases","Advanced models show LLM evaluation biases"]},"model":"grok-4.3","cost_usd":0.006559,"raw_usage":{"total_tokens":2944,"prompt_tokens":587,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":65590500,"prompt_tokens_details":{"text_tokens":587,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2298,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":587,"tokens_out":59,"duration_ms":29313,"temperature":1.0,"reasoning_tokens":2298,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-15T19:56:39.492700+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Repeating the CALM measurements on the same model and inputs but obtaining substantially different bias scores when a different set of guiding principles is used would indicate that the isolation procedure does not reliably separate the biases.","supporting_citations":[],"review_version":1}