{"id":"c2d222f1-a969-4e8a-8b1e-a89c4452cab6","arxiv_id":"2607.11871","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LLM-as-judge scoring biases concentrate in low-dimensional, type-specific activation subspaces that support bidirectional causal steering and cross-domain failure prediction.","lead":"LLM judges shift scores on surface cues like prestige or bandwagon notes; this paper shows those biases live as low-dimensional directions in the model's hidden states. Steering those directions can both induce and cancel unfair scores, and a simple projection predicts failures on unseen benchmarks.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged construct-validity premise; residual risk is ordinary for this class of empirical MI work.","rationale":"The central claim is a representation-level account of LLM-as-judge scoring bias with three interlocking parts: a depth-sharpening, type-specific bias subspace recovered by independent estimator families; bidirectional causal control along that subspace far above random and swap controls; and a linear projection that anticipates degradation on unseen benchmarks. The paper supplies multi-judge, multi-bias, multi-benchmark evidence plus the right controls (random direction, type-swap, matched-budget text, CV defense, TOST human equivalence). The single softest joint is exactly the operationalization the reader already flagged: surface-cue + score-shift case-control substrate. Because the authors already test construct validity and out-of-substrate transfer, and because no stronger technical inconsistency appears (e.g., no contradiction between geometry and causality, no unacknowledged ceiling artifact, no estimator-family disagreement that would collapse the subspace claim), the appropriate stress-test outcome is non-finding. Verdict remains CONDITIONAL pending public code/data; confidence and correctness_risk stay as the reader set them. The proposed concrete_test is a direct falsifier of the case-control premise without requiring new data collection.","tokens_in":52489,"tokens_out":768,"duration_ms":6071,"concrete_test":"Re-estimate the three retained bias directions (Geometric Median, PCA, Classifier) on the full D_neg without the δ_s=2 / 90th-percentile Mahalanobis core filter, then re-run the calibrated attack/defense protocol of Table 19 and the linear-projection outcome predictor of Section 4.5 on the same held-out trio. If within-type W1 falls below 2× the matched-norm random baseline or cross-domain AUC drops below ~0.70, the case-control substrate is load-bearing in a way that weakens the fairness claim; otherwise the geometry generalizes beyond the strong-shift tail.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption correctly isolates the load-bearing premise: bias is operationalized as a semantics-preserving surface-cue transformation that produces a score shift of at least δ_s=2 (direction fitting) or δ_o=1 (outcome prediction), with directions estimated only on the effective-bias / biased-core case-control subset (Sections 3.1–3.3, Appendix F). If those cues are rationally quality-relevant to the judge, or if the threshold selects a non-representative tail, the recovered geometry and its causal/predictive utility would not underwrite the broader fairness claim. The paper already partially closes this with TOST human equivalence tests on bit-identical-body and prose-rewrite types (Appendix C/C.1), matched-budget text-attack comparisons (Appendix J.3), random-direction and bias-type-swap controls (J.1–J.2), 5-fold CV defense retaining ≥80% of in-sample W1-reduction (K.1), and cross-domain prediction on three held-out benchmarks (4.5, M.7). No stronger internal inconsistency or unaddressed technical flaw is evident in the three-part claim (geometry, bidirectional steering, operational prediction). Residual risks (white-box coverage limited to three mid-scale judges; artifacts promised but not yet hashed) are ordinary for large-scale LLM empirical work and do not overturn the stated claims.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that LLM-as-judge scoring bias is a representation-level phenomenon: baseline activations form a tight manifold while biased inputs are displaced along a low-dimensional, type-specific subspace that sharpens with depth. Across seven judges, seven bias types, and nine benchmarks, the authors (i) document strong behavioral asymmetry (negative surface cues penalize more than positive ones reward), (ii) recover the bias subspace with directional and discriminative estimators that agree within family and partially across architectures, (iii) show bidirectional activation steering along that subspace (attack on clean inputs, defense on biased ones) far above matched-norm random and bias-type-swap controls, and (iv) train a linear projection onto the same features that predicts score degradation on three entirely held-out benchmarks (AUC ~0.82 vs ~0.63 text baseline). The contribution is a unified geometric–causal–operational account rather than a new mitigation recipe.","tokens_in":52828,"tokens_out":1303,"duration_ms":12362,"significance":"If the three-part claim holds, the paper supplies a mechanistic account of a widely used evaluation primitive that currently sits inside RLHF and benchmark pipelines. The combination of multi-estimator geometry, bidirectional causal control with random and swap ablations, and a simple transferable linear predictor is stronger than typical input–output bias catalogs and is operationally useful for white-box judges. Strengths include the multi-judge behavioral replication, the TOST human equivalence checks on both prose-rewrite and bit-identical-body perturbations, the matched-budget text-attack comparison, 5-fold CV defense retaining ≥80% of in-sample W1 reduction, and cross-domain prediction sensitivity analysis. These make the result a solid contribution to mechanistic interpretability of evaluators and to fairness auditing of LLM judges.","major_comments":[{"comment":"Sections 3.1–3.3 and Appendix F: direction estimation is restricted to the effective-bias / biased-core case-control subset (score shift ≥ δ_s = 2 and Mahalanobis 90th percentile). The paper correctly treats null-shift samples as null observations and evaluates causal/predictive claims outside the fitting substrate, but the fairness claim still depends on the premise that the surface cues are not rationally quality-relevant. Appendix C/C.1 TOST tests and the matched-budget text comparison (J.3) substantially mitigate this; the manuscript should state more explicitly in the main text (not only the appendix) that the geometry is conditioned on the established surface-cue operationalization of the LLM-as-judge literature, and report a brief sensitivity of recovered directions to δ_s ∈ {1,2,3} so readers can see how much the subspace depends on the strong-shift tail.","section":"§3.1–3.3, App. F, C/C.1"},{"comment":"Section 4.4 and Appendices J.1–J.2: bidirectional steering is presented as interventional sufficiency, not unique natural pathway—this scoping is appropriate. The random-direction control is strong (order-of-magnitude gap); the bias-type-swap control correctly shows an intermediate effect (shared subspace + type-specific component). For the central causal claim, the paper should either (a) add a short path-patching / mediation sketch on one bias type and one layer, or (b) keep the current scoping but move the “we do not claim unique pathway” language into the main-text causal paragraph rather than only the discussion, so readers do not over-read the attack/defense results as full causal identification.","section":"§4.4, App. J.1–J.2"}],"minor_comments":[{"comment":"Table 2 vs Table 4 / Figure 3: the positive-aggregate convention (five score-inflating types vs all seven) is explained in Appendix I.1 but should be flagged once in the main-text caption of Table 2 to avoid confusion with the full-pool +0.07 figure mentioned in §4.2.","section":"Table 2, §4.2"},{"comment":"Figure 1 MDS and Figure 2 Δh MDS: axis scales differ across panels; a shared color legend and a note that MDS is used for visualization only (not for the estimators) would help non-MI readers.","section":"Fig. 1–2"},{"comment":"Appendix H limitations: white-box coverage is limited to three mid-scale judges; the cross-architecture cosine band [0.47, 0.62] is useful—consider promoting one sentence of that partial-transfer result into the main-text geometry section.","section":"App. H, I.3"},{"comment":"Reproducibility: seed, split, and ~1400 A100-hour budget are stated; ensure the promised code release includes the exact α-search hyperparameters (Algorithm 1) and the nested question-ID split files so the cross-domain AUC can be regenerated.","section":"App. G"},{"comment":"Minor notation: ℳ_base is introduced as a manifold but used as an empirical cluster; a one-line clarification that it is the empirical support of H_base^(l) would avoid geometric overclaim.","section":"§3.2"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a strong empirical MI paper on an important applied setting. The construct-validity premise is the only load-bearing soft spot, and the authors have already done more to close it (TOST, matched-budget text, CV defense, held-out domains) than is typical. I would not require path-patching for acceptance; a clearer main-text scoping sentence and a δ_s sensitivity check would suffice. Fit for a serious ML / interpretability venue is good; novelty relative to SteerFair and LAGER is adequately differentiated in Appendix A.3."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean empirical MI paper that actually delivers the three things it promises: a typed, depth-sharpening bias subspace in judge residual streams; bidirectional activation steering that beats matched-norm random and type-swap controls; and a linear projection that predicts score degradation on three held-out benchmarks (AUC ~0.82 vs text ~0.63).\n\nWhat is new is not the toolkit—steering vectors and linear concept directions are established—but the systematic application to the evaluator role across seven bias types, seven judges, and nine benchmarks, with honest differentiation from SteerFair and LAGER. The behavioral asymmetry (large negative penalties, weak/heterogeneous positives) is replicated across models and prompt configs; geometry is recovered by independent estimator families with high within-family agreement; and the causal loop is closed with attack/defense plus random and swap controls. The operational predictor is the part that transfers, and they correctly anchor the claim to the simple linear projection rather than the overfit GBDT.\n\nSoft spots are real but proportionate. Direction fitting is restricted to the effective-bias / biased-core case-control subset defined by score shift (δ_s=2). That is the load-bearing premise: if those surface cues are rationally quality-relevant, or if the threshold selects a non-representative tail, the fairness claim weakens. They partially address this with TOST human equivalence on bit-identical and prose-rewrite types, matched-budget text attacks, CV defense retaining ≥80% of in-sample W1 reduction, and prediction on benchmarks never used for vector estimation. White-box coverage is only three mid-scale open models; artifacts are promised but not yet hashed. Free parameters (thresholds, α*, layer) are standard for this class of work and do not invent circularity.\n\nMath and citation pattern look solid; related work is fair. This is for people who care about judge reliability inside RLHF and eval pipelines, and for MI readers who want a non-generative application of representation engineering. It deserves a serious referee. I would engage with it and expect it to survive review with ordinary revisions (code release, sensitivity tables).","headline":"Solid three-part MI account of LLM-as-judge bias (geometry, bidirectional steering, cross-domain prediction) that earns referee time; main residual is the surface-cue/score-shift operationalization of bias, which the authors partially close with human TOST and out-of-substrate tests.","tokens_in":53466,"tokens_out":544,"would_cite":true,"duration_ms":8121,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"LLM judge bias lives as a low-dimensional direction in the model’s hidden state, and steering along it both creates and cancels unfair scores.","keywords":["LLM-as-judge","scoring bias","mechanistic interpretability","activation steering","bias subspace","representation engineering","outcome prediction"],"falsifier":"Show that matched-norm random directions or swapped bias-type directions produce score shifts as large as the recovered bias direction under the same validity and rank-preservation constraints, or that a linear projection onto those features fails to beat a strong text baseline on truly held-out domains.","tokens_in":53341,"feed_emoji":"⚖️","tokens_out":657,"duration_ms":5272,"temperature":0.7,"pith_summary":"When large language models are used as automatic judges of answers, their scores move with surface cues that have nothing to do with answer quality—prestige labels, peer consensus notes, length, tone, identity claims. Prior work mostly treated that as black-box input-output noise and tried to fix it with better prompts. This paper argues the same bias has a clear internal geometry. Clean judging inputs sit in a tight activation manifold; biased inputs are displaced along a low-dimensional, type-specific subspace that becomes sharper in deeper layers and is recovered by several independent estimators. Steering the hidden state along that subspace bidirectionally controls scores: adding the direction makes a clean answer look biased, subtracting it restores fair scoring on a biased answer, while matched-norm random directions barely move the score. The same direction features also let a simple linear predictor flag judge failures on three benchmarks never seen during training, beating text-only detectors. If the account holds, bias is no longer only a catalog of prompt tricks but a representation object that can be measured, steered, and predicted.","feed_headline":"Judge bias sits on a steerable direction inside the model","feed_subtitle":"Add it to fake unfair scores, subtract it to restore fair ones; the same features predict failures on new domains.","key_machinery":"The bias direction (or low-dimensional bias subspace) recovered from effective bias samples via directional-change and discriminative-boundary estimators; unit-normalized and used both for activation steering (add or subtract at mid-to-late layers) and for linear outcome prediction.","core_discovery":"LLM-as-judge scoring bias admits a representation-level account: baseline activations form a tight manifold while biased inputs are displaced along a low-dimensional, type-specific subspace that sharpens with depth; that same subspace is an interventional handle that bidirectionally controls scores and supplies features that predict judge degradation on held-out domains.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLM judge bias is a steerable low-dim activation subspace","Steer one bias direction to fake or restore fair judge scores","Biased inputs displace judges along a type-specific subspace","Linear projection on bias features predicts new-domain failures","Same subspace controls scores and flags judge degradation"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"The paper treats score shifts caused by carefully constructed surface-cue edits as bias, and fits the direction only on the strong-shift subset of those edits; if many of those cues are rationally quality-relevant or the strong-shift tail is unrepresentative, the geometry may not underwrite the broader fairness claim.","fun_headline_variants_meta":{"raw":{"variants":["LLM judge bias is a steerable low-dim activation subspace","Steer one bias direction to fake or restore fair judge scores","Biased inputs displace judges along a type-specific subspace","Linear projection on bias features predicts new-domain failures","Same subspace controls scores and flags judge degradation"]},"model":"grok-4.5","effort":"low","cost_usd":0.004856,"raw_usage":{"total_tokens":1392,"prompt_tokens":778,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":48560000,"prompt_tokens_details":{"text_tokens":778,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":534,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":778,"tokens_out":80,"duration_ms":4492,"temperature":1.0,"reasoning_tokens":534,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T02:32:31.489105+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Show that matched-norm random directions or swapped bias-type directions produce score shifts as large as the recovered bias direction under the same validity and rank-preservation constraints, or that a linear projection onto those features fails to beat a strong text baseline on truly held-out domains.","supporting_citations":[],"review_version":1}