{"id":"8f0637c4-5713-4ce6-9a7a-b16302d97d9e","arxiv_id":"2607.04590","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Limited attention makes pairwise preference labels non-identifiable for reward, can reverse Bradley-Terry rankings, and bounds learning by attended information rather than raw label count.","lead":"Human preference labels used to train AI are not clean measurements of reward: limited attention can reverse true rankings and hide large quality gaps. This matters because RLHF pipelines treat near-even votes as indifference when they may instead mean the evaluator failed to see the difference.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the channel-adequacy caveat already flagged.","rationale":"The reader’s strongest_claim accurately summarizes what the paper proves under Definition 1, and the weakest_assumption correctly locates the only material contingency (channel adequacy for real feedback). No derivation gap, misstated theorem, or unacknowledged empirical overclaim was found that would move the verdict. CONDITIONAL remains appropriate: accept-shaped theory with complete proofs and supportive public-data case studies, contingent on treating Definition 1 as a working measurement model and on releasing the reproduction script. Confidence stays high because every formal claim is checkable from the text and appendix.","tokens_in":20607,"tokens_out":497,"duration_ms":23139,"concrete_test":"Independently recompute Example 1’s population FOCs at ε=0.4 (and at the Arena-scale max |ℓ|≈1.8 with the same β ratios) by maximizing L(r;ℓ) in Eq. (10) with equal weights; confirm the B≻A ranking and report the exact r† values. Separately re-run the Arena bootstrap cyclic-energy test after treating ties as half-wins and at the 50/200 vote thresholds already claimed; if either the ranking reversal or the p<0.01 cyclic excess fails, the concrete illustrations weaken.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The theoretical core (cycle obstruction Prop. 1, local projection Prop. 2, ranking reversal Example 1, non-identification Prop. 3, Fisher/KL/Fano bounds Props. 4–6 and Thm. 1, cyclic KL price Prop. 7) follows rigorously from the reduced-form channel in Definition 1; appendix proofs are complete and elementary. The single contingency that could prevent the claims from transferring to real RLHF is whether ℓ_z = η_z + β_z Δ*_z adequately describes alignment-relevant human labels (Section 2.2, Remark 1). That is already the reader’s weakest_assumption; no sharper internal inconsistency, hidden assumption in the derivations, or empirical contradiction was found. Arena cyclic energy rejects every scalar score but is only consistent with (not unique to) heterogeneous attention; the perceptual study measures channel ingredients directly but is not preferential. These are scope limits the paper itself states, not load-bearing breaks.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that standard Bradley–Terry reward learning from pairwise human comparisons systematically misreads attention-limited labels. Motivated by Shannon rational inattention, it posits a reduced-form channel ℓ_z = η_z + β_z Δ*_z in which observed log odds mix deliberative reward gaps with query-specific attention and defaults. From this channel it derives: a cycle criterion for scalar representability (Prop. 1); a local projection characterization of the population BT fit that can reverse deliberative rankings even when every pairwise majority is correct (Prop. 2, Example 1); observational non-identification of reward, attention, and defaults from passive labels (Prop. 3); entropy–information separation and rank-one Fisher matrices (Props. 4–5); KL/Fano sample-complexity bounds governed by total attended information rather than label count (Prop. 6, Thm. 1); and a one-eighth-law KL price of cyclic energy discarded by any scalar reward (Prop. 7). Two case studies check signatures: excess cyclic energy in Chatbot Arena votes, and response-time/gaze information about gap magnitude absent from labels in a perceptual task.","tokens_in":20861,"tokens_out":1662,"duration_ms":27696,"significance":"If the reduced-form channel is a useful description of alignment-relevant feedback, the results are load-bearing for RLHF practice: they separate hard-because-close from hard-because-hidden comparisons, show that BT can recover the wrong ranking under heterogeneous attention, and reframe sample complexity as attended information rather than annotation volume. The theoretical core is cleanly executed: appendix proofs for Props. 1–7, Lemma 1, and Theorem 1 are elementary and complete (cycle telescoping, IFT for the local projection, outer-product Fisher matrices, Pinsker/Fano, Hodge decomposition). The one-eighth law is a parameter-free second-order prediction that matches Arena misfit to within about 10% even outside the small-ε regime. The paper is explicit about scope (Remark 1, Section 5 caveats) and states falsifiable geometric and process signatures rather than fitting the channel to produce the reversal. That combination of identification theory, information bounds, and priced cyclic loss is a genuine contribution to the foundations of preference-based alignment.","major_comments":[{"comment":"Definition 1 / Remark 1: The entire identification, projection, and information analysis is conditional on the reduced-form marginal channel ℓ_z = η_z + β_z Δ*_z. Lemma 1 only yields a conditional logit given the realized evidence state; Remark 1 correctly notes that the marginal label law equals this form only in a special environment. For alignment-relevant queries, deliberative Δ* typically requires integrating over unresolved uncertainty, so the gap between conditional and marginal can be material. The manuscript should either (i) state sufficient conditions under which the reduced form is a good approximation to the marginal choice probability for the intended RLHF setting, or (ii) more sharply bound which of Props. 1–7 survive under qualitatively different attention models (e.g., feature-based or threshold attention). Without that, the transfer claim in the abstract and conclusion","section":"Section 2.2, Definition 1, Remark 1"},{"comment":"Section 5.1 / Abstract: The Arena analysis cleanly rejects every scalar score (bootstrap p = 0.008; LR p = 0.007) and prices the misfit with Prop. 7 to ratio 0.89. The paper itself notes that aggregation over prompts and annotators can generate cyclic energy without per-query attention heterogeneity. That alternative is not quantified (e.g., by prompt-stratified cycle energy or annotator-level Hodge residuals). Given that the abstract presents the result as exhibiting the predicted attention signature, the manuscript should either add a simple stratification check or rephrase the abstract/conclusion so that the finding is stated as rejection of scalar representability consistent with, but not diagnostic of, heterogeneous attention—matching the more careful language already in the last paragraph of §5.1.","section":"Section 5.1, Abstract"}],"minor_comments":[{"comment":"Figure 1 caption and Example 1: the figure uses R*_A(2), R*_B(1), R*_C(0) while the text of Example 1 uses the same values; the figure also shows β_BC = 5ε producing ℓ_BC = 5ε, which is consistent, but the cycle sum annotation 'ε + 5ε − ε ≠ 0' would be clearer if the oriented edges were labeled with signs matching the chosen orientation in the incidence matrix.","section":"Figure 1, Example 1"},{"comment":"Proposition 2 is stated for ∥ℓ∥_∞ ≤ ε with remainder O(ε³). Example 1 notes that the exact FOCs preserve the reversal at ε = 0.4. A short remark on the range of ε where the ranking of r† remains reversed (or a one-line numerical check) would help readers judge relevance outside the asymptotic regime.","section":"Proposition 2, Example 1"},{"comment":"Section 5.2: the mutual-information numbers (0.33 bits for sign, 0.0001 for magnitude from labels; 0.035 bits from RT) are useful; please state the estimator (e.g., histogram / KSG) and the permutation procedure more explicitly so the 0.035 figure is reproducible from the public aDDM data alone.","section":"Section 5.2"},{"comment":"Related work: the connection to combinatorial Hodge ranking (Jiang et al.) is well credited; a brief pointer to recent preference-game / von Neumann-winner work (already cited as [11,21,32]) on how cyclic energy from attention differs operationally from cyclic energy from genuine intransitive preferences would help practitioners choose between scalar and non-scalar pipelines.","section":"Related work"},{"comment":"Typos / notation: 'deliberatively' is used consistently; in Eq. (10) the objective L(r;ℓ) uses σ(ℓ_e) as the true probability, which is fine, but a one-word reminder that this is the population BT likelihood under the human channel would reduce confusion with the usual correctly-specified BT likelihood.","section":"Equation (10)"}],"recommendation":"minor_revision","confidential_remarks":"The theoretical core is stronger than the empirical packaging. I would not reject on the channel-adequacy caveat alone—the paper is unusually careful about reduced-form status—but I would push the authors to tighten the abstract's Arena claim so that referees in the alignment community do not over-read cyclic energy as a unique fingerprint of attention. Fit for a theory-leaning ML/AI journal is good; for a pure empirical HCI venue it would be thin. No citation or novelty concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple. Under a reduced-form channel ℓ = η + βΔ*, passive pairwise labels cannot separate reward from attention and defaults, and heterogeneous attention can make the population Bradley–Terry fit reverse the deliberative ranking even when every pairwise majority is correct. Recovery is governed by total attended information scaling like β^{2}, not by raw label count.\n\nWhat is new is the packaging and the consequences for reward learning. Rational inattention, Hodge decomposition, and Fisher/Fano bounds are known; the paper turns them into an explicit measurement model for RLHF-style comparisons and derives the cycle obstruction, the local projection target, the three-item reversal, the rank-one Fisher matrix, and the one-eighth KL price of the cyclic component. Appendix proofs are complete and elementary. The Arena analysis rejects every scalar score with a bootstrap null and prices the misfit with the one-eighth law to within about ten percent; the perceptual study shows response times carry magnitude information that labels lack. Both use public data and state their caveats.\n\nThe soft spot is the channel itself. Definition 1 is motivated by a Shannon first-order condition but is exact only in a special conditional environment. If real alignment labels are generated differently, the transfer is not automatic. Arena cyclic energy is consistent with heterogeneous attention but not unique to it (aggregation can produce cycles too). The perceptual task is not preferential. These are scope limits the paper already flags, not internal breaks. Free parameters (β_z, η_z) are explicit; the theory is not circular.\n\nThis is for people who care about the measurement model behind preference learning and scalable oversight. The math is checkable; the empirics are supportive rather than decisive. I would send it to peer review. Engage with it if you work on reward models or oversight protocols; the non-identification and attended-information bounds are the parts worth carrying forward.","headline":"Clean theory of attention-limited pairwise labels: non-identification, ranking reversal under heterogeneous β, and β^{2} sample complexity, with solid proofs and two supportive public-data case studies.","tokens_in":21454,"tokens_out":508,"would_cite":true,"duration_ms":5105,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Human pairwise preferences under limited attention cannot generally reveal true reward rankings, and learning is limited by attended information rather than label count.","keywords":["reward learning","Bradley-Terry","rational inattention","RLHF","pairwise comparisons","attention-limited measurement","cyclic preference data","scalable oversight"],"falsifier":"Find a large human preference graph in which the cyclic energy of the log-odds field is statistically indistinguishable from sampling noise under a fitted Bradley-Terry model, or a controlled comparison task in which response times and gaze carry no additional information about gap magnitude once the binary labels are known.","tokens_in":21451,"feed_emoji":"👁️","tokens_out":616,"duration_ms":5024,"temperature":0.7,"pith_summary":"Standard reward learning from human comparisons treats each vote as a noisy readout of a true value difference. This paper models each vote instead as the output of a low-capacity attention channel: the observed log-odds equal a default tendency plus an attention multiplier times the deliberative reward gap. Under that channel, passive labels cannot separate reward from attention or defaults, and when attention varies across pairs the usual Bradley-Terry fit can reverse the true ranking even though every pairwise majority is correct. What determines sample complexity is not the number of labels but the total attended information they carry, which scales with the square of the attention multipliers. Case studies on Chatbot Arena votes and on a perceptual comparison task show the predicted cyclic signature that no scalar reward can represent, and that response times and gaze carry gap information that the binary labels themselves do not. The practical claim is that weak preference signals may mean evaluation difficulty rather than genuine indifference.","feed_headline":"Limited attention can reverse true rankings from pairwise votes","feed_subtitle":"Passive comparison data cannot separate reward from attention; learning scales with attended information, not label count.","key_machinery":"The attention-scaled comparison channel: observed log-odds equal a default term plus an attention multiplier times the deliberative reward gap. This reduced form produces cycle obstructions to scalar representability, a weighted projection limit for Bradley-Terry fitting, rank-one local information matrices, and sample-complexity lower bounds controlled by attended information.","core_discovery":"Passive pairwise comparison data cannot generally distinguish reward, attention, and default tendencies; heterogeneous attention can make the population Bradley-Terry solution reverse the deliberative ranking even when every observed majority points the right way; and reward recovery is governed by total attended information, which scales as the square of the attention multipliers, not by the raw number of labels.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Heterogeneous attention can reverse true rankings from pairwise votes","Passive comparisons cannot separate reward from attention limits","Bradley-Terry recovers misleading ranks under limited attention","Reward recovery scales with attended info not raw label count","Pairwise votes confound reward attention and default tendencies"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The claim rests on the premise that real human comparison labels are well described by a reduced-form channel that multiplies the true reward gap by a query-specific attention weight and adds a default term.","fun_headline_variants_meta":{"raw":{"variants":["Heterogeneous attention can reverse true rankings from pairwise votes","Passive comparisons cannot separate reward from attention limits","Bradley-Terry recovers misleading ranks under limited attention","Reward recovery scales with attended info not raw label count","Pairwise votes confound reward attention and default tendencies"]},"model":"grok-4.5","effort":"low","cost_usd":0.005848,"raw_usage":{"total_tokens":1568,"prompt_tokens":794,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":58480000,"prompt_tokens_details":{"text_tokens":794,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":700,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":794,"tokens_out":74,"duration_ms":5089,"temperature":1.0,"reasoning_tokens":700,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T16:46:29.106854+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Find a large human preference graph in which the cyclic energy of the log-odds field is statistically indistinguishable from sampling noise under a fitted Bradley-Terry model, or a controlled comparison task in which response times and gaze carry no additional information about gap magnitude once the binary labels are known.","supporting_citations":[],"review_version":1}