{"id":"93466f84-59be-48bf-8c0f-22bd74fdcf4e","arxiv_id":"2607.02672","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Local pairwise comparisons are provably blind to perfectly inseparable priorities and can distort conflicted preferences, but allowing indecision reports speeds up simulated preference learning.","lead":"This paper builds a formal model of “internal pluralism”—one person holding several moral priorities—and shows that standard local pairwise comparison questions can erase global priorities like fairness and hide internal conflict. It then argues, via simulations, that letting people report indecision instead of forcing a choice can make preference learning faster and more faithful.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Rule-first permutation-invariant aggregation (Thm 3.3) is unvalidated; under priority-first or default-background aggregation, perfectly inseparable priorities need not cancel, so non-identifiability may fail.","rationale":"The reader's weakest_assumption correctly identifies the rule-first permutation-invariant aggregation as the load-bearing premise. My independent pass confirms that Theorem 3.3's proof turns entirely on Lemma B.3, which is a statement about permutation-invariance of the rule aggregator. The paper's own text (Section 2.2) acknowledges the order is assumed and defers alternatives. The computational argument (tracking |F| vs. m quantities) is an argument for tractability, not for descriptive adequacy; for a model whose stated purpose is to characterize what human values can be learned from forced comparisons, this is a gap. I considered other potential concerns: the Section 4 simulation's active-learning dependence and the assumption of accurate indecision reporting are explicitly flagged as limitations in Section 5 and are secondary; the interview coding lacking inter-rater reliability weakens but does not invalidate the formal results. The formal model appears internally consistent, so I do not see a fatal flaw. The right response is to treat the paper's core negative result as conditional on the stated aggregation architecture, exactly as the reader did. Hence no change to the verdict.","tokens_in":60128,"tokens_out":14588,"duration_ms":148072,"concrete_test":"Recompute the response distributions in Example 3 (Egalitarianism + Family, two inputs) under priority-first aggregation with the same linear priority weights and a max rule aggregator, or under a single default-background-rule model. If the zero-threshold response distribution R^0_{h_beta;0,0}(q; M_insep) differs from R^0_{h_beta;0,0}(q; M_family) for any beta > 0 (e.g., nonzero KL divergence), then Theorem 3.3's cancellation fails under this alternative order, and the non-identifiability result does not generalize. This check isolates the aggregation-order assumption from all other features of the model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.2 fixes aggregation as rules-first, then priorities, with a permutation-invariant, unanimous rule aggregator. Theorem 3.3's cancellation (Lemma B.3) uses this: for a perfectly inseparable priority, the multiset of projection gaps for y vs. y' equals the multiset for y' vs. y up to permutation, so permutation-invariance forces equal directional evidence, and the priority drops out of kappa. If the individual instead aggregates priorities first, or conditions on a single default/status-quo background rule, the cancellation need not occur. For example, with Egalitarianism (perfectly inseparable) and Family (perfectly separable), priority-first aggregation with a max rule aggregator yields kappa(q) = max_F [omega_egal * Delta_F^egal + omega_fam * Delta_F^fam], which changes when Egalitarianism is removed; the priority leaves a detectable trace. The paper explicitly defers other aggregation orders to future work, yet the 'erased without a trace' result and Corollary 3.5 non-identifiability are stated as properties of local pairwise comparisons. The sole justification for rule-first is computational ('natural'), not behavioral; no evidence shows humans aggregate counterfactual evidence this way. Since the model aims to establish limits of pairwise elicitation for real human priorities, this unvalidated architectural assumption is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a formal model of \"internal pluralism\" in which an individual's preferences over decision rules are represented by multiple weighted priorities, each a complete transitive utility over the rule space. It defines how these rule-level priorities generate local pairwise comparison responses through two aggregation steps: first over background rules within each priority, then over priorities, yielding directional evidence scores and latent states of decisiveness, conflict, and indifference. The contributions are: (i) a characterization of score-based random utility models (S-RUMs), including Bradley-Terry, as exactly the perfectly separable, zero-threshold special case of this model (Theorem 2.3); (ii) results showing that perfectly inseparable priorities—such as certain formalizations of egalitarianism, proportionality, and equal treatment—are erased from local pairwise comparison data and are not identifiable, while general inseparable priorities can be misinterpreted by separable-consistent decoders (Theorem 3.3, Corollaries 3.4–3.5, Example 4); and (iii) simulations in linear priority models showing that forced decisive responses under latent indecision cause weight-estimation and worst-case regret losses, and that explicitly allowing indecision reports substantially accelerates Bayesian active learning. The paper concludes by proposing \"priority-aware learning\" as a more faithful elicitation paradigm.","tokens_in":60538,"tokens_out":6220,"duration_ms":70585,"significance":"If the results are taken at face value, this is a significant conceptual contribution to preference learning and AI alignment. The formal equivalence between S-RUMs and the zero-threshold separable case is clean and useful, and the non-identification theorem for perfectly inseparable priorities is a genuine formal obstruction to a widely used elicitation pipeline. The simulation study is carefully designed, with generative data, hyperparameter settings, multiple response models, standard errors, and an explicit active-learning procedure; the finding that indecision reports contain useful information beyond avoiding forced responses is plausible and practically relevant. The paper is also commendable for including detailed appendix proofs, explicit assumptions, and a candid discussion of limitations. However, the breadth of the title and abstract exceeds what the formal model actually establishes, because the central erasure and non-identification results are contingent on a specific, behaviorally unvalidated aggregation architecture. The paper is therefore more a rigorous conditional result than a general impossibility theorem about local pairwise comparisons.","major_comments":[{"comment":"The core non-identification result is derived under the assumption that aggregation happens over background rules first, then over priorities, with a permutation-invariant and unanimous rule aggregator. The authors explicitly defer other aggregation orders to future work, but the manuscript's framing—\"erased without a trace\" and \"not identifiable by local pairwise comparisons at all\"—presents this as a property of local pairwise comparisons themselves, not of a particular model architecture. Under a plausible alternative, priority-first aggregation with a max aggregator, a perfectly inseparable egalitarian priority does not cancel: the decisiveness score kappa(q) can depend on the egalitarian weight, so the priority leaves a detectable trace and the conclusion of Theorem 3.3 fails. Since no behavioral evidence is offered for rule-first aggregation, the load-bearing claim is conditional o","section":""},{"comment":"The general-inseparability example relies on an additional \"balanced sign-responsive\" condition on the rule aggregator, but this condition is not stated in the main text and is not satisfied by all aggregators permitted by Definition 2.3 (it excludes, for example, lower-percentile aggregators). The formal construction in Appendix B.6 correctly proves the claim under this extra assumption, but the main-text Example 4 reads as a general demonstration of misinterpretation of inseparable priorities. This is a load-bearing point for the paper's broader narrative that general inseparable priorities are systematically distorted by separable-consistent decoders. The condition should be stated in the main text, and the scope of the claim should be narrowed to aggregators satisfying it, or the authors should explain why the restriction is harmless.","section":""}],"minor_comments":[{"comment":"Typo: \"assumpions\" should be \"assumptions\".","section":""},{"comment":"Heading \"Eqal Treatment\" should be \"Equal Treatment\".","section":""},{"comment":"The quantifier structure in scale-free indistinguishability is ambiguous: it should be made explicit whether the rescaling beta' is allowed to depend on the query q or must be uniform across all queries. Lemma B.4 requires a uniform rescaling, and the proof of Theorem 3.3 constructs one, but the definition as written can be read otherwise.","section":""},{"comment":"The paper shows that egalitarianism can be perfectly inseparable, but only in a two-recipient, two-input toy example. The abstract's broad statement that \"egalitarianism ... local comparisons may fail to capture them\" is supported by the weaker inseparability result (Proposition 3.1), but the strong erasure result (Corollary 3.4) applies only to perfect inseparability. A sentence clarifying the gap between the general notion and the toy instances would help readers avoid overgeneralizing.","section":""},{"comment":"The simulation treats inputs as drawn from a continuous distribution despite the model's finite-X assumption; the authors note this and suggest discretization, but a more careful statement of how the reported numbers map to the formal model would improve reproducibility.","section":""}],"recommendation":"major_revision","confidential_remarks":"This is a well-executed paper with detailed proofs and careful simulations, but the headline claims overstate the generality of the formal results. The rule-first aggregation assumption is the key risk: if that assumption is not behaviorally grounded, the central erasure and non-identification results are conditional on a specific model class. I recommend major revision, with the expectation that the authors either provide behavioral support for the aggregation order, analyze alternatives, or reframe the conclusions as properties of this model class. The paper is not fatally flawed internally—the logic is sound—but its contribution as a statement about the limits of pairwise comparisons requires closer alignment between the formal model and the paper's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a serious formal contribution. It builds a pluralistic priority model, proves S-RUMs are a zero-threshold separable special case, and shows that perfectly inseparable priorities (egalitarianism, proportionality, equal treatment) can be erased and made unidentifiable from local pairwise comparisons. The simulations are careful and suggest allowing indecision reports speeds learning substantially. That is the real news.\n\nWhat is new is the formal apparatus: directional evidence scores, latent conflict vs indifference, the S-RUM equivalence theorem (Thm 2.3), and the non-identifiability results. The proofs are detailed and the appendices carry through. The simulation design is sound: 40 random weight vectors, standard errors, active learning with BALD, and convergence checks. The interview coding is exploratory but honest.\n\nThe main soft spot is the one the stress-test flags: Theorem 3.3 depends on aggregation over background rules first, then priorities, with a permutation-invariant unanimous rule aggregator. The paper calls this “natural” for computational reasons, but offers no behavioral evidence that humans aggregate this way. If people instead aggregate priorities first, or anchor on a single default background rule, perfectly inseparable priorities leave a detectable trace and the non-identifiability result fails. The paper acknowledges other aggregation orders are future work, but the abstract and contributions state the erasure result as a property of local comparisons, not as a property of one aggregation order. That is a real caveat, not a fatal flaw. The formal results are valid as theorems; their reach to real human elicitation is conditional on an unvalidated architectural assumption.\n\nAlso minor: simulations don't release code or data, and the interview coding lacks inter-rater reliability. The paper itself flags both limitations.\n\nThis is for preference-learning and AI-alignment researchers who want a formal way to talk about internal pluralism and indecision. It deserves a serious referee. The reviewer should ask for code/data, a behavioral or empirical argument for the aggregation order, and a sharper statement of which results are model-dependent.\n\nI would send it to review rather than desk reject. My own verdict would be “accept with major revisions”: the core formal work stands, but the load-bearing assumption needs to be either validated or the claims narrowed.","headline":"A clean formal model showing local pairwise comparisons can miss global priorities, but its central erasure theorem rests on an untested assumption about how people aggregate across background rules.","tokens_in":60948,"tokens_out":2207,"would_cite":true,"duration_ms":24601,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Local pairwise comparisons—the standard tool for learning what people want—cannot detect global priorities like egalitarianism, and can be systematically distorted by internal conflict; allowing people to report indecision substantially acc","keywords":["internal pluralism","pairwise comparisons","decision rules","random utility models","preference learning","inseparability","indecision","active learning"],"falsifier":"Take the two-recipient, two-input allocation example from the paper (Proposition 3.2), where egalitarianism is perfectly inseparable. Ask participants a set of local pairwise comparisons, fit a Bradley-Terry model, and then ask them to choose between complete decision rules (e.g., a rule that alternates between the two recipients vs. a rule that always gives to one). If participants' global rule choices are predicted by the local-fit model, the erasure phenomenon is behaviorally irrelevant; if the global choices reveal a preference for balanced rules that the local comparisons could not have p","tokens_in":59981,"feed_emoji":"⚖️","tokens_out":5090,"duration_ms":50407,"temperature":0.7,"pith_summary":"The paper claims that the standard way of learning human preferences—forcing people to choose between two options in a single case—rests on two assumptions that fail under internal pluralism, the idea that one person weighs several authoritative priorities at once. It builds a formal model in which each priority is a preference ordering over whole decision rules, and shows that standard score-based models like Bradley-Terry are exactly the special case where priorities are perfectly separable and indecision is impossible. Once those restrictions are lifted, global priorities such as egalitarianism or proportionality can be perfectly inseparable: they give equal and opposite evidence on every local question, so they cancel out and are erased without a trace by any learner that assumes separability. Even when priorities are separable, forcing a decisive answer when two strong priorities conflict distorts responses, and the paper's simulations show that letting people report conflict or indifference—even as a single generic 'indecision' button—cuts the number of queries needed for accurate learning by a large factor.","feed_headline":"Pairwise questions erase global priorities, model shows","feed_subtitle":"Forced local comparisons are a special case of the new model; indecision reports cut queries dramatically.","key_machinery":"The load-bearing object is the pluralistic priority model: a weighted collection of 'priorities', each a complete, transitive preference relation over the space of decision rules, together with a two-stage aggregation. For a local query (y,y';x), each priority first aggregates its 'projection gaps' across all background rules via a permutation-invariant, unanimous rule aggregator; the priorities' weighted evidence is then tallied into two directional scores s+ and s-, whose sum and difference give valence and decisiveness. Thresholds on these quantities produce four latent states: decisive preference either way, indifference (too little total evidence), and conflict (strong evidence both way","core_discovery":"On the paper's own terms, the central discovery is Theorem 3.3: any model that contains a nonzero number of perfectly inseparable priorities is behaviorally indistinguishable, over local pairwise comparisons, from the same model with those priorities deleted. Because a permutation-invariant rule aggregator sees perfectly inseparable priorities as contributing identical evidence for both options at every query, their weight cancels from the decisiveness signal up to a scale factor that is absorbed by the noise temperature. Consequently, any decoder that fits a perfectly separable model (the class that includes Bradley-Terry) will erase these priorities without a trace on exhaustive data (Coro","pith_inferences":["If local comparisons cannot identify inseparable priorities, then eliciting priorities directly—for example, asking people to state their criteria in text and then weighting those criteria—becomes not just a convenience but a necessary complement to pairwise data; the paper sketches this 'priority-aware learning' direction.","The rule-first aggregation assumption (aggregate evidence over background rules before combining priorities) is what makes the perfect-inseparability cancellation go through; if individuals instead weight background rules by plausibility or context, inseparable priorities would leave a detectable trace. Testing this aggregation order behaviorally is a natural next step.","The paper's formal distinction between indifference and conflict maps onto behavioral findings that people defer choice under conflict but not under indifference; a testable extension would compare learning speed when the indecision option is framed as 'no strong feeling' versus 'genuinely torn'.","Because the indecision benefit persists even when the learner does not know the response thresholds (learned jointly), the acceleration is not an artifact of privileged information; this suggests practical systems can offer an indecision button without a separate calibration phase."],"forward_implications":["Standard preference-learning pipelines built on Bradley-Terry or other score-based random utility models are, by Theorem 2.3, restricted to perfectly separable priorities with no indecision; the paper's results therefore delimit exactly when those pipelines are justified.","Perfectly inseparable priorities—formalized versions of egalitarianism, proportionality, and equal treatment—can be erased without a trace, so a learner may converge to a rule that is near-worst according to the individual's true priorities even with unlimited data.","Because perfectly inseparable weights are non-identifiable from local comparisons, no amount of additional local data can recover them; richer query formats are necessary.","Allowing people to report indecision (even a single generic option conflating conflict and indifference) substantially reduces the number of queries needed to learn accurate priority weights, with gains concentrated in the tens-of-queries regime that is practical for elicitation.","Learning weights is not the same as learning rules: in the forced-response simulations, weight errors were large (14–24% of worst case) while average regret was modest, but worst-case regret reached up to 39% of the utility range."],"fun_headline_variants":["Pairwise comparisons erase inseparable priorities","Tied priorities vanish from pairwise data","Indecision cuts queries in preference learning","Global priorities defeat local pairwise tests","Forced comparisons hide pluralistic priorities"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claim depends on the assumption that, when answering a local question, a person summarizes each priority by aggregating its implications across all possible 'background' behaviors of the rule symmetrically (permutation-invariantly), before combining priorities—if instead people weight some background scenarios as more plausible or salient, perfectly inseparable priorities would not cancel and could be learned from local comparisons.","fun_headline_variants_meta":{"raw":{"variants":["Pairwise comparisons erase inseparable priorities","Tied priorities vanish from pairwise data","Indecision cuts queries in preference learning","Global priorities defeat local pairwise tests","Forced comparisons hide pluralistic priorities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000289,"raw_usage":{"total_tokens":1527,"prompt_tokens":738,"completion_tokens":789,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":482,"completion_tokens_details":{"reasoning_tokens":729}},"tokens_in":482,"tokens_out":789,"duration_ms":7440,"temperature":1.0,"reasoning_tokens":729,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T08:56:49.620224+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the two-recipient, two-input allocation example from the paper (Proposition 3.2), where egalitarianism is perfectly inseparable. Ask participants a set of local pairwise comparisons, fit a Bradley-Terry model, and then ask them to choose between complete decision rules (e.g., a rule that alternates between the two recipients vs. a rule that always gives to one). If participants' global rule choices are predicted by the local-fit model, the erasure phenomenon is behaviorally irrelevant; if the global choices reveal a preference for balanced rules that the local comparisons could not have p","supporting_citations":[],"review_version":2}