{"id":"2b287190-5f28-459d-a96f-7b1d90588220","arxiv_id":"2607.11432","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A Bradley-Terry-style rationality model with an incomparability score based on utility-difference standard deviation recovers multi-dimensional rewards and Pareto frontiers from trajectory comparisons that include incomparability labels.","lead":"The paper introduces comparison-based RL (CbRL) that treats incomparability of trajectories as rational multi-objective feedback, plus a multi-objective Bradley-Terry model that recovers vector rewards and Pareto policies from such labels. This bridges preference-based RL and multi-objective RL without requiring dense rewards or scalarization weights from the expert.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Synthetic expert is drawn from the same MOBT family that is recovered, so empirical success mainly verifies self-consistency rather than external validity of the incomparability score.","rationale":"The reader correctly isolates the synthetic-expert circularity as the weakest assumption. The theory (desiderata compliance, non-convexity impossibility, KL bounds for both global and local optima) is carefully derived and internally consistent; no mathematical error is apparent. The only load-bearing empirical gap is precisely the one the reader flags: reconstruction success under a self-generated expert does not establish that the proposed rationality model captures real human incomparability. Because the authors themselves acknowledge the absence of human data (Section 5) and the verdict is already CONDITIONAL, no further downgrade is warranted. The concrete test above would settle the issue without requiring a full-scale human study.","tokens_in":31233,"tokens_out":541,"duration_ms":6971,"concrete_test":"Collect a modest human-labeled set (e.g., 200–500 trajectory pairs from GridWorld or MO-Hopper) that includes free-form “incomparable” responses; fit MOBT and a simple alternative (e.g., independent per-objective BT + hard threshold on conflicting signs). If MOBT’s held-out log-likelihood or induced hypervolume ratio is not statistically better than the alternative, the empirical support for the std-based score collapses.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central empirical claim (Section 4) is that MOBT reconstructs a multi-dimensional reward whose induced Pareto front approaches the true front (hypervolume ratio 0.96 at 5 k pairs in LQR) and yields low test KL. All labels, however, are generated by a synthetic expert that itself samples from the MOBT softmax of Eqs. 7–10 (Appendix D.1: “we have modeled our synthetic expert to use our MOBT as its rationality model”). Consequently the optimizer is recovering parameters of the identical generative family that produced the data. This does not test whether real human incomparability is well-described by the standard-deviation score h_∥ = √d · std(δ) + β, nor whether the four mode-desiderata of Definition 2.2 are the right inductive bias for human experts. The theoretical guarantees (Lemma 3.1, Thms. 3.2–3.3) remain valid inside the model class, but the paper’s claim to “recover the Pareto frontier of policies” from comparison feedback rests on an untested modeling assumption about human rationality.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper formalizes comparison-based RL (CbRL) via the Markov Decision Process with Comparisons (MDPC), allowing four outcomes (direct preference, inverse preference, indifference, incomparability). It states four mode-desiderata for a rationality model (Definition 2.2), proves that no model satisfying them can yield a convex negative log-likelihood (Proposition 2.1), and introduces the multi-objective Bradley-Terry (MOBT) model with scores (7)–(10). Lemma 3.1 shows compliance under α>0, α>β; Theorems 3.2–3.3 give KL sample-complexity bounds of order RΛdk √(log(1/δ)/N) for the global optimum and an additive P(∥) term for local optima under linear features and boundedness. Experiments on synthetic GridWorld, MO-Hopper and LQR data (labels drawn from MOBT itself) report low test KL, multi-dimensional reward recovery, and LQR hypervolume ratios approaching 0.96 at 5k pairs, plus a robustness sweep over mistake probability ε.","tokens_in":31570,"tokens_out":1267,"duration_ms":13178,"significance":"If the modeling assumptions hold, the work cleanly bridges PbRL and MORL by treating incomparability as a rational signal rather than noise, and supplies the first sample-complexity guarantees for a multi-dimensional rationality model under offline comparisons. The impossibility of convex NLL (Prop. 2.1), the explicit desiderata, the reduction to classical BT when d=1 (Remark 3.1), and the local-optimum bound that isolates the incomparability probability are technically solid contributions. The LQR closed-form Pareto evaluation and the total-variation comparison against BT/RK/Davidson baselines (Appendix D.4) strengthen the empirical case inside the model class. The main limitation is that all labels are generated by the same MOBT family that is recovered, so external validity of the standard-deviation incomparability score remains untested; the theoretical results themselves do not depend on that loop.","major_comments":[{"comment":"Section 4 and Appendix D.1: every experimental label is drawn from the MOBT softmax of Eqs. (7)–(10) itself (“we have modeled our synthetic expert to use our MOBT as its rationality model”). Consequently the reported KL values, reward matrices (Fig. 3) and LQR hypervolume ratios (0.86–0.96) demonstrate self-consistency of the optimizer rather than that real human incomparability obeys h_∥=√d·std(δ)+β. The central claim that the model “recovers the Pareto frontier of policies” from comparison feedback therefore rests on an untested inductive bias. At minimum the paper should (i) state this limitation prominently in the abstract and Section 4, and (ii) either supply a non-MOBT synthetic expert (e.g., Thurstone-style multi-objective noise or a lexicographic rule) or a small human pilot that records genuine incomparability labels.","section":null},{"comment":"Theorem 3.3 / Eq. (13): the additive error term is proportional to the marginal probability of incomparability P_θ*(∥). In the multi-objective regimes the paper targets this probability is expected to be non-negligible (and is the source of non-convexity). The bound therefore does not guarantee that a local optimum recovered by ADAM is close to the expert’s distribution when conflict is high. The manuscript should either (a) quantify how large P(∥) can be under the desiderata before the additive term dominates, or (b) provide empirical evidence that the local optima found in practice remain useful for Pareto recovery even when the incomparability ratio reaches the 0.4–0.6 range examined in Table 4 of the appendix.","section":null}],"minor_comments":[{"comment":"Definition 2.2, Eq. (4): the limit is written “lim_δ→+∞ t” with t∈{−1,1}^d∖{1_d,−1_d}; a short clarifying sentence that the limit is taken along the ray c·t, c→+∞, would remove ambiguity.","section":null},{"comment":"Figure 1 caption and surrounding text: the 2-D illustration is helpful but the axes are labeled only δ1, δ2; adding the four mode regions explicitly in the figure legend would improve readability.","section":null},{"comment":"Assumption 3.1: the feature map φ is assumed known. A brief remark on how one would estimate or over-estimate d and the feature dimension in practice (the paper already notes that d can be overestimated) would help practitioners.","section":null},{"comment":"Table 2: report the corresponding train/test split sizes and the number of random seeds more prominently; the 95 % C.I. notation is clear but the absolute number of runs is easy to miss.","section":null},{"comment":"Appendix B derivation of h_∥: the projection argument is correct, yet the final step equates the Euclidean distance to √d·std; a one-line identity ||x−x̄1||_2 = √d·std(x) would make the algebra self-contained.","section":null},{"comment":"Typographical: “thereinforcement learning” (Abstract), “asincomparable” (Abstract), and occasional missing spaces after commas in the arXiv text should be cleaned.","section":null}],"recommendation":"major_revision","confidential_remarks":"The theoretical core is publishable and the impossibility/convexity result is neat. The synthetic-expert loop is the single load-bearing empirical weakness; if the authors add even a modest non-MOBT baseline or a human pilot, the paper becomes a clear accept for a solid ML venue. Without that, major revision is the appropriate bar."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real novelty here is the MDPC framework plus the four-class MOBT model that treats incomparability as a first-class signal rather than noise. They give clean desiderata (Def. 2.2), prove the model meets them (Lemma 3.1), show that any model satisfying those desiderata cannot have convex NLL (Prop. 2.1), and then give KL sample-complexity bounds for both global and local optima under linear features (Thms. 3.2–3.3). That package is new relative to standard BT/Rao-Kupper/Davidson work and to the few multi-objective preference papers that still force a known scalarization.\n\nWhat they do well: the impossibility result is useful and correctly derived; the local-optimum bound that isolates the incomparability probability as the extra error term is honest about non-convexity; the LQR Pareto-front recovery (hypervolume ratio climbing to ~0.96) and the GridWorld reward-matrix visualizations are clear; the robustness-to-ϵ plots are sensible. Citations cover the right PbRL, MORL, and decision-theory sources without obvious gaps.\n\nSoft spots, in proportion. The central empirical loop is synthetic: labels are drawn from the same MOBT softmax that is later recovered (Appendix D.1). So the low test KL and the recovered fronts mainly verify that the optimizer can fit its own generative family, not that real human incomparability follows √d·std(δ)+β. The authors flag the missing human data in Section 5, so this is a limitation of scope rather than a hidden flaw. The linear-feature assumption is standard but restrictive; free parameters α, β, W are estimated, not free-floating. No critical math errors.\n\nThis is for people working on multi-objective alignment, robotics preference learning, or anyone who has ever thrown away “I can’t decide” labels. It deserves a serious referee. I would engage with the theory and the modeling idea; I would not yet treat the empirical Pareto claims as external validation.","headline":"Clean theoretical bridge from PbRL to multi-objective settings via an explicit incomparability model; the math holds, the synthetic experiments mainly check self-consistency.","tokens_in":32114,"tokens_out":504,"would_cite":true,"duration_ms":8938,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"When experts label trajectories incomparable, a multi-objective Bradley-Terry model recovers multi-dimensional rewards and the Pareto frontier of policies.","keywords":["preference-based reinforcement learning","multi-objective RL","Bradley-Terry model","incomparability","Pareto frontier","human feedback","rationality models","comparison-based RL"],"falsifier":"Collect real human comparison labels (including free-form “cannot compare”) on a multi-objective control task whose ground-truth objectives are known, fit MOBT, and test whether the recovered reward matrices and the hypervolume of the induced Pareto front match the known objectives within the reported KL and hypervolume ratios.","tokens_in":32149,"feed_emoji":"⚖️","tokens_out":951,"duration_ms":21243,"temperature":0.7,"pith_summary":"Preference-based reinforcement learning has treated an expert’s refusal to rank two trajectories as noise or indifference. This paper argues the refusal is often rational: it signals that multiple conflicting objectives make neither trajectory dominate the other. The authors formalize Markov decision processes with four-way comparisons (prefer, prefer-inverse, indifferent, incomparable), list four mode-desiderata any rationality model must satisfy, and introduce a multi-objective Bradley-Terry model whose scores use the average utility difference for clear preferences and the standard deviation of the difference vector for incomparability. They prove that the non-convex negative log-likelihood still admits sample-complexity guarantees on the KL divergence between true and estimated comparison distributions, and they show in simulation that the recovered multi-dimensional reward produces policies whose Pareto front approaches the true front. A sympathetic reader cares because the approach turns an everyday human response into a signal that lets multi-objective reinforcement learning run on ordinary pairwise feedback instead of dense vector rewards.","feed_headline":"Incomparability labels recover multi-objective rewards","feed_subtitle":"A Bradley-Terry extension turns expert 'can't decide' into the Pareto frontier of policies.","key_machinery":"The multi-objective Bradley-Terry model (MOBT): a four-class softmax whose incomparability score is defined as √d times the standard deviation of the utility-difference vector; this single geometric quantity makes incomparability the modal outcome precisely on the non-standard diagonals where objectives conflict.","core_discovery":"The multi-objective Bradley-Terry (MOBT) model—softmax of four scores that are the signed average utility difference for direct and inverse preference, a constant for indifference, and the standard deviation of the utility-difference vector plus bias for incomparability—satisfies the four natural mode conditions of Definition 2.2, falls back to ordinary Bradley-Terry when the problem is one-dimensional, and can be learned from an offline dataset so that the induced comparison distribution is close in KL to the expert’s.","pith_inferences":["Allowing an explicit “cannot compare” answer may lower cognitive load relative to forcing experts to state scalarization weights or multi-criteria scores.","High estimated incomparability bias or frequent incomparability labels could serve as a diagnostic that a preference dataset is multi-objective rather than noisy.","An online active-learning variant that chooses trajectory pairs expected to reduce incomparability uncertainty would be a direct algorithmic extension.","The same score construction could be ported to ranking or social-choice settings where partial orders arise from conflicting criteria."],"forward_implications":["Offline datasets that retain “incomparable” labels can reconstruct multi-dimensional rewards without ever receiving dense multi-objective reward vectors.","Standard single-objective preference models that discard or re-label incomparabilities recover only one scalarization and cannot traverse the Pareto front.","KL error between true and estimated comparison distributions scales as O(R Λ d k √(log(1/δ)/N)) under linear utility features.","Even a local minimum of the non-convex likelihood still yields KL error controlled by the marginal probability of incomparability.","The recovered multi-dimensional reward can be handed to any multi-objective RL solver to obtain the Pareto set of policies."],"fun_headline_variants":["Incomparability labels recover multi-objective rewards","MOBT model turns 'can't decide' into multi-dim rewards","Expert incomparability recovers Pareto policies from prefs","Bradley-Terry extension learns multi-reward from offline pairs","Incomparable trajectory labels yield multi-objective RL rewards"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Every experimental label is generated by a synthetic expert that itself follows the proposed MOBT model, so reconstruction success mainly verifies recovery of parameters from the same family that produced the data.","fun_headline_variants_meta":{"raw":{"variants":["Incomparability labels recover multi-objective rewards","MOBT model turns 'can't decide' into multi-dim rewards","Expert incomparability recovers Pareto policies from prefs","Bradley-Terry extension learns multi-reward from offline pairs","Incomparable trajectory labels yield multi-objective RL rewards"]},"model":"grok-4.5","effort":"low","cost_usd":0.00423,"raw_usage":{"total_tokens":1247,"prompt_tokens":717,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":42300000,"prompt_tokens_details":{"text_tokens":717,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":465,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":717,"tokens_out":65,"duration_ms":5237,"temperature":1.0,"reasoning_tokens":465,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T05:37:47.568399+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Collect real human comparison labels (including free-form “cannot compare”) on a multi-objective control task whose ground-truth objectives are known, fit MOBT, and test whether the recovered reward matrices and the hypervolume of the induced Pareto front match the known objectives within the reported KL and hypervolume ratios.","supporting_citations":[],"review_version":1}