{"id":"c0ed4757-5a91-4bd5-9fa3-47d9dadea645","arxiv_id":"2506.21887","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Active-MoSH interactively learns a decision maker's soft and hard bounds on objectives and actively samples Pareto-optimal points, with a sensitivity analysis module intended to build trust.","lead":"This paper presents a system that helps people choose between options with competing goals, such as cancer treatment plans, by adjusting soft targets and hard limits during an interactive search. It aims to help decision makers find better tradeoffs and gain confidence that they have not missed a superior option.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The simulated DM's 5% larger bound adjustments under T-MoSH inject the global component's advantage, so the claimed trust/efficiency benefit is not yet independently established.","rationale":"The paper has real strengths: it extends a coherent SHF utility model to interactive settings, provides submodular sparsification guarantees via existing theory, and includes a genuine user study with 21 MTurk workers, IRB approval, screening, counterbalanced tasks, and multiple-comparison correction. These features support a conditional rather than summary rejection. However, the central comparative claim is about efficiency and trust. The synthetic evaluation is the main quantitative evidence, and Appendix A.4.1 gives the T-MoSH arm a 5% larger adjustment magnitude whenever T-MoSH finds an improved point. That is a direct injection of the effect under test; it does not follow from the sensitivity analysis itself. Additionally, the simulated Active-MoSH DM is constructed to rank displayed points by proximity to the adjusted bound, so the Section 3.2 ranking assumption cannot be validated by the simulation, only instantiated. The user study provides some external evidence for convergence (H1 significant vs. pairwise and full ranking at p about 0.05; not vs. partial ranking at p = 0.07), but the trust metric ISE is not significant (adj. p = 0.15), and the significant trust finding comes from self-reported Likert responses. Given this, a conditional verdict is appropriate until a neutral simulation protocol, especially one removing the T-MoSH magnitude boost, and ideally code release confirm the advantage. If the boost is removed and the advantage disappears, the central claim would need to be weakened to the user-study result alone.","tokens_in":34244,"tokens_out":4788,"duration_ms":57193,"concrete_test":"Rerun the Section 5.3 and 5.5 comparisons with the T-MoSH magnitude boost removed: set the 5% increase in Appendix A.4.1 to 0 for Active-T-MoSH, holding all other simulation parameters fixed, and report Active-T-MoSH versus Active-MoSH and all baselines in SHF Utility Ratio. If the gap collapses or reverses, the global component's claimed benefit is an artifact of the injected adjustment magnitude. As a secondary check, run the same comparison with an alternative simulated DM that adjusts bounds using global utility differences rather than proximity to displayed points; if Active-MoSH's convergence advantage over pairwise and ranking disappears, the Section 3.2 ranking interpretation is doing the work rather than the feedback mechanism itself.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The most load-bearing weakness is the simulation protocol in Appendix A.4.1, not the modeling formalism itself. The appendix sets default bound-adjustment magnitudes of 0.3 and 0.4, then states that when y* is outside the bounds and T-MoSH 'promotes an improved point,' the simulated DM is assumed to gain enhanced confidence and increase the magnitude of the bound adjustment by 5%. This directly injects the effect being tested: Active-T-MoSH receives larger moves toward the known ground truth than Active-MoSH. Since the only difference between those two arms is this magnitude boost plus the sensitivity-analysis samples, the T-MoSH advantage in Figures 3 and 4 cannot be attributed to the information T-MoSH provides. The same section also implements the Active-MoSH DM by selecting the displayed point closest to y* and ranking points by proximity to the adjusted bound, which is exactly the Plackett-Luce likelihood in Eq. 5; thus the synthetic experiments cannot independently validate the Section 3.2 feedback interpretation. With the primary trust metric ISE not reaching significance in the user study (adj. p = 0.15 vs. pairwise), the central efficiency/trust claim currently rests on a simulation whose key comparison is favorably biased.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Active-MoSH, an interactive framework for multi-objective optimization in which a decision-maker (DM) iteratively adjusts soft and hard bounds on objectives. The local component models DM preferences and bounds with Bayesian posteriors, interprets bound adjustments as implicit rankings via a Plackett-Luce likelihood (Eq. 5), and uses an active sampling strategy to present small, Pareto-optimal query sets. The global component T-MoSH uses multi-objective sensitivity analysis to propose points that may indicate overlooked improvements, aiming to build DM trust. The authors evaluate Active-MoSH against pairwise, ranking, and uniform feedback baselines in simulations on synthetic and real-world problems (including brachytherapy) and in a user study on AI-generated image selection. They claim improved convergence efficiency and trust, measured by SHF Utility Ratio and an Iteration Stop Efficiency metric.","tokens_in":34535,"tokens_out":7295,"duration_ms":67251,"significance":"The problem is important: interactive multi-objective decision-making with expressive, dual-level preferences is relevant to high-stakes applications, and the paper is among the first to formalize iterative soft/hard bound adjustments probabilistically. The framework is principled, and the sparse query selection inherits a submodularity guarantee (Theorem 1). The paper also attempts a human-subject validation, which is commendable. However, the main validation is undermined by a simulation protocol that injects the T-MoSH advantage (A.4.1) and by a user study whose primary trust metric was not significant. As currently presented, the evidence does not yet establish the claimed efficiency and trust benefits.","major_comments":[{"comment":"The simulation protocol directly injects the benefit of T-MoSH: when y* is outside the bounds and T-MoSH promotes an improved point, the simulated DM is assumed to gain enhanced confidence and increase the magnitude of the bound adjustment by 5%. Since this magnitude boost is the only systematic difference between the Active-MoSH and Active-T-MoSH arms in the synthetic experiments, the higher SHF Utility Ratios for Active-T-MoSH in Figures 3, 4, 5, and 8 cannot be attributed to the information T-MoSH provides. The authors should either re-run the comparison without this assumption or justify the 5% boost with human behavioral data.","section":"Appendix A.4.1 (simulation setup)"},{"comment":"The Section 3.2 feedback interpretation is load-bearing: the posterior updates in Eq. (3) rely on a Plackett-Luce likelihood over rankings induced by proximity to the adjusted bound. The simulated DM in A.4.1 is implemented to choose the point in Ym closest to y* as the reference and to rank points by distance to the new bound, which is exactly the same rule as Eq. (5). Consequently, the synthetic experiments cannot independently validate the ranking interpretation; they only show the method converges when the DM behaves as modeled. The paper needs an external test of this assumption, for example a user study that directly measures how users rank displayed points after a bound adjustment, or a comparison with alternative feedback models on real user data.","section":"Section 3.2 and Appendix A.4.1"},{"comment":"The primary behavioral metric for trust, Iteration Stop Efficiency (ISE), did not reach statistical significance against any baseline (adj. p = 0.15 vs. pairwise, 0.18 vs. partial ranking, 0.29 vs. full ranking). The significant result on the Likert trust question is a self-report that may be confounded with perceived expressiveness or novelty, and the user study did not include an Active-MoSH (without T-MoSH) condition, so the specific contribution of T-MoSH to trust is not isolated. The abstract's claim that the framework 'enhance[s] DM trust' is therefore not supported by the behavioral evidence.","section":"Section 6.3, Eq. (10)"},{"comment":"The evaluation metric SHF Utility Ratio (Eq. 2) and the feedback likelihood (Eq. 5) are both defined via the same SHF utility function, and the simulation ground truth y* is sampled to lie in the high-utility regions of the chosen soft and hard bounds (A.4.1). This creates a circular evaluation: the simulation rewards methods that align with the SHF model that the method itself assumes. To establish external validity, the authors should include at least one evaluation that does not depend on SHF utilities, such as distance to the known Pareto point under an independent utility model, or behavioral outcome measures from the user study.","section":"Section 2.2, Eq. (2) and Eq. (5); Appendix A.4.1"}],"minor_comments":[{"comment":"The x-axis of Figures 3--5 is not explicitly described in the text or captions; please add axis labels and state whether the horizontal axis is feedback units or iterations.","section":"Section 5.2 and Figure 3"},{"comment":"The Metropolis-Hastings implementation uses only 20 burn-in steps for the posterior over the preference vector; please provide diagnostics or a sensitivity analysis for this choice.","section":"Appendix A.2.1"},{"comment":"There are several incomplete references, e.g., 'Ziebart et al., Shaikh et al., 2024' in the related work section; please correct the citation entries.","section":"References"},{"comment":"The user study measures self-reported trust with a single Likert item; consider using a validated trust scale to improve reliability.","section":"Section 6.1"},{"comment":"When each feedback instance is assigned a single unit, the comparison at '10 units' means different numbers of feedback instances per method; please clarify this interpretation in the caption.","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The simulation bias in A.4.1 is a fundamental issue that must be addressed before publication. I also suggest the authors temper the abstract's significance claims until the ISE result is strengthened or the behavioral trust evidence is more direct."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Eddie,\n\nHere's my read on arXiv:2506.21887 (Active-MoSH). The novel piece is the interactive loop: treating a DM's bound adjustment as an implicit Plackett-Luce ranking over displayed points, and updating posteriors over both preference weights and soft/hard bounds. T-MoSH's sensitivity component—EI-style screening under bound constraints—is a reasonable way to surface \"have I missed something?\" candidates. The paper is clearly written, the appendix is unusually transparent, and the user study (21 MTurk participants, four mechanisms, Likert plus behavioral metrics) is a genuine effort, even if underpowered.\n\nWhere it wobbles: the simulation protocol in A.4.1 injects the very effect it claims to test. When T-MoSH promotes a point, the simulated DM is granted 5% larger bound adjustments. That's effectively the only difference between the Active-MoSH and Active-T-MoSH arms producing the advantage in Figures 3 and 4. The synthetic experiments therefore cannot independently establish that T-MoSH's information, rather than the larger simulated moves, drives the improvement. The same appendix generates ground truth y* inside high-utility regions of the sampled SHF, and the ranking likelihood (Eq. 5) is built from the same SHF utilities used in the evaluation metric (Eq. 2), so the synthetic results are partly circular. The primary trust metric, ISE, didn't reach significance in the user study (p=0.15 vs pairwise), and the trust claim ultimately leans on a Likert item that is significant but self-reported. Ten trials per simulation is thin. The feedback interpretation—that bound adjustments imply a ranking—is a reasonable assumption but untested against alternatives.\n\nNone of this kills the paper. The user study does show a significant convergence advantage at the second iteration for Active-T-MoSH against pairwise and full ranking (adj. p ≈ 0.05), and the expressiveness and mental-effort results are coherent with the design. The authors are upfront about the ISE null result and the cognitive-load tradeoff. The main unresolved question is whether T-MoSH's benefit is real or an artifact of the simulation's confidence boost.\n\nWho should read it: anyone working on interactive multi-objective optimization or preference elicitation with bound-style feedback. It's a solid extension of MoSH with a clear experimental structure, but the T-MoSH claim needs a cleaner protocol—equal adjustment magnitudes, or a user study that separates trust from familiarity—before I'd take it as established.\n\nI'd send it to review. The flaws are fixable, the contribution is incrementally real, and the user study infrastructure is reusable.","headline":"A genuinely interactive extension of MoSH with an honest user study, but the T-MoSH advantage is partly injected by the simulation protocol, so the central trust/efficiency claim needs a cleaner test before it lands.","tokens_in":35071,"tokens_out":2520,"would_cite":false,"duration_ms":25905,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Active-MoSH claims that learning from iterative soft-hard bound adjustments gets decision makers to their ideal Pareto-optimal point in fewer interactions than standard feedback, with stronger confidence.","keywords":["multi-objective optimization","preference learning","soft-hard bounds","interactive decision-making","Bayesian inference","active query sampling","sensitivity analysis","user study"],"falsifier":"Record what users actually do: give a participant the Active-MoSH interface, have them adjust a hard or soft bound, then separately ask them to rank the points shown. If the stated rankings disagree with distance-to-the-new-bound ordering across users and iterations, the Plackett-Luce likelihood in Eq. (5) is misspecified and the posterior updates are not tracking real preferences.","tokens_in":34052,"feed_emoji":"🎯","tokens_out":9429,"duration_ms":100809,"temperature":0.7,"pith_summary":"High-stakes multi-objective choices—a cancer treatment plan, a structural design, a cover image—force a decision maker to balance competing goals with expensive evaluations. This paper proposes Active-MoSH, which turns the way experts already talk about goals—aspirational soft bounds and non-negotiable hard bounds per objective—into a Bayesian preference-learning signal. Each adjustment of a bound is read as an implicit ranking of the displayed candidate points, and the system actively selects the next small query set while a global sensitivity-analysis component flags potentially overlooked high-value regions. If the framework works as claimed, decision makers reach their ideal Pareto-optimal point in fewer interactions, with more confidence, than with pairwise, ranking, or uniform feedback; the paper offers simulations, a brachytherapy case, and an image-selection user study as evidence.","feed_headline":"Soft-hard bound tweaks guide users to the ideal tradeoff faster","feed_subtitle":"In a user study, it converged faster and earned higher trust ratings than ranking or pairwise feedback.","key_machinery":"The central object is the soft-hard utility function ($u_\\alpha$), a bounded utility that rewards objective values near the aspirational soft bound and forbids values beyond the non-negotiable hard bound, combined across objectives by a scalarization $s_\\lambda(u_\\alpha(x))$. The load-bearing mechanism is the interpretation of a bound adjustment as an implicit ranking of the displayed points, scored by a Plackett-Luce likelihood, which turns expert constraint-talk into a full preference signal. Active sampling then maximizes the expected SHF utility ratio over posterior draws of $\\lambda$ and $\\alpha$, while T-MoSH maximizes expected improvement with the perturbed hard bound slightly relaxed to identify potentially overlooked, high-value regions.","core_discovery":"On its own terms, the discovery is that the familiar expert act of restating targets and limits—'keep tumor coverage above 95% if possible, never below 90%, and keep bladder dose under 601 cGy'—is enough to drive a preference-learning loop. Active-MoSH treats each bound adjustment as an implicit ranking of the K displayed Pareto points, with points closest to the newly adjusted bound ranked highest, and updates posteriors over the hidden preference vector $\\lambda$ and hidden ideal bounds $\\alpha$ through a Plackett-Luce likelihood. It then samples new queries by maximizing the expected SHF utility ratio, using Gaussian-process surrogates and random scalarizations from the posteriors, and sparsifies them to a small set. The global component T-MoSH maximizes expected improvement under a slightly relaxed hard bound to surface high-value points outside the current region. The paper asserts that this combined loop finds a decision maker's ideal point with fewer interaction units than pairwise, ranking, or uniform feedback, and that the sensitivity-based global view increases confidence; the user study is presented as evidence of both.","pith_inferences":["The paper leaves implicit that T-MoSH's expected-improvement screen can double as a stopping rule: when no candidate outside the current bounds has positive expected improvement, the system has a defensible answer to 'have I missed anything?'.","A testable extension is to learn the ranking interpretation itself, letting bound adjustments and explicit rankings coexist so the likelihood can be corrected if users adjust bounds for global rather than local reasons.","The interaction-unit accounting could be validated directly by measuring task-completion time or physiological effort, which would sharpen the efficiency claim beyond the paper's assumption-heavy unit assignments."],"forward_implications":["With the same number of interaction units, Active-MoSH reaches a higher SHF utility ratio than pairwise, full ranking, partial ranking, and uniform feedback on the Branin-Currin, Four Bar Truss, and brachytherapy benchmarks.","The ablation results show that dropping the Plackett-Luce preference update or replacing active sampling with random sampling degrades convergence, so the paper's efficiency claim depends on both components.","Active-T-MoSH, which adds T-MoSH sensitivity analysis, generally converges at least as fast and with lower variance than Active-MoSH alone.","In the user study, participants reached higher utility at the second feedback iteration with Active-T-MoSH than with pairwise and full-ranking feedback, and they rated it more expressive and more trustworthy, while also rating it more mentally demanding."],"supporting_citations":[{"why":"Defines soft-hard utility functions and the MoSH-Dense/MoSH-Sparse algorithms that Active-MoSH adapts for interactive preference queries.","marker":"Chen et al. [2024]"},{"why":"Supplies random scalarizations and the UCB acquisition heuristic used to sample points from the expensive black-box objectives.","marker":"Paria et al. [2019]"},{"why":"Provides the pairwise, ranking, and information-gain active querying baselines against which Active-MoSH is evaluated.","marker":"Bıyık et al. [2019b]"},{"why":"Introduces the Plackett-Luce model used as the likelihood for bound-induced rankings in Eq. (5).","marker":"Luce [1959]"},{"why":"Provides robust submodular observation selection and the GPC algorithm that backs MoSH-Sparse's query-set guarantee.","marker":"Krause et al. [2008]"},{"why":"Supplies expected improvement, the criterion T-MoSH maximizes to flag potentially overlooked high-value points.","marker":"Jones et al. [1998]"},{"why":"Provides the epsilon-constraint approach and clinical data used to generate the brachytherapy tradeoff surface.","marker":"Deufel et al. [2020]"},{"why":"Supplies the Bayesian Plackett-Luce inference used by the partial-ranking baseline.","marker":"Guiver and Snelson [2009]"}],"fun_headline_variants":["Bound tweaks guide interactive learning to ideal tradeoff faster","Fewer queries to find ideal tradeoff with soft-hard bounds","Active sampling plus sensitivity increases confidence in decisions","Learn preferences from bound adjustments, not just rankings","Interactive framework speeds Pareto selection in brachytherapy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking interpretation is the load-bearing premise: when a decision maker moves a soft or hard bound, the framework assumes the displayed points are ranked by distance to the new bound, and if real users adjust bounds for global reasons rather than to signal which displayed points they like, the posteriors can converge to the wrong preferences.","fun_headline_variants_meta":{"raw":{"variants":["Bound tweaks guide interactive learning to ideal tradeoff faster","Fewer queries to find ideal tradeoff with soft-hard bounds","Active sampling plus sensitivity increases confidence in decisions","Learn preferences from bound adjustments, not just rankings","Interactive framework speeds Pareto selection in brachytherapy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000571,"raw_usage":{"total_tokens":2748,"prompt_tokens":1041,"completion_tokens":1707,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":657,"completion_tokens_details":{"reasoning_tokens":1632}},"tokens_in":657,"tokens_out":1707,"duration_ms":13794,"temperature":1.0,"reasoning_tokens":1632,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:16:12.446933+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record what users actually do: give a participant the Active-MoSH interface, have them adjust a hard or soft bound, then separately ask them to rank the points shown. If the stated rankings disagree with distance-to-the-new-bound ordering across users and iterations, the Plackett-Luce likelihood in Eq. (5) is misspecified and the posterior updates are not tracking real preferences.","supporting_citations":[],"review_version":1}