{"id":"6825b84c-3deb-4b98-9938-7b58c8378d9f","arxiv_id":"2411.11182","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Combining CMA-ES sampling with information gain query selection improved perceived behavioral adaptation and overall preference in a 14-person study, though ease-of-use gains over pure information gain were not significant.","lead":"This paper combines two existing ways of asking people to rank robot behaviors into one algorithm, CMA-ES-IG, and tests it on a robot arm and a social robot. In a small user study, participants rated it as more adaptive and preferred it over both prior methods, though its ease-of-use advantage over one baseline was not statistically significant.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The bridge from simulation to user study is unvalidated: Section 4.1 replaces sampled feature queries with nearest real trajectories, so CMA-ES-IG's perceived benefits may be artifacts of dataset coverage rather than the algorithm.","rationale":"The paper's central contribution is an algorithm that improves user experience by generating more appealing and distinguishable queries. The mechanistic justification is Table 1: CMA-ES-IG matches IG in alignment and beats both baselines in query quality. But Table 1 is computed on synthetic feature points sampled from a normal distribution, not on the physical trajectories users actually ranked. Section 4.1 explicitly inserts a nearest-neighbor retrieval step, so every property of the theoretical query (information gain, quality, distinctness) is only as good as the retrieval. With 1000 or 1500 trajectories in 4- or 6-dimensional autoencoder features, the nearest neighbor to an arbitrary sample can be far; the paper gives no measure of this distance. The user study itself only measures subjective ratings, not objective alignment, so it cannot independently validate the mechanism. This is the weakest link because if retrieval distorts queries, all experimental outcomes—higher perceived behavioral adaptation, preference, and ease of use—could be driven by dataset artifacts or by the retrieval rule rather than by CMA-ES-IG's sampling. This is a correctable issue: the released code makes a dataset-constrained simulation straightforward. The other weaknesses (small sample, uncorrected multiple comparisons, overclaimed ease-of-use) are secondary and do not by themselves invalidate the central preference claim. For these reasons, the conditional verdict remains appropriate.","tokens_in":11687,"tokens_out":4828,"duration_ms":48432,"concrete_test":"Using the released code, rerun the Section 3.1 simulation in the actual learned feature spaces and datasets: for each CMA-ES-IG sampled query, replace each feature point by its nearest dataset trajectory exactly as done in Section 4.1, then recompute the AUC for alignment and quality shown in Table 1. Also report the distribution of nearest-neighbor distances in feature space. If the quality advantage of CMA-ES-IG over IG and the alignment values persist under this dataset-constrained retrieval, the concern is resolved; if they degrade substantially, the user-study results cannot be attributed to the algorithm.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Algorithm 1 (lines 4-6) samples feature points from N(mu,C), selects a query by maximizing information gain, and updates the belief from the ranking. In the user study, however, Section 4.1 states: \"We then identified the closest trajectories in the dataset to the sampled trajectory features to present to the participant.\" Thus users never see the sampled feature points; they see nearest neighbors from a fixed dataset of 1000 handover or 1500 gesture trajectories. The theoretical argument and the Section 3.1 simulation evaluate information gain and quality on continuous sampled feature points, with no finite-dataset or nearest-neighbor retrieval step. If the learned autoencoder features (d=4, d=6) are sparsely covered by the dataset, the retrieved queries can be far from the intended points: their information gain can be lower, their average reward can differ, and the distinctness that makes ranking easy can be degraded. The user study does not report nearest-neighbor distances or compare the actual presented trajectories against the intended ones. Consequently, the subjective preference for CMA-ES-IG, and the claim that it maintains preference-learning alignment while increasing average reward, are not directly supported by the data as presented; they could reflect the retrieval rule or dataset instead of the CMA-ES-IG sampling mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CMA-ES-IG, an algorithm that combines CMA-ES adaptive sampling with an information-gain query selection criterion for preference-based reward learning from trajectory rankings. The central claim is that CMA-ES-IG preserves the preference-learning alignment of information-gain querying while increasing the average reward of the trajectories presented to users, and that users perceive the robot as adapting without sacrificing the distinguishability of the ranking options. The authors support this with a simulated-preference study across feature spaces of dimension 8, 16, and 32, and with a within-subjects user study (n=14) in which participants ranked JACO arm handover trajectories and Blossom robot gestures under IG, CMA-ES, and CMA-ES-IG. The user study reports that CMA-ES-IG received the highest ease-of-use rating, significantly higher than CMA-ES but not significantly higher than IG, and the highest perceived behavioral adaptation, significantly higher than both baselines, and was ranked as the preferred algorithm.","tokens_in":11958,"tokens_out":4179,"duration_ms":41613,"significance":"If the central claim holds, the contribution is valuable to human-robot interaction and personalization: it reframes query generation as a user-experience design problem, and it is one of the few papers to combine an explicit Bayesian reward model with black-box trajectory optimization and to evaluate the combination on both physical and social robot tasks. The user study uses validated Likert scales with high internal consistency (Cronbach's alpha .89 and .97), appropriate non-parametric repeated-measures tests, and a counterbalanced within-subjects design. The authors also provide code and a hyperparameter sensitivity analysis, which strengthens reproducibility. The main risk is whether the user-study results can be attributed to the CMA-ES-IG mechanism, because the presented trajectories are retrieved from a fixed dataset rather than being the sampled feature-space queries that the simulation and algorithm description analyze.","major_comments":[{"comment":"The user study replaces the feature-space query produced by Algorithm 1 with the nearest real trajectory from a pre-collected dataset, as stated in §4.1: 'We then identified the closest trajectories in the dataset to the sampled trajectory features to present to the participant.' However, the simulation in §3.1 and the theoretical motivation in §3 evaluate information gain and quality on continuous sampled feature points, with no finite-dataset retrieval step. The paper does not report nearest-neighbor distances, the coverage of the learned autoencoder feature space, or whether the retrieved trajectories preserve the information gain, average reward, and pairwise distinctness of the intended queries. Without this, the significant user-experience results in §4.3 could be driven by the dataset construction or the retrieval rule rather than by CMA-ES-IG's sampling and IG mechanism. Please add a validation of this bridge, for example by reporting nearest-neighbor distance distributions, recomputing information gain and quality on the actually presented queries, or running the simulation with the retrieval step included.","section":"§4.1, Algorithm 1"},{"comment":"The abstract claims that users find CMA-ES-IG 'more intuitive and easier to use than previous approaches,' but the ease-of-use analysis in §4.3 only shows a significant improvement over CMA-ES (W=5.5, p=.016); the comparison with IG (M=5.50 vs 5.13) is described as 'empirically easier' without a reported significance test. Since IG is one of the previous approaches, the stronger claim in the abstract and conclusion is not supported by the reported statistics. Please report the IG comparison p-value or qualify the claim.","section":"§4.3 and Abstract"},{"comment":"With n=14 and multiple pairwise repeated-measures tests across ease of use, behavioral adaptation, and overall ranking, the p-values are reported without correction for multiple comparisons. For example, the behavioral-adaptation comparison against CMA-ES (p=.033) would not survive a Holm-Bonferroni correction over six tests. This does not invalidate the results, but the paper should report adjusted p-values or explicitly frame the comparisons as exploratory, especially given the small sample size.","section":"§4.3, Figs. 6–8"},{"comment":"The quality metric used in the simulation is the average reward of the presented trajectories, which is precisely the quantity that CMA-ES and CMA-ES-IG are designed to increase. Consequently, the quality results do not independently establish that users perceive higher appeal, and the Fig. 3 caption's claim that 'CMA-ES-IG performing significantly better' is not supported by the reported statistics, since no tests, effect sizes, or confidence intervals are given for the simulated AUC comparisons. Please provide inferential statistics for the simulated comparisons or soften the significance claim.","section":"§3.1, Table 1 and Fig. 3"}],"minor_comments":[{"comment":"The symbol D is used both for the trajectory dataset and for the number of samples drawn from the CMA-ES distribution; please rename one of them for clarity.","section":"§3, Algorithm 1"},{"comment":"The text cites 'Habiban et al.' but the reference list gives the correct spelling 'Habibian'; please correct the in-text citation.","section":"§2"},{"comment":"The conclusion contains a typo: 'prefered' should be 'preferred.'","section":"§4.4"},{"comment":"Each algorithm was paired with a specific handover object or gesture emotion within a domain; although the algorithm order was counterbalanced, please clarify how the assignment of algorithms to tasks was randomized and whether any task-specific effects were examined.","section":"§4.2"},{"comment":"The sample count D for the CMA-ES sampling and the posterior belief, as well as several autoencoder and CMA-ES hyperparameters beyond the initial step size, are not reported; please list them in the appendix or supplement.","section":"§4.1"},{"comment":"The information-gain objective is approximated by medoid selection, but the exact procedure used for the IG baseline is not specified; please state how the IG baseline generates its candidate set and whether it samples uniformly from the feature space.","section":"§3, Equation (3)"}],"recommendation":"major_revision","confidential_remarks":"I see this as a meaningful contribution with a clear algorithmic idea and a genuine user study, so rejection is not warranted. The load-bearing issue is the unvalidated nearest-trajectory retrieval step: the simulation and the algorithm operate on continuous feature-space samples, while the user study presents retrieved real trajectories, and the paper does not show that retrieval preserves the properties that drive the claimed benefits. The requested retrieval analysis is feasible within the scope of the manuscript, which is why I recommend major revision rather than reject. In addition, the abstract's ease-of-use claim goes beyond the significant results, and the multiple-comparison issue should be addressed for the user-study statistics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real contribution is the hybrid itself: CMA-ES-IG uses CMA-ES to shift query sampling toward higher-reward regions while keeping the information-gain objective so queries stay distinguishable. That is a natural combination, and the simulation supports it — CMA-ES-IG matches IG on alignment and beats both baselines on the quality of presented trajectories. The user study is well structured for an HRI paper: within-subjects, two domains, validated Likert scales, appropriate non-parametric tests, and code is released. That is real work and worth taking seriously.\n\nThe soft spots are fixable but important. The biggest one is the gap between Algorithm 1 and what participants actually saw. In the user study, the algorithm samples feature points, but the paper then says \"we identified the closest trajectories in the dataset to the sampled trajectory features to present to the participant.\" So users never rank the sampled queries; they rank nearest neighbors from a fixed dataset of 1000 or 1500 trajectories. The simulation, however, evaluates directly on sampled feature points with no retrieval step. The paper never reports nearest-neighbor distances or checks how far the presented trajectories are from the intended ones. This means the user-study results — the main evidence for the title claim — could reflect dataset coverage or the retrieval rule rather than what CMA-ES-IG actually does. The comparison across algorithms is likely still fair since all three use the same retrieval, but the claim that CMA-ES-IG improves user experience because it presents better queries is not directly supported.\n\nSecond, the abstract overstates ease of use. The EOU difference between CMA-ES-IG and IG was not significant (M=5.50 vs 5.13, no p reported for that comparison). Only the CMA-ES comparison was significant. The behavioral adaptation finding is stronger — significant against both baselines — so the abstract should be softened, not the whole paper.\n\nThird, n=14 with multiple uncorrected pairwise tests is thin. The effect sizes are moderate to large, so the results are suggestive, but not conclusive. The authors should report whether the comparisons survive a Holm-Bonferroni correction or at least present adjusted values.\n\nDespite these issues, the paper is honest and the logic is clear. I would send it to peer review with a request for the authors to address the retrieval gap: report nearest-neighbor distances, rerun the simulation with the retrieval step, or use a dense trajectory set. A revised version could be a solid contribution to preference-based HRI. I would not desk-reject it.","headline":"A sensible hybrid algorithm with a promising user study, but the missing bridge between sampled features and retrieved real trajectories leaves the main claim about user experience unproven.","tokens_in":12476,"tokens_out":1745,"would_cite":true,"duration_ms":20011,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CMA-ES-IG combines an evolution strategy with an information-gain objective to generate ranking queries that improve in reward while staying easy to distinguish, and users rate it the most adaptive and preferred way to teach a robot.","keywords":["preference-based reward learning","human-robot interaction","CMA-ES","information gain","active learning","user experience","assistive robots","trajectory ranking"],"falsifier":"Collect a trajectory dataset with deliberately sparse coverage of the feature space, run the same user study, and measure whether CMA-ES-IG still beats the baselines on perceived behavioral adaptation and ease of use; if its advantage disappears or reverses, the reported effect comes from the retrieval rule or dataset, not from the sampling-and-selection algorithm. Alternatively, in simulation, compute the information gain of the actual nearest-neighbor queries and compare it with the information gain of the raw CMA-ES samples: if they diverge substantially, the algorithm is not delivering the queries it claims to.","tokens_in":11464,"feed_emoji":"🤖","tokens_out":7013,"duration_ms":61443,"temperature":0.7,"pith_summary":"Preference-based robot teaching asks users to rank candidate behaviors, but the two standard query-generation strategies pull in opposite directions: information-gain selection gives options that are easy to rank yet do not visibly get better, while CMA-ES gives options that improve but become too similar to distinguish. This paper tries to establish that a hybrid, CMA-ES-IG, gets both: it samples candidates from CMA-ES's adaptive distribution and then picks the subset with the highest expected information gain. The simulations show the hybrid keeps preference-learning alignment while markedly increasing the average reward of the presented trajectories. The user study, across a physical JACO arm and a social Blossom robot, reports that participants rated CMA-ES-IG highest on ease of use and on perceived behavioral adaptation, significantly above the pure CMA-ES baseline for both and above the pure information-gain baseline for adaptation, and ranked it as their preferred method. If true, this would mean that how much the robot appears to adapt during teaching is a designable property, not a side effect, of the query-selection algorithm.","feed_headline":"Users rate hybrid CMA-ES-IG as the most adaptive robot-teaching method","feed_subtitle":"It pairs information-gain query selection with CMA-ES sampling, so users see the robot improve while still giving clear feedback.","key_machinery":"Algorithm 1: sample D trajectory-feature vectors from a CMA-ES multivariate normal distribution parameterized by mean µ and covariance C; sample a set of belief trajectories from the current posterior over reward weights; select the |Q| queries that maximize expected information gain, approximated by finding |Q| medoids among the samples; collect the user's ranking; update the posterior over weights with the Bradley-Terry model; update µ and C with the standard CMA-ES rule. The combination is the carrier: CMA-ES moves the sampling distribution toward high-reward regions so presented trajectories improve, while the information-gain selection keeps the presented options far apart in feature space so users can still rank them.","core_discovery":"The paper argues that the two dominant approaches to preference-based robot teaching each sacrifice something users care about: pure information-gain (IG) query selection produces options that are easy to distinguish but do not visibly improve, while CMA-ES produces visibly improving options that become too similar to rank. CMA-ES-IG combines them by sampling trajectory features from CMA-ES's adaptive distribution and then selecting the subset with maximum expected information gain. In simulation across 8-, 16-, and 32-dimensional feature spaces, it matches or exceeds the baselines on learned-preference alignment while producing queries with substantially higher average reward. In a within-subjects study with 14 participants teaching a physical JACO arm and a social Blossom robot, CMA-ES-IG received the highest ease-of-use ratings and the highest perceived behavioral adaptation, significantly above both baselines on adaptation, and was ranked the preferred algorithm.","pith_inferences":["The user-study advantage may depend on dataset coverage: because participants rank the nearest real trajectories to the sampled feature vectors, a sparse trajectory dataset could sever the link between the CMA-ES sample and what the user actually sees; an ablation with denser versus sparser datasets would test whether the perceived adaptation comes from the algorithm or from the retrieval rule.","If the perceived-adaptation effect is real, it suggests that presenting visible learning progress—not just arriving at a good final policy—drives user trust; measuring trust and willingness to keep teaching over longer interactions would likely show stronger retention for CMA-ES-IG than for IG.","The medoid approximation trades exact information gain for tractability; a version that directly optimizes the continuous query set rather than selecting from a finite sample might further increase both alignment and quality.","The framework could extend beyond rankings to pairwise choices or 'approximately equal' responses, using the same query-selection principle."],"forward_implications":["If correct, assistive robots can be taught through rankings that users perceive as responsive, which may increase adoption and sustained engagement with personalization.","The approach suggests that 'query quality' (reward of presented options) and 'query informativeness' are not mutually exclusive; future active-learning objectives can include perceptual experience as an explicit optimization criterion.","The simulation results imply the benefit grows in higher-dimensional feature spaces, where CMA-ES-IG improves alignment while IG degrades.","The same combination strategy could be applied to other black-box optimization settings where humans evaluate candidates, such as haptic texture design or exoskeleton assistance, where candidate quality matters during training.","Since all methods took under one second per query, CMA-ES-IG is deployable in real-time interaction loops."],"supporting_citations":[{"why":"Supplies the information-gain query-selection objective that CMA-ES-IG builds on, along with the user-friendly 'easy questions' framing.","marker":"[6]"},{"why":"Provides the CMA-ES algorithm whose adaptive sampling drives the increasing reward of presented trajectories.","marker":"[25]"},{"why":"Gives the medoid-based batch approximation used to make the information-gain selection computationally tractable.","marker":"[7]"},{"why":"Defines the Bradley-Terry rational-choice model used to compute ranking probabilities and update the belief over reward weights.","marker":"[8]"},{"why":"Supplies the active preference-based learning formulation and the linear-reward framework that the simulation and belief updates rely on.","marker":"[40]"},{"why":"Provides the CMA-ES update rules referenced in Algorithm 1 for updating the mean and covariance.","marker":"[23]"}],"fun_headline_variants":["CMA-ES-IG ranked most intuitive for teaching robots","Users prefer CMA-ES-IG for adaptive robot learning","CMA-ES-IG makes robot preference learning feel better","New method improves user experience in robot teaching","For assistive robots, CMA-ES-IG wins on ease of use"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim depends on the assumption that the real trajectory closest to a CMA-ES-sampled feature point keeps the sampled query's information gain and quality, so that what users actually rank matches what the algorithm believes it is presenting.","fun_headline_variants_meta":{"raw":{"variants":["CMA-ES-IG ranked most intuitive for teaching robots","Users prefer CMA-ES-IG for adaptive robot learning","CMA-ES-IG makes robot preference learning feel better","New method improves user experience in robot teaching","For assistive robots, CMA-ES-IG wins on ease of use"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000436,"raw_usage":{"total_tokens":2185,"prompt_tokens":882,"completion_tokens":1303,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1221}},"tokens_in":498,"tokens_out":1303,"duration_ms":9472,"temperature":1.0,"reasoning_tokens":1221,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:49:18.425789+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a trajectory dataset with deliberately sparse coverage of the feature space, run the same user study, and measure whether CMA-ES-IG still beats the baselines on perceived behavioral adaptation and ease of use; if its advantage disappears or reverses, the reported effect comes from the retrieval rule or dataset, not from the sampling-and-selection algorithm. Alternatively, in simulation, compute the information gain of the actual nearest-neighbor queries and compare it with the information gain of the raw CMA-ES samples: if they diverge substantially, the algorithm is not delivering the queries it claims to.","supporting_citations":[{"cited_title":"In: Conference on robot learning","cited_arxiv_id":null,"evidence_quote":"Gives the medoid-based batch approximation used to make the information-gain selection computationally tractable."},{"cited_title":"the method of paired comparisons","cited_arxiv_id":null,"evidence_quote":"Defines the Bradley-Terry rational-choice model used to compute ranking probabilities and update the belief over reward weights."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the active preference-based learning formulation and the linear-reward framework that the simulation and belief updates rely on."}],"review_version":1}