{"id":"1f0b2b6a-5c9d-4e79-8651-9265799b34c6","arxiv_id":"2501.11829","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Motion cues in VR air taxi simulations reduce passengers' trust, understanding, and acceptance, while leaving most optimized interface parameters unchanged.","lead":"This study tested whether adding physical motion to a VR-simulated air taxi flight changes passengers' trust and their preferred interface. It found that motion cues lower trust, understanding, and acceptance, while changing only a few interface choices. A generalist might read it to see how much simulation realism matters when designing and testing interfaces for future air taxis.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Non-independence of Pareto-front ratings inflates the reported Bayes factors; the 'extreme evidence' claim is not established by the analysis as presented.","rationale":"The reader's weakest assumption — non-independence of Pareto-front observations nested within participants — is the most load-bearing concern, and I agree it should be fixed before the extreme-evidence claims are accepted. The paper's central contribution is the claim that motion fidelity decreases trust, understanding, and acceptance. That claim rests on three Bayes factors computed from pooled Pareto points (Section 3.7.4), with each participant contributing multiple points from one adaptive trajectory. The pooled analysis has two distinct problems: (1) it treats repeated within-participant measurements as independent, inflating effective sample size; and (2) because the optimizer adapts the UI to the participant's own ratings over 30 runs, successive ratings are serially correlated and the Pareto points are not exchangeable draws. A participant-level or multilevel analysis addresses both. I also agree with the reader's secondary concern about the confound between motion and optimized UI parameters: the two groups' optimized designs differ (boundary box, chevron size), so the rating comparison mixes the effect of motion with the effect of the UI the participant ended up evaluating. However, I would not call this an alternative weakest assumption; it is a consequence of the same between-subjects optimization design and is partially acknowledged in the discussion. The paper has real strengths: a concrete, reproducible apparatus; a clear between-subjects manipulation; a plausible direction of effect; and honest reporting of null results for immersion. The concern is not that the effect is fabricated but that the reported strength of evidence is not supported by the statistical procedure as described, and the preprint itself does not provide the participant-level data needed to verify the claim. The proposed test is feasible with the raw data and would definitively settle whether the evidence survives the appropriate unit of analysis. The verdict should remain CONDITIONAL rather than REJECT because the issue is fixable and the qualitative direction of the finding may survive the reanalysis; but the 'extreme evidence' language should be revised if the participant-level BF does not replicate.","tokens_in":23296,"tokens_out":2061,"duration_ms":20649,"concrete_test":"Re-run the Section 3.7.4 comparisons using participant-level aggregation: compute each participant's mean rating over their Pareto-front points, then run a Bayesian t-test or a multilevel model with random intercepts for participants on the resulting n=20 vs n=20 values. If the trust Bayes factor drops from 3970 to below 30 (or the credible interval crosses zero), the extreme-evidence claim is an artifact of pooling and should be downgraded. As a secondary check, recompute the boundary-box comparison using participant-level proportions of Pareto points with value >= 0.5 rather than pooled means; if the difference is no longer extreme, the binary threshold artifact is confirmed.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (Section 4.1: \"strong to extreme evidence that motion fidelity reduces users' trust, understanding, and acceptance\") rests on independent Bayesian t-tests over pooled Pareto-front observations (Section 3.7.4), with n=48 and n=42 Pareto points rather than n=20 participants per group. Each participant contributed multiple Pareto points from a single 30-run adaptive loop in which the UI was updated based on that participant's own ratings, so observations are nested within participants and serially correlated across runs. Treating them as independent overstates the effective sample size and can massively inflate Bayes factors, exactly where the paper reports BF=3970 (trust), BF=68 (acceptance), and BF=57 (understanding). The same pooled-Pareto-point procedure underlies the headline design-parameter differences (Other Chevron Size BF=16.5 and Boundary Box BF=32474 in Table 1), compounded by the binary 0.5 threshold: the no-Motion boundary-box median (0.45) is effectively on the threshold and its IQR crosses it, so the extreme BF for a binary parameter is especially sensitive to how Pareto-selected values are pooled. A participant-level analysis or a multilevel Bayes factor is required before the extreme-evidence language can be taken at face value. A related confound: because the optimized UI parameters differ between groups (boundary box, chevron size), the motion-vs-no-motion comparison of ratings is not a pure manipulation of motion; the effect could be partly mediated by the different UIs the optimizer converged to, as the paper's own discussion of boundary boxes acknowledges.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a between-subjects VR study (N=40, n=20 per group) that uses multi-objective Bayesian optimization (MOBO) to personalize a 12-parameter air-taxi interface for each participant, with one group experiencing 3-DoF motion cues via a chair and the other VR only. The authors analyze only the Pareto-front designs extracted per participant, then use independent Bayesian t-tests to compare the two groups on six subjective objectives and on each design parameter. They report strong to extreme evidence that motion fidelity lowers trust, understanding, and acceptance, and strong/extreme evidence of group differences in other-chevron size and boundary-box visibility, while interpreting most remaining parameter differences as evidence for personalization rather than group effects.","tokens_in":23510,"tokens_out":3751,"duration_ms":39527,"significance":"If the central claim is supported, the result is practically important for UAM simulation methodology: it would indicate that motion fidelity systematically lowers subjective ratings and that such fidelity should be considered when generalizing from VR-only studies to real flights. The human-in-the-loop MOBO pipeline is a novel and timely methodological contribution, and the paper is transparent about many of its design choices. However, the strength of the central claim rests on a statistical analysis whose assumptions about independence are not met, and the group comparison is confounded by the very design parameters the optimization produced. The evidentiary basis therefore needs reworking before the headline conclusions can be accepted.","major_comments":[{"comment":"The central claim of strong to extreme evidence (BF=3970.14 for trust, 67.92 for acceptance, 56.85 for understanding) rests on independent Bayesian t-tests applied to pooled Pareto-front observations. Section 3.7.1 reports n=48 and n=42 Pareto points from only 20 participants per group, meaning each participant contributes multiple observations from a single 30-run adaptive loop in which the UI was updated based on that participant's own ratings. These observations are nested within participants and serially correlated, so the effective sample size is smaller than the reported number of Pareto points. Treating them as independent can substantially inflate Bayes factors. A participant-level summary (e.g., one value per participant, such as the mean or median over that participant's Pareto front) or a multilevel Bayes factor is required before the 'strong to extreme evidence' language in Section 4.1 can be taken at face value.","section":"3.7.1 and 3.7.4"},{"comment":"The group comparison of questionnaire ratings is confounded with the optimized UI design. Table 1 reports strong evidence of a difference in Other Chevron Size and extreme evidence of a difference in Boundary Box between the motion and no-motion groups. Because the participants in the two groups rated different UI configurations, the observed differences in trust, understanding, and acceptance are not attributable solely to motion fidelity; they could partly reflect the differing chevron sizes or boundary-box visibility. This is especially important because Section 4.2 argues that boundary boxes were preferred in the motion condition as a compensatory response to reduced trust. The paper should either acknowledge and discuss this confound explicitly or provide an analysis that controls for the design configuration (for example, by comparing ratings on a common, non-personalized design).","section":"3.7.3 and 4.1"},{"comment":"The extreme Bayes factor for Boundary Box (BF=32473.69) is not robust. Boundary Box is a binary parameter defined by a 0.5 threshold on a continuous MOBO value, and the no-motion median is 0.45 with an IQR of (0.35, 0.63), so the median lies on one side of the threshold while the IQR crosses it. The reported extreme evidence is therefore heavily dependent on the arbitrary threshold and on pooling multiple Pareto points per participant. A sensitivity analysis around the threshold, or a model that treats the binary parameter directly rather than thresholding the pooled continuous values, is needed before this parameter can be claimed as an extreme group difference.","section":"Table 1 and Section 3.7.3"}],"minor_comments":[{"comment":"The abstract states that 'minimal evidence was found for differences or equality in the optimized interface designs,' but Table 1 reports strong evidence for Other Chevron Size and extreme evidence for Boundary Box; the wording should be aligned with the actual Bayes factors.","section":"Abstract and Section 4.2"},{"comment":"The descriptions of the binary parameters contain a copy-paste error: for 'Additional Information on Display', the parenthetical says '(< 0.5 no boundary box' instead of 'no additional information'; the same style of error appears in the description of the 'Map Display' and 'Boundary Box' parameters.","section":"Section 3.3"},{"comment":"The text reports means and SDs for the two groups, while Table 1 reports medians and IQRs; please clarify which summary is used for the Bayesian t-tests and keep the presentation consistent.","section":"Section 3.7.3"},{"comment":"The trust panel is annotated with 'BF > 100' rather than the exact reported value of 3970.14; using the exact value (or a consistent descriptive label) would improve accuracy.","section":"Figure 6(a)"},{"comment":"No data or analysis code availability statement is provided. Given that the analysis is central to the claims, making the de-identified data and analysis scripts available would strengthen reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a promising empirical contribution with an interesting methodological angle, but the statistical analysis currently overstates the evidence for the main claim. The non-independence of Pareto-front observations and the confound between motion fidelity and optimized UI parameters are load-bearing issues that can be addressed by reanalysis (for example, participant-level summaries or a multilevel model, plus a sensitivity analysis for the binary threshold). These fixes are within the scope of a revision rather than requiring a new study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the Fly Away paper. What you should know: it's a genuine first—nobody has combined multi-objective Bayesian optimization with a motion-fidelity manipulation in simulated urban air mobility. The empirical result that motion cues lower trust, understanding, and acceptance is plausible and worth taking seriously, but the \"extreme evidence\" framing (BF=3970 for trust) is not supported by the analysis as presented.\n\nWhat's good: the study is carefully run. N=40, between-subjects, 30 optimization runs per participant, validated questionnaires for trust, understanding, mental demand, and perceived safety. The MOBO setup with BoTorch and qEHVI is standard and competently described. The paper is honest about its own limitations—it flags the immersion null, the lack of simulator sickness measures, and the redundancy of objectives. The finding that most UI parameters did not differ between motion conditions, while two did (boundary box and other chevron size), is a useful nuance. The practical guidance—VR suffices for relative UI comparisons, motion matters for absolute ratings—is actionable.\n\nThe soft spots are real but fixable. The Bayesian t-tests in Section 3.7.4 pool Pareto-front observations across runs and participants (n=48 vs n=42), but each participant contributed multiple Pareto points from a single 30-run adaptive loop. Those observations are nested and serially correlated. Treating them as independent inflates the effective sample size and can massively inflate Bayes factors—exactly where the paper reports its most dramatic numbers. A participant-level aggregation or multilevel Bayes factor is needed before the extreme-evidence language is warranted.\n\nSecond, the group comparison is not a pure motion manipulation. Because the optimizer adapted the UI to each participant, the motion and no-motion groups ended up with different optimized UIs on two parameters. The rating differences could be partly mediated by those design differences, not just by motion. The paper acknowledges this in the discussion of boundary boxes, but the causal claim in Section 4.1 is stated more strongly than the design supports.\n\nThird, the binary threshold at 0.5 drives one of the headline design differences. The no-motion boundary-box median is 0.45 with an IQR crossing the threshold; the extreme BF=32473 is especially sensitive to how Pareto-selected values are pooled and thresholded. A sensitivity analysis on the threshold would strengthen the claim.\n\nWho is this for? Researchers in UAM HCI and anyone designing simulator-based studies of automated vehicles. It deserves a serious referee—the combination is novel and the empirical groundwork is solid—but the statistical analysis needs revision before the central claim is taken at face value. If I were handling it, I'd send it to review and ask for a revision addressing non-independence and the confound. Worth citing for the MOBO-plus-motion setup, not for the BF numbers.","headline":"A genuine first combining MOBO with a motion-fidelity manipulation in simulated air taxis, but the headline Bayes factors rest on non-independent Pareto-front observations and a group comparison partly confounded by the optimized UIs themselves.","tokens_in":24121,"tokens_out":3015,"would_cite":true,"duration_ms":30619,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding physical motion to a VR air-taxi simulation makes passengers report lower trust, understanding, and acceptance of the interface.","keywords":["urban air mobility","virtual reality","motion fidelity","multi-objective Bayesian optimization","trust in automation","air taxi interface","Pareto front","human-in-the-loop optimization"],"falsifier":"Run the same procedure with each participant rating one randomly assigned UI design exactly once, with no adaptive optimization, and compare motion versus no-motion; if trust ratings no longer separate the groups, the reported effect depends on the adaptive loop rather than on motion itself. Alternatively, re-analyze the Pareto-front ratings with a model that treats participant as a random effect; if the trust Bayes factor falls below the extreme threshold, the independence assumption carried the result.","tokens_in":1576,"feed_emoji":"🛩️","tokens_out":1573,"duration_ms":72573,"temperature":0.7,"pith_summary":"This paper asks whether the physical motion fidelity of a simulator changes what passengers feel and what interface they prefer in automated air taxis. The authors ran 40 participants through a VR air-taxi flight, half with a three-degree-of-freedom motion chair and half without, while a Bayesian optimizer repeatedly adjusted twelve interface parameters to maximize six subjective goals. They report strong-to-extreme Bayesian evidence that adding motion lowers trust, understanding, and acceptance, and moderate evidence it lowers perceived safety and aesthetics. The optimized interfaces themselves differed mainly in two parameters: passengers who felt motion preferred smaller chevrons on other air taxis and wanted boundary boxes around them, while passengers without motion preferred larger chevrons and no boxes. If this holds, UAM researchers need motion cues when they want realistic absolute ratings, but simple VR may suffice for comparing interface variants.","feed_headline":"Motion in VR air-taxi sims cuts passenger trust","feed_subtitle":"Adding a motion chair lowered trust, understanding, and acceptance in a VR air-taxi study.","key_machinery":"The central mechanism is Multi-Objective Bayesian Optimization (MOBO), a human-in-the-loop procedure that maps twelve design parameters (trajectory lengths, transparencies, chevron sizes, map and box toggles) to six subjective objectives (trust, perceived safety, mental demand, understanding, acceptance, and aesthetics), and proposes the design expected to improve the Pareto trade-off most. Each participant went through 30 iterations, and only the non-dominated Pareto-optimal designs were analyzed. The comparison across motion conditions then rests on independent Bayesian t-tests whose Bayes factors quantify evidence for difference versus equality.","core_discovery":"The paper claims that adding physical motion cues to a VR air-taxi simulation does not simply make the experience feel more real; it changes the user's subjective response to the interface. On Pareto-optimal designs, the motion group rated trust lower (BF = 3970.14), understanding lower (BF = 56.85), and acceptance lower (BF = 67.92) than the no-motion group. The optimized interfaces differed between the groups mainly in two ways: without motion, participants preferred larger chevrons marking other air taxis and no boundary boxes; with motion, they preferred smaller chevrons and boundary boxes around other air taxis. Other design parameters showed no strong group difference, and the authors interpret the large spread of Pareto-front designs as evidence against a single best air-taxi interface.","pith_inferences":["One plausible reading, not tested by the authors: the motion-driven drop in trust may be a familiarity effect, since no participant had flown in an air taxi; repeated exposure to the motion condition could attenuate the gap.","The comparison pools Pareto-front points across participants even though each participant's optimization loop personalized the interface, so the group-level optimized UI is a statistical construct; a design that is best for the average may not be best for any given passenger.","A testable extension would vary motion amplitude (none, 3-DoF, 6-DoF) to see whether trust decreases monotonically with fidelity or shows a threshold.","Because boundary boxes appeared compensatory in the motion condition, a design implication is that trust-reducing contexts may call for reassurance elements that are unnecessary in calmer conditions."],"forward_implications":["Studies that need realistic absolute ratings of air-taxi interfaces should include motion cues, because no-motion ratings overstate trust, understanding, and acceptance.","Comparative UI studies can probably use VR or even monitor setups without motion, since motion changed only two of twelve optimized design parameters.","Boundary boxes around other air taxis should be included in air-taxi simulations regardless of motion, because the motion condition preferred them.","A one-size-fits-all optimized interface is unlikely to satisfy passengers; personalization appears necessary.","Reducing the objective set from six to fewer dimensions (for instance, dropping understanding because it tracks trust, or aesthetics because it tracks acceptance) would make future optimization simpler."],"supporting_citations":[{"why":"Supplies the earlier finding that path visualizations increase trust and perceived safety, and provides the chevron and boundary-box design vocabulary.","marker":"[17]"},{"why":"Provides the validated trust-in-automation questionnaire used to measure both trust and understanding.","marker":"[42]"},{"why":"Compares simulators across fidelity levels and motivates the hypothesis that motion fidelity changes user perception in VR.","marker":"[83]"},{"why":"Supplies the human-in-the-loop Bayesian optimization procedure that the study adapts to air-taxi UI design.","marker":"[10]"},{"why":"Provides the optimization software implementation used to run the multi-objective Bayesian optimizer.","marker":"[4]"},{"why":"Provides the trust-in-automation theory used to interpret why UI improvements cannot fully compensate for contextual trust deficits.","marker":"[45]"},{"why":"Shows that path visualizations beat alternatives across flight phases and visibility conditions, informing the choice of design parameters.","marker":"[76]"},{"why":"Supplies the Bayes-factor interpretation thresholds used to label evidence as moderate, strong, or extreme.","marker":"[46]"},{"why":"Provides the NASA-TLX mental workload subscale used to measure mental demand.","marker":"[30]"},{"why":"Supplies the acceptance scale items adapted to measure acceptance of the visualizations.","marker":"[77]"}],"fun_headline_variants":["Motion chair in air-taxi VR lowers trust","VR air-taxi sims: motion reduces trust","Physical motion hurts trust in VR taxi study","Air-taxi VR: motion changes interface needs","Less trust with motion in VR air-taxi test"],"cache_read_input_tokens":26240,"weakest_assumption_plain":"The headline evidence counts each Pareto-optimal rating as an independent data point, but all ratings from one participant came from a single session in which the interface kept adapting to that person's own preferences.","fun_headline_variants_meta":{"raw":{"variants":["Motion chair in air-taxi VR lowers trust","VR air-taxi sims: motion reduces trust","Physical motion hurts trust in VR taxi study","Air-taxi VR: motion changes interface needs","Less trust with motion in VR air-taxi test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000147,"raw_usage":{"total_tokens":1151,"prompt_tokens":879,"completion_tokens":272,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":200}},"tokens_in":495,"tokens_out":272,"duration_ms":2987,"temperature":1.0,"reasoning_tokens":200,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:48:26.298069+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same procedure with each participant rating one randomly assigned UI design exactly once, with no adaptive optimization, and compare motion versus no-motion; if trust ratings no longer separate the groups, the reported effect depends on the adaptive loop rather than on motion itself. Alternatively, re-analyze the Pareto-front ratings with a model that treats participant as a random effect; if the trust Bayes factor falls below the extreme threshold, the independence assumption carried the result.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the validated trust-in-automation questionnaire used to measure both trust and understanding."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that path visualizations beat alternatives across flight phases and visibility conditions, informing the choice of design parameters."},{"cited_title":"Lee and E.J","cited_arxiv_id":null,"evidence_quote":"Supplies the Bayes-factor interpretation thresholds used to label evidence as moderate, strong, or extreme."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the NASA-TLX mental workload subscale used to measure mental demand."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the acceptance scale items adapted to measure acceptance of the visualizations."}],"review_version":1}