{"id":"012ef860-050f-46d7-9827-fc3a5f1ce5e0","arxiv_id":"2412.14982","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An MPC-based motion planner recreates on-road acceleration profiles on a compact test track, and a 47-participant study found no significant subjective motion sickness difference between track and road, though objective exposure was 12% lower.","lead":"This paper develops a model-predictive-control method to reproduce the accelerations of an on-road drive on a small 70 by 175 meter test track, then tests with 47 people whether motion sickness matches. The goal is to let automated-vehicle comfort researchers run repeatable, safe motion sickness experiments off public roads.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The main threat is not unmodeled vertical/roll/pitch (Table III shows their MSDVt contribution is small) but the invalid statistics behind 'no difference': a paired Mann-Whitney U, inconsistent ANOVA df, and no equivalence test despite a significant 12% lower objective exposure.","rationale":"I disagree with the reader's choice of weakest assumption. The measured MSDVz, roll, and pitch differences in Table III are large in percentage but small in absolute terms (MSDVz 0.45 vs 1.52, MSDVphi 0.03 vs 0.07, MSDVtheta 0.04 vs 0.08) compared with MSDVx/y around 7-21, and the paper's combined MSDVt shows only about 1% difference for the generated path and 17% for the measured track data for that participant, with a 12% average reduction across participants. Thus the unmodeled channels are not the decisive threat. The decisive threat is the inference from the human data: the statistical tests are invalid or misreported, and the observed trend is toward lower sickness on the track. A null result from an underpowered, mis-specified test cannot establish the recreation claim; equivalence testing or a confidence interval is required. This concern is consistent with the reader's overall CONDITIONAL verdict, so I recommend no change to that verdict, but the condition should explicitly require reanalysis with proper paired and equivalence tests.","tokens_in":16222,"tokens_out":7805,"duration_ms":50385,"concrete_test":"Obtain the raw paired per-participant MISCmax and MSDV data for both conditions, and rerun the comparison with a Wilcoxon signed-rank test and a two one-sided test (TOST) for equivalence using a pre-registered margin such as ±1 MISC point or ±20% of the on-road mean. Also recompute the repeated-measures ANOVA with correct within-subject degrees of freedom. If the 90% confidence interval for the mean difference excludes the equivalence margin, the claim that the track recreates on-road sickness is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the compact track recreates on-road motion sickness rests on the subjective null result in Section VI.B, and that null result is not statistically grounded. Three concrete problems: (1) The maximum-MISC comparison uses a Mann-Whitney U test on data from a within-subject design; the observations are paired, so the independence assumption is violated and the test has reduced power. (2) The repeated-measures ANOVA is reported as F(2,48)=70.0 for time and F(2,48)=0.79 for condition; with 47 participants and repeated measures, these degrees of freedom are not plausible unless the design was collapsed, and the text does not explain how. (3) No equivalence margin or power analysis is given; 'p=0.53' is treated as proof of no difference, but the mean MISCmax is 15% lower on the track (2.69 vs 2.29) and objective MSDV is about 12% lower and significant (p<.001). The Limitations section acknowledges the lower objective exposure but still asserts the non-significant subjective result, which is exactly the unsupported step. The unmodeled vertical/roll/pitch channels are a secondary issue: Table III indicates their combined contribution to MSDVt is below 10%, so even if these channels are not tracked, the paper's own data suggest they are unlikely to overturn the objective comparison.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an MPC-based motion planner that maps on-road longitudinal and lateral acceleration profiles onto a compact 70 m by 175 m test track, and validates the approach with a 47-participant within-subject experiment comparing subjective motion sickness (MISC) between an on-road drive and a test-track drive. The authors report that the test-track motion closely matches the reference accelerations, that objective MSDV is about 12% lower on the track, and that the subjective MISC response shows no statistically significant difference between conditions. They conclude that the method enables reproducible, safer motion sickness experiments for automated vehicles.","tokens_in":16533,"tokens_out":3474,"duration_ms":32471,"significance":"If the statistical support were solid, this would be a practically useful contribution: it offers a way to run controlled, reproducible motion sickness studies on a small test track instead of public roads, which is valuable for AV development. The experimental design is a genuine strength: 47 participants, within-subject cross-randomized design, a controlled non-driving task, and both objective and subjective outcome measures. The MPC formulation is also not circular in the load-bearing sense, because the human MISC data were not used to tune the controller weights. However, the central claim of \"no difference\" between conditions currently rests on a null result obtained with inappropriate statistical tests and without an equivalence framework, so the main conclusion is not yet established.","major_comments":[{"comment":"The maximum-MISC comparison uses a Mann-Whitney U test on paired within-subject data. Each participant contributes one observation per condition, so the two observations are dependent; the independence assumption of the Mann-Whitney U test is violated, and the test has reduced power to detect a within-subject difference. The reported p=0.53 is therefore not a valid basis for the claim of no difference. A paired test (e.g., Wilcoxon signed-rank test) or a mixed-effects model with participant as a random effect should be reported instead.","section":"Section VI.B, MISCmax comparison"},{"comment":"The reported F(2,48)=70.0 for time and F(2,48)=0.79 for condition are not consistent with a simple 47-participant, two-condition repeated-measures design; the degrees of freedom suggest a collapsed or aggregated analysis, but the text does not state how time points were binned, how missing or aborted trials (MISC=6 terminations) were handled, whether sphericity corrections were applied, or whether the ANOVA was actually performed on condition means across participants. Please specify the exact model, the factors, and the df, and verify that the reported statistics correspond to the design.","section":"Section VI.B, repeated-measures ANOVA"},{"comment":"The central conclusion \"no difference\" is an equivalence claim, but no equivalence margin, confidence interval, or power analysis is provided. The descriptive statistics show a 15% lower mean MISCmax on the track (2.69 vs. 2.29) and a 12% lower objective MSDV that is significant (p<.001, Figure 8). A non-significant p-value cannot support equivalence; the Limitations section repeats the assertion that \"there was no significant difference\" without addressing this. Please report a pre-specified equivalence bound and a confidence interval for the within-subject difference, or reframe the claim as \"no detected difference\" with appropriate caution.","section":"Section VI.B and Section VII (Limitations)"},{"comment":"The paper states in Section VI.A that the contribution of vertical, yaw, roll, and pitch to the total MSDV is less than 10%, which is consistent with Table III and reduces the concern about unmodeled motion channels. However, the same section describes the track motion as \"perfectly aligned\" with the generated path while also reporting a significant 12% reduction in MSDV relative to on-road. These statements should be reconciled, and the 95% confidence interval for the objective MSDV difference should be reported so the reader can judge whether the objective reduction is compatible with the subjective null result.","section":"Section V.C / Table III and Figure 8"}],"minor_comments":[{"comment":"In the cost function, the terms wdδ dδ_k and wdax da_x,k appear without squares; if these are intended as L1 penalties on control rates, please state so, and if they are intended as quadratic penalties, the squares appear to be missing.","section":"Equation (10)"},{"comment":"The notation for the adaptive weights is inconsistent: Eq. (14) defines i = {X,Y}, but Eq. (15) uses j = {x,y}; please harmonize the indices and define i_norm explicitly.","section":"Equations (14)-(17)"},{"comment":"The procedure says the NDRT video contained 20 countable events, but Section VI.B reports counts \"out of 16 in total.\" Please correct this inconsistency.","section":"Section IV.E and Section VI.B"},{"comment":"The individual-level correspondence is described as \"good,\" but the fit has adjusted R²=0.45 with y=a*x; this is a modest correlation. Please temper the wording or add additional agreement metrics (e.g., limits of agreement or ICC).","section":"Figure 10 and surrounding text"},{"comment":"The on-road path in Figure 4a is plotted over a much larger coordinate range than the test-track path in Figure 4b; this is understandable, but the figures would benefit from a shared scale or an explicit note about the different extents.","section":"Section V.A / Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The core idea and experimental effort are valuable, and the paper could become acceptable after a rigorous reanalysis of the subjective outcome. The main issue is not the MPC method but the statistical support for the equivalence claim; the authors should be asked to reanalyze the paired MISC data with appropriate methods and to provide an equivalence margin or, failing that, to soften the central claim to 'no detected difference.'"},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: the method here is real and worth engaging with, but the central claim—that the compact track recreates on-road motion sickness—is not actually established by the statistics as reported. The MPC formulation that maps on-road ax/ay to a 70 by 175 m track is a sensible extension of simulator washout work, and the 47-participant within-subject experiment is a serious effort. The paper deserves peer review, but it needs a statistical rewrite before the conclusion can be trusted.\n\nWhat is new: applying MPC-based motion cueing to a real vehicle on a compact test track, with adaptive boundary weights, and validating it against human MISC responses. That is a useful contribution to the AV motion comfort community. The tracking analysis is honest: the authors show longitudinal acceleration deviations up to about 25%, increased yaw rate, and they quantify the objective MSDV differences. Table III is helpful—it shows vertical, roll, and pitch contribute less than 10% to total MSDV, so the worry about unmodeled channels is a minor issue, not the main threat.\n\nThe soft spots are in Section VI. The maximum-MISC comparison uses a Mann-Whitney U test on paired within-subject data, which violates independence and loses power. The repeated-measures ANOVA is reported as F(2,48)=70.0 for time and F(2,48)=0.79, p=.75 for condition; with 47 participants those degrees of freedom are not explained, and p=.75 does not match an F of 0.79 with those df. The objective MSDV is about 12% lower on the track and significant (p<.001), yet the paper treats p=.53 on MISC as proof of equivalence without an equivalence margin or power analysis. The Limitations section acknowledges the lower objective exposure but still asserts the null result. That is the load-bearing flaw. There is also a small reporting inconsistency: the NDRT count is described as 20 in the procedure and 16 in the results.\n\nThat said, the core method is not circular—the MPC weights are tuned to track accelerations, and the human MISC data are independent. The direction is right. For a motion sickness researcher, this is a useful step toward reproducible test-track studies. With corrected statistics (paired tests, proper ANOVA, and ideally an equivalence test), the conclusion could hold. As written, I would send it to review with major revision required.","headline":"A genuinely useful MPC-based method for recreating on-road motion sickness on a compact track, backed by a serious human experiment, but the 'no difference' claim rests on statistical tests that do not survive scrutiny.","tokens_in":17071,"tokens_out":2419,"would_cite":true,"duration_ms":21617,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A compact 70 m by 175 m test track can reproduce on-road motion sickness: in 47 participants, a model predictive controller that matches recorded longitudinal and lateral accelerations produced statistically indistinguishable subjective…","keywords":["motion sickness","automated vehicles","model predictive control","test track","motion sickness dose value","within-subject experiment","MISC scale","acceleration replication"],"falsifier":"Run the same within-subject protocol on a deliberately rough track with speed bumps, an undulating surface, or a slalom so that vertical, roll, and pitch accelerations differ strongly between track and road while matched longitudinal and lateral MSDV remain similar; if mean MISC or the proportion reaching high MISC differs significantly between conditions, the assumption that only longitudinal and lateral acceleration drive sickness is false.","tokens_in":16049,"feed_emoji":"🚗","tokens_out":6071,"duration_ms":47573,"temperature":0.7,"pith_summary":"Motion sickness experiments for automated vehicles usually require either realistic but unrepeatable public-road drives or large, expensive test tracks. This paper proposes that a model predictive control planner can compress an on-road drive into a 70 m by 175 m test track by tracking only the longitudinal and lateral accelerations of the original ride. In a within-subject trial with 47 participants watching a video, riders reported statistically indistinguishable motion sickness levels on the track and on the road, with mean maximum MISC scores of 2.29 versus 2.69. The paper argues that this makes human-in-the-loop motion sickness assessment simpler, safer, and reproducible, because every participant on the track experiences nearly the same applied motion.","feed_headline":"Compact test track reproduces on-road motion sickness","feed_subtitle":"MPC replays road accelerations at lower speed; riders reported no significant difference in sickness.","key_machinery":"The load-bearing object is an offline model predictive controller built on a linear bicycle model, using steering rate and longitudinal jerk as control inputs. The cost function penalizes squared errors between the generated and recorded longitudinal and lateral accelerations, position offsets from the track centre, and control effort, with adaptive weights that reduce acceleration-tracking priority near the track edges and pull the vehicle back toward the centre. Standstill events from the on-road drive are inserted manually by reconstructing constant-acceleration deceleration and acceleration phases. The objective comparison relies on ISO-2631 frequency-weighted motion sickness dose values computed for the tracked axes.","core_discovery":"The central discovery is that subjective motion sickness, not just objective acceleration statistics, transfers from road to track when the replanned trajectory matches the on-road longitudinal and lateral accelerations. Across 47 participants, repeated-measures ANOVA found that time significantly increased MISC but condition did not (F(2,48)=0.79, p=0.75), and individual maximum MISC values correlated between conditions with fitted slope 1.03 and adjusted R-squared 0.45 (p<0.001). The objective motion sickness dose value was about 12% lower on the track (p<0.001), and unmodeled axes differed substantially (up to 71% for vertical MSDV and about 47 to 62% for roll and pitch), yet these differences did not translate into a detectable subjective difference at the low-to-moderate sickness levels studied. The authors conclude that a compact track can stand in for on-road driving in motion sickness studies.","pith_inferences":["The paper only establishes equivalence for low-to-moderate sickness levels (mean maximum MISC around 2.3 to 2.7 with termination at 6), so extrapolating to severe nausea is an untested extension.","The same MPC replay idea could be extended to include vertical, roll, and pitch reference accelerations if the track and vehicle actuation allow, which would directly test whether the unmodeled axes become important on rougher roads or at higher sickness levels.","The manual standstill insertion is a practical workaround; automating stop-and-go events would make the method a turnkey tool for urban driving scenarios that involve frequent stops.","The moderate individual-level correlation (R-squared 0.45) suggests the method reproduces group averages and extreme responses well, but researchers should still expect residual individual variability from personality or environment, as the authors themselves flag."],"forward_implications":["Motion sickness studies can move from public roads to small controlled areas, removing traffic risk and day-to-day variability.","Because the applied track motion is essentially identical for every participant, within- and between-group comparisons become cleaner and more reproducible.","The reference can be any recorded on-road drive, so the method can replay different road types or candidate automated-vehicle control strategies without building a full route.","The roughly 12% lower objective MSDV on the track is a measurable offset; if subjective equivalence holds across scenarios, it can be accounted for in dose-based protocols.","Early-stage prototype vehicles that lack public-road approval could still be tested for motion comfort during development on a compact track."],"supporting_citations":[{"why":"Defines the ISO-2631 frequency-weighted motion sickness dose values used to objectively compare road and track motion.","marker":"[13]"},{"why":"Supplies the optimization solver that carries out the MPC replanning of the road accelerations.","marker":"[16]"},{"why":"Basis for tracking only longitudinal and lateral accelerations, by showing rotational and vertical body motions contribute little to objective motion sickness.","marker":"[21]"},{"why":"Earlier moving-base simulator comparison that found on-road sickness more severe, motivating a track-based replication instead.","marker":"[25]"},{"why":"Used to screen participants by susceptibility so the sample excludes extremes and is representative for comparisons.","marker":"[29]"},{"why":"Provides the MISC scale used for minute-by-minute subjective motion sickness ratings.","marker":"[30]"},{"why":"Suggests a frequency-splitting MPC direction that the paper identifies for improving longitudinal replication in future work.","marker":"[32]"}],"fun_headline_variants":["Small track recreates on-road motion sickness","MPC test track replicates road sickness","Compact track gives same motion sickness as road","Road nausea reproduced on 70-by-175m track","Test track matches road for motion sickness"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that vertical, roll, pitch, and yaw accelerations contribute so little to motion sickness that matching only longitudinal and lateral accelerations is enough, even though the paper's own measured data show vertical MSDV differs by about 71% and roll and pitch by about 47 to 62% between track and road.","fun_headline_variants_meta":{"raw":{"variants":["Small track recreates on-road motion sickness","MPC test track replicates road sickness","Compact track gives same motion sickness as road","Road nausea reproduced on 70-by-175m track","Test track matches road for motion sickness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1373,"prompt_tokens":972,"completion_tokens":401,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":334}},"tokens_in":588,"tokens_out":401,"duration_ms":3698,"temperature":1.0,"reasoning_tokens":334,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:43:45.983538+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same within-subject protocol on a deliberately rough track with speed bumps, an undulating surface, or a slalom so that vertical, roll, and pitch accelerations differ strongly between track and road while matched longitudinal and lateral MSDV remain similar; if mean MISC or the proportion reaching high MISC differs significantly between conditions, the assumption that only longitudinal and lateral acceleration drive sickness is false.","supporting_citations":[{"cited_title":"Mechanical vibration and shock— evaluation of human exposure to whole body vibration. part 1: General requirements,","cited_arxiv_id":null,"evidence_quote":"Defines the ISO-2631 frequency-weighted motion sickness dose values used to objectively compare road and track motion."},{"cited_title":"Forces nlp: an eﬀicient implementation of interior-point methods for multi- stage nonlinear nonconvex programs,","cited_arxiv_id":null,"evidence_quote":"Supplies the optimization solver that carries out the MPC replanning of the road accelerations."},{"cited_title":"The impact of body and head dynamics on motion comfort assessment","cited_arxiv_id":"2307.03608","evidence_quote":"Basis for tracking only longitudinal and lateral accelerations, by showing rotational and vertical body motions contribute little to objective motion sickness."},{"cited_title":"Validation of a moving base driving simulator for motion sickness research,","cited_arxiv_id":null,"evidence_quote":"Earlier moving-base simulator comparison that found on-road sickness more severe, motivating a track-based replication instead."},{"cited_title":"Evaluation multimodaler physiologischer merkmale zur objektiven detektion von kinetose im pkw,","cited_arxiv_id":null,"evidence_quote":"Used to screen participants by susceptibility so the sample excludes extremes and is representative for comparisons."},{"cited_title":"How feelings of unpleasantness develop during the progression of motion sickness symptoms,","cited_arxiv_id":null,"evidence_quote":"Provides the MISC scale used for minute-by-minute subjective motion sickness ratings."},{"cited_title":"Motion Cueing Algorithm for Effective Motion Perception: A frequency-splitting MPC Approach","cited_arxiv_id":"2309.01689","evidence_quote":"Suggests a frequency-splitting MPC direction that the paper identifies for improving longitudinal replication in future work."}],"review_version":1}