REVIEW 3 major objections 2 minor 3 references
Symbolic machine learning ranks rotation and acceleration as top cybersickness triggers in VR flight games compared to racing games.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-27 08:17 UTC pith:B372332A
load-bearing objection With n=37 and no reported validation, the symbolic ML rankings of rotation and acceleration as cybersickness triggers are likely unstable associations rather than solid causal factors. the 3 major comments →
Identifying cybersickness causes in virtual reality games using symbolic machine learning algorithms
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The authors claim that symbolic machine learning applied to participant responses from a flight game and a race game can estimate and rank the impact of potential cybersickness causes, showing rotation and acceleration as more frequent triggers in the flight game, less experienced VR users as more prone overall with the effect stronger in the race game due to its greater user control, and distinct sets of causes arising under short-term versus long-term VR exposures.
What carries the argument
Symbolic machine learning algorithms applied to user response data to estimate and rank the relative impact of factors such as rotation, acceleration, and prior VR experience on cybersickness.
Load-bearing premise
The 37 valid samples from two games allow symbolic machine learning to identify true causal factors of cybersickness rather than correlations, and the findings extend beyond these specific titles and protocols.
What would settle it
Re-running the protocols with a larger participant pool or additional game types and obtaining substantially different ranked causes for cybersickness would show the original rankings do not hold.
If this is right
- Rotation and acceleration should be limited in flight-style VR games to reduce the frequency of cybersickness.
- Racing-style VR games require designs that accommodate users with varying levels of prior experience to limit discomfort.
- Mitigation approaches must be adjusted separately for short-term and long-term VR exposure scenarios.
- Symbolic machine learning offers a method to compare and prioritize causes across different VR game mechanics.
Where Pith is reading between the lines
- The same ranking technique could be tested on VR applications outside entertainment, such as training or education, to identify context-specific triggers.
- If the game-type differences persist, VR systems might benefit from built-in adjustments that scale movement based on detected user experience and session duration.
- The emphasis on controller freedom in the race game suggests that greater user agency in movement options can shift which physical factors dominate discomfort.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports an experimental study applying symbolic machine learning (decision lists and rule induction) to data from 37 valid participants (out of 88 volunteers) who played a flight game and a race game under six protocols. It claims to rank causes of cybersickness, concluding that rotation and acceleration trigger symptoms more frequently in the flight game than the race game, that lower VR experience increases susceptibility (especially in the race game), and that distinct causes emerge for short- versus long-term exposures, with suggested mitigation strategies.
Significance. If the symbolic rankings prove stable and the causal interpretation holds, the work could inform VR game design guidelines for reducing cybersickness. The use of symbolic methods offers potential interpretability advantages over black-box models, and the comparison across game types and exposure durations addresses a practical question in HCI. However, the small effective sample and absence of validation details limit immediate impact.
major comments (3)
- [Abstract and Results] Abstract and Results section: the central ranking of rotation/acceleration as more frequent triggers in the flight game (versus race game) is derived from symbolic ML on only 37 valid samples with no reported cross-validation, permutation testing, or stability analysis across participant splits; without these, the rankings may reflect sampling artifacts rather than reliable associations.
- [Methods and Results] Methods and Results: the leap from learned rules to 'causes that trigger cybersickness' and 'more prone to feel discomfort' is unsupported by any causal identification strategy, confounding controls, or comparison against classical ML feature importances with statistical tests; the paper provides no evidence that the symbolic output distinguishes causation from correlation.
- [Methods] Participant recruitment: the reduction from 88 volunteers to 37 valid samples is not accompanied by analysis of selection bias, exclusion criteria effects, or power calculations, which directly affects the generalizability claim for experience level and exposure duration effects.
minor comments (2)
- [Methods] Clarify the exact symbolic algorithms employed (e.g., specific decision-list learner or rule-induction method) and any hyperparameter choices.
- [Results] Add a table or figure showing the top-ranked rules with their coverage and accuracy metrics on the 37 samples.
Simulated Author's Rebuttal
We thank the referee for the thoughtful and constructive comments, which have helped us clarify the scope and limitations of our work. We address each major comment below and indicate where revisions have been made to the manuscript.
read point-by-point responses
-
Referee: [Abstract and Results] Abstract and Results section: the central ranking of rotation/acceleration as more frequent triggers in the flight game (versus race game) is derived from symbolic ML on only 37 valid samples with no reported cross-validation, permutation testing, or stability analysis across participant splits; without these, the rankings may reflect sampling artifacts rather than reliable associations.
Authors: We acknowledge that the absence of explicit stability checks leaves the rankings vulnerable to sampling variability. Symbolic rule induction was chosen precisely because it can surface interpretable patterns from modest sample sizes, yet we agree that additional safeguards are warranted. In the revised manuscript we have added (i) 5-fold cross-validation of the induced rule sets and (ii) a bootstrap stability analysis across 100 random participant splits, reporting the frequency with which each top-ranked condition (rotation, acceleration, prior experience) reappears. These results are now summarized in a new subsection of Results and the corresponding tables have been updated. revision: yes
-
Referee: [Methods and Results] Methods and Results: the leap from learned rules to 'causes that trigger cybersickness' and 'more prone to feel discomfort' is unsupported by any causal identification strategy, confounding controls, or comparison against classical ML feature importances with statistical tests; the paper provides no evidence that the symbolic output distinguishes causation from correlation.
Authors: The referee correctly identifies that our observational design cannot support causal claims. The original wording occasionally used “causes” and “trigger” in a colloquial sense; this was imprecise. We have revised the manuscript throughout to replace such language with “associated factors,” “predictive indicators,” and “frequently co-occurring conditions.” A new paragraph in the Discussion explicitly states the correlational nature of the findings, notes the lack of randomized controls or instrumental-variable strategies, and contrasts the symbolic rules with permutation-based feature importances obtained from a random-forest baseline (now reported in an appendix). No causal identification is claimed or performed. revision: yes
-
Referee: [Methods] Participant recruitment: the reduction from 88 volunteers to 37 valid samples is not accompanied by analysis of selection bias, exclusion criteria effects, or power calculations, which directly affects the generalizability claim for experience level and exposure duration effects.
Authors: We have expanded the Methods section with a full accounting of the 51 exclusions (motion-sickness history, incomplete sessions, technical failures, and voluntary withdrawal) together with a table comparing demographic and baseline SSQ scores between included and excluded participants. A brief discussion of possible selection effects on the experience-level findings has been added. Because the study is retrospective, formal a-priori power calculations cannot be supplied; we now report post-hoc observed power for the main contrasts and list the modest sample size as an explicit limitation in both the Abstract and Discussion. revision: partial
Circularity Check
Empirical ML ranking study with no circular derivations
full rationale
The paper is an observational study applying symbolic machine learning (decision lists, rule induction) to 37 valid participant samples across two VR games and six protocols. No equations, parameter-fitting derivations, or self-citation chains appear in the provided text; claims about rotation/acceleration frequency, experience level, and exposure duration are presented as direct outputs of the learned rules on the collected data. The derivation chain is therefore self-contained and does not reduce any result to its inputs by construction.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption Symbolic machine learning algorithms can identify and rank causal factors of cybersickness from user-reported data
read the original abstract
Virtual reality (VR) and head-mounted displays are constantly gaining popularity in various fields such as education, military, entertainment, and health. Although such technologies provide a high sense of immersion, they can also trigger symptoms of discomfort. This condition is called cybersickness (CS) and is quite popular in recent virtual reality publications. This work proposes a novel experimental analysis using symbolic machine learning to rank potential causes of CS in VR games. We estimate CS causes and rank them according to their impact using classical machine learning. Experiments are performed using two virtual reality games and 6 experimental protocols along with 37 valid samples from a total of 88 volunteers. Our results show that rotation and acceleration triggered cybersickness more frequently in a flight game in contrast to a race game. We could also observe that subjects that are less experienced with VR are more prone to feel discomfort. Former experience plays a more important role on the race game, as this game provides more liberty to the user in terms of controllers, more displacement alternatives and a more user-controlled acceleration. Furthermore, different causes that trigger discomfort arise based on short or long term VR exposures. We suggest strategies for mitigating CS for these two scenarios: short and long term exposure experiences and compare the two highlighted scenarios (race and flight).
Figures
Reference graph
Works this paper leans on
-
[1]
M. Calvelo, Á. Pifieiro, R. Garcia-Fandino, An immersive journey to the molecular structure of sars-cov-2: Virtual reality in covid-19, Comput. Struct. Biotechnol. J. A. Statista, The statistics portal, Web site: https://www.statista.com/statistics/ 591181/global-augmented-virtual-reality-market-size/. B.G. Studios, The elder scrolls v: Skyrim, Bethesda G...
work page internal anchor Pith review Pith/arXiv arXiv 2015
-
[2]
Padmanaban, R
N. Padmanaban, R. Konrad, T. Stramer, E.A. Cooper, G. Wetzstein, Optimizing virtual reality for all users through gaze-contingent and adaptive focus displays, Proceedings of the National Academy of Sciences (2017) 201617251, J. Van Waveren, The asynchronous time warp for virtual reality on consumer hardware, in: Proc, in: 22nd ACM Conference on Virtual Re...
2017
-
[3]
Rodrigues, A
E. Rodrigues, A. Conci, L. Panos, Morphological classifiers, Pattern Recogn. 84 (2018) 82-96. E. Frank, M.A. Hall, LH. Witten, Data Mining: Practical Machine Learning Tools and Techniques, 4th Edition, Morgan Kaufmann, 2016. P. Flach, Machine Learning: The Art and Science of Algorithms That Make Sense of Data, Cambridge, 2012, F. Pedregosa, G. Varoquaux, ...
2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.