{"id":"903ad6ad-65fe-4709-926b-ed2e55c8bdae","arxiv_id":"2411.13327","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Reinforcement-learning fine-tuning on gameplay data improved EMG-based finger-movement decoding in 15 able-bodied users, with a 39% gain in a separate motion test.","lead":"The authors fine-tune a pretrained EMG hand-movement classifier using reinforcement learning on muscle data recorded while people play a guitar-hero style game. The fine-tuned policy roughly doubled game accuracy and improved a separate motion test by 39% across 15 able-bodied participants.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that RL (not just extra gameplay data and human adaptation) drives the gains is untested: Section 7.1 admits the game and RL contributions cannot be separated, and no supervised-fine-tuning control on the same gameplay data is run.","rationale":"The reader's CONDITIONAL verdict is appropriate. I agree with the conditionality but locate the load-bearing gap differently: the reader named the Eq. (4) ideal-intention/label-alignment assumption as the weakest premise; I see that as acknowledged in the text and empirically supported by the independent Motion Test. The more serious gap is that no control separates the AWAC update from the mere availability of additional gameplay data and human co-adaptation. The paper itself flags this in Section 7.1. Since the paired design, randomized Motion Test order, and independent evaluation are solid evidence for the combined intervention, the result should not be rejected; but the headline causal attribution to RL is not yet isolated. A supervised fine-tuning control on the same dataset is the minimal experiment that would settle it. Hence no verdict change from conditional acceptance.","tokens_in":19259,"tokens_out":5750,"duration_ms":74477,"concrete_test":"Add a matched control condition (same n, same pretrained π0, same accumulated Dn and model-selection protocol) in which the classifier is fine-tuned supervisedly on the song labels a*_t from the same gameplay data, using the same loss, architecture, and model selection as pretraining. Compare the supervised-fine-tuned policy against π8 on the independent Motion Test EMR. If the supervised arm reproduces a statistically indistinguishable improvement over π0, the RL-specific attribution fails; if π8 is significantly better, the 'RL' element is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing unsecured premise is the attribution of the observed EMR gains to the AWAC/RL update. The experimental design compares π0 (pretrained on D0 only) with π8 (trained on D0 plus eight gameplay episodes via AWAC). These differ by both the added usage data and the learning rule, so the paired comparison establishes only that the combined game-plus-RL intervention improves decoding. Section 7.1 says explicitly: 'A limitation of our study is that we are unable to independently assess the distinct contributions of our game environment and RL training procedure.' Without a supervised fine-tuning (or behavior-cloning) arm on the same accumulated dataset Dn, the same Motion Test EMR gain could plausibly be obtained by simply retraining the classifier on the song-labeled gameplay data; the paper's conclusion that RL is the effective ingredient is therefore not yet supported. The Eq. (4) ideal-intention assumption is acknowledged and empirically plausible, but it is secondary here: even if song labels are occasionally wrong, the independent Motion Test still shows the combined pipeline works. The missing control targets the stronger, causal 'through RL' claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes an approach to myoelectric control in which a supervised pretrained EMG classifier is fine-tuned with the AWAC offline RL algorithm using data collected while the user plays a Guitar Hero-style game. The game's song provides movement labels that define a reward function (Eq. 4), and the policy is iteratively fine-tuned over eight repeated play-throughs of the same song. In a 15-subject real-time human-in-the-loop experiment, the final RL policy π8 significantly outperforms the initial supervised policy π0 in gameplay EMR (0.36 vs. 0.78), gameplay F1 macro (0.55 vs. 0.75), Motion Test EMR (0.43 vs. 0.60), and Motion Test F1 macro (0.53 vs. 0.71), all by Wilcoxon signed-rank tests. The authors additionally analyze mutual information and population stability to explain inter-subject variability and to identify outlier participants. The central claim is that RL fine-tuning on usage-based data improves decoding accuracy and robustness for simultaneous finger movement control.","tokens_in":19477,"tokens_out":11606,"duration_ms":125879,"significance":"The paper addresses an important problem in myoelectric control: reducing the labeled-data burden and closing the offline-online gap. Its strengths include a real-time human-in-the-loop evaluation with 15 participants, a paired design with randomized test order for the Motion Test, the reuse of the initial policy after training (repetition 9) to separate human adaptation from policy improvement, and an independent Motion Test that is not used during training or model selection. The MI-based analysis yields a falsifiable diagnostic, namely a low-MI interval [0.28, 0.35] below which fine-tuning appears to fail. The empirical result that the combined game-plus-RL fine-tuning procedure improves decoding is credible and well supported by the Motion Test data. However, the stronger claim that RL specifically, rather than the addition of usage data under any fine-tuning rule, is the effective ingredient, is not established by the present design, and the abstract and title attribute the gains to RL. The authors are explicit about this gap in Section 7.1, which is commendable, but the attribution remains load-bearing for the paper's stated conclusion.","major_comments":[{"comment":"The experimental design cannot support the claim that RL is the cause of the observed improvements. The paired comparison of π0 and π8 varies two factors simultaneously: the training data (D0 versus D0∪...∪D8) and the learning rule (supervised RMSE minimization versus AWAC). Section 7.1 explicitly states that the distinct contributions of the game environment and the RL training procedure cannot be independently assessed. Because the title and abstract attribute the gains specifically to reinforcement learning, a control arm is needed in which the same accumulated gameplay dataset Dn is used to fine-tune the pretrained policy with supervised learning or behavior cloning. Without such a control, the Motion Test results support only the combined intervention, not the causal role of RL. I consider this the key missing experiment: it is feasible within the same experimental setup and would directly resolve whether the 'through RL' claim is warranted.","section":"Sections 5 and 7.1; Table 1"},{"comment":"The gameplay results for π8 are evaluated on data that were used for both training and model selection. Section 6.3 selects the deployed policy as the one with the highest simulated episodic return computed on the recorded game data accumulated so far, and the repetition-8 gameplay, whose EMR and normalized return are reported in Table 1 and Fig. 5, is part of that same dataset. The gameplay EMR for π8 is therefore optimistically biased, and the headline 'more than two-fold increase in decoding accuracy during gameplay' is not a clean out-of-sample measurement. The Motion Test is the only fully independent evaluation, as it is a different task not used in training or selection, and the paper should present it as the primary evidence for the method. At minimum, the gameplay evaluation should be performed on a held-out song or on a post-training replay that does not enter the training set; the Appendix A remark that using training data for model selection is 'not ideal' understates this concern.","section":"Section 6.3 and Table 1"},{"comment":"The reward in Eq. (4) assumes that the song-provided label a*_t equals the participant's true intended movement at each 200 ms step. The paper acknowledges this assumption, but its validity is load-bearing for the quality of the RL training signal, and Section 7.3 shows that its validity plausibly varies across participants: the two participants with the lowest mutual information are exactly those for whom fine-tuning fails. Fig. 10 further shows that MI depends on the deployed policy, since it drops when π0 is replayed at repetition 9, indicating that participants adapt their muscle activations to the policy. This interaction between human adaptation and the assumed labels is not captured by Eq. (4). Since the paper already computes MI as a post-hoc diagnostic, it would strengthen the work to present label-alignment quality as a per-participant diagnostic and to discuss how reward misattribution arises when participants lag, anticipate, or execute a different movement than the note requires.","section":"Section 4.2, Eq. (4), and Section 7.3"}],"minor_comments":[{"comment":"The abstract highlights the gameplay accuracy gain as a two-fold increase, but that metric is evaluated on data used for training and model selection (see major comment 2); the independent Motion Test result (39% EMR improvement) is the stronger evidence and should be foregrounded.","section":"Abstract"},{"comment":"The action-randomization procedure (ϵ = 0.9, replacing negative-reward samples with uniformly random movements) is unusual and should be justified: it injects many samples with randomly relabeled actions into the training set, and the paper does not state how the new reward is computed for the selected random movement or why a uniform distribution over movements is appropriate.","section":"Section 5.2"},{"comment":"Multiple Wilcoxon signed-rank tests are reported (gameplay and Motion Test, overall and per-DOF in Fig. 6); please state whether any correction for multiple comparisons was applied.","section":"Section 6.4"},{"comment":"The claim that 'every measure in all scenarios improves with RL' mixes descriptive and inferential statements, since several per-DOF comparisons in Fig. 6 are not marked as statistically significant; please qualify the claim accordingly.","section":"Section 7.2 and Fig. 6"},{"comment":"The email given for the first author appears as 'tamino@chalmers.se', which likely belongs to a different author; please correct the contact address.","section":"Author block"},{"comment":"Reference [20] lists the author as 'C. labs at Reality Labs'; this should read 'CTRL-labs at Meta Reality Labs'.","section":"Reference [20]"},{"comment":"The introduction claims that gameplay data collection 'helps reduce the length of the initial recording session,' but no experiment quantifies this reduction; please qualify or support the claim.","section":"Section 1"},{"comment":"The caption is difficult to parse regarding which repetitions use π_i and which use π0; a concise statement of the repetition structure and of which repetitions produce training data would improve readability.","section":"Figure 5 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and the combined-pipeline result is credible and well executed methodologically. The main gap is the attribution of the gains to RL: the design confounds the learning rule with the quantity of usage data, and the authors themselves flag this in Section 7.1. I would regard the contribution as acceptable after either (a) the addition of a supervised-fine-tuning/behavior-cloning control arm on the same accumulated gameplay data, or (b) a substantial reframing of the title and abstract so that the claims are restricted to the combined game-plus-RL intervention, with the Motion Test positioned as the primary outcome. I did not find evidence of citation or novelty problems; the related-work coverage is adequate. No code or data release is indicated, which would otherwise strengthen reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: this is a well-run pilot that demonstrates a real, measurable improvement in myoelectric decoding after game-based fine-tuning. What it does not demonstrate is that the improvement comes from reinforcement learning per se.\n\nWhat's new and good: as far as I know, no one has applied offline RL (AWAC) to fine-tune a multi-label EMG finger-movement classifier on data collected in a guitar-hero style game. That combination is novel at the application level. The experimental design is above average for this area: 15 able-bodied subjects, paired comparisons, randomized test order for the Motion Test, and replaying the song with the original policy (π0) at the end to separate human learning from policy learning. The results are internally consistent, the analysis of failure cases (MI, PSI, SNR) is honest and useful, and the independent Motion Test gain (0.43 to 0.60 EMR) is real evidence that the combined game-plus-RL pipeline works.\n\nThe soft spot is the central attribution. π8 differs from π0 in two ways: it was trained on additional gameplay data, and it was trained with AWAC. Without a supervised fine-tuning or behavior-cloning arm on the same accumulated dataset Dn, you cannot say the gain comes from RL. The paper actually admits this in Section 7.1: 'we are unable to independently assess the distinct contributions of our game environment and RL training procedure.' The stress-test note on this is exactly right. A supervised arm would be easy to add and would settle it. As it stands, the title and abstract frame the result as 'through RL,' which the experiment does not isolate.\n\nA secondary, minor concern: the reward uses song-provided labels as ground truth, which assumes ideal participant intention. The paper acknowledges this, and I don't think it's fatal — the Motion Test gain holds regardless. But noisy labels would be another reason to be cautious about the RL-specific claim.\n\nNo code, data, or game artifacts are released, which limits reproducibility, though this is a preliminary study.\n\nWho this is for: researchers in myoelectric control or human-in-the-loop RL will get value as a proof-of-concept and a template for game-based data collection. I'd send it to peer review, but with a clear request: either add a supervised fine-tuning control or substantially soften the causal language. As is, it's a solid pilot that overclaims its mechanism.","headline":"A competent able-bodied pilot showing game-based fine-tuning improves myoelectric decoding, but the specific 'through RL' claim is not isolated from the added data.","tokens_in":20025,"tokens_out":2613,"would_cite":true,"duration_ms":30139,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A rhythm game plus reinforcement learning doubles muscle-signal decoding accuracy for prosthetic control.","keywords":["Reinforcement learning","Electromyography","Myoelectric control","Motor intent decoding","Prosthetic limbs","Serious games","Human-in-the-loop","Multi-label classification"],"falsifier":"Re-run the same eight-repetition fine-tuning protocol on new participants but delay the note-to-label alignment by one 200 ms step, so the reward is computed against the previous note's movement; if the method's gains persist under this systematic misalignment, the reward assumption is not what drives the improvement, and if they collapse, the assumption is confirmed as load-bearing.","tokens_in":19026,"feed_emoji":"🎮","tokens_out":4593,"duration_ms":46842,"temperature":0.7,"pith_summary":"The paper claims that fine-tuning a pretrained electromyography (EMG) classifier with reinforcement learning on data collected while people play a rhythm game yields large, statistically significant gains in online motor-intent decoding. Across 15 able-bodied participants, exact-match accuracy during gameplay rose from 0.36 to 0.78, and on a separate prompted motion test it rose from 0.43 to 0.60. The idea is that gameplay produces usage-based EMG data with an automatic reward signal, so the controller can keep improving outside a long supervised recording session. If correct, this closes part of the gap between offline classifier accuracy and real-time prosthetic control.","feed_headline":"Reinforcement learning doubles muscle-signal decoding accuracy","feed_subtitle":"A guitar-hero game supplies training data that lifts a hand controller's accuracy from 36% to 78%.","key_machinery":"The load-bearing mechanism is a Markov decision process whose reward is $r(s_t,a_t)=1$ for a correct non-rest prediction, $0$ for correctly predicting rest, and $-1$ otherwise, combined with Advantage Weighted Actor-Critic (AWAC), an off-policy actor-critic algorithm that keeps the policy close to the data seen so far. The environment is a Guitar Hero-style game in which notes specify which finger movement should occur at each step, so gameplay supplies both the usage data and the labels $a^*_t$ used to compute reward. AWAC lets the pretrained policy be fine-tuned on this dataset without the distribution shift that pure offline RL would suffer, and the game's timing demands mimic the temporal precision needed in daily prosthetic use.","core_discovery":"On its own terms, the paper establishes that an RL fine-tuning loop can replace part of the labeled-data burden in myoelectric control. Starting from a supervised policy trained on static recordings of thirteen finger movements (including simultaneous ones), the authors apply the Advantage Weighted Actor-Critic (AWAC) algorithm to data gathered while each participant plays an eight-repetition rhythm game whose notes define the desired movement at each 200 ms window. The final policy outperforms the initial one on every reported metric, with the largest gains in single-degree-of-freedom movements and in prediction stability, and the improvement transfers to a separate motion test that resembles the original recording protocol. The authors attribute the gains to RL aligning the policy with usage data rather than to human learning alone, since replaying the game with the original supervised policy yields markedly lower scores.","pith_inferences":["A practical screening rule suggested by the paper's correlation results: compute the mutual information between a user's EMG features and the ideal labels during a short gameplay session, and only invest in RL fine-tuning above a threshold; the authors note a possible threshold around MI 0.28-0.35.","The method's reliance on song-provided labels means it can only reward movements the game asks for; extending to free-form daily activity would require a different reward source, such as task success criteria, which the paper leaves open.","Because the authors could not separate the contribution of the game environment from that of RL, one testable extension is to fine-tune on gameplay data with the same AWAC update but a shuffled or unrelated reward; if improvement persists, gamification rather than reward-driven learning is doing the work."],"forward_implications":["Gameplay data can substitute for part of the supervised recording session, shortening initial calibration for multi-degree-of-freedom controllers.","Because the method works with any reward signal derivable from a task, the same fine-tuning loop could be applied to other serious games or daily-use tasks.","Improvements concentrate in single-degree-of-freedom movements, which had the lowest initial accuracy, so RL fine-tuning acts partly as a corrective for weak classes.","The number of prediction changes during gameplay drops consistently, meaning users experience a more stable and predictable controller.","A gameplay mutual-information measure $I(s; a^*)$ predicts which participants will benefit, offering a pre-training diagnostic."],"supporting_citations":[{"why":"Supplies the AWAC algorithm that carries the online fine-tuning update.","marker":"[36]"},{"why":"Provides the Hudgins time-domain features used as the policy's input representation.","marker":"[9]"},{"why":"Defines the network architecture and online evaluation setup the study builds on.","marker":"[19]"},{"why":"Defines the Motion Test protocol used to evaluate transfer beyond gameplay.","marker":"[39]"},{"why":"Earlier RL myoelectric controller with human-in-the-loop reward that motivates this approach.","marker":"[27]"},{"why":"Earlier RL system decoding three simultaneous degrees of freedom, the direct predecessor to this task.","marker":"[28]"}],"fun_headline_variants":["RL fine-tuning doubles muscle-signal decoding accuracy","Game-based RL boosts bionic hand accuracy twofold","Rhythm game trains AI to read muscle signals twice as well","Guitar-hero data sharpens myoelectric control via RL","Two-fold accuracy gain: RL fine-tuning with game data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The training signal assumes that whenever a song note appears, the participant is actually trying to perform exactly that movement at that moment; if a user lags, anticipates, or intentionally deviates, the reward punishes the wrong behavior and can reinforce mistaken mappings.","fun_headline_variants_meta":{"raw":{"variants":["RL fine-tuning doubles muscle-signal decoding accuracy","Game-based RL boosts bionic hand accuracy twofold","Rhythm game trains AI to read muscle signals twice as well","Guitar-hero data sharpens myoelectric control via RL","Two-fold accuracy gain: RL fine-tuning with game data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000465,"raw_usage":{"total_tokens":2325,"prompt_tokens":955,"completion_tokens":1370,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":1290}},"tokens_in":571,"tokens_out":1370,"duration_ms":11969,"temperature":1.0,"reasoning_tokens":1290,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:33:22.380576+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same eight-repetition fine-tuning protocol on new participants but delay the note-to-label alignment by one 200 ms step, so the reward is computed against the previous note's movement; if the method's gains persist under this systematic misalignment, the reward assumption is not what drives the improvement, and if they collapse, the assumption is confirmed as load-bearing.","supporting_citations":[{"cited_title":"Hudgins, P","cited_arxiv_id":null,"evidence_quote":"Provides the Hudgins time-domain features used as the policy's input representation."},{"cited_title":"Zbinden, J","cited_arxiv_id":null,"evidence_quote":"Defines the network architecture and online evaluation setup the study builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Motion Test protocol used to evaluate transfer beyond gameplay."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier RL myoelectric controller with human-in-the-loop reward that motivates this approach."},{"cited_title":"Vasan and P","cited_arxiv_id":null,"evidence_quote":"Earlier RL system decoding three simultaneous degrees of freedom, the direct predecessor to this task."}],"review_version":1}