{"id":"e13247bc-72a0-485c-9256-a3dd356aab85","arxiv_id":"2411.16273","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A CNN trained on combined EMG and IMU signals classifies five daily movements for ankle exoskeleton control with 96.5% offline accuracy, though new-user accuracy is much lower before fine-tuning and real-time operation is not evaluated.","lead":"A CNN that fuses muscle and motion sensor data classifies five daily movements with 96.5% accuracy in offline tests, and adapts to new users with only ten samples per movement. The work adds a public dataset for ankle exoskeleton control, relevant to assistive robotics and fall prevention.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 96.5% accuracy is measured on complete 5-second trials from seen subjects under a random split, so it does not establish real-time or new-user performance; the paper's own cross-subject result (69.0%) shows the gap.","rationale":"The reader's conditional verdict is appropriate. I find no internal inconsistency in the core training pipeline beyond the evaluation-protocol gap the reader identified; the reported numbers are plausible for whole-trial, within-subject classification, and the release of code and data is a real plus. The decisive issue is not whether 96.5% can be reproduced on the same split—it likely can—but whether that number measures the thing the abstract promises. The paper's own transfer-learning experiment makes this concrete: unseen-subject accuracy without fine-tuning is 69.0±7.6%, a 27.5-point drop from the headline. That gap, plus the absence of any sliding-window or latency analysis, means the abstract's 'real-time prediction' and 'new users' framing is not demonstrated. The paper could be made acceptable by reporting a deployment-matching evaluation (subject-independent, causal windows) and by tightening the abstract. Since the reader already conditioned the verdict on these same issues, I do not recommend a different verdict.","tokens_in":11293,"tokens_out":3728,"duration_ms":36719,"concrete_test":"Recompute the CNN result under a deployment-matching protocol: leave-one-subject-out cross-validation with no fine-tuning, and separately with causal sliding windows of 250, 500, and 1000 ms (reporting accuracy and decision latency per window) on the held-out subject. If the best windowed, cross-subject accuracy is materially below 96.5%, or if the latency exceeds the exoskeleton control budget, the abstract's real-time and new-user claims should be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's motivating claim is that the sensor-processing pipeline supports accurate real-time prediction of user intent for ankle exoskeletons. The headline 96.5±0.8% does not support that deployment claim as reported. The input to both CNN and LSTM is one complete 5-second trial (5000 samples at 1 kHz) already segmented and labeled; the 80/20 split is by trial, not by subject, so trials from all three subjects appear in both training and testing. This protocol answers 'can the network label a complete, already-segmented movement from a subject it has trained on?' It does not answer 'can the controller act on a short causal window from a new user before or during the movement?' No sliding-window evaluation, latency analysis, or causal masking is reported in Methods or Results. The relevance of the gap is quantified by the paper itself: when the CNN is trained on two subjects and tested on the third without fine-tuning, accuracy drops to 69.0±7.6%. The transfer-learning section shows that 10 samples per class plus frozen convolutional layers recovers 89.7±3.7%, but Figure 4's caption says 50 samples were used for fine-tuning, so the exact calibration cost is unclear. Because the abstract's real-time and generalizability claims rest on the 96.5% number, the evaluation protocol is the load-bearing assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a deep-learning pipeline for classifying five lower-limb locomotion tasks (turn left/right, pick up object, walk forwards/backwards) from eight surface EMG channels and three IMUs, for potential real-time ankle exoskeleton control. The authors compare CNNs and LSTMs on IMU-only, EMG-only, and fused data; report a best CNN accuracy of 96.5±0.8% on an 80/20 random split of 1,504 trials from three subjects; assess transfer learning by pre-training on two subjects and fine-tuning on the third; and simulate sensor failure by zeroing test channels. They release the dataset and code. The abstract claims real-time prediction and new-user generalization with only ten calibration samples per class.","tokens_in":11545,"tokens_out":2978,"duration_ms":30417,"significance":"The dataset release and reproducible processing code are useful community contributions, and the comparison of textile-electrode EMG with IMU fusion for multi-class locomotion is a relevant step for exoskeleton control. If the headline results were reproduced under a subject-independent split and with causal sliding-window inputs, the system would be an important benchmark. The transfer-learning idea is promising and the sensor-failure analysis is a pragmatic safety check. However, the current evaluation does not support the abstract's real-time or new-user claims: the 96.5% figure is an offline, trial-level, same-subject result, and the paper's own cross-subject accuracy before fine-tuning is 69.0±7.6%. The contribution is therefore a well-executed offline benchmark with a valuable dataset, rather than a demonstrated deployment-ready controller.","major_comments":[{"comment":"The headline accuracy of 96.5±0.8% is obtained from an 80/20 random split of individual trials, so trials from all three subjects appear in both training and test sets. This estimates how well the network labels a complete, pre-segmented 5-second trial from a subject it has already trained on; it does not estimate performance for a new user or for real-time control. The paper's own transfer-learning experiment quantifies the gap: training on two subjects and testing on the third without fine-tuning yields only 69.0±7.6% accuracy. To support the abstract's claims of 'real-time prediction' and generalization to new users, the authors should report a subject-independent split (e.g., leave-one-subject-out) and a causal sliding-window evaluation with prediction latency, or clearly restrict the claims to offline same-subject classification.","section":"Methods, Machine learning models; Results, Motion classification using deep learning; Transfer learning for model…"},{"comment":"There is a factual inconsistency about the calibration cost. The text and abstract state that fine-tuning used '10 samples per category' (and 'ten samples per class'), while the caption of Figure 4(a) says the model was 'fine-tuned using 50 samples'. This directly affects the central transfer-learning claim and must be corrected in the manuscript.","section":"Transfer learning for model deployment; Figure 4"},{"comment":"The study includes only three subjects, and the transfer-learning evaluation is performed on a single held-out subject. The reported standard deviations are computed over five random seeds, not over subjects, so they do not capture between-subject variability. The Discussion's statement that 'the transfer learning results indicate good generalizability' is stronger than the evidence supports; the authors should report per-subject results and clearly acknowledge that the target-subject sample size is n=1.","section":"Human subject study; Transfer learning for model deployment"}],"minor_comments":[{"comment":"The caption contains a typo: 'transfer leaning' should be 'transfer learning'.","section":"Figure 4(a) caption"},{"comment":"The layer names are written 'Cov-1D' instead of 'Conv-1D' throughout the table.","section":"Table 3"},{"comment":"Equation (2) defines accuracy as TC/(TC+FC) without explaining how TC and FC are computed for a multi-class problem; rewriting it as the fraction of correctly classified samples over the total number of samples would avoid ambiguity.","section":"Statistical analysis, Eq. (2)"},{"comment":"The terminology for the electrodes is inconsistent: the abstract refers to 'towel electrodes', while the Methods and Discussion describe a graphene/PEDOT:PSS textile composite; using one consistent term (e.g., 'textile electrodes') would improve clarity.","section":"Abstract; Discussion; Methods"},{"comment":"The electrode schematic is reproduced from reference [25]; the authors should confirm that reproduction permission is obtained or that the figure is original.","section":"Figure 1(a)"}],"recommendation":"major_revision","confidential_remarks":"The core offline benchmarking result is plausible and the dataset/code availability is a genuine strength, but the paper's framing substantially overstates real-time and new-user performance. The evaluation split and the 10-vs-50-sample discrepancy are fixable with additional experiments or careful re-writing. I would be comfortable with a major revision rather than a rejection, provided the authors add (or clearly scope) the subject-independent and causal evaluation and reconcile the transfer-learning numbers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the dataset is the contribution. If you need a public EMG+IMU benchmark for ankle exoskeleton locomotion classification, this is a useful one. The 96.5% accuracy, however, is a within-subject whole-trial number, so treat the real-time and new-user claims with skepticism.\n\nThe paper does several things well. It provides a new public dataset of 1,504 trials across five daily tasks with IMUs and sEMG from three subjects, along with code. The textile/towel electrode integration is a genuine practical step. The transfer-learning experiment—train on two subjects, fine-tune on the third—is honest and shows a plausible calibration path, though note the abstract says 10 samples per class while Figure 4 says 50 and the text says 10; that discrepancy needs fixing. The sensor-failure robustness test is a good safety-oriented addition that most papers skip.\n\nThe main soft spot is the evaluation protocol. The headline 96.5±0.8% comes from an 80/20 random split of trials, not a subject split. Since all three subjects appear in both training and test, this number reflects the network's ability to label a complete, already-segmented trial from a person it has trained on. It does not estimate performance for a new user. The paper's own cross-subject result, 69.0±7.6% before fine-tuning, shows the gap. Relatedly, the abstract promises real-time prediction, but the models ingest the full 5-second trial. There is no sliding-window evaluation, causal masking, or latency analysis. That is a load-bearing mismatch between framing and evidence.\n\nNothing about the core modeling is broken. The CNN and LSTM architectures are standard, the comparison is clean, and the training details are reproducible. The sensor-fusion advantage over single-modality is real but modest (96.5 vs roughly 93.5). The small number of subjects (three, all young and healthy) is an acknowledged limitation, and the authors do list several limitations in the Discussion.\n\nWho gets value: researchers working on exoskeleton control or wearable motion classification who want a shared dataset and a baseline to beat. For that purpose, the paper deserves a serious referee. The claims, though, need to be scaled back to match the protocol, or the experiments need to be extended to subject-independent and real-time settings.","headline":"Useful public dataset and fabric-electrode integration, but the 96.5% accuracy is measured on seen subjects' whole trials and does not support the real-time or new-user claims as reported.","tokens_in":12129,"tokens_out":2399,"would_cite":true,"duration_ms":22627,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A convolutional network fusing EMG and IMU signals classifies five daily-living motions with 96.5±0.8% accuracy.","keywords":["ankle exoskeleton","motion intention prediction","surface EMG","IMU","convolutional neural network","transfer learning","sensor fusion","wearable sensors"],"falsifier":"Run a leave-one-subject-out evaluation and stream the sensors continuously, requiring a classification decision every 200 ms instead of over a whole five-second trial. If a new subject's per-decision accuracy stays near the 69% seen without fine-tuning, or if sliding-window accuracy falls below 80% on continuous data, the paper's real-time deployability claim is contradicted; if ten calibration samples per class fails to reproduce the 89.7% fine-tuned accuracy on a larger, more diverse subject pool, the transfer-learning claim is weakened.","tokens_in":11077,"feed_emoji":"🦿","tokens_out":9197,"duration_ms":75405,"temperature":0.7,"pith_summary":"This paper argues that a single convolutional neural network, fed with time-aligned signals from three inertial measurement units and eight surface-EMG electrodes, can recognise the everyday ankle movements an exoskeleton needs to assist: walking forwards, walking backwards, turning left, turning right, and squatting to pick up an object. On a dataset of 1,504 five-second trials from three healthy volunteers, the fused-signal CNN reaches 96.5±0.8% accuracy on held-out trials, outperforming an LSTM (87.5±2.9%) and either sensor modality alone. The paper also claims that the model can adapt to a new user from just ten samples per class, reaching 89.7±3.7% after fine-tuning, and that it remains above 80% accurate when individual sensors are silenced. The reason to care is that real-time and reliable motion-intention classification is the step that lets an ankle exoskeleton aid older adults safely instead of hindering them.","feed_headline":"CNN fusing EMG and IMU hits 96.5% on five exoskeleton motions","feed_subtitle":"A new user needs only ten calibration samples, and accuracy stays above 80% when a sensor fails.","key_machinery":"The machinery is a one-dimensional convolutional network that operates directly on a five-second, 1000 Hz multi-channel recording: eight surface-EMG channels (tibialis anterior, gastrocnemius medial and lateral, and soleus, bilaterally) plus three IMUs (left shank, right shank, right foot), synchronised by upsampling the IMU stream. The network stacks three 1D convolution, batch-normalisation, ReLU, and max-pooling stages followed by dropout and a fully connected softmax layer (21,655 trainable parameters), and is trained with cross-entropy loss and the ADAM optimiser. The same input pipeline feeds a two-layer LSTM for comparison. This shared, time-aligned representation is what lets the paper attribute differences in accuracy to model architecture and sensor modality rather than to preprocessing.","core_discovery":"The central discovery is that fusing kinematic (IMU) and muscular (sEMG) signals into one multi-channel CNN input supports higher motion-classification accuracy than either signal type alone, or than an LSTM on the same inputs, for a realistic set of daily-living movements. With all channels, the CNN classifies five movement classes with 96.5±0.8% accuracy; EMG-only and IMU-only versions reach 93.9±1.6% and 93.3±1.6%, and a single-leg version reaches 92.9±1.6%. The remaining misclassifications concentrate between directionally mirrored actions: turning left versus turning right, and walking forwards versus walking backwards. When the model is pre-trained on two subjects and fine-tuned with ten samples per class from a third subject, accuracy reaches 89.7±3.7%; without fine-tuning it is 69.0±7.6%. Zeroing all channels of one sensor at test time, simulating a failed sensor, still leaves accuracy above 80%, with the foot IMU the most important single sensor at 82.8±2.9%.","pith_inferences":["The 96.5% figure is trial-level, not decision-level: the model sees a complete five-second recording, so a continuous real-time controller would need a sliding-window variant, and the paper does not report per-decision latency; that measurement is the next test.","With only three young healthy participants, the transfer-learning result shows feasibility, not population generality; the open question is whether ten samples per class still suffice for older adults, people with gait impairment, or day-to-day variations in electrode placement.","Silencing a sensor to zero is an extreme but clean test; realistic failures may be partial, intermittent, or noisy, which could degrade performance differently than the paper's zero-signal simulation."],"forward_implications":["If the accuracy holds in real use, a CNN can support an ankle exoskeleton through the five motions needed to navigate a barrier-free environment, with misclassifications mostly between directionally mirrored movements that an exoskeleton could handle cautiously.","A new user could be fitted in minutes: ten labelled samples per class are enough to fine-tune the pre-trained model to 89.7±3.7% accuracy, making per-person retraining unnecessary.","Sensor redundancy is a safety feature: with any single IMU or one leg's EMG silenced, classification accuracy remains above 80%, so the device can stay safe until the failed sensor is replaced.","Because EMG-only and IMU-only models both reach roughly 93-94%, users with weak or degraded muscle signals could still be served by the IMU channel alone.","The released dataset and code give other groups a public benchmark for five-class daily-living motion classification with synchronised EMG and IMU signals."],"supporting_citations":[{"why":"Supplies the survey of sensor modalities and algorithms for lower-limb exoskeleton locomotion detection that frames the accuracy gap the paper targets.","marker":"[20]"},{"why":"Li et al.'s CNN-LSTM ankle-movement classifier is the closest comparable high-accuracy result; the paper distinguishes its own full-body daily-living motions from this ankle-only work.","marker":"[22]"},{"why":"Kim et al.'s CNN on EMG and IMU data for daily-living motions is the main baseline at 88%, which the paper compares its 96.5% against.","marker":"[24]"},{"why":"Provides the textile electrode fabrication method, the CNN architecture this paper adapts, and the transfer-learning-with-few-samples result it extends.","marker":"[25]"},{"why":"Si et al.'s EMG CNN at 95.5% is a further comparison point, with movements the paper notes differ from daily life.","marker":"[28]"},{"why":"Supports the transfer-learning premise that fine-tuning with a small number of samples yields high-accuracy EMG classification.","marker":"[30]"}],"fun_headline_variants":["Exoskeleton CNN fuses EMG+IMU to 96.5% motion accuracy","Ten calibration samples fine-tune exoskeleton CNN to 89.7%","Exoskeleton CNN survives sensor loss, stays above 80% accuracy","CNN beats LSTM for exoskeleton motion: 96.5% vs 87.5%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline accuracy comes from a random 80/20 split of trials in which the same three people contribute to both training and testing, so the claim assumes this predicts performance for a new user in real time; the paper's own unseen-subject test drops to 69.0±7.6% before fine-tuning.","fun_headline_variants_meta":{"raw":{"variants":["Exoskeleton CNN fuses EMG+IMU to 96.5% motion accuracy","Ten calibration samples fine-tune exoskeleton CNN to 89.7%","Exoskeleton CNN survives sensor loss, stays above 80% accuracy","CNN beats LSTM for exoskeleton motion: 96.5% vs 87.5%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001525,"raw_usage":{"total_tokens":6148,"prompt_tokens":1026,"completion_tokens":5122,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":5029}},"tokens_in":642,"tokens_out":5122,"duration_ms":30965,"temperature":1.0,"reasoning_tokens":5029,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:19:18.869212+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a leave-one-subject-out evaluation and stream the sensors continuously, requiring a classification decision every 200 ms instead of over a whole five-second trial. If a new subject's per-decision accuracy stays near the 69% seen without fine-tuning, or if sliding-window accuracy falls below 80% on continuous data, the paper's real-time deployability claim is contradicted; if ten calibration samples per class fails to reproduce the 89.7% fine-tuned accuracy on a larger, more diverse subject pool, the transfer-learning claim is weakened.","supporting_citations":[{"cited_title":"Sensors and algorithms for locomotion intention detection of lower limb exoskeletons.Medical Engineering &amp; Physics, 113:103960, March 2023","cited_arxiv_id":null,"evidence_quote":"Supplies the survey of sensor modalities and algorithms for lower-limb exoskeleton locomotion detection that frames the accuracy gap the paper targets."},{"cited_title":"Soft wearable flexible bioelectronics integrated with an ankle-foot exoskeleton for estimation of metabolic costs and physical effort.npj Flexible Electronics, 7(1), January 2023","cited_arxiv_id":null,"evidence_quote":"Kim et al.'s CNN on EMG and IMU data for daily-living motions is the main baseline at 88%, which the paper compares its 96.5% against."},{"cited_title":"Recognition of lower limb movements baesd on electromyogra- phy (emg) texture maps","cited_arxiv_id":null,"evidence_quote":"Si et al.'s EMG CNN at 95.5% is a further comparison point, with movements the paper notes differ from daily life."}],"review_version":1}