{"id":"f3fde6a5-2352-4960-8af1-fc7d030e34e9","arxiv_id":"2604.18920","paper_version":2,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"SPARC articulatory features predict sEMG signals more accurately than phoneme features across aloud, mimed, and subvocal speech, with consistent anatomical patterns and above-chance performance even in silent mode.","lead":"The study tests whether articulatory features from SPARC can predict muscle signals from the face and neck better than simple sound-based phoneme features during normal, mimed, and silent speech in 24 people. A general reader might care because this could help build better devices that let people communicate without speaking out loud.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"SPARC may outperform phoneme one-hot due to higher feature dimensionality/flexibility in elastic-net mTRF rather than superior encoding.","rationale":"Reader's weakest assumption correctly flags the linear mTRF + CV adequacy. The dimensionality concern is a more specific, testable instantiation of that assumption focused on the feature comparison itself. It does not overturn the empirical design but qualifies how strongly the accuracy numbers support SPARC as the preferred intermediate representation. No other internal inconsistency (e.g., circularity or missing controls) is evident from the described methods.","tokens_in":1648,"tokens_out":368,"duration_ms":31831,"concrete_test":"Extract and report the input feature dimensionality for SPARC vs. phoneme one-hot, plus the mean number of non-zero mTRF coefficients per electrode (across CV folds and subjects) for both representations. If SPARC models retain >2x more active features on average and the accuracy gap shrinks after matching effective feature count (e.g., via explicit L0 or post-hoc pruning), the representational-superiority interpretation weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that SPARC's higher linear prediction accuracy reflects better articulatory representation of sEMG, not merely greater model capacity. SPARC consists of continuous multi-dimensional articulatory parameters while phoneme one-hot is sparse and bounded by phoneme inventory size (~40-60). Elastic-net regularization mitigates but does not eliminate differences in effective degrees of freedom; if the same alpha/lambda schedule is applied without explicit capacity matching, SPARC can achieve higher r-values by fitting more parameters to noise, especially in lower-SNR subvocal conditions. Variance partitioning would then attribute unique variance to SPARC partly because of its richer basis rather than unique neural information.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.3","summary":"The manuscript compares Speech Articulatory Coding (SPARC) features against phoneme one-hot encodings for linearly predicting sEMG envelopes in aloud, mimed, and subvocal speech. Using elastic-net regularized mTRF models with sentence-level cross-validation across 24 subjects, it reports higher prediction accuracy for SPARC on nearly all electrodes and modes, substantial unique variance from SPARC in partitioning analyses, minimal unique variance from phonemes, and anatomically interpretable mTRF weights consistent across modes. The work positions SPARC as a robust intermediate representation for sEMG-based silent speech modeling.","tokens_in":1800,"tokens_out":585,"duration_ms":32775,"significance":"If robust after addressing dimensionality confounds, the results would support SPARC as a more effective and interpretable articulatory target than phoneme encodings for sEMG interfaces, particularly for subvocal speech where above-chance encoding is shown. The cross-mode consistency and variance partitioning provide useful empirical data on articulatory feature encoding from muscle signals.","major_comments":[{"comment":"Methods section on mTRF modeling and feature sets: The comparison applies the same elastic-net regularization schedule to SPARC (continuous, multi-dimensional articulatory parameters) and phoneme one-hot (sparse, bounded by ~40-60 phonemes). No explicit control for effective degrees of freedom or feature dimensionality is described, raising the possibility that SPARC's higher capacity contributes to elevated r-values and unique variance rather than superior representation of sEMG. A matched-dimensionality control or reporting of effective df would be required to support the central claim.","section":"Methods (mTRF and feature comparison)"},{"comment":"Results section on variance partitioning: The reported unique contribution of SPARC may partly reflect its richer basis set rather than unique neural information. It is unclear whether the partitioning isolates representation quality after accounting for the continuous vs. discrete nature of the features; subsampling SPARC to phoneme dimensionality or adding a capacity-matched baseline would test this.","section":"Results (variance partitioning)"}],"minor_comments":[{"comment":"Abstract: No quantitative accuracy values (e.g., mean r or percentage improvement), SPARC extraction details, electrode montage, or statistical test descriptions are provided, limiting immediate assessment of effect sizes.","section":"Abstract"},{"comment":"Results: The claim that subvocal speech remains above chance would be strengthened by explicit reporting of the chance-level baseline, exact p-values, and correction for multiple comparisons across electrodes.","section":"Results"}],"recommendation":"major_revision","confidential_remarks":"The work is timely for silent-speech BCI but the related-work section appears thin on prior sEMG encoding literature; a broader citation of comparable mTRF studies would improve context."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments, which identify a valid methodological concern about potential dimensionality confounds in our feature comparison. We address each point below and commit to revisions that directly test this issue.","responses":[{"response":"We acknowledge that the continuous, multi-dimensional nature of SPARC versus the discrete phoneme one-hot encoding could introduce a capacity difference, even under elastic-net regularization. Although the L1/L2 penalties and sentence-level cross-validation are intended to limit effective model complexity, we agree that an explicit control is needed to isolate representation quality. In the revised manuscript, we will add a matched-dimensionality analysis by randomly subsampling SPARC features to approximately 50 dimensions (matching the phoneme set) and re-evaluate both prediction accuracies and unique variances. We will also report effective degrees of freedom derived from the regularization paths for both feature sets.","revision_made":"yes","referee_comment":"Methods section on mTRF modeling and feature sets: The comparison applies the same elastic-net regularization schedule to SPARC (continuous, multi-dimensional articulatory parameters) and phoneme one-hot (sparse, bounded by ~40-60 phonemes). No explicit control for effective degrees of freedom or feature dimensionality is described, raising the possibility that SPARC's higher capacity contributes to elevated r-values and unique variance rather than superior representation of sEMG. A matched-dimensionality control or reporting of effective df would be required to support the central claim."},{"response":"We agree that the variance partitioning results require additional controls to rule out capacity effects. To directly address this, the revised manuscript will include a capacity-matched baseline in which SPARC features are subsampled to phoneme dimensionality before repeating the unique-variance analysis. This will clarify whether SPARC's unique contribution persists after dimensionality matching, thereby strengthening the claim that it captures superior articulatory information for sEMG encoding.","revision_made":"yes","referee_comment":"Results section on variance partitioning: The reported unique contribution of SPARC may partly reflect its richer basis set rather than unique neural information. It is unclear whether the partitioning isolates representation quality after accounting for the continuous vs. discrete nature of the features; subsampling SPARC to phoneme dimensionality or adding a capacity-matched baseline would test this."}],"tokens_in":1381,"tokens_out":490,"duration_ms":34985,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's central result is that Speech Articulatory Coding features outperform phoneme one-hot encodings when used to predict sEMG envelopes via elastic-net mTRF, and this holds for aloud, mimed, and subvocal speech in 24 subjects. Variance partitioning attributes most of the explained variance to SPARC with little unique contribution from phonemes, and the mTRF weights show anatomically plausible patterns that stay consistent across modes. Subvocal performance stays above chance, which is useful for silent-speech applications. The design uses sentence-level cross-validation, which is a reasonable step against leakage, and the focus stays on encoding rather than end-to-end decoding. That keeps the claims grounded. The work is incremental but cleanly executed on the comparison it sets out to do. The main soft spot is the dimensionality difference: SPARC supplies continuous multi-dimensional articulatory parameters while phoneme one-hots are sparse and limited by inventory size. Elastic-net regularization reduces but does not remove the risk that SPARC simply has more degrees of freedom to fit noise, especially in lower-SNR subvocal conditions. The abstract gives no numbers on feature counts, exact regularization schedule, or any explicit capacity-matched control, so a referee would need to see those details to judge whether the unique SPARC variance truly reflects better articulatory information. The citation pattern looks standard for the subfield and does not appear to over-claim prior results. This is the kind of targeted empirical study that matters for the BCI and speech-neuroscience crowd working on intermediate representations. Readers who care about sEMG encoding or silent-speech interfaces will find the head-to-head and the anatomical weight maps worth their time. The paper is coherent on its own terms and shows clear thinking about the comparison, so it deserves a serious referee even if the effect sizes turn out modest and the dimensionality concern needs addressing in revision.","headline":"SPARC features give higher linear prediction accuracy for sEMG than phoneme one-hots across modes, but the gain could partly trace to feature richness rather than pure representational superiority.","tokens_in":2313,"tokens_out":452,"would_cite":false,"duration_ms":25992,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"SPARC articulatory features predict sEMG envelopes more accurately than phoneme representations across all tested speech modes.","keywords":["sEMG encoding","SPARC","articulatory features","phoneme features","silent speech","mTRF","speech modes"],"falsifier":"Finding no significant accuracy advantage for SPARC over phonemes when using a nonlinear decoder on the same data, or observing subvocal prediction accuracy drop to chance levels in an independent replication.","tokens_in":2566,"feed_emoji":"","tokens_out":633,"duration_ms":29891,"temperature":0.7,"pith_summary":"The paper examines whether Speech Articulatory Coding features can linearly predict surface electromyography signals from speech muscles in aloud, mimed, and subvocal conditions. Using a regularized linear model on data from 24 subjects, it shows SPARC outperforms simple phoneme codes on nearly every electrode and in every mode, including when speech is silent. Subvocal speech remains predictable above chance level, and the model weights point to consistent links between electrodes and specific articulatory movements. This matters because it identifies a robust intermediate representation that could support silent speech interfaces without relying on audible output or phoneme-based assumptions.","feed_headline":"Articulatory codes predict muscle signals better than phonemes in silent speech","feed_subtitle":"SPARC features achieve higher accuracy across aloud, mimed and subvocal modes, with signals still detectable when no sound is made.","key_machinery":"Speech Articulatory Coding (SPARC) features as the central representation in elastic-net regularized multivariate temporal response function (mTRF) models for predicting sEMG envelopes.","core_discovery":"SPARC features yield higher prediction accuracy than phoneme one-hot representations on nearly all electrodes and in all speech modes. Aloud and mimed speech perform comparably, subvocal speech remains above chance, variance partitioning shows substantial unique contribution from SPARC, and mTRF weight patterns reveal anatomically interpretable relationships consistent across modes. This supports SPARC as a robust intermediate target for sEMG-based silent-speech modeling.","pith_inferences":["These findings suggest SPARC could serve as a target for training decoders in practical silent speech applications.","The consistency across modes implies potential for models trained on audible speech to generalize to silent conditions.","Extending the analysis to real-time decoding scenarios could test whether the linear advantage holds under streaming constraints."],"forward_implications":["Aloud and mimed speech show comparable encoding performance using SPARC.","Subvocal speech exhibits detectable articulatory activity above chance levels.","SPARC contributes uniquely to predictions beyond what phoneme features provide.","Anatomically interpretable mTRF weights remain consistent across speech modes."],"fun_headline_variants":["SPARC predicts sEMG better than phonemes across all speech modes","SPARC features achieve higher sEMG accuracy than phonemes in every mode","Articulatory SPARC uniquely contributes to sEMG prediction over phonemes","Consistent articulatory patterns link electrodes to speech movements across modes"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The assumption that a linear model with elastic-net regularization and sentence-level cross-validation adequately captures the true encoding relationship without overfitting or missing important nonlinear dynamics.","fun_headline_variants_meta":{"raw":{"variants":["SPARC predicts sEMG better than phonemes across all speech modes","SPARC features achieve higher sEMG accuracy than phonemes in every mode","Articulatory SPARC uniquely contributes to sEMG prediction over phonemes","Consistent articulatory patterns link electrodes to speech movements across modes"]},"model":"grok-4.3","cost_usd":0.008686,"raw_usage":{"total_tokens":3893,"prompt_tokens":622,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":86862000,"prompt_tokens_details":{"text_tokens":622,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3199,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":622,"tokens_out":72,"duration_ms":56378,"temperature":1.0,"reasoning_tokens":3199,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-10T02:46:48.043335+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Finding no significant accuracy advantage for SPARC over phonemes when using a nonlinear decoder on the same data, or observing subvocal prediction accuracy drop to chance levels in an independent replication.","supporting_citations":[],"review_version":1}