{"id":"fa2b65c5-ea77-4288-a3f5-4a90937eeeb6","arxiv_id":"2504.19046","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A TCN plus scaled dot-product attention model reproduces ACE electrodograms with STOI 0.6031 versus 0.6126 for ACE, slightly underperforming its own training target on a 20-file TIMIT test set.","lead":"This paper trains a neural network combining temporal convolutions with attention to generate cochlear implant electrodograms that imitate the standard ACE coding strategy. On a 20-file test set the model reaches a speech intelligibility score (STOI) of 0.6031, slightly below the 0.6126 scored by ACE, while the claimed flexibility advantages are never measured.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline rests on a single pair of 20-file STOI averages with no per-file variance; before the 'closely approximating' claim can be accepted, the model-vs-ACE difference must be shown to be within a pre-specified equivalence margin, not merely close in point estimate.","rationale":"The reader's strongest claim correctly places the load-bearing content in the two aggregate STOI numbers. My review found no independent support beyond those numbers: no code, no per-file table, no error bars, no ablation of the TCN or attention components, and no human-listening validation. The ACE-as-training-target circularity is real, but the narrowest claim is only that the learned model closely approximates ACE's intelligibility; that claim could survive the circularity objection if the measurement were statistically solid. What undermines the claim as written is the absence of any uncertainty quantification around the 0.0095 difference with n=20. Without per-file paired data, the degree of approximation cannot be assessed, and the 'slightly outperforms' wording is equally unsupported. A paired equivalence analysis is the minimal check that would turn the headline into a supported result, so the reader's CONDITIONAL verdict remains appropriate.","tokens_in":7852,"tokens_out":7907,"duration_ms":87981,"concrete_test":"Re-evaluate the same 20 test files per file for both the proposed model and the ACE reference, computing paired STOI differences; report mean, SD, and 95% CI. Run a two one-sided tests (TOST) equivalence procedure with a pre-specified margin (e.g., ±0.02 STOI). If the CI for the mean difference falls inside the margin, the 'closely approximating' claim is quantitatively supported; if it is wider than the margin, the claim is unsupported and should be reworded or reevaluated on more files.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central load-bearing measurement is the pair of aggregate STOI scores in the Abstract and Section V: 0.6031 (model) vs 0.6126 (ACE). No per-file scores, standard deviation, confidence interval, or significance/equivalence test is reported. With only 20 test files, the standard error of a STOI mean for vocoded TIMIT sentences can be large enough that the observed 0.0095 gap could be within noise; the aggregate could also hide opposite-sign per-file differences. The prose 'closely approximating' is therefore a statistical assertion made without the statistics needed to support it. Because the model is trained to regress ACE electrodograms, the two arms are not independent, which makes a paired uncertainty analysis more important, not less. The paper itself concedes that cross-study STOI comparisons are confounded by dataset and vocoder differences, but the within-dataset claim still lacks any variance estimate.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a deep learning approach for cochlear implant (CI) signal coding, combining temporal convolutional networks (TCNs) with scaled dot-product attention to generate electrodograms. The model is trained on TIMIT audio files whose target outputs are electrodograms produced by the ACE strategy implemented in Matlab. The predicted electrodograms are vocoded back to audio and evaluated with the short-time objective intelligibility (STOI) metric. The authors report an average STOI of 0.6031 for their model versus 0.6126 for the ACE strategy over a 20-file test set, and conclude that the model closely approximates ACE while offering potential advantages in flexibility and adaptability. The paper includes architecture details, training curves, and a qualitative electrodogram comparison, but provides no statistical analysis of the reported STOI scores.","tokens_in":7904,"tokens_out":2655,"duration_ms":29084,"significance":"If the quantitative claim were properly supported, the work would provide a useful data point on learned alternatives to traditional CI coding strategies: the architecture is simple, the training pipeline is clearly described, and the use of a standard objective metric (STOI) makes the result comparable to prior work. The manuscript also correctly acknowledges that cross-dataset STOI comparisons are confounded. However, the central result rests on a single pair of aggregate STOI means with no variance or significance assessment, and the training objective is to imitate ACE itself, so the 'advanced alternative' framing is not currently supported. The study does not provide reproducible code or sufficiently detailed experimental settings, which limits its immediate value to the community.","major_comments":[{"comment":"The headline claim that the model's STOI of 0.6031 'closely approximates' the ACE score of 0.6126 is a statistical assertion, but no variance, confidence interval, or significance or equivalence test is reported. With only 20 test files, the standard error of the mean STOI for vocoded speech can easily exceed the observed 0.0095 gap, and the aggregate means could conceal opposite-sign per-file differences. The authors should report per-file STOI scores, the mean and standard deviation for each condition, and a paired analysis appropriate to the matched test set; ideally they should specify a pre-defined equivalence margin. Without this, the central claim is not distinguishable from a chance difference.","section":"Section V (Experimental Results), Abstract"},{"comment":"The experimental design makes the comparison to ACE partly self-referential: the model is trained to regress ACE-generated electrodograms, and Section V explicitly states that the ACE electrodogram 'serves as the ground truth' when assessing the model's output. Consequently, the model's attainable STOI is bounded by the fidelity of ACE itself, and the observed proximity in STOI is an expected consequence of the imitation objective rather than evidence of an independent improvement. To support the 'advanced alternative' claim, the authors should either compare against a non-ACE target, evaluate with a vocoder or listener test that does not derive its reference from ACE, or explicitly reframe the contribution as 'a learned model that can reproduce ACE-like performance' and discuss what is gained (e.g., computational cost, personalization) beyond imitation.","section":"Section IV-B (Audio material), Section V"},{"comment":"The dataset description is underspecified. The text says 'The dataset consisted of 100 audio files for both training and validation' — this is ambiguous between 100 files total and 100 per split. There is no mention of the number of speakers, sentence duration, sampling rate, gender balance, or whether the 20 test files are disjoint from the training/validation speakers. These details are necessary for interpreting the intelligibility scores and for any future replication. The authors should provide a precise breakdown of the TIMIT subsets used.","section":"Section IV-B (Audio material)"},{"comment":"The comparison with existing studies (TMHINT-based scores from Refs. [20] and [11]) is presented as contextualization, but the authors themselves note the datasets and vocoders differ. As such, the sentence 'our model demonstrates competitive performance within the context of CI coding strategies' is not supported by the evidence. Since the model and ACE are evaluated on the same TIMIT test set, the only valid comparison is the within-dataset model-vs-ACE contrast, which currently lacks statistical backing. The cross-study numbers should be removed or clearly labeled as non-comparable.","section":"Section V (Experimental Results)"}],"minor_comments":[{"comment":"The phrase 'artificial intelligent (AI)' in the Abstract should be corrected to 'artificial intelligence (AI)'.","section":"Section I (Introduction), Section III-B"},{"comment":"The architecture diagram is hard to read: the connections between the TCN, attention, and decoder modules are not clearly labeled, and the meaning of the 'M x 1 Conv.' and '1 x L Conv.' annotations is not explained in the text. A short description of tensor shapes would improve reproducibility.","section":"Section IV-A (DL model), Fig. 4"},{"comment":"The notation in Eq. (2) states Q∈ R^{n×dk}, K∈ R^{m×dk}, V∈ R^{m×dv}, but the text and the standard formulation in Ref. [15] require QK^T to be compatible, which implies Q should be R^{n×dk} and K^T is R^{dk×m}; this is correct as written, but the dimensions of V and the output are not specified. Please clarify the output dimensions and the relationship between n and the sequence length.","section":"Section III-B (Eq. 2)"},{"comment":"The training curves in Fig. 5 appear to be smoothed or downsampled, and the y-axis label 'Loss' is not specific about which loss (MSE, BCE, or combined) is plotted. Please label the curve with the exact loss function and report the final validation loss.","section":"Section V (Experimental Results)"},{"comment":"The spectrograms in Fig. 6 are mentioned in the text but are not clearly visible or labeled in the figure; please ensure the figure contains both the electrodograms and the spectrograms with distinct captions or subplot labels.","section":"Section V (Experimental Results), Fig. 6"},{"comment":"Several references are incomplete or informally cited; for example, Ref. [15] is listed as 'Attention is all you need' without authors, venue, or year. Please provide complete bibliographic information for all references.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is plausible as a short conference contribution, but the central quantitative claim is statistically unsupported and the experimental framing conflates imitation of ACE with an 'advanced alternative' to ACE. The authors should be given the opportunity to add proper statistical analysis and to reframe or strengthen their claims. I would not recommend reject, as the underlying idea—learning CI electrodograms with a TCN-attention model—is worthy of further scrutiny. The lack of code or detailed implementation details is also a concern for reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"One thing to know: this is a modest paper that does one new thing—it trains a TCN plus scaled dot-product attention network to generate cochlear-implant electrodograms, and reports a STOI of 0.6031 versus ACE's 0.6126 on a 20-file TIMIT test set. The narrow claim is plausible and the paper is honest about ACE slightly outperforming it. But the quantitative support is thin, and the design makes the result partly circular. What's actually new: the specific architecture stack for electrodogram generation. Huang et al. already replaced ACE stages with DNN/CNN/LSTM, so the attention-plus-TCN combination is a routine extension, but it hasn't been tried in this application before. The STOI numbers on the authors' TIMIT split are new too, though only in the sense that no one else has measured them. What the paper does well: it reports the raw numbers without hiding the gap, and it explicitly concedes that cross-dataset comparisons are confounded. The training/validation curves look plausible, and the description of the loss (MSE + BCE) is sensible. The citations to prior DL-CI work are adequate, and the heavy self-citation isn't a real problem since the cited papers are on-point. Soft spots, in proportion: the headline claim rests entirely on two point estimates from 20 test files. No standard deviation, no confidence interval, no significance or equivalence test. The stress-test note is right: 0.0095 could easily be within noise at n=20, and the paired nature of the comparison makes a variance estimate more important, not less. The deeper issue is that the model is trained to regress ACE electrodograms and then evaluated against ACE as ground truth. So the result measures imitation fidelity, not whether the model is a better or more flexible strategy. The claimed advantages—flexibility, adaptability, personalization—are never measured. Also, no code, data, or architecture details are provided, so replication from the text alone is impossible. One minor technical point: the STOI formula in Eq. (3) is a simplification of the actual STOI algorithm, which involves intermediate intelligibility measures and clipping; this doesn't change the headline but shows imprecision. Who is this for? Someone in the CI signal-processing community who wants a quick look at whether attention-based networks are worth pursuing. It's a conference-level contribution, not a definitive study. Recommendation: yes, engage with it—send it to peer review if it comes your way. The idea is reasonable and the result, while under-supported, is not wrong. Require code/data, per-file STOI scores with variance, and ideally a non-circular evaluation. With those, it could become a solid short paper.","headline":"A small, honest empirical paper: TCN+attention can mimic ACE electrodograms nearly enough to match its STOI on 20 files, but the headline 'closely approximating' rests on no statistics and a self-referential training target.","tokens_in":782,"tokens_out":2992,"would_cite":false,"duration_ms":37331,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One deep-learning model—a temporal convolutional network plus scaled dot-product attention—can generate cochlear-implant electrodograms whose vocoded speech is nearly as intelligible as the standard ACE strategy.","keywords":["cochlear implant","electrodogram","ACE strategy","deep learning","scaled dot-product attention","temporal convolutional network","STOI","speech intelligibility"],"falsifier":"Compute per-file STOI on the same 20 test clips and report the distribution with a paired confidence interval; if the mean gap of 0.0095 lies within the noise of the measurement, or if removing the attention layer produces the same 0.6031 mean, the claim that this attention-based model approximates ACE is not supported.","tokens_in":7507,"feed_emoji":"🦻","tokens_out":10524,"duration_ms":92874,"temperature":0.7,"pith_summary":"The paper is trying to establish that a temporal convolutional network (TCN) combined with scaled dot-product attention can learn to convert speech into cochlear-implant electrode stimulation patterns (electrodograms) well enough that reconstructed audio is nearly as intelligible as the standard ACE strategy. It reports a mean STOI of 0.6031 for the model versus 0.6126 for ACE on 20 TIMIT test files, and takes the small gap as evidence that a learned coding strategy can preserve the speech cues ACE preserves. The payoff, if true, is a differentiable, data-driven alternative to hand-tuned ACE processing that could be retrained or personalized for individual users. The claim is about approximation, not superiority: the model is trained to imitate ACE electrodograms.","feed_headline":"Attention network nears ACE speech intelligibility","feed_subtitle":"A deep-learning model reaches 0.6031 STOI versus ACE's 0.6126 on TIMIT test speech.","key_machinery":"The central machinery is the scaled dot-product attention mechanism, $\\mathrm{Attention}(Q,K,V) = \\mathrm{softmax}(QK^{\\top}/\\sqrt{d_k})V$, placed on top of a temporal convolutional network. The TCN's causal dilated convolutions supply long-range temporal context, and the attention layer lets the model weight the most relevant parts of that context when deciding stimulation levels for each electrode band. The training objective couples mean squared error for continuous electrodogram values with binary cross entropy for band-selection decisions, and the whole pipeline is supervised by ACE electrodograms, so the architecture is doing imitation learning of ACE rather than discovering a new coding strategy from first principles.","core_discovery":"The paper's central claim is that an encoder–decoder network built from a temporal convolutional backbone and a scaled dot-product attention layer can reproduce ACE electrodograms closely enough that a sine-vocoder reconstruction of the predicted patterns reaches a mean STOI of 0.6031, within 0.0095 of the 0.6126 obtained from the ACE strategy on the same 20-file TIMIT test set. The authors present this as evidence that deep learning can serve as a viable alternative to conventional cochlear-implant coding, retaining the essential spectral and temporal cues while offering the flexibility of a learned, adaptable system. The model is trained on 100 TIMIT utterances whose electrodograms were generated by a Matlab implementation of ACE, optimized with a combined mean-squared-error and binary-cross-entropy loss, and evaluated by vocoding its predicted electrodograms back into audio and computing STOI against the original speech.","pith_inferences":["Because the training target is ACE itself, the model's intelligibility is capped by ACE; training on a perceptual objective such as STOI could in principle push the model past 0.6126, a route the paper does not explore.","The 0.0095 gap is small enough that it may fall within per-sentence and per-listener variability; a paired listening test or confidence intervals on per-file STOI would show whether the difference is perceptible at all.","The attention weights could be inspected to identify which electrodes or frequency bands carry intelligibility-critical information, turning the model into an exploratory tool for cochlear-implant coding rather than only an ACE mimic.","A natural next experiment is to train the same architecture in noisy conditions, since TCNs are known for noise robustness; if attention helps there, the model's advantage over ACE may be larger than the clean-speech numbers suggest."],"forward_implications":["If the network really matches ACE within 0.0095 STOI, a learned front-end could replace the hand-designed ACE stages without a measurable loss of intelligibility on clean speech.","Because the model is differentiable, the same architecture could be fine-tuned on individual listeners' feedback or on noisy speech, which is where ACE's fixed mapping is least flexible.","The reported gap sets a concrete target for further work: an architecture that closes or reverses the 0.0095 gap would have a demonstrable intelligibility advantage over ACE.","The training recipe—ACE electrodograms as targets, a combined MSE/BCE loss, and STOI evaluation—provides an objective loop for testing other neural cochlear-implant coding strategies."],"supporting_citations":[{"why":"Defines the STOI metric used to compare the model's vocoded output with ACE; without it there is no quantitative intelligibility claim.","marker":"[19]"},{"why":"Establishes the ElectrodeNet line of deep-learning ACE emulation and supplies the closest prior STOI baselines this paper extends with a TCN and attention.","marker":"[11]"},{"why":"Provides the scaled dot-product attention operation and the sqrt(d_k) scaling used in the model's attention layer.","marker":"[15]"},{"why":"Demonstrates that an end-to-end deep network can replace a cochlear-implant coding strategy, the premise for training a network to generate electrodograms.","marker":"[9]"},{"why":"Shows a denoising deep-learning coding strategy optimized with a CI-specific loss and evaluated with objective measures, the methodological template for the proposed model.","marker":"[10]"},{"why":"Supplies the ACE processing block diagram that defines the processing stages the network is trained to imitate.","marker":"[8]"}],"fun_headline_variants":["Deep learning CI coding close to ACE on TIMIT","Attention layer brings CI coding within 0.01 of ACE","AI-generated electrodograms rival ACE strategy","DL model achieves near-ACE speech intelligibility for CIs","Attention network for CIs: 0.6031 STOI vs ACE 0.6126"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central comparison rests on treating the average STOI over 20 test files, with no error bars or significance testing, as enough to show that 0.6031 closely approximates 0.6126; the model's ceiling is also fixed by the ACE electrodograms used as training targets.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning CI coding close to ACE on TIMIT","Attention layer brings CI coding within 0.01 of ACE","AI-generated electrodograms rival ACE strategy","DL model achieves near-ACE speech intelligibility for CIs","Attention network for CIs: 0.6031 STOI vs ACE 0.6126"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000348,"raw_usage":{"total_tokens":1877,"prompt_tokens":890,"completion_tokens":987,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":899}},"tokens_in":506,"tokens_out":987,"duration_ms":9705,"temperature":1.0,"reasoning_tokens":899,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:02:55.900764+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute per-file STOI on the same 20 test clips and report the distribution with a paired confidence interval; if the mean gap of 0.0095 lies within the noise of the measurement, or if removing the attention layer produces the same 0.6031 mean, the claim that this attention-based model approximates ACE is not supported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the STOI metric used to compare the model's vocoded output with ACE; without it there is no quantitative intelligibility claim."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the ElectrodeNet line of deep-learning ACE emulation and supplies the closest prior STOI baselines this paper extends with a TCN and attention."},{"cited_title":"Gajecki, W","cited_arxiv_id":null,"evidence_quote":"Demonstrates that an end-to-end deep network can replace a cochlear-implant coding strategy, the premise for training a network to generate electrodograms."},{"cited_title":"Gajecki, Y","cited_arxiv_id":null,"evidence_quote":"Shows a denoising deep-learning coding strategy optimized with a CI-specific loss and evaluated with objective measures, the methodological template for the proposed model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ACE processing block diagram that defines the processing stages the network is trained to imitate."}],"review_version":1}