{"id":"db56a164-eda2-4351-87ed-fb2981b13a63","arxiv_id":"2501.01673","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"IFENet combines dual-path Mamba for speech and Kolmogorov-Arnold Networks for EEG to extract the attended speaker, achieving state-of-the-art SI-SDR on the KUL and AVED datasets.","lead":"A new speech extraction network, IFENet, uses a Mamba-based speech encoder and a KAN-based EEG encoder to pull out the voice a listener is attending to, improving SI-SDR by 36% and 29% over an earlier model on two datasets. The method could improve hearing aids and cochlear implants by using brain signals to identify the target speaker.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'state-of-the-art' claim is under-supported: Table I omits NeuroHeed/NeuroHeed+, the strongest published baselines the paper itself cites, and reports no significance information.","rationale":"The reader's weakest assumption concerns per-subject training and the unbenchmarked AVED dataset. That is a fair concern about generalization, but per-subject training is standard in EEG-based auditory attention decoding, and the KUL dataset is a public, widely used benchmark. The more directly load-bearing weakness for the paper's central claim is the baseline selection: the paper calls IFENet state-of-the-art but Table I omits the NeuroHeed and NeuroHeed+ systems that the introduction itself identifies as prior neuro-steered extraction models. Without those comparisons, the 36%/29% figures over MSFNet do not establish SOTA status. This is a concrete, testable omission rather than a disagreement with field consensus. The paper has independent support in the form of clear ablations showing both proposed modules contribute, especially EEGKAN, and the KUL numbers are internally consistent with the reported relative improvements. I therefore do not recommend changing the CONDITIONAL verdict; the condition should be that the authors compare against NeuroHeed/NeuroHeed+ and report statistical variability before the SOTA claim is accepted.","tokens_in":8449,"tokens_out":4591,"duration_ms":48464,"concrete_test":"Obtain or re-implement the official NeuroHeed+ system (ref. [24]) and evaluate it on the identical per-subject 80/10/10 splits of the KUL dataset used for IFENet. If NeuroHeed+ attains an SI-SDR of at least 6.85 dB, or if its difference from IFENet is within the run-to-run variability, the 'outperforms state-of-the-art' claim fails. Report per-subject results and standard deviations across at least five training seeds for both models to test whether the 0.55 dB gap over BASEN* on KUL is statistically meaningful.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that IFENet outperforms the state-of-the-art model, and Table I is the only evidence offered. However, Table I compares IFENet only with BASEN, BASEN*, and MSFNet; it does not compare against NeuroHeed or NeuroHeed+, which the introduction explicitly cites as existing neuro-steered speaker extraction systems. The claimed 36% and 29% relative improvements are computed against MSFNet, not against the best available published system, so the headline 'state-of-the-art' conclusion does not follow from the presented comparison. This is not merely a citation issue: if NeuroHeed+ performs at or above IFENet under the same protocol, the central claim is false. The concern is sharpened by the absence of standard deviations, confidence intervals, or repeated-run results; on KUL, IFENet (6.85 dB) is only 0.55 dB above BASEN* (6.30 dB), and it is unclear whether this gap is reproducible. The paper's own ablation shows EEGKAN drives most of the gain, but that does not establish superiority over the strongest prior systems.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes IFENet, a time-domain neural network for EEG-guided target speaker extraction. The architecture uses a SpeechBiMamba encoder based on dual-path Mamba for long-range speech modeling, and an EEGKAN encoder that combines multi-head attention with Kolmogorov-Arnold Networks for EEG feature extraction. The model is evaluated on the KUL dataset and a newly introduced AVED dataset against BASEN, BASEN*, and MSFNet, reporting relative SI-SDR improvements of 36% and 29% over MSFNet, with ablations showing that EEGKAN contributes most of the gain.","tokens_in":8709,"tokens_out":5201,"duration_ms":46882,"significance":"If the reported results hold, IFENet offers a competitive architecture for neuro-steered speaker extraction, and the combination of dual-path Mamba for speech and KAN-based attention for EEG is a sensible design direction. The ablation study in Table II clearly isolates the contribution of EEGKAN and demonstrates that it is the dominant factor (KUL: 6.85 vs. 3.89 SI-SDR without EEGKAN). The introduction of the Mandarin AVED dataset is a potentially useful new resource for the community. However, the headline claim of state-of-the-art performance is undercut by the omission of the strongest published baselines (NeuroHeed, NeuroHeed+) from the comparison and by the absence of any statistical significance or variance information.","major_comments":[{"comment":"The claim that IFENet 'outperforms the state-of-the-art model' is not supported by the presented comparison. The Introduction cites NeuroHeed [23] and NeuroHeed+ [24] as existing neuro-steered speaker extraction systems, but Table I compares IFENet only with BASEN, BASEN*, and MSFNet. The headline relative improvements of 36% and 29% are computed against MSFNet, which is not established as the best prior method. Without including NeuroHeed/NeuroHeed+ under the same protocol, the central SOTA claim does not follow from the evidence.","section":"Section IV.A, Table I"},{"comment":"The evaluation reports a single run per condition with no standard deviations, confidence intervals, or significance tests. In Table I, on KUL, IFENet's SI-SDR (6.85 dB) is only 0.55 dB above BASEN* (6.30 dB); without repeated runs or statistical testing, the reader cannot judge whether this gap is reproducible. Given that the central claim is superiority over prior systems, measures of variability are essential.","section":"Section III.B, Section IV.A"},{"comment":"The text misreports the AVED results: it states that IFENet outperforms MSFNet by '0.2 in PESQ, 0.3 in STOI, and 0.6 in ESTOI', but the differences in Table I are 0.20, 0.03, and 0.06, respectively. This factor-of-ten error obscures the actual magnitude of the improvements and must be corrected and checked against the raw results.","section":"Section IV.A"},{"comment":"The abstract says the results are achieved 'under an open evaluation condition', but this term is never defined in the experimental setup or results. If it refers to the per-subject 80/10/10 trial split described in Section III.B, this should be stated explicitly; if it refers to a different protocol, the description in Section III.B does not currently match it. As written, the headline condition is unverifiable.","section":"Abstract, Section III.B"}],"minor_comments":[{"comment":"There are repeated typographical issues: 'Datesets' should be 'Datasets', and 'A VED' appears with a space in many places (e.g., Table I, Section IV).","section":"Section III.A"},{"comment":"The AVED dataset description lacks basic acoustic and task details, such as signal-to-noise ratio of the mixtures, trial duration, speech segment lengths, and how the two talkers are paired. This makes it difficult to compare the new dataset with existing benchmarks or to interpret the absolute SI-SDR values, which are noticeably higher for all methods on AVED than on KUL.","section":"Section III.A.2"},{"comment":"The citation for MSFNet is incomplete: 'in ACM Multimedia. in Proc.' appears truncated and the publication year is missing.","section":"Reference [25]"},{"comment":"The SI-SDR formula is typeset ambiguously in the provided text; the numerator and denominator need clear delimiters so the projection term is unambiguous.","section":"Section II.D, Eq. (3)"},{"comment":"The description of SpeechBiMamba does not explain how the 'flip' operation is implemented (e.g., time reversal) or how the outputs of the forward and backward Mamba passes are combined before the concatenation step.","section":"Section II.B"},{"comment":"The text says EEGKAN consists of a multi-head attention block, three KAN layers, and dropouts, but Fig. 1 labels the EEG encoder block with '×5'; the relationship between this factor and the number of KAN layers/number of layers in the EEGKAN module is not clarified.","section":"Section II.C, Fig. 1/Fig. 2"}],"recommendation":"major_revision","confidential_remarks":"The omission of NeuroHeed and NeuroHeed+ is the most serious issue: these are the strongest published baselines in the neuro-steered extraction literature, and the paper cites them in the Introduction. The authors should be asked either to include them under the same protocol or to justify explicitly why MSFNet is the appropriate SOTA reference. Also note that MSFNet appears to be from the same research group, so the self-comparison should be supplemented with independent baselines."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: the paper is a real engineering contribution—SpeechBiMamba + EEGKAN is a sensible combination and the ablation cleanly shows EEGKAN drives the gain—but the \"state-of-the-art\" claim doesn't hold. Table I compares only against BASEN, BASEN*, and MSFNet, omitting NeuroHeed and NeuroHeed+, which the paper itself cites in the introduction and which are the strongest published neuro-steered extractors. So the 36%/29% numbers are relative to MSFNet, not to the best available system. The stress-test note is right; I don't think it's contrived.\n\nWhat's new: applying dual-path Mamba to speech encoding and KAN to EEG encoding in this task. The ablation is well done: removing EEGKAN drops SI-SDR from 6.85 to 3.89 on KUL, so the attention-plus-KAN encoder is doing real work. The AVED dataset is a resource, but it's new, self-collected, and not yet benchmarked elsewhere, so results on it are less informative.\n\nSoft spots in proportion: (1) Missing baselines is the big one—the SOTA claim is unsupported as written. (2) \"Open evaluation condition\" is undefined; the per-subject 80/10/10 split is described, but what \"open\" means is not. (3) No variance, confidence intervals, or repeated-run data. On KUL, IFENet (6.85) is only 0.55 dB above BASEN* (6.30), and on typical datasets that difference may not be reproducible. (4) Minor: the paper says \"interoperability\" where it means \"interpretability\" for KAN. (5) The new dataset needs external validation.\n\nThat said, the method is plausible and the paper ships enough detail to reproduce. If the authors add NeuroHeed/NeuroHeed+ and BESD/U-BESD under the same protocol, report variance, and define the evaluation condition, the engineering claim would be much firmer. As is, I'd treat it as a qualified contribution, not a SOTA result.\n\nFor peer review: yes, send it out. A serious referee can help fix the evaluation. For a reading group, maybe; it's a good case study in baseline selection. I wouldn't cite the SOTA claim, but I'd cite the architecture if my work touched this niche.","headline":"Real architecture, clean ablation, but the SOTA claim is under-supported because Table I skips the strongest baselines.","tokens_in":9253,"tokens_out":1939,"would_cite":true,"duration_ms":18729,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims IFENet, combining a dual-path Mamba speech encoder with a KAN-based EEG encoder, outperforms the previous best model, with relative SI-SDR gains of 36% on KUL and 29% on AVED.","keywords":["target speaker extraction","EEG","auditory attention decoding","Mamba","Kolmogorov-Arnold Networks","SI-SDR","speech separation","brain-computer interface"],"falsifier":"Re-run IFENet and MSFNet under leave-one-subject-out cross-validation on the KUL dataset (train on 15 subjects, test on the 16th) and compare SI-SDR; if the relative gain disappears or reverses, the claim that the modules generalize beyond per-subject training is undermined. An external subject-independent benchmark on AVED would test the same point for the 29% gain.","tokens_in":8261,"feed_emoji":"🧠","tokens_out":7935,"duration_ms":72689,"temperature":0.7,"pith_summary":"The paper sets out to improve neuro-oriented target speaker extraction—pulling the attended talker out of a mixed recording using the listener's EEG as the clue to who is being attended. It proposes IFENet, a time-domain network whose speech encoder uses dual-path Mamba (SpeechBiMamba) to model local and global structure in long speech sequences, and whose EEG encoder uses Kolmogorov-Arnold layers inside an attention block (EEGKAN) to identify the attended speaker. On the KUL and newly introduced AVED datasets, IFENet reports relative SI-SDR gains of 36% and 29% over the MSFNet baseline, together with higher PESQ, STOI, and ESTOI scores. The ablation study shows that the EEGKAN module is the larger contributor, indicating that the quality of EEG feature extraction, not just speech sequence modeling, drives the improvement.","feed_headline":"EEG-guided speaker extraction beats prior best by 36 percent","feed_subtitle":"A dual-path Mamba speech encoder plus a KAN-based EEG encoder raises extracted speech quality on two datasets.","key_machinery":"The machinery is a pair of feature-extraction modules inserted into a Conv-TasNet-style pipeline. SpeechBiMamba is a dual-path Mamba: it scans the speech representation through selective state-space blocks once forward and once backward, over local segments and over the whole sequence, so long-range dependencies are modeled in linear time. EEGKAN is an attention-based EEG encoder in which the feed-forward MLPs are replaced by Kolmogorov-Arnold Networks, giving the EEG branch learnable univariate activation functions that the paper argues better capture speaker-related attention information; a convolutional multi-layer cross-attention (CMCA) module then fuses the speech and EEG embeddings before mask estimation.","core_discovery":"The paper reports that IFENet—a time-domain, end-to-end network built on the Conv-TasNet encoder–mask–decoder structure—extracts the attended speaker from a mixture using EEG as the only cue. Its speech encoder, SpeechBiMamba, applies dual-path Mamba blocks forward and backward over both local segments and the full sequence to capture long-range speech structure; its EEG encoder, EEGKAN, replaces MLP layers with Kolmogorov-Arnold Networks inside a multi-head attention block. On the KUL dataset IFENet reaches an SI-SDR of 6.85 dB versus MSFNet's 5.05 dB, a 36% relative improvement, and on the AVED dataset 8.76 dB versus 6.78 dB, a 29% relative improvement, with higher PESQ, STOI, and ESTOI in both cases. The ablation study attributes the larger share of the gain to EEGKAN: removing it drops KUL SI-SDR to 3.89 dB, whereas removing SpeechBiMamba leaves 6.64 dB.","pith_inferences":["Beyond the paper, the larger ablation penalty for removing EEGKAN (KUL SI-SDR falls from 6.85 to 3.89 dB) suggests that a simpler speech encoder paired with a strong EEG front-end could reproduce most of the gain; testing that configuration would separate the two modules' contributions.","Beyond the paper, because models are trained and tested per subject, the reported gains may reflect subject-specific EEG signatures rather than a general attention decoder; a leave-one-subject-out experiment on KUL would show which.","Beyond the paper, the newly introduced AVED dataset is Mandarin and recorded in-lab, so releasing it and having independent groups benchmark on it would establish whether the 29% relative improvement transfers across recording setups."],"forward_implications":["The 36% and 29% relative SI-SDR gains over MSFNet, if reproduced, mean the architecture is a strong candidate for replacing CNN-only feature extractors in neuro-steered speech extraction.","The ablation result implies that EEG feature extraction carries most of the improvement, so future work that strengthens the EEG branch rather than the speech branch should be prioritized within this architecture.","Because the system runs entirely in the time domain and needs no enrollment utterance from the target speaker, it is compatible with hearing-assist scenarios where only the listener's EEG and the mixed audio are available.","The consistent gains across the English KUL and Mandarin AVED recordings suggest the modules are not tied to one language or stimulus format."],"supporting_citations":[{"why":"MSFNet is the state-of-the-art baseline whose SI-SDR, STOI, ESTOI, and PESQ scores IFENet is directly compared against and improved by 36% and 29%.","marker":"[25]"},{"why":"BASEN is a prior brain-assisted enhancement network baseline, and its convolutional multi-layer cross-attention module is reused in IFENet for speech–EEG fusion.","marker":"[22]"},{"why":"The KUL dataset supplies one of the two evaluation benchmarks, including the 16-subject, 64-channel EEG recording protocol used in the experiments.","marker":"[28]"},{"why":"Mamba provides the selective state-space model backbone that SpeechBiMamba builds on for long-sequence speech modeling.","marker":"[32]"},{"why":"Dual-Path Mamba supplies the short- and long-term bidirectional sequence modeling structure adapted by SpeechBiMamba.","marker":"[40]"},{"why":"Kolmogorov-Arnold Networks define the learnable-activation building block that EEGKAN substitutes for MLP layers in the EEG encoder.","marker":"[42]"},{"why":"NeuroHeed is the attention-based neuro-steered extraction model cited as motivation for using attention mechanisms in the EEG encoder.","marker":"[23]"},{"why":"Conv-TasNet supplies the end-to-end time-domain encoder–mask–decoder backbone on which IFENet is structured.","marker":"[31]"},{"why":"The SI-SDR metric is both the evaluation metric and the training loss function used to optimize and compare the extraction models.","marker":"[29]"}],"fun_headline_variants":["EEG-guided speech extraction beats prior best by 36%","Dual-path Mamba + EEG-KAN sharpens target speech extraction","Brain-wave cues help Mamba-KAN model extract target speaker","IFENet: EEG-attended speaker extraction gains 36% SI-SDR","Mamba and KAN team up for EEG-driven speaker extraction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that per-subject training and the newly recorded AVED dataset represent real-world use; if the method must generalize to unseen subjects, or if AVED is not a fair proxy, the reported 36% and 29% gains may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["EEG-guided speech extraction beats prior best by 36%","Dual-path Mamba + EEG-KAN sharpens target speech extraction","Brain-wave cues help Mamba-KAN model extract target speaker","IFENet: EEG-attended speaker extraction gains 36% SI-SDR","Mamba and KAN team up for EEG-driven speaker extraction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000854,"raw_usage":{"total_tokens":3719,"prompt_tokens":959,"completion_tokens":2760,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":2668}},"tokens_in":575,"tokens_out":2760,"duration_ms":17943,"temperature":1.0,"reasoning_tokens":2668,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:22:23.883912+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run IFENet and MSFNet under leave-one-subject-out cross-validation on the KUL dataset (train on 15 subjects, test on the 16th) and compare SI-SDR; if the relative gain disappears or reverses, the claim that the modules generalize beyond per-subject training is undermined. An external subject-independent benchmark on AVED would test the same point for the 29% gain.","supporting_citations":[{"cited_title":"MSFNet: Multi-scale fusion network for brain-controlled speaker extraction,","cited_arxiv_id":null,"evidence_quote":"MSFNet is the state-of-the-art baseline whose SI-SDR, STOI, ESTOI, and PESQ scores IFENet is directly compared against and improved by 36% and 29%."},{"cited_title":"BASEN: Time-domain brain-assisted speech enhancement network with convolutional cross attention in multi-talker conditions,","cited_arxiv_id":null,"evidence_quote":"BASEN is a prior brain-assisted enhancement network baseline, and its convolutional multi-layer cross-attention module is reused in IFENet for speech–EEG fusion."},{"cited_title":"Neuroheed: Neuro- steered speaker extraction using EEG signals,","cited_arxiv_id":null,"evidence_quote":"NeuroHeed is the attention-based neuro-steered extraction model cited as motivation for using attention mechanisms in the EEG encoder."},{"cited_title":"Conv-TasNet: Surpassing ideal time- frequency magnitude masking for speech separation,","cited_arxiv_id":null,"evidence_quote":"Conv-TasNet supplies the end-to-end time-domain encoder–mask–decoder backbone on which IFENet is structured."},{"cited_title":"SDR-half-baked or well done?","cited_arxiv_id":null,"evidence_quote":"The SI-SDR metric is both the evaluation metric and the training loss function used to optimize and compare the extraction models."}],"review_version":1}