{"id":"5cf03cd3-8050-422c-82e2-873a730e317d","arxiv_id":"2501.16762","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Transfer entropy and an upper-bound estimate of directed redundancy between EEG electrodes are inversely related to reconstruction distortion for attended speech, but not for distracting speech.","lead":"This paper tests whether information-theoretic rates measured between speech stimuli and EEG reconstructions relate to how well the brain tracks attended speech. It finds that for the attended talker, higher transfer entropy and higher estimated directed redundancy between electrode signals accompany lower reconstruction distortion, but not for the distracting talker.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim conflates the directed-redundancy upper bound R in Eq. (12) with the redundancy itself; R is a min over single-electrode transfer entropies, so the observed relation with distortion may reflect the weakest electrode, not shared information.","rationale":"The reader's weakest assumption concerned whether the speech envelope S can serve as the hidden redundancy process phi, worrying that volume conduction or unrelated neural sources could make R meaningless. My concern is related but distinct and arguably more load-bearing: even if S is exactly phi, Eq. (12) is only an upper bound on I_red, and it is constructed as a minimum over individual single-electrode transfer entropies. Such a minimum is not a measure of redundant information shared among electrodes; it is a weakest-link statistic. Therefore, the observed decreasing trend of D with R for attended speech does not logically imply that the true directed redundancy is inversely proportional to distortion. This is the step that connects the empirical plots to the abstract's central claim. The paper's own language acknowledges the upper-bound status ('The directed redundancy ... is then upper bounded by R'), but the framing in the abstract and conclusions treats R as the directed redundancy itself. A concrete check is available from the existing data: identify which term in Eq. (12) attains the minimum, and compare R against a true redundancy estimate. That check is inexpensive and would settle whether the min-of-singletons construction drives the result. Because the issue is addressable by re-analysis rather than a fundamental mathematical contradiction, the appropriate verdict remains conditional, matching the reader's assessment. I partially agree with the reader because both critiques target the validity of R as a measure of redundancy, but my critique focuses on the min-of-singletons/upper-bound conflation rather than the choice of phi.","tokens_in":12870,"tokens_out":9477,"duration_ms":95397,"concrete_test":"For every trial and distortion-rate bin, record which of the three components in Eq. (12) attains the minimum. If the active term is RS->E or RE->S_hat in the large majority of bins, then re-run the Fig. 3a analysis replacing R with a genuine multi-electrode redundancy estimate, e.g., the minimum over all pairwise shared-information terms from a PID decomposition or the redundancy lower bound of [33]. If the negative slope disappears or reverses, the central claim is an artifact of the min-of-singletons construction; if the monotone relationship survives, then R is a faithful proxy in this dataset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In 'Rate Redundancy in EEG Signals', the paper defines R = min(RS->S_hat, RE->S_hat, RS->E), where RE->S_hat = min_j TE(E_j -> S_hat) and RS->E = min_j TE(S -> E_j). Lemma 1 states that this quantity is an upper bound on the true directed redundancy I_red, not an estimator of it. Because the minimum is taken over six electrodes individually, R is controlled by the weakest single link: one electrode with small TE(S -> E_j) or TE(E_j -> S_hat) forces R down, and R can be large even when the six electrodes carry no shared stimulus information (for example, if each electrode encodes a different part of S). Thus the Fig. 3a regression of distortion D on R does not establish that redundant information among the electrodes is proportional to the stimulus-reconstruction correlation; it may only show that the weakest electrode's stimulus tracking improves as reconstruction quality improves. The abstract's statement that 'the directed redundancy is proportional to the correlation' and the conclusion's 'greater amount of redundant information' go beyond what Eq. (12) supports unless R is shown to be tight or at least monotonically related to the true I_red. The paper never reports which term attains the minimum, so the reader cannot tell whether R reflects a multi-source shared-information term or simply the smallest single-electrode transfer entropy.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes neural speech tracking under a competing-talker EEG paradigm and proposes a rate-distortion interpretation. For six left-temporal electrodes, the authors compute transfer-entropy rates from the speech envelope S to the reconstructed stimulus Ŝ, from S to each electrode, and from each electrode to Ŝ, and combine them via the directed-redundancy upper bound R = min(RS→Ŝ, RE→Ŝ, RS→E) from Eq. (12). They then plot the distortion D = 1 − |ρ| against these rates, fit linear regressions, and report that for attended stimuli both transfer entropy and the directed redundancy are significantly related to distortion, while no such relationship holds for distracting stimuli. The central theoretical apparatus is the directed-redundancy framework of Ref. [34], with the hidden redundancy process φ identified with the speech envelope S.","tokens_in":13210,"tokens_out":3471,"duration_ms":33328,"significance":"If the claims are sustained, the paper would provide a rate-distortion perspective on cortical speech tracking and connect an information-theoretic measure of directed redundancy to EEG electrode correlations. The empirical use of real EEG data, the standard mTRF decoder, and the strong transfer-entropy result for the attended stimulus (p = 6.1e-06 in Table 1) are concrete strengths. However, the headline redundancy claim rests on an upper-bound proxy R, not on the exact directed redundancy, and the statistical evidence for that claim is marginal. The significance is therefore conditional on validating that R behaves as a faithful indicator of the true redundancy in this EEG setting.","major_comments":[{"comment":"Eq. (12) defines R as min(RS→Ŝ, RE→Ŝ, RS→E), and Lemma 1 states that this quantity is only an upper bound on the directed redundancy I^red, not an estimate or measurement of it. The abstract's claim that 'the directed redundancy is proportional to the correlation' and the conclusion's 'greater amount of redundant information' therefore go beyond what the reported analysis supports: the regressions in Fig. 3 and Table 1 are against R, not against I^red. Because the minimum is taken over six individual electrodes, R can be controlled by the single weakest electrode's transfer entropy, so the observed R–D relationship may reflect the weakest single-electrode tracking performance rather than shared or redundant information among electrodes. The paper does not report which term attains the minimum, nor any evidence that R is tight or monotonically related to I^red in this application. Please provide such evidence, or explicitly soften the claims to refer to an upper-bound proxy.","section":"Rate Redundancy in EEG Signals, Eq. (12)"},{"comment":"The identification of the speech envelope S with the hidden redundancy process φ is assumed without validation. The directed-redundancy framework of [34] requires that φ causally drives the redundant information in the source processes, but EEG electrode correlations are also driven by volume conduction and by shared neural sources that may not be causally linked to S. If φ is not S, then R does not measure redundancy about the stimulus, and the interpreted meaning of the R–D relationship is lost. A control analysis, for example using a surrogate φ unrelated to S or comparing with a non-causal redundancy measure, is needed to support the mapping between the model in Fig. 1b and the EEG setup.","section":"Fig. 1b and Section 'Rate Redundancy in EEG Signals'"},{"comment":"The p-values for R are marginal (0.0177 for attended, 0.0416 for distractor) and are reported without any multiple-comparison correction across the four rate measures and two attention conditions. In addition, the support threshold 'pdf > 0.01' and the 0.005-bit bin width are introduced after inspecting the data, and no confidence intervals or effect sizes are reported for the regression slopes. These choices make the statistical case for the redundancy–distortion proportionality fragile. The strong RS→Ŝ result (p = 6.1e-06) can support the transfer-entropy claim, but the redundancy claim needs a pre-specified or corrected inference procedure, or at minimum a report of confidence intervals for the slopes.","section":"Table 1 and Fig. 3"}],"minor_comments":[{"comment":"The sentence describing transfer entropy refers to 'the target Y', but Eq. (1) uses Z as the target; please align the notation.","section":"Introduction, after Eq. (1)"},{"comment":"The label 'RR→Ŝ' in Fig. 2(c) appears to be a typo for 'RE→Ŝ' as defined in Eq. (9); please correct it for consistency.","section":"Simulation Study, Fig. 2"},{"comment":"The sentence 'I(En(1),...,En(|LT|); Ŝn) = 0 or ∞' is confusing, since for continuous variables mutual information with a deterministic function is generally infinite, not zero; the intended statement needs clarification.","section":"Section 'Rate Redundancy in EEG Signals'"},{"comment":"The operational distortion-rate curves would benefit from error bars or confidence bands; as presented, the visual trend in Fig. 3b is difficult to assess without accompanying uncertainty information.","section":"Fig. 3"},{"comment":"There is a typo, 'phenomonen' for 'phenomenon', in the Introduction; the paper also uses 'directed redundancy' and 'rate redundancy' interchangeably, which should be defined consistently.","section":"Abstract and Introduction"}],"recommendation":"major_revision","confidential_remarks":"The central theoretical quantity R is inherited from the authors' own prior work ([33], [34]) and is used as a proxy for directed redundancy without independent validation in this EEG setting. Given that the empirical claims are framed around this quantity, I would recommend that a reviewer with expertise in information decomposition specifically assess whether Lemma 1's upper bound is appropriate when the sources are six electrodes and the target is a deterministic function of those electrodes. The paper's scope fits a signal-processing/information-theory venue, but the match with the journal's interests in rate-distortion theory should be judged accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. The attended transfer-entropy result is real and worth knowing. The redundancy headline is not yet supported: R in Eq. (12) is an upper bound on directed redundancy, not the quantity itself, and because it is a min over single-electrode transfer entropies it can be controlled by the weakest electrode rather than by shared information. The abstract and conclusion say “directed redundancy is proportional” to correlation, but the analysis actually regresses distortion on this upper bound.\n\nWhat is genuinely new: this is the first application of the directed redundancy measure from [34] to a real EEG competing-talker dataset. Linear stimulus reconstruction and transfer entropy are standard, but the combination is new and the attended-versus-distractor contrast is a reasonable empirical question. The attended RS→Ŝ result is statistically strong (p = 6.1e-06) and consistent with the known phenomenon of cortical speech tracking. The distractor finding—no monotone relationship, and an actually positive slope for R—is interesting and could be informative about how attention changes redundancy.\n\nNow the soft spots, in rough proportion. First and most important: R is a minimum over three single-electrode transfer entropies. The paper never reports which term attains the minimum, and it offers no evidence that R is tight or even monotonically related to the true directed redundancy. So the regression of distortion on R may simply reflect the weakest electrode’s tracking quality, not shared information across electrodes. That is a load-bearing interpretive gap for the paper’s central claim. Second, the TE estimator and its parameters are not specified for the EEG analysis; the only estimator named appears in an example from the authors’ prior work. Without that detail, the numbers are not reproducible. Third, the p-values for R are marginal (0.0177 attended, 0.0416 distractor) and no multiple-comparison correction is applied; with four rates and two conditions, one of those p-values is expected by chance. Fourth, the support threshold (pdf > 0.01) is post hoc, and the overlapping rate bins smooth the curves in ways that can make trends look cleaner than they are. The assumption that the speech envelope S is the hidden redundancy process φ is reasonable and clearly stated in the model, but unvalidated; I would want that acknowledged as an assumption rather than treated as fact. Self-citation to [33, 34] is not a problem here because the measure is genuinely theirs.\n\nWho is this for? Researchers in auditory attention decoding and anyone working on information-theoretic measures of neural redundancy. The attended transfer-entropy relationship alone justifies a serious referee. The directed redundancy claim needs more work before it can be accepted. If the authors report which terms bind in Eq. (12), specify the TE estimator, add bootstrap confidence intervals and multiple-comparison corrections, and soften the language from “redirected redundancy” to “upper bound on directed redundancy,” I would be happy to see it published. I would not desk-reject it.","headline":"A useful first application of directed redundancy to real EEG speech-tracking data, but the main claim treats an upper bound R as the quantity itself, and the redundancy result rests on marginal statistics.","tokens_in":13702,"tokens_out":3322,"would_cite":true,"duration_ms":32779,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A17","94A34"],"pacs":[],"model":"deepseek-v4-flash","headline":"For attended speech, more redundant EEG information means more accurate neural tracking.","keywords":["directed redundancy","transfer entropy","rate-distortion","neural speech tracking","auditory attention decoding","electroencephalography (EEG)","cocktail party","hidden redundancy process"],"falsifier":"Shuffle or time-shift the six electrode signals relative to the speech envelopes to break the causal $S \\to$ electrode link while preserving the marginal statistics of each signal; if the attended distortion $D$ still decreases with $R$ at comparable strength, the causal-redundancy interpretation is not supported. A second check is to repeat the analysis on electrodes from a non-auditory scalp region: a preserved attended trend would show the effect is not specific to the speech-tracking pathway.","tokens_in":12692,"feed_emoji":"🧠","tokens_out":10930,"duration_ms":87878,"temperature":0.7,"pith_summary":"This paper tries to establish that neural tracking of attended speech obeys a rate-distortion relation: for the talker a subject is listening to, the transfer entropy from the speech envelope to the EEG-reconstructed envelope, and the directed redundancy of the left-temporal electrode signals, both grow as the reconstruction distortion $D = 1 - |\\rho|$ shrinks. The authors compute these quantities on a competing-talker EEG dataset and fit linear models to the distortion-rate points, finding statistically significant negative slopes for the attended condition for the rate $R_{S\\to\\hat S}$ and for the redundancy bound $R$. For the distracting talker the relationship is absent or reversed, with the redundancy slope positive and significant. If the claim holds, attention can be read out information-theoretically: better tracking of the attended talker is accompanied by more causal information flow and more redundancy across electrodes, and this signature is specific to the attended stream.","feed_headline":"Attended speech obeys an EEG rate-distortion law","feed_subtitle":"More neural redundancy means better tracking of the talker you attend to — and not the ignored one.","key_machinery":"The carrying object is the directed-redundancy upper bound of Eq. (12), $R = \\min(R_{S\\to\\hat S}, R_{E\\to\\hat S}, R_{S\\to E})$, where each $R_{\\cdot\\to\\cdot}$ is a transfer entropy and $R_{E\\to\\hat S} = \\min_j TE(E_n(j) \\to \\hat S_n)$ runs over the six left-temporal electrodes. This bound comes from Lemma 1 of [34], which replaces the causal minimal sufficient statistics in the definition of directed redundancy by the raw source processes, so that the min over the weakest causal link in the chain $S \\to$ electrodes $\\to \\hat S$ upper-bounds the redundant information about the speech envelope that the electrodes share with the reconstruction. The distortion measure $D = 1 - |\\rho|$ turns higher Pearson correlation between stimulus and reconstruction into lower distortion, and the linear regressions of $D$ (in dB) on the rates provide the reported significance tests.","core_discovery":"The central claim is that, for attended speech, the operational distortion $D = 1 - |\\rho|$ between the speech envelope and the envelope reconstructed from six left-temporal EEG electrodes is linearly related to two information measures: the transfer entropy $R_{S\\to\\hat S}$ from the stimulus to the reconstruction, and the directed-redundancy upper bound $R = \\min(R_{S\\to\\hat S}, R_{E\\to\\hat S}, R_{S\\to E})$. Higher values of either quantity go with lower distortion, i.e. with a higher correlation between the attended stimulus and the neural reconstruction; the fitted slopes are negative and the $p$-values are $6.1\\times10^{-6}$ for $R_{S\\to\\hat S}$ and $0.0177$ for $R$. For the distractor, the same relation is not observed: the redundancy slope is positive and significant ($p=0.0416$), so more redundant electrode activity is associated with worse tracking of the ignored talker. The paper reads this as evidence that attention concentrates rate and redundancy in the neural pathway that tracks the attended speech.","pith_inferences":["A testable extension would be to use $R$ or $R_{S\\to\\hat S}$ directly as an auditory attention decoder, comparing its trial-level accuracy against the standard correlation-based decoder, especially in low-correlation trials where the linear relationship may be easier to detect.","The reversed distractor slope hints that $R$ could separate speech-driven shared information from volume-conduction shared activity; a source-localization or MEG study could test whether the distractor redundancy originates from non-auditory generators.","Because $R$ is an upper bound from Lemma 1 rather than the true directed redundancy, the linear relation could partly be an artifact of the bound's slack; estimating the causal minimal sufficient statistics directly would show whether the proportionality survives.","The analysis uses linear reconstruction; whether the same rate-distortion relation holds for nonlinear decoders or for spectro-temporal features beyond the envelope is left open by the paper."],"forward_implications":["For attended speech, the stimulus-to-reconstruction transfer entropy can be used as a predictor of reconstruction quality: larger $R_{S\\to\\hat S}$ implies smaller $D$.","The directed-redundancy bound $R$ across left-temporal electrodes is likewise a linear predictor of distortion in the attended condition, so redundancy in the attended pathway is informative rather than wasted.","For the distractor, the same monotone relation fails, and the redundancy slope reverses sign, indicating that redundant electrode activity in the ignored pathway does not reflect faithful tracking.","The results suggest that operational rate-distortion thinking applies to cortical speech tracking: the attended stream is encoded with higher rate and higher redundancy at lower distortion, matching a rate-distortion tradeoff at the scalp level."],"supporting_citations":[{"why":"defines the directed redundancy measure and Lemma 1, the upper bound R that the paper computes.","marker":"[34]"},{"why":"supplies the competing-talker EEG paradigm and the linear stimulus-reconstruction approach used to form the reconstructed envelope.","marker":"[23]"},{"why":"provides the decoder-estimation software used to fit the linear reconstruction coefficients g.","marker":"[38]"},{"why":"supplies the EEG dataset of 15 subjects with competing male and female talkers used in the simulation study.","marker":"[39]"},{"why":"defines transfer entropy, the building block of all rates and of the redundancy bound.","marker":"[35]"}],"fun_headline_variants":["Attention boosts neural redundancy only for the talker you hear","EEG redundancy tracks attended speech, not the ignored talker","Rate-distortion holds only for attended speech, not distractors","More EEG redundancy means better tracking of attended speech only"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the speech envelope $S$ is the hidden redundancy process $\\phi$ of the directed-redundancy framework, so that $R$ in Eq. (12) is a valid upper bound on redundant speech information actually carried by the electrodes; if the shared electrode activity is dominated by volume conduction or by neural sources unrelated to $S$, the observed relation between $R$ and distortion would not have the interpretation the paper gives it.","fun_headline_variants_meta":{"raw":{"variants":["Attention boosts neural redundancy only for the talker you hear","EEG redundancy tracks attended speech, not the ignored talker","Rate-distortion holds only for attended speech, not distractors","More EEG redundancy means better tracking of attended speech only"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000573,"raw_usage":{"total_tokens":2701,"prompt_tokens":931,"completion_tokens":1770,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":1702}},"tokens_in":547,"tokens_out":1770,"duration_ms":12033,"temperature":1.0,"reasoning_tokens":1702,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T10:54:25.402099+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Shuffle or time-shift the six electrode signals relative to the speech envelopes to break the causal $S \\to$ electrode link while preserving the marginal statistics of each signal; if the attended distortion $D$ still decreases with $R$ at comparable strength, the causal-redundancy interpretation is not supported. A second check is to repeat the analysis on electrodes from a non-auditory scalp region: a preserved attended trend would show the effect is not specific to the speech-tracking pathway.","supporting_citations":[{"cited_title":"Directed redundancy in time series,","cited_arxiv_id":null,"evidence_quote":"defines the directed redundancy measure and Lemma 1, the upper bound R that the paper computes."},{"cited_title":"Attentional selection in a cocktail party environment can be decoded from single-trial eeg,","cited_arxiv_id":null,"evidence_quote":"supplies the competing-talker EEG paradigm and the linear stimulus-reconstruction approach used to form the reconstructed envelope."},{"cited_title":"The multivariate temporal response function (mtrf) toolbox: a matlab toolbox for relating neural signals to continuous stimuli,","cited_arxiv_id":null,"evidence_quote":"provides the decoder-estimation software used to fit the linear reconstruction coefficients g."},{"cited_title":"Noise-robust cortical tracking of attended speech in real-life environments,","cited_arxiv_id":null,"evidence_quote":"supplies the EEG dataset of 15 subjects with competing male and female talkers used in the simulation study."},{"cited_title":"Measuring information transfer,","cited_arxiv_id":null,"evidence_quote":"defines transfer entropy, the building block of all rates and of the redundancy bound."}],"review_version":1}