REVIEW 3 major objections 5 minor 39 references
Rate-Distortion under Neural Tracking of Speech: A Directed Redundancy Approach
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read For attended speech, more redundant EEG information means more accurate neural tracking.
desk verdict A useful first application of directed redundancy to real EEG speech-tracking data, but the main claim treats an upper bound R as the quantity itself, and the redundancy result rests on marginal statistics. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the directed-redundancy upper bound of Eq. (12), $R = \min(R_{S\to\hat S}, R_{E\to\hat S}, R_{S\to E})$, where each $R_{\cdot\to\cdot}$ is a transfer entropy and $R_{E\to\hat S} = \min_j TE(E_n(j) \to \hat S_n)$ runs over the six left-temporal electrodes. This bound comes from Lemma 1 of [34], which replaces the causal minimal sufficient statistics in the definition of directed redundancy by the raw source processes, so that the min over the weakest causal link in the chain $S \to$ electrodes $\to \hat S$ upper-bounds the redundant information about the speech envelope that the electrodes share with the reconstruction. The distortion measure $D = 1 - |\rho|$ turns higher Pearson correlation between stimulus and reconstruction into lower distortion, and the linear regressions of $D$ (in dB) on the rates provide the reported significance tests.
What would settle it
Shuffle or time-shift the six electrode signals relative to the speech envelopes to break the causal $S \to$ electrode link while preserving the marginal statistics of each signal; if the attended distortion $D$ still decreases with $R$ at comparable strength, the causal-redundancy interpretation is not supported. A second check is to repeat the analysis on electrodes from a non-auditory scalp region: a preserved attended trend would show the effect is not specific to the speech-tracking pathway.
Extended reading notes
Core claim
The central claim is that, for attended speech, the operational distortion $D = 1 - |\rho|$ between the speech envelope and the envelope reconstructed from six left-temporal EEG electrodes is linearly related to two information measures: the transfer entropy $R_{S\to\hat S}$ from the stimulus to the reconstruction, and the directed-redundancy upper bound $R = \min(R_{S\to\hat S}, R_{E\to\hat S}, R_{S\to E})$. Higher values of either quantity go with lower distortion, i.e. with a higher correlation between the attended stimulus and the neural reconstruction; the fitted slopes are negative and the $p$-values are $6.1\times10^{-6}$ for $R_{S\to\hat S}$ and $0.0177$ for $R$. For the distractor, the same relation is not observed: the redundancy slope is positive and significant ($p=0.0416$), so more redundant electrode activity is associated with worse tracking of the ignored talker. The paper reads this as evidence that attention concentrates rate and redundancy in the neural pathway that tracks the attended speech.
Load-bearing premise
The argument assumes that the speech envelope $S$ is the hidden redundancy process $\phi$ of the directed-redundancy framework, so that $R$ in Eq. (12) is a valid upper bound on redundant speech information actually carried by the electrodes; if the shared electrode activity is dominated by volume conduction or by neural sources unrelated to $S$, the observed relation between $R$ and distortion would not have the interpretation the paper gives it.
Editorial extensions
If this is right
- For attended speech, the stimulus-to-reconstruction transfer entropy can be used as a predictor of reconstruction quality: larger $R_{S\to\hat S}$ implies smaller $D$.
- The directed-redundancy bound $R$ across left-temporal electrodes is likewise a linear predictor of distortion in the attended condition, so redundancy in the attended pathway is informative rather than wasted.
- For the distractor, the same monotone relation fails, and the redundancy slope reverses sign, indicating that redundant electrode activity in the ignored pathway does not reflect faithful tracking.
- The results suggest that operational rate-distortion thinking applies to cortical speech tracking: the attended stream is encoded with higher rate and higher redundancy at lower distortion, matching a rate-distortion tradeoff at the scalp level.
Reading between the lines
- A testable extension would be to use $R$ or $R_{S\to\hat S}$ directly as an auditory attention decoder, comparing its trial-level accuracy against the standard correlation-based decoder, especially in low-correlation trials where the linear relationship may be easier to detect.
- The reversed distractor slope hints that $R$ could separate speech-driven shared information from volume-conduction shared activity; a source-localization or MEG study could test whether the distractor redundancy originates from non-auditory generators.
- Because $R$ is an upper bound from Lemma 1 rather than the true directed redundancy, the linear relation could partly be an artifact of the bound's slack; estimating the causal minimal sufficient statistics directly would show whether the proportionality survives.
- The analysis uses linear reconstruction; whether the same rate-distortion relation holds for nonlinear decoders or for spectro-temporal features beyond the envelope is left open by the paper.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes neural speech tracking under a competing-talker EEG paradigm and proposes a rate-distortion interpretation. For six left-temporal electrodes, the authors compute transfer-entropy rates from the speech envelope S to the reconstructed stimulus Ŝ, from S to each electrode, and from each electrode to Ŝ, and combine them via the directed-redundancy upper bound R = min(RS→Ŝ, RE→Ŝ, RS→E) from Eq. (12). They then plot the distortion D = 1 − |ρ| against these rates, fit linear regressions, and report that for attended stimuli both transfer entropy and the directed redundancy are significantly related to distortion, while no such relationship holds for distracting stimuli. The central theoretical apparatus is the directed-redundancy framework of Ref. [34], with the hidden redundancy process φ identified with the speech envelope S.
Significance. If the claims are sustained, the paper would provide a rate-distortion perspective on cortical speech tracking and connect an information-theoretic measure of directed redundancy to EEG electrode correlations. The empirical use of real EEG data, the standard mTRF decoder, and the strong transfer-entropy result for the attended stimulus (p = 6.1e-06 in Table 1) are concrete strengths. However, the headline redundancy claim rests on an upper-bound proxy R, not on the exact directed redundancy, and the statistical evidence for that claim is marginal. The significance is therefore conditional on validating that R behaves as a faithful indicator of the true redundancy in this EEG setting.
major comments (3)
- [Rate Redundancy in EEG Signals, Eq. (12)] Eq. (12) defines R as min(RS→Ŝ, RE→Ŝ, RS→E), and Lemma 1 states that this quantity is only an upper bound on the directed redundancy I^red, not an estimate or measurement of it. The abstract's claim that 'the directed redundancy is proportional to the correlation' and the conclusion's 'greater amount of redundant information' therefore go beyond what the reported analysis supports: the regressions in Fig. 3 and Table 1 are against R, not against I^red. Because the minimum is taken over six individual electrodes, R can be controlled by the single weakest electrode's transfer entropy, so the observed R–D relationship may reflect the weakest single-electrode tracking performance rather than shared or redundant information among electrodes. The paper does not report which term attains the minimum, nor any evidence that R is tight or monotonically related to I^red in this application. Please provide such evidence, or explicitly soften the claims to refer to an upper-bound proxy.
- [Fig. 1b and Section 'Rate Redundancy in EEG Signals'] The identification of the speech envelope S with the hidden redundancy process φ is assumed without validation. The directed-redundancy framework of [34] requires that φ causally drives the redundant information in the source processes, but EEG electrode correlations are also driven by volume conduction and by shared neural sources that may not be causally linked to S. If φ is not S, then R does not measure redundancy about the stimulus, and the interpreted meaning of the R–D relationship is lost. A control analysis, for example using a surrogate φ unrelated to S or comparing with a non-causal redundancy measure, is needed to support the mapping between the model in Fig. 1b and the EEG setup.
- [Table 1 and Fig. 3] The p-values for R are marginal (0.0177 for attended, 0.0416 for distractor) and are reported without any multiple-comparison correction across the four rate measures and two attention conditions. In addition, the support threshold 'pdf > 0.01' and the 0.005-bit bin width are introduced after inspecting the data, and no confidence intervals or effect sizes are reported for the regression slopes. These choices make the statistical case for the redundancy–distortion proportionality fragile. The strong RS→Ŝ result (p = 6.1e-06) can support the transfer-entropy claim, but the redundancy claim needs a pre-specified or corrected inference procedure, or at minimum a report of confidence intervals for the slopes.
minor comments (5)
- [Introduction, after Eq. (1)] The sentence describing transfer entropy refers to 'the target Y', but Eq. (1) uses Z as the target; please align the notation.
- [Simulation Study, Fig. 2] The label 'RR→Ŝ' in Fig. 2(c) appears to be a typo for 'RE→Ŝ' as defined in Eq. (9); please correct it for consistency.
- [Section 'Rate Redundancy in EEG Signals'] The sentence 'I(En(1),...,En(|LT|); Ŝn) = 0 or ∞' is confusing, since for continuous variables mutual information with a deterministic function is generally infinite, not zero; the intended statement needs clarification.
- [Fig. 3] The operational distortion-rate curves would benefit from error bars or confidence bands; as presented, the visual trend in Fig. 3b is difficult to assess without accompanying uncertainty information.
- [Abstract and Introduction] There is a typo, 'phenomonen' for 'phenomenon', in the Introduction; the paper also uses 'directed redundancy' and 'rate redundancy' interchangeably, which should be defined consistently.
Circularity Check
No significant circularity: the empirical rate–distortion relationship is computed from independent functionals, and the imported directed-redundancy measure is not fitted to the present data.
full rationale
The derivation chain is not circular. R in Eq. (12) is defined as min(RS→Ŝ, RE→Ŝ, RS→E), each term being a transfer entropy estimated from time series, while D in Eq. (8) is 1 − |ρ|, where ρ is the Pearson correlation between the reconstructed envelope and the attended stimulus. These are different functionals of the data, and no equation in the paper sets R equal to a function of D or fits R to match D. The observed inverse relationship in Fig. 3a and Table 1 is an empirical regression of computed distortion points on computed rate/redundancy points, not a prediction derived from fitted parameters. The directed-redundancy construction is imported from the authors' prior mathematical work ([33], [34]), but those results are stated as definitions and lemmas (e.g., Definition 1 and Lemma 1 quoted in the paper) with explicit assumptions and do not themselves assert any EEG relationship, so the self-citations are not load-bearing in a way that forces the present conclusion. Two interpretive choices—identifying the hidden redundancy process φ with the speech envelope S, and reporting the Lemma 1 upper bound R as "the directed redundancy"—are assumptions or overstatements that affect validity, but they do not constitute a circular reduction because the plotted relationship could have failed to appear and, indeed, is reported to be absent for the distractor stimuli.
Assumptions & free parameters
free parameters (4)
- regularization lambda in mTRF decoder =
not reported
- support threshold (pdf > 0.01) =
not reported numerically
- rate bin width =
0.005 bits
- TE estimator parameters =
not reported
assumptions (4)
- domain assumption The speech envelope S is the hidden redundancy process φ driving the EEG electrode signals.
- standard math Lemma 1's min-of-transfer-entropy bound is a valid upper bound for six electrodes when taking minima over electrodes.
- domain assumption The linear decoder in Eq (5) captures the stimulus-response mapping sufficiently for reconstruction.
- ad hoc to paper The support-threshold filtering does not bias the rate-distortion trend.
invented entities (1)
-
hidden redundancy process φ (assumed equal to speech envelope S)
independent evidence
Cite this review
Pith. "Pith review of Rate-Distortion under Neural Tracking of Speech: A Directed Redundancy Approach." pith.science (2026). https://pith.science/paper/ZVWVAH7C
@misc{pith2026250116762,
author = {Pith},
title = {Pith review of: Rate-Distortion under Neural Tracking of Speech: A Directed Redundancy Approach},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZVWVAH7C}},
note = {Machine review of arXiv:2501.16762}
}
read the original abstract
The data acquired at different scalp EEG electrodes when human subjects are exposed to speech stimuli are highly redundant. The redundancy is partly due to volume conduction effects and partly due to localized regions of the brain synchronizing their activity in response to the stimuli. In a competing talker scenario, we use a recent measure of directed redundancy to assess the amount of redundant information that is causally conveyed from the attended stimuli to the left temporal region of the brain. We observe that for the attended stimuli, the transfer entropy as well as the directed redundancy is proportional to the correlation between the speech stimuli and the reconstructed signal from the EEG signals. This demonstrates that both the rate as well as the rate-redundancy are inversely proportional to the distortion in neural speech tracking. Thus, a greater rate indicates a greater redundancy between the electrode signals, and a greater correlation between the reconstructed signal and the attended stimuli. A similar relationship is not observed for the distracting stimuli.
Figures
Reference graph
Works this paper leans on
-
[34]
Directed redundancy in time series,
J. Østergaard, “Directed redundancy in time series,” in IEEE International Symposium on Information Theory (ISIT), 2024
work page 2024
-
[1]
A mathematical theory of communication,
C. E. Shannon, “A mathematical theory of communication,” The Bell System Technical Journal, vol. 27, 1948
work page 1948
-
[2]
Information theory and cognition: A review,
K. Sayood, “Information theory and cognition: A review,” Entropy, vol. 20, no. 9, 2018
work page 2018
-
[3]
Broadbent, Perception and communication, Pergamon Press, 1958
D. Broadbent, Perception and communication, Pergamon Press, 1958
work page 1958
-
[4]
H. B. Barlow, Possible principles underlying the transformations of sensory messages, In W. Rosenblith (Ed.), Sensory communication, MA: MIT Press, 1961
work page 1961
-
[5]
Information through a spiking neuron,
C Stevens and A. Zador, “Information through a spiking neuron,” in Advances in neural information processing systems, 1995
work page 1995
-
[6]
Time resolution dependence of information measures for spiking neurons: scaling and universality,
S. E. Marzen, M. R. DeWeese, and J. P. Crutchfield, “Time resolution dependence of information measures for spiking neurons: scaling and universality,” Frontiers in Computational Neuroscience, vol. 9, 2015
work page 2015
-
[7]
On the upper bound of the information capacity in neuronal synapses,
M. Veleti´ c, P. A. Floor, Y. Chahibi, and I. Balasingham, “On the upper bound of the information capacity in neuronal synapses,” IEEE Transactions on Communications, vol. 64, no. 12, pp. 5025–5036, 2016
work page 2016
Show all 39 references
-
[8]
Axonal channel capacity in neuro-spike communication,
K. Aghababaiyan, V. Shah-Mansouri, and B. Maham, “Axonal channel capacity in neuro-spike communication,” IEEE Transactions on NanoBioscience, vol. 17, no. 1, pp. 78–87, 2018
2018
-
[9]
Capacity and error proba- bility analysis of neuro-spike communication exploiting temporal modulation,
K. Aghababaiyan, V. Shah-Mansouri, and B. Maham, “Capacity and error proba- bility analysis of neuro-spike communication exploiting temporal modulation,” IEEE Transactions on Communications, vol. 68, no. 4, pp. 2078–2089, 2020
2020
-
[10]
Theoretical understanding of the early visual processes by data com- pression and data selection,
L. Zhaoping, “Theoretical understanding of the early visual processes by data com- pression and data selection,” Network: Computation in Neural Systems, vol. 17, pp. 301 – 334, 2006
2006
-
[11]
Adaptive allocation of human visual working memory capacity during statistical and categorical learning,
C. J. Bates, R. A. Lerch, C. R. Sims, and R. A. Jacobs, “Adaptive allocation of human visual working memory capacity during statistical and categorical learning,” Journal of Vision, vol. 19, 2019
2019
-
[12]
The informational capacity of the human ear,
H. Jacobson, “The informational capacity of the human ear,” Science (New York, N.Y.), vol. 112, 1950
1950
-
[13]
Information loss in the human auditory system,
M. Z. Jahromi, A. Zahedi, J. Jensen, and J. Østergaard, “Information loss in the human auditory system,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 27, no. 3, pp. 472–481, 2019
2019
-
[14]
Rate-distortion theory of neural coding and its implications for working memory,
A. M. V. Jakob and S. J. Gershman, “Rate-distortion theory of neural coding and its implications for working memory,” eLife, vol. 12, pp. e79450, jul 2023
2023
-
[15]
Efficient data compression in perception and perceptual memory,
C. J. Bates and R. A. Jacobs, “Efficient data compression in perception and perceptual memory,” Psychological Review, vol. 127, pp. 891 – 917, 2020
2020
-
[16]
Auditory representations of acoustic signals,
X. Yang, K. Wang, and S.A. Shamma, “Auditory representations of acoustic signals,” IEEE Transactions on Information Theory, vol. 38, no. 2, pp. 824–839, 1992
1992
-
[17]
A mathematical theory of energy efficient neural compu- tation and communication,
T. Berger and W. B. Levy, “A mathematical theory of energy efficient neural compu- tation and communication,” IEEE Transactions on Information Theory, vol. 56, no. 2, pp. 852–874, 2010
2010
-
[18]
Nonnegative decomposition of multivariate informa- tion,
P. L. Williams and R. D. Beer, “Nonnegative decomposition of multivariate informa- tion,” CoRR, vol. abs/1004.2515, 2010
2010 arXiv
-
[19]
Bits and pieces: Understanding infor- mation decomposition from part-whole relationships and formal logic,
A. J. Gutknecht, M. Wibral, and A. Makkeh, “Bits and pieces: Understanding infor- mation decomposition from part-whole relationships and formal logic,” Proceedings of the Royal Society A, vol. 477, 2021
2021
-
[20]
Selective cortical representation of attended speaker in multi-talker speech perception,
N. Mesgarani and E. F. Chang, “Selective cortical representation of attended speaker in multi-talker speech perception,” Nature, vol. 485, no. 7397, pp. 233–236, 2012
2012
-
[21]
Robust cortical entrainment to the speech envelope relies on the spectro-temporal fine structure,
N. Ding, M. Chatterjee, and J. Z. Simon, “Robust cortical entrainment to the speech envelope relies on the spectro-temporal fine structure,” NeuroImage, vol. 88, pp. 41–46, 2014
2014
-
[22]
On the speech envelope in the cortical tracking of speech,
M. F. Issa, I. Khan, M. Ruzzoli, N. Molinaro, and M. Lizarazu, “On the speech envelope in the cortical tracking of speech,” NeuroImage, vol. 297, pp. 120675, 2024
2024
-
[23]
Attentional selection in a cocktail party environment can be decoded from single-trial eeg,
J.A. O’Sullivan, A.J. Power, N. Mesgarani, and et al., “Attentional selection in a cocktail party environment can be decoded from single-trial eeg,” Cereb Cortex, vol. 25, 2015
2015
-
[24]
Cortical auditory attention decod- ing during music and speech listening,
A. Simon, G. Loquet, J. Østergaard, and S. Bech, “Cortical auditory attention decod- ing during music and speech listening,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, June 2023
2023
-
[25]
Selective cortical representation of attended speaker in multi-talker speech perception,
N. Mesgarani and E. F. Chang, “Selective cortical representation of attended speaker in multi-talker speech perception,” Nature, vol. 485, 2012
2012
-
[26]
A comparison of regularization methods in forward and backward models for auditory attention decoding,
D. D. E. Wong, S. A. Fuglsang, J. Hjortkjær, E. Ceolini, M. Slaney, and A. de Cheveign´ e, “A comparison of regularization methods in forward and backward models for auditory attention decoding,” Frontiers in Neuroscience, vol. 12, 2018
2018
-
[27]
A tutorial on auditory attention identification methods,
E. Alickovic, T. Lunner, F. Gustafsson, and L. Ljung, “A tutorial on auditory attention identification methods,” Frontiers in Neuroscience, vol. 13, 2019
2019
-
[28]
Machine learning for decoding listen- ers’ attention from electroencephalography evoked by continuous speech,
T. de Taillez, B. Kollmeier, and B. T. Meyer, “Machine learning for decoding listen- ers’ attention from electroencephalography evoked by continuous speech,” European Journal of Neuroscience, vol. 51, no. 5, pp. 1234–1241, 2020
2020
-
[29]
Neural tracking of the speech envelope is differentially modulated by attention and language experience,
R. Reetzke, G.N. Gnanateja, and B. Chandrasekaran, “Neural tracking of the speech envelope is differentially modulated by attention and language experience,” Brain Lang., vol. 213, 2021
2021
-
[30]
Information-theoretic limits on the performance of auditory attention decoders,
R. Abeysekara, C. J. Smalt, I. M. Dushyanthi Karunathilake, J. Z. Simon, and B. Babadi, “Information-theoretic limits on the performance of auditory attention decoders,” in 2023 57th Asilomar Conference on Signals, Systems, and Computers, 2023, pp. 1479–1483
2023
-
[31]
T. M. Cover and J. A. Thomas, Elements of Information Theory, John Wiley & Sons, 2012
2012
-
[32]
A first course in information theory,
R. Yeung, “A first course in information theory,” Springer, 2002
2002
-
[33]
Higher-order common information,
J. Østergaard, “Higher-order common information,” Electronically available at: https://arxiv.org/abs/2406.02001, 2024
2024
-
[35]
Measuring information transfer,
T. Schreiber, “Measuring information transfer,” Phys. Rev. Lett., vol. 85, 2000
2000
-
[36]
On the speech envelope in the cortical tracking of speech,
M. F. Issa, I. Khan, M. Ruzzoli, N. Molinaro, and M. Lizarazu, “On the speech envelope in the cortical tracking of speech,” NeuroImage, vol. 297, 2024
2024
-
[37]
Niedermeyer and F
E. Niedermeyer and F. L. da Silva, Electroencephalography: Basic Principles, Clinical Applications, and Related Fields, Lippincott Williams & Wilkins, 2004
2004
-
[38]
The multivariate temporal response function (mtrf) toolbox: a matlab toolbox for relating neural signals to continuous stimuli,
M. J. Crosse, G. M. Di Liberto, A. Bednar, and E. C. Lalor, “The multivariate temporal response function (mtrf) toolbox: a matlab toolbox for relating neural signals to continuous stimuli,” Frontiers in Human Neuroscience, vol. 10, pp. 604, 2016
2016
-
[39]
Noise-robust cortical tracking of attended speech in real-life environments,
S. A. Fuglsang, T. Dau, and J. Hjortkjær, “Noise-robust cortical tracking of attended speech in real-life environments,” NeuroImage, vol. 156, pp. 435 – 444, 2017
2017
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.