REVIEW 3 major objections 5 minor 59 references
Under a fixed observation budget, learnable acoustic front ends form more informative low-rate samples for speech reconstruction than direct waveform sampling.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 05:55 UTC pith:ODCWU7KE
load-bearing objection Solid empirical front-end for low-rate speech reconstruction; the headline “informativeness” claim is confounded by unmatched baselines, but the method and ablations still stand. the 3 major comments →
Semantic Sampling via Learnable Observation Front Ends
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Under a matched scalar observation budget, finite-dimensional observations formed from learned acoustic responses—via a semantic feature filterbank, a magnitude-constrained signed observation matrix, and low-rate temporal readout—retain more reconstruction-relevant information than direct low-rate waveform samples or the predetermined low-rate inputs used by neural audio restoration systems. On speech reconstruction at 4 kHz, 2 kHz, and 1 kHz observation rates, this semantic sampling front end improves waveform fidelity, spectral consistency, and perceptual quality relative to uniform-sinc sampling and representative neural restoration baselines.
What carries the argument
The learnable observation front end: a causal damped-cosine filterbank that maps the waveform to multi-channel acoustic responses, a fixed-magnitude signed mixing matrix that collapses those responses into a few complementary observation channels, and a windowed temporal integral that produces the low-rate finite samples. These three stages jointly determine the information content of every observation value before reconstruction.
Load-bearing premise
The central claim rests on treating an end-to-end trained front-end plus reconstructor as a fair comparison against public pretrained restoration models that never see the same training speech and that always start from ordinary downsampled waveforms.
What would settle it
Retrain the neural restoration baselines from scratch on the same speech data and loss while still giving them only uniformly downsampled waveforms at the matched rates; if their waveform, spectral, and intelligibility scores then match or beat the proposed front end at 1–4 kHz, the claim that the learned observations themselves are more informative would not hold.
If this is right
- Under a fixed observation budget, reconstruction quality can be raised by redesigning the sampling stage rather than only by enlarging the reconstructor.
- Low-rate acoustic acquisition can target content-related spectral-temporal structure instead of raw bandwidth or sparsity alone.
- The same front-end design can be trained for objectives beyond waveform fidelity, such as intelligibility or linguistic content recovery.
- Cross-dataset transfer without fine-tuning already improves waveform and spectral metrics over fixed low-rate baselines, with further gains after target-domain fine-tuning.
Where Pith is reading between the lines
- If the front end truly shapes what is recoverable, analog or hardware filter–mix–integrate stages could cut digital sampling rates for speech devices without sacrificing intelligibility.
- A matched retrain of the neural restorers on the same low-rate waveforms would test whether the reported gains come mainly from observation formation or from joint end-to-end optimization.
- Similar response-space measurements may help other structured acoustic signals under tight budgets, not only speech.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes semantic sampling: finite-dimensional observations for acoustic reconstruction are formed not by direct low-rate waveform sampling, but by a learnable front end consisting of a causal damped-cosine filterbank (Eq. 4), a fixed-magnitude signed observation matrix B (Eqs. 7–9), and a low-rate temporal readout (Eqs. 10–11). A time-domain encoder–BiLSTM–decoder then reconstructs the waveform from the resulting P×Ns observations. Under matched scalar observation rates (4/2/1 kHz on 16 kHz speech), the method is reported to outperform Uniform-Sinc and publicly released neural restoration models (AudioUNet, NU-Wave 2, AudioSR) on SI-SDR, SNR, MR-STFT, LSD, STOI, and PESQ (Table II), with supporting ablations on K/P/Lr (Table III), qualitative waveform/spectrogram figures, and cross-dataset transfer to AISHELL-1 without fine-tuning (Table IV).
Significance. If the observation-informativeness claim holds under a properly controlled comparison, the work is a useful contribution to acoustic sampling and low-rate speech reconstruction: it reframes the sampling stage as a structured, learnable response-analysis problem rather than fixed waveform decimation, and shows that filter–mix–readout design can matter under a fixed scalar budget. Strengths include a clearly structured front end, multi-rate evaluation, systematic ablations of channel vs. temporal allocation, and cross-corpus results without target-domain training. The paper does not provide theory or machine-checked guarantees; its value is empirical and architectural. The main significance risk is that current gains may partly reflect end-to-end co-adaptation and baseline mismatch rather than observation quality alone.
major comments (3)
- §IV-A3 and Table II: the central claim that the front end yields more informative observations under a matched budget is not cleanly isolated. AudioUNet, NU-Wave 2, and AudioSR are used as publicly released pretrained models without retraining on LibriSpeech, while the proposed system jointly optimizes filterbank, B, readout, and a custom reconstructor on the same data and loss (Eqs. 18–22). This confounds observation formation with reconstructor co-adaptation and architecture. A load-bearing control is needed: train the same reconstructor (or an identical capacity model) on ordinary downsampled waveforms at the same f_obs, and/or retrain the neural baselines on LibriSpeech under the same protocol. Without that, superiority cannot be attributed solely to semantic sampling.
- §III-B and abstract: the term “semantic” is used for reconstruction-oriented acoustic responses trained only with waveform/spectral/energy losses (Eqs. 19–22). No linguistic, phonetic, or task-level semantic labels or objectives appear. Either justify the term against the semantic-communication literature cited in §II-C, or replace it with a more precise label (e.g., reconstruction-oriented / content-aware acoustic sampling) so the claim does not over-reach the training objective.
- Table II vs. Uniform-Sinc: Uniform-Sinc is the only fully matched classical baseline, and the gap there is large, which supports the idea that structured observations help. However, the paper’s strongest narrative repeatedly groups fixed sampling with neural restoration methods. Please separate (i) fixed sampling + same reconstructor, (ii) fixed sampling + strong pretrained restorers, and (iii) learned front end + proposed reconstructor, so readers can see how much of the gain is front-end design versus training regime.
minor comments (5)
- §III-B1, Eq. (4): implementation details of finite-length truncation, mean correction, and energy normalization are mentioned only briefly; a short appendix with filter length and initialization of fk, bk would aid reproducibility.
- §III-D: regularization weights βh, βb, βu and the exact forms of Rh, Rb, Ru are not specified numerically; please report them or the selection procedure.
- Fig. 4: the signed matrix visualization is helpful; adding a brief note on how many filters are effectively used (row/column diversity) would strengthen the ablation discussion in §IV-D.
- Table IV PESQ: without fine-tuning, pretrained restorers sometimes win PESQ while the proposed method wins SI-SDR/STOI/LSD; a short discussion of this metric split would avoid over-claiming uniform perceptual superiority.
- Typos/formatting: “PREPRINT” headers and occasional spacing issues (e.g., “RelatedWork”, “QuantitativeResults”) should be cleaned for journal submission.
Circularity Check
No circular derivation: end-to-end joint training is standard empirical learning, not a tautology; claims rest on held-out metrics versus external baselines.
full rationale
The paper defines a learnable front end A_ϑA (filterbank + constrained matrix B + readout) and jointly optimizes it with reconstructor D_ϑD by minimizing reconstruction loss L_rec plus front-end regularizers on training data (Eqs. 18–22). Reconstruction quality on held-out LibriSpeech and cross-dataset AISHELL-1 is then reported as evidence that the formed observations are more informative under a matched scalar budget f_obs. This is ordinary supervised end-to-end learning; the test metrics are not forced by construction from the training objective, nor do any equations reduce SI-SDR/SNR/MR-STFT/etc. to a fitted constant. Baselines (Uniform-Sinc, pretrained AudioUNet/NU-Wave 2/AudioSR) supply independent external comparisons, and ablations vary K/P/L_r while holding the budget fixed. Related-work self-citations on semantic communication are background only and do not underwrite uniqueness or the experimental claim. No self-definitional loop, fitted-input-as-prediction, load-bearing self-citation chain, or renamed known result appears in the derivation or results.
Axiom & Free-Parameter Ledger
free parameters (4)
- filterbank size K
- observation channels P and readout interval Lr
- loss weights λt, λf, λe and regularization coefficients βh, βb, βu
- learning rates (5e-5 front-end, 2e-4 reconstructor) and AdamW schedule
axioms (3)
- domain assumption Reconstruction quality under a fixed scalar observation budget is a valid proxy for the information content of the formed observations.
- ad hoc to paper Causal damped-cosine filters (Eq. 4) plus fixed-magnitude signed mixing (Eq. 9) are sufficiently expressive to capture reconstruction-relevant speech structure.
- ad hoc to paper Public pretrained restoration models evaluated on downsampled LibriSpeech constitute fair baselines for observation-quality comparison.
invented entities (2)
-
semantic sampling / learnable observation front end
no independent evidence
-
constrained semantic observation matrix B
no independent evidence
read the original abstract
Sampling determines the form of information available to downstream reconstruction systems. Conventional lowrate sampling forms finite-dimensional observations directly from the raw waveform, with the sampling rule mainly guided by bandwidth, sparsity, or fixed signal-level structures. For acoustic signals such as speech, however, reconstruction-relevant information is often expressed through content-related spectral-temporal structures rather than waveform samples alone. This paper proposes semantic sampling via learnable observation front ends, where finite-dimensional observations are generated from learned signal responses instead of directly subsampled waveform points. The proposed front end consists of a semantic feature filterbank, a constrained semantic observation matrix, and a low-rate readout module. The filterbank maps the input waveform into multiple acoustic response channels, the observation matrix combines these responses into a small number of observation channels, and the readout module produces low-rate finite-dimensional samples. A reconstruction network is then used to recover the signal from the resulting observations. Experiments on low-rate speech reconstruction show that, under the same observation budget, the proposed semantic sampling front end provides more informative observations than fixed low-rate sampling and neural restoration methods based on predetermined low-rate waveforms. The improvements in waveform fidelity, spectral consistency, and perceptual quality show that learnable observation front ends preserve more useful information for acoustic signal reconstruction under the same observation budget.
Figures
Reference graph
Works this paper leans on
-
[1]
Certain topics in telegraph transmission theory,
H. Nyquist, “Certain topics in telegraph transmission theory,”Transac- tions of the American Institute of Electrical Engineers, vol. 47, no. 2, pp. 617–644, 1928
1928
-
[2]
Communication in the presence of noise,
C. E. Shannon, “Communication in the presence of noise,”Proceedings of the IRE, vol. 37, no. 1, pp. 10–21, 1949
1949
-
[3]
Sampling-50 years after shannon,
M. Unser, “Sampling-50 years after shannon,”Proceedings of the IEEE, vol. 88, no. 4, pp. 569–587, 2000
2000
-
[4]
Generalized sampling expansion,
A. Papoulis, “Generalized sampling expansion,”IEEE Transactions on Circuits and Systems, vol. 24, no. 11, pp. 652–654, 1977
1977
-
[5]
Sampling signals with finite rate of innovation,
M. Vetterli, P. Marziliano, and T. Blu, “Sampling signals with finite rate of innovation,”IEEE Transactions on Signal Processing, vol. 50, no. 6, pp. 1417–1428, 2002
2002
-
[6]
Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency informa- tion,
E. J. Cand `es, J. Romberg, and T. Tao, “Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency informa- tion,”IEEE Transactions on Information Theory, vol. 52, no. 2, pp. 489–509, 2006
2006
-
[7]
Compressed sensing,
D. L. Donoho, “Compressed sensing,”IEEE Transactions on Informa- tion Theory, vol. 52, no. 4, pp. 1289–1306, 2006
2006
-
[8]
Stable signal recovery from incomplete and inaccurate measurements,
E. J. Candes, J. K. Romberg, and T. Tao, “Stable signal recovery from incomplete and inaccurate measurements,”Communications on Pure and Applied Mathematics: A Journal Issued by the Courant Institute of Mathematical Sciences, vol. 59, no. 8, pp. 1207–1223, 2006
2006
-
[9]
The restricted isometry property and its implications for compressed sensing,
E. J. Candes, “The restricted isometry property and its implications for compressed sensing,”Comptes rendus mathematique, vol. 346, no. 9-10, pp. 589–592, 2008
2008
-
[10]
Model- based compressive sensing,
R. G. Baraniuk, V . Cevher, M. F. Duarte, and C. Hegde, “Model- based compressive sensing,”IEEE Transactions on Information Theory, vol. 56, no. 4, pp. 1982–2001, 2010
1982
-
[11]
Audio super resolution using neural networks,
V . Kuleshov, S. Z. Enam, and S. Ermon, “Audio super resolution using neural networks,”arXiv preprint arXiv:1708.00853, 2017
Pith/arXiv arXiv 2017
-
[12]
Nu-wave: A diffusion probabilistic model for neural audio upsampling,
J. Lee and S. Han, “Nu-wave: A diffusion probabilistic model for neural audio upsampling,”arXiv preprint arXiv:2104.02321, 2021
Pith/arXiv arXiv 2021
-
[13]
Nu-wave 2: A general neural audio upsampling model for various sampling rates,
S. Han and J. Lee, “Nu-wave 2: A general neural audio upsampling model for various sampling rates,”arXiv preprint arXiv:2206.08545, 2022
Pith/arXiv arXiv 2022
-
[14]
V oicefixer: Toward general speech restoration with neural vocoder,
H. Liu, Q. Kong, Q. Tian, Y . Zhao, D. Wang, C. Huang, and Y . Wang, “V oicefixer: Toward general speech restoration with neural vocoder,” arXiv preprint arXiv:2109.13731, 2021
Pith/arXiv arXiv 2021
-
[15]
Audiosr: Versatile audio super-resolution at scale,
H. Liu, K. Chen, Q. Tian, W. Wang, and M. D. Plumbley, “Audiosr: Versatile audio super-resolution at scale,” inICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 1076–1080
2024
-
[16]
Towards high-quality and efficient speech bandwidth extension with parallel amplitude and phase prediction,
Y .-X. Lu, Y . Ai, H.-P. Du, and Z.-H. Ling, “Towards high-quality and efficient speech bandwidth extension with parallel amplitude and phase prediction,”IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 236–250, 2025
2025
-
[17]
Train- able frontend for robust and far-field keyword spotting,
Y . Wang, P. Getreuer, T. Hughes, R. F. Lyon, and R. A. Saurous, “Train- able frontend for robust and far-field keyword spotting,” in2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2017, pp. 5670–5674
2017
-
[18]
Learning filterbanks from raw speech for phone recogni- tion,
N. Zeghidour, N. Usunier, I. Kokkinos, T. Schaiz, G. Synnaeve, and E. Dupoux, “Learning filterbanks from raw speech for phone recogni- tion,” in2018 IEEE international conference on acoustics, speech and signal Processing (ICASSP). IEEE, 2018, pp. 5509–5513
2018
-
[19]
Speaker recognition from raw waveform with sincnet,
M. Ravanelli and Y . Bengio, “Speaker recognition from raw waveform with sincnet,” in2018 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2018, pp. 1021–1028
2018
-
[20]
Leaf: A learnable frontend for audio classification,
N. Zeghidour, O. Teboul, F. D. C. Quitry, and M. Tagliasacchi, “Leaf: A learnable frontend for audio classification,”arXiv preprint arXiv:2101.08596, 2021
Pith/arXiv arXiv 2021
-
[21]
Necessary density conditions for sampling and interpola- tion of certain entire functions,
H. J. Landau, “Necessary density conditions for sampling and interpola- tion of certain entire functions,”Acta Mathematica, vol. 117, pp. 37–52, 1967
1967
-
[22]
Near-optimal signal recovery from random projections: Universal encoding strategies?
E. J. Candes and T. Tao, “Near-optimal signal recovery from random projections: Universal encoding strategies?”IEEE Transactions on In- formation Theory, vol. 52, no. 12, pp. 5406–5425, 2006
2006
-
[23]
From theory to practice: Sub-nyquist sampling of sparse wideband analog signals,
M. Mishali and Y . C. Eldar, “From theory to practice: Sub-nyquist sampling of sparse wideband analog signals,”IEEE Journal of Selected Topics in Signal Processing, vol. 4, no. 2, pp. 375–391, 2010
2010
-
[24]
Recon- net: Non-iterative reconstruction of images from compressively sensed measurements,
K. Kulkarni, S. Lohit, P. Turaga, R. Kerviche, and A. Ashok, “Recon- net: Non-iterative reconstruction of images from compressively sensed measurements,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 449–458
2016
-
[25]
Deepcodec: Adaptive sensing and recovery via deep convolutional neural networks,
A. Mousavi, G. Dasarathy, and R. G. Baraniuk, “Deepcodec: Adaptive sensing and recovery via deep convolutional neural networks,”arXiv preprint arXiv:1707.03386, 2017
Pith/arXiv arXiv 2017
-
[26]
Learning a compressed sensing measurement matrix via gradient unrolling,
S. Wu, A. Dimakis, S. Sanghavi, F. Yu, D. Holtmann-Rice, D. Storcheus, A. Rostamizadeh, and S. Kumar, “Learning a compressed sensing measurement matrix via gradient unrolling,” inInternational Conference on Machine Learning. PMLR, 2019, pp. 6828–6839
2019
-
[27]
Compressed sensing using generative models,
A. Bora, A. Jalal, E. Price, and A. G. Dimakis, “Compressed sensing using generative models,” inInternational conference on machine learn- ing. PMLR, 2017, pp. 537–546
2017
-
[28]
Deep compressed sensing,
Y . Wu, M. Rosca, and T. Lillicrap, “Deep compressed sensing,” in International Conference on Machine Learning. PMLR, 2019, pp. 6850–6860
2019
-
[29]
Scalable deep compressive sensing,
Z. Zhang, Y . Liu, X. Cao, F. Wen, and C. Zhu, “Scalable deep compressive sensing,”arXiv preprint arXiv:2101.08024, 2021
Pith/arXiv arXiv 2021
-
[30]
Content-aware scalable deep compressed sens- ing,
B. Chen and J. Zhang, “Content-aware scalable deep compressed sens- ing,”IEEE Transactions on Image Processing, vol. 31, pp. 5412–5426, 2022
2022
-
[31]
Learning-based compressive mri,
B. G ¨ozc¨u, R. K. Mahabadi, Y .-H. Li, E. Ilıcak, T. Cukur, J. Scarlett, and V . Cevher, “Learning-based compressive mri,”IEEE Transactions on Medical Imaging, vol. 37, no. 6, pp. 1394–1406, 2018
2018
-
[32]
Learning-based optimization of the under-sampling pattern in mri,
C. D. Bahadir, A. V . Dalca, and M. R. Sabuncu, “Learning-based optimization of the under-sampling pattern in mri,” ininternational conference on information processing in medical imaging. Springer, 2019, pp. 780–792
2019
-
[33]
Auto-encoding variational bayes,
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”arXiv preprint arXiv:1312.6114, 2013
Pith/arXiv arXiv 2013
-
[34]
Neural discrete representation learning,
A. Van Den Oord, O. Vinyalset al., “Neural discrete representation learning,”Advances in neural information processing systems, vol. 30, 2017
2017
-
[35]
Representation learning with contrastive predictive coding,
A. v. d. Oord, Y . Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”arXiv preprint arXiv:1807.03748, 2018
Pith/arXiv arXiv 2018
-
[36]
High fidelity neural audio compression,
A. D ´efossez, J. Copet, G. Synnaeve, and Y . Adi, “High fidelity neural audio compression,”arXiv preprint arXiv:2210.13438, 2022
Pith/arXiv arXiv 2022
-
[37]
Deep learning enabled semantic communication systems,
H. Xie, Z. Qin, G. Y . Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,”IEEE Transactions on Signal Pro- cessing, vol. 69, pp. 2663–2675, 2021
2021
-
[38]
Semantic communications: Principles and challenges,
Z. Qin, X. Tao, J. Lu, W. Tong, and G. Y . Li, “Semantic communications: Principles and challenges,”arXiv preprint arXiv:2201.01389, 2022
Pith/arXiv arXiv 2022
-
[39]
Beyond transmitting bits: Context, semantics, and task-oriented communications,
D. G ¨und¨uz, Z. Qin, I. E. Aguerri, H. S. Dhillon, Z. Yang, A. Yener, K. K. Wong, and C.-B. Chae, “Beyond transmitting bits: Context, semantics, and task-oriented communications,”IEEE Journal on Selected Areas in Communications, vol. 41, no. 1, pp. 5–41, 2023
2023
-
[40]
A theory of semantic communication,
Y . Shao, Q. Cao, and D. G¨und¨uz, “A theory of semantic communication,” IEEE Transactions on Mobile Computing, vol. 23, no. 12, pp. 12 211– 12 228, 2024
2024
-
[41]
Task-oriented explainable semantic communications,
S. Ma, W. Qiao, Y . Wu, H. Li, G. Shi, D. Gao, Y . Shi, S. Li, and N. Al-Dhahir, “Task-oriented explainable semantic communications,” IEEE Transactions on Wireless Communications, vol. 22, no. 12, pp. 9248–9262, 2023
2023
-
[42]
Rate-distortion-perception theory for semantic communication,
J. Chai, Y . Xiao, G. Shi, and W. Saad, “Rate-distortion-perception theory for semantic communication,” in2023 IEEE 31st International Conference on Network Protocols (ICNP). IEEE, 2023, pp. 1–6. PREPRINT 13
2023
-
[43]
Classification-oriented semantic wireless com- munications,
E. Kutay and A. Yener, “Classification-oriented semantic wireless com- munications,” inICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 9096–9100
2024
-
[44]
Generative ai driven task-oriented adaptive semantic communications,
Y . Fu, W. Cheng, J. Wang, L. Yin, and W. Zhang, “Generative ai driven task-oriented adaptive semantic communications,”IEEE Transactions on Wireless Communications, vol. 25, pp. 9078–9093, 2025
2025
-
[45]
Bandwidth extension on raw audio via generative adversarial networks,
S. Kim and V . Sathe, “Bandwidth extension on raw audio via generative adversarial networks,”arXiv preprint arXiv:1903.09027, 2019
Pith/arXiv arXiv 1903
-
[46]
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” inICASSP 2020-2020 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 6199–6203
2020
-
[47]
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”Advances in neural information processing systems, vol. 33, pp. 17 022–17 033, 2020
2020
-
[48]
Dif- fwave: A versatile diffusion model for audio synthesis,
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “Dif- fwave: A versatile diffusion model for audio synthesis,”arXiv preprint arXiv:2009.09761, 2020
Pith/arXiv arXiv 2009
-
[49]
Long short-term memory,
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997
1997
-
[50]
Bidirectional recurrent neural net- works,
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural net- works,”IEEE Transactions on Signal Processing, vol. 45, no. 11, pp. 2673–2681, 1997
1997
-
[51]
Librispeech: an asr corpus based on public domain audio books,
V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2015, pp. 5206–5210
2015
-
[52]
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,” in2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA). IEEE, 2017, pp. 1–5
2017
-
[53]
Sdr–half-baked or well done?
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “Sdr–half-baked or well done?” inICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, pp. 626–630
2019
-
[54]
Distance measures for speech processing,
A. Gray and J. Markel, “Distance measures for speech processing,”IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 24, no. 5, pp. 380–391, 1976
1976
-
[55]
S. R. Quackenbush, T. P. Barnwell, and M. A. Clements,Objective Measures of Speech Quality. Englewood Cliffs, NJ: Prentice Hall, 1988
1988
-
[56]
An algorithm for intelligibility prediction of time–frequency weighted noisy speech,
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time–frequency weighted noisy speech,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, no. 7, pp. 2125–2136, 2011
2011
-
[57]
Perceptual evaluation of speech quality (pesq): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,
I.-T. Recommendation, “Perceptual evaluation of speech quality (pesq): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,”Rec. ITU-T P . 862, 2001
2001
-
[58]
Decoupled weight decay regularization,
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2017
Pith/arXiv arXiv 2017
-
[59]
Sgdr: Stochastic gradient descent with warm restarts,
——, “Sgdr: Stochastic gradient descent with warm restarts,”arXiv preprint arXiv:1608.03983, 2016
Pith/arXiv arXiv 2016
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.