Pith. sign in

REVIEW 3 major objections 5 minor 59 references

Under a fixed observation budget, learnable acoustic front ends form more informative low-rate samples for speech reconstruction than direct waveform sampling.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 05:55 UTC pith:ODCWU7KE

load-bearing objection Solid empirical front-end for low-rate speech reconstruction; the headline “informativeness” claim is confounded by unmatched baselines, but the method and ablations still stand. the 3 major comments →

arxiv 2607.11260 v1 pith:ODCWU7KE submitted 2026-07-13 eess.AS

Semantic Sampling via Learnable Observation Front Ends

classification eess.AS
keywords semantic samplinglearnable observation front endlow-rate observationspeech reconstructionacoustic signal processingfilterbankobservation matrix
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that for speech, what you keep under a tight observation budget depends on how the finite samples are formed, not only how many there are. Conventional low-rate sampling takes points from the raw waveform and can discard content-related spectral-temporal structure before any reconstructor sees the data. The authors replace that step with a learnable front end: a filterbank that produces multi-channel acoustic responses, a constrained mixing matrix that combines them into a few observation channels, and a temporal readout that yields the low-rate samples. Jointly trained with a reconstruction network, this front end keeps more recoverable speech structure than fixed downsampling or neural restorers that start from predetermined low-rate waveforms. A sympathetic reader cares because the result reframes sampling itself as an optimizable stage that can preserve intelligibility and spectral detail when every scalar observation is expensive.

Core claim

Under a matched scalar observation budget, finite-dimensional observations formed from learned acoustic responses—via a semantic feature filterbank, a magnitude-constrained signed observation matrix, and low-rate temporal readout—retain more reconstruction-relevant information than direct low-rate waveform samples or the predetermined low-rate inputs used by neural audio restoration systems. On speech reconstruction at 4 kHz, 2 kHz, and 1 kHz observation rates, this semantic sampling front end improves waveform fidelity, spectral consistency, and perceptual quality relative to uniform-sinc sampling and representative neural restoration baselines.

What carries the argument

The learnable observation front end: a causal damped-cosine filterbank that maps the waveform to multi-channel acoustic responses, a fixed-magnitude signed mixing matrix that collapses those responses into a few complementary observation channels, and a windowed temporal integral that produces the low-rate finite samples. These three stages jointly determine the information content of every observation value before reconstruction.

Load-bearing premise

The central claim rests on treating an end-to-end trained front-end plus reconstructor as a fair comparison against public pretrained restoration models that never see the same training speech and that always start from ordinary downsampled waveforms.

What would settle it

Retrain the neural restoration baselines from scratch on the same speech data and loss while still giving them only uniformly downsampled waveforms at the matched rates; if their waveform, spectral, and intelligibility scores then match or beat the proposed front end at 1–4 kHz, the claim that the learned observations themselves are more informative would not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Under a fixed observation budget, reconstruction quality can be raised by redesigning the sampling stage rather than only by enlarging the reconstructor.
  • Low-rate acoustic acquisition can target content-related spectral-temporal structure instead of raw bandwidth or sparsity alone.
  • The same front-end design can be trained for objectives beyond waveform fidelity, such as intelligibility or linguistic content recovery.
  • Cross-dataset transfer without fine-tuning already improves waveform and spectral metrics over fixed low-rate baselines, with further gains after target-domain fine-tuning.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the front end truly shapes what is recoverable, analog or hardware filter–mix–integrate stages could cut digital sampling rates for speech devices without sacrificing intelligibility.
  • A matched retrain of the neural restorers on the same low-rate waveforms would test whether the reported gains come mainly from observation formation or from joint end-to-end optimization.
  • Similar response-space measurements may help other structured acoustic signals under tight budgets, not only speech.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes semantic sampling: finite-dimensional observations for acoustic reconstruction are formed not by direct low-rate waveform sampling, but by a learnable front end consisting of a causal damped-cosine filterbank (Eq. 4), a fixed-magnitude signed observation matrix B (Eqs. 7–9), and a low-rate temporal readout (Eqs. 10–11). A time-domain encoder–BiLSTM–decoder then reconstructs the waveform from the resulting P×Ns observations. Under matched scalar observation rates (4/2/1 kHz on 16 kHz speech), the method is reported to outperform Uniform-Sinc and publicly released neural restoration models (AudioUNet, NU-Wave 2, AudioSR) on SI-SDR, SNR, MR-STFT, LSD, STOI, and PESQ (Table II), with supporting ablations on K/P/Lr (Table III), qualitative waveform/spectrogram figures, and cross-dataset transfer to AISHELL-1 without fine-tuning (Table IV).

Significance. If the observation-informativeness claim holds under a properly controlled comparison, the work is a useful contribution to acoustic sampling and low-rate speech reconstruction: it reframes the sampling stage as a structured, learnable response-analysis problem rather than fixed waveform decimation, and shows that filter–mix–readout design can matter under a fixed scalar budget. Strengths include a clearly structured front end, multi-rate evaluation, systematic ablations of channel vs. temporal allocation, and cross-corpus results without target-domain training. The paper does not provide theory or machine-checked guarantees; its value is empirical and architectural. The main significance risk is that current gains may partly reflect end-to-end co-adaptation and baseline mismatch rather than observation quality alone.

major comments (3)
  1. §IV-A3 and Table II: the central claim that the front end yields more informative observations under a matched budget is not cleanly isolated. AudioUNet, NU-Wave 2, and AudioSR are used as publicly released pretrained models without retraining on LibriSpeech, while the proposed system jointly optimizes filterbank, B, readout, and a custom reconstructor on the same data and loss (Eqs. 18–22). This confounds observation formation with reconstructor co-adaptation and architecture. A load-bearing control is needed: train the same reconstructor (or an identical capacity model) on ordinary downsampled waveforms at the same f_obs, and/or retrain the neural baselines on LibriSpeech under the same protocol. Without that, superiority cannot be attributed solely to semantic sampling.
  2. §III-B and abstract: the term “semantic” is used for reconstruction-oriented acoustic responses trained only with waveform/spectral/energy losses (Eqs. 19–22). No linguistic, phonetic, or task-level semantic labels or objectives appear. Either justify the term against the semantic-communication literature cited in §II-C, or replace it with a more precise label (e.g., reconstruction-oriented / content-aware acoustic sampling) so the claim does not over-reach the training objective.
  3. Table II vs. Uniform-Sinc: Uniform-Sinc is the only fully matched classical baseline, and the gap there is large, which supports the idea that structured observations help. However, the paper’s strongest narrative repeatedly groups fixed sampling with neural restoration methods. Please separate (i) fixed sampling + same reconstructor, (ii) fixed sampling + strong pretrained restorers, and (iii) learned front end + proposed reconstructor, so readers can see how much of the gain is front-end design versus training regime.
minor comments (5)
  1. §III-B1, Eq. (4): implementation details of finite-length truncation, mean correction, and energy normalization are mentioned only briefly; a short appendix with filter length and initialization of fk, bk would aid reproducibility.
  2. §III-D: regularization weights βh, βb, βu and the exact forms of Rh, Rb, Ru are not specified numerically; please report them or the selection procedure.
  3. Fig. 4: the signed matrix visualization is helpful; adding a brief note on how many filters are effectively used (row/column diversity) would strengthen the ablation discussion in §IV-D.
  4. Table IV PESQ: without fine-tuning, pretrained restorers sometimes win PESQ while the proposed method wins SI-SDR/STOI/LSD; a short discussion of this metric split would avoid over-claiming uniform perceptual superiority.
  5. Typos/formatting: “PREPRINT” headers and occasional spacing issues (e.g., “RelatedWork”, “QuantitativeResults”) should be cleaned for journal submission.

Circularity Check

0 steps flagged

No circular derivation: end-to-end joint training is standard empirical learning, not a tautology; claims rest on held-out metrics versus external baselines.

full rationale

The paper defines a learnable front end A_ϑA (filterbank + constrained matrix B + readout) and jointly optimizes it with reconstructor D_ϑD by minimizing reconstruction loss L_rec plus front-end regularizers on training data (Eqs. 18–22). Reconstruction quality on held-out LibriSpeech and cross-dataset AISHELL-1 is then reported as evidence that the formed observations are more informative under a matched scalar budget f_obs. This is ordinary supervised end-to-end learning; the test metrics are not forced by construction from the training objective, nor do any equations reduce SI-SDR/SNR/MR-STFT/etc. to a fitted constant. Baselines (Uniform-Sinc, pretrained AudioUNet/NU-Wave 2/AudioSR) supply independent external comparisons, and ablations vary K/P/L_r while holding the budget fixed. Related-work self-citations on semantic communication are background only and do not underwrite uniqueness or the experimental claim. No self-definitional loop, fitted-input-as-prediction, load-bearing self-citation chain, or renamed known result appears in the derivation or results.

Axiom & Free-Parameter Ledger

4 free parameters · 3 axioms · 2 invented entities

The work is empirical; the central claim rests on architectural choices, loss weights, and the fairness of the baseline protocol rather than on unproved mathematical axioms. Free parameters are the usual hyper-parameters of a deep audio model; invented entities are the named modules of the front end.

free parameters (4)
  • filterbank size K
    Number of learnable filters; ablated over {16,32,64,128} and selected by validation performance.
  • observation channels P and readout interval Lr
    Trade-off parameters that fix the observation rate fobs = P f0 / Lr; chosen by ablation.
  • loss weights λt, λf, λe and regularization coefficients βh, βb, βu
    Hand-set coefficients that balance time, frequency, energy and front-end regularizers (Eqs. 20–21).
  • learning rates (5e-5 front-end, 2e-4 reconstructor) and AdamW schedule
    Optimizer hyper-parameters selected for stable joint training.
axioms (3)
  • domain assumption Reconstruction quality under a fixed scalar observation budget is a valid proxy for the information content of the formed observations.
    Stated throughout Introduction and Section IV; underpins all claims that better metrics imply more informative sampling.
  • ad hoc to paper Causal damped-cosine filters (Eq. 4) plus fixed-magnitude signed mixing (Eq. 9) are sufficiently expressive to capture reconstruction-relevant speech structure.
    Architectural choice introduced in Section III-B without independent theoretical guarantee.
  • ad hoc to paper Public pretrained restoration models evaluated on downsampled LibriSpeech constitute fair baselines for observation-quality comparison.
    Protocol of Section IV-A3; not a standard mathematical axiom.
invented entities (2)
  • semantic sampling / learnable observation front end no independent evidence
    purpose: Name the overall process of forming finite observations from learned acoustic responses rather than raw waveform points.
    Terminological framing introduced by the authors; no independent physical existence outside the proposed architecture.
  • constrained semantic observation matrix B no independent evidence
    purpose: Linearly combine K filter responses into P observation channels under fixed-magnitude signed entries.
    Defined by Eq. 9; a design choice of the paper.

pith-pipeline@v1.1.0-grok45 · 22977 in / 2782 out tokens · 31554 ms · 2026-07-14T05:55:47.057136+00:00 · methodology

0 comments
read the original abstract

Sampling determines the form of information available to downstream reconstruction systems. Conventional lowrate sampling forms finite-dimensional observations directly from the raw waveform, with the sampling rule mainly guided by bandwidth, sparsity, or fixed signal-level structures. For acoustic signals such as speech, however, reconstruction-relevant information is often expressed through content-related spectral-temporal structures rather than waveform samples alone. This paper proposes semantic sampling via learnable observation front ends, where finite-dimensional observations are generated from learned signal responses instead of directly subsampled waveform points. The proposed front end consists of a semantic feature filterbank, a constrained semantic observation matrix, and a low-rate readout module. The filterbank maps the input waveform into multiple acoustic response channels, the observation matrix combines these responses into a small number of observation channels, and the readout module produces low-rate finite-dimensional samples. A reconstruction network is then used to recover the signal from the resulting observations. Experiments on low-rate speech reconstruction show that, under the same observation budget, the proposed semantic sampling front end provides more informative observations than fixed low-rate sampling and neural restoration methods based on predetermined low-rate waveforms. The improvements in waveform fidelity, spectral consistency, and perceptual quality show that learnable observation front ends preserve more useful information for acoustic signal reconstruction under the same observation budget.

Figures

Figures reproduced from arXiv: 2607.11260 by Guangming Shi, Pengfei He, Shuai Ma, Xiang Cheng, Yuxuan Liu.

Figure 1
Figure 1. Figure 1: Overview of the proposed semantic sampling framework. The input waveform [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Waveform reconstruction results at different observation rates. Each row corresponds to one observation rate. The left column shows the reference [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Spectrogram reconstruction results at different observation rates. Each row corresponds to one observation rate. The left column shows the reference [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Visualization of the semantic sampling front end at the [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

59 extracted references · 15 linked inside Pith

  1. [1]

    Certain topics in telegraph transmission theory,

    H. Nyquist, “Certain topics in telegraph transmission theory,”Transac- tions of the American Institute of Electrical Engineers, vol. 47, no. 2, pp. 617–644, 1928

  2. [2]

    Communication in the presence of noise,

    C. E. Shannon, “Communication in the presence of noise,”Proceedings of the IRE, vol. 37, no. 1, pp. 10–21, 1949

  3. [3]

    Sampling-50 years after shannon,

    M. Unser, “Sampling-50 years after shannon,”Proceedings of the IEEE, vol. 88, no. 4, pp. 569–587, 2000

  4. [4]

    Generalized sampling expansion,

    A. Papoulis, “Generalized sampling expansion,”IEEE Transactions on Circuits and Systems, vol. 24, no. 11, pp. 652–654, 1977

  5. [5]

    Sampling signals with finite rate of innovation,

    M. Vetterli, P. Marziliano, and T. Blu, “Sampling signals with finite rate of innovation,”IEEE Transactions on Signal Processing, vol. 50, no. 6, pp. 1417–1428, 2002

  6. [6]

    Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency informa- tion,

    E. J. Cand `es, J. Romberg, and T. Tao, “Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency informa- tion,”IEEE Transactions on Information Theory, vol. 52, no. 2, pp. 489–509, 2006

  7. [7]

    Compressed sensing,

    D. L. Donoho, “Compressed sensing,”IEEE Transactions on Informa- tion Theory, vol. 52, no. 4, pp. 1289–1306, 2006

  8. [8]

    Stable signal recovery from incomplete and inaccurate measurements,

    E. J. Candes, J. K. Romberg, and T. Tao, “Stable signal recovery from incomplete and inaccurate measurements,”Communications on Pure and Applied Mathematics: A Journal Issued by the Courant Institute of Mathematical Sciences, vol. 59, no. 8, pp. 1207–1223, 2006

  9. [9]

    The restricted isometry property and its implications for compressed sensing,

    E. J. Candes, “The restricted isometry property and its implications for compressed sensing,”Comptes rendus mathematique, vol. 346, no. 9-10, pp. 589–592, 2008

  10. [10]

    Model- based compressive sensing,

    R. G. Baraniuk, V . Cevher, M. F. Duarte, and C. Hegde, “Model- based compressive sensing,”IEEE Transactions on Information Theory, vol. 56, no. 4, pp. 1982–2001, 2010

  11. [11]

    Audio super resolution using neural networks,

    V . Kuleshov, S. Z. Enam, and S. Ermon, “Audio super resolution using neural networks,”arXiv preprint arXiv:1708.00853, 2017

  12. [12]

    Nu-wave: A diffusion probabilistic model for neural audio upsampling,

    J. Lee and S. Han, “Nu-wave: A diffusion probabilistic model for neural audio upsampling,”arXiv preprint arXiv:2104.02321, 2021

  13. [13]

    Nu-wave 2: A general neural audio upsampling model for various sampling rates,

    S. Han and J. Lee, “Nu-wave 2: A general neural audio upsampling model for various sampling rates,”arXiv preprint arXiv:2206.08545, 2022

  14. [14]

    V oicefixer: Toward general speech restoration with neural vocoder,

    H. Liu, Q. Kong, Q. Tian, Y . Zhao, D. Wang, C. Huang, and Y . Wang, “V oicefixer: Toward general speech restoration with neural vocoder,” arXiv preprint arXiv:2109.13731, 2021

  15. [15]

    Audiosr: Versatile audio super-resolution at scale,

    H. Liu, K. Chen, Q. Tian, W. Wang, and M. D. Plumbley, “Audiosr: Versatile audio super-resolution at scale,” inICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 1076–1080

  16. [16]

    Towards high-quality and efficient speech bandwidth extension with parallel amplitude and phase prediction,

    Y .-X. Lu, Y . Ai, H.-P. Du, and Z.-H. Ling, “Towards high-quality and efficient speech bandwidth extension with parallel amplitude and phase prediction,”IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 236–250, 2025

  17. [17]

    Train- able frontend for robust and far-field keyword spotting,

    Y . Wang, P. Getreuer, T. Hughes, R. F. Lyon, and R. A. Saurous, “Train- able frontend for robust and far-field keyword spotting,” in2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2017, pp. 5670–5674

  18. [18]

    Learning filterbanks from raw speech for phone recogni- tion,

    N. Zeghidour, N. Usunier, I. Kokkinos, T. Schaiz, G. Synnaeve, and E. Dupoux, “Learning filterbanks from raw speech for phone recogni- tion,” in2018 IEEE international conference on acoustics, speech and signal Processing (ICASSP). IEEE, 2018, pp. 5509–5513

  19. [19]

    Speaker recognition from raw waveform with sincnet,

    M. Ravanelli and Y . Bengio, “Speaker recognition from raw waveform with sincnet,” in2018 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2018, pp. 1021–1028

  20. [20]

    Leaf: A learnable frontend for audio classification,

    N. Zeghidour, O. Teboul, F. D. C. Quitry, and M. Tagliasacchi, “Leaf: A learnable frontend for audio classification,”arXiv preprint arXiv:2101.08596, 2021

  21. [21]

    Necessary density conditions for sampling and interpola- tion of certain entire functions,

    H. J. Landau, “Necessary density conditions for sampling and interpola- tion of certain entire functions,”Acta Mathematica, vol. 117, pp. 37–52, 1967

  22. [22]

    Near-optimal signal recovery from random projections: Universal encoding strategies?

    E. J. Candes and T. Tao, “Near-optimal signal recovery from random projections: Universal encoding strategies?”IEEE Transactions on In- formation Theory, vol. 52, no. 12, pp. 5406–5425, 2006

  23. [23]

    From theory to practice: Sub-nyquist sampling of sparse wideband analog signals,

    M. Mishali and Y . C. Eldar, “From theory to practice: Sub-nyquist sampling of sparse wideband analog signals,”IEEE Journal of Selected Topics in Signal Processing, vol. 4, no. 2, pp. 375–391, 2010

  24. [24]

    Recon- net: Non-iterative reconstruction of images from compressively sensed measurements,

    K. Kulkarni, S. Lohit, P. Turaga, R. Kerviche, and A. Ashok, “Recon- net: Non-iterative reconstruction of images from compressively sensed measurements,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 449–458

  25. [25]

    Deepcodec: Adaptive sensing and recovery via deep convolutional neural networks,

    A. Mousavi, G. Dasarathy, and R. G. Baraniuk, “Deepcodec: Adaptive sensing and recovery via deep convolutional neural networks,”arXiv preprint arXiv:1707.03386, 2017

  26. [26]

    Learning a compressed sensing measurement matrix via gradient unrolling,

    S. Wu, A. Dimakis, S. Sanghavi, F. Yu, D. Holtmann-Rice, D. Storcheus, A. Rostamizadeh, and S. Kumar, “Learning a compressed sensing measurement matrix via gradient unrolling,” inInternational Conference on Machine Learning. PMLR, 2019, pp. 6828–6839

  27. [27]

    Compressed sensing using generative models,

    A. Bora, A. Jalal, E. Price, and A. G. Dimakis, “Compressed sensing using generative models,” inInternational conference on machine learn- ing. PMLR, 2017, pp. 537–546

  28. [28]

    Deep compressed sensing,

    Y . Wu, M. Rosca, and T. Lillicrap, “Deep compressed sensing,” in International Conference on Machine Learning. PMLR, 2019, pp. 6850–6860

  29. [29]

    Scalable deep compressive sensing,

    Z. Zhang, Y . Liu, X. Cao, F. Wen, and C. Zhu, “Scalable deep compressive sensing,”arXiv preprint arXiv:2101.08024, 2021

  30. [30]

    Content-aware scalable deep compressed sens- ing,

    B. Chen and J. Zhang, “Content-aware scalable deep compressed sens- ing,”IEEE Transactions on Image Processing, vol. 31, pp. 5412–5426, 2022

  31. [31]

    Learning-based compressive mri,

    B. G ¨ozc¨u, R. K. Mahabadi, Y .-H. Li, E. Ilıcak, T. Cukur, J. Scarlett, and V . Cevher, “Learning-based compressive mri,”IEEE Transactions on Medical Imaging, vol. 37, no. 6, pp. 1394–1406, 2018

  32. [32]

    Learning-based optimization of the under-sampling pattern in mri,

    C. D. Bahadir, A. V . Dalca, and M. R. Sabuncu, “Learning-based optimization of the under-sampling pattern in mri,” ininternational conference on information processing in medical imaging. Springer, 2019, pp. 780–792

  33. [33]

    Auto-encoding variational bayes,

    D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”arXiv preprint arXiv:1312.6114, 2013

  34. [34]

    Neural discrete representation learning,

    A. Van Den Oord, O. Vinyalset al., “Neural discrete representation learning,”Advances in neural information processing systems, vol. 30, 2017

  35. [35]

    Representation learning with contrastive predictive coding,

    A. v. d. Oord, Y . Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”arXiv preprint arXiv:1807.03748, 2018

  36. [36]

    High fidelity neural audio compression,

    A. D ´efossez, J. Copet, G. Synnaeve, and Y . Adi, “High fidelity neural audio compression,”arXiv preprint arXiv:2210.13438, 2022

  37. [37]

    Deep learning enabled semantic communication systems,

    H. Xie, Z. Qin, G. Y . Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,”IEEE Transactions on Signal Pro- cessing, vol. 69, pp. 2663–2675, 2021

  38. [38]

    Semantic communications: Principles and challenges,

    Z. Qin, X. Tao, J. Lu, W. Tong, and G. Y . Li, “Semantic communications: Principles and challenges,”arXiv preprint arXiv:2201.01389, 2022

  39. [39]

    Beyond transmitting bits: Context, semantics, and task-oriented communications,

    D. G ¨und¨uz, Z. Qin, I. E. Aguerri, H. S. Dhillon, Z. Yang, A. Yener, K. K. Wong, and C.-B. Chae, “Beyond transmitting bits: Context, semantics, and task-oriented communications,”IEEE Journal on Selected Areas in Communications, vol. 41, no. 1, pp. 5–41, 2023

  40. [40]

    A theory of semantic communication,

    Y . Shao, Q. Cao, and D. G¨und¨uz, “A theory of semantic communication,” IEEE Transactions on Mobile Computing, vol. 23, no. 12, pp. 12 211– 12 228, 2024

  41. [41]

    Task-oriented explainable semantic communications,

    S. Ma, W. Qiao, Y . Wu, H. Li, G. Shi, D. Gao, Y . Shi, S. Li, and N. Al-Dhahir, “Task-oriented explainable semantic communications,” IEEE Transactions on Wireless Communications, vol. 22, no. 12, pp. 9248–9262, 2023

  42. [42]

    Rate-distortion-perception theory for semantic communication,

    J. Chai, Y . Xiao, G. Shi, and W. Saad, “Rate-distortion-perception theory for semantic communication,” in2023 IEEE 31st International Conference on Network Protocols (ICNP). IEEE, 2023, pp. 1–6. PREPRINT 13

  43. [43]

    Classification-oriented semantic wireless com- munications,

    E. Kutay and A. Yener, “Classification-oriented semantic wireless com- munications,” inICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 9096–9100

  44. [44]

    Generative ai driven task-oriented adaptive semantic communications,

    Y . Fu, W. Cheng, J. Wang, L. Yin, and W. Zhang, “Generative ai driven task-oriented adaptive semantic communications,”IEEE Transactions on Wireless Communications, vol. 25, pp. 9078–9093, 2025

  45. [45]

    Bandwidth extension on raw audio via generative adversarial networks,

    S. Kim and V . Sathe, “Bandwidth extension on raw audio via generative adversarial networks,”arXiv preprint arXiv:1903.09027, 2019

  46. [46]

    Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,

    R. Yamamoto, E. Song, and J.-M. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” inICASSP 2020-2020 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 6199–6203

  47. [47]

    Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

    J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”Advances in neural information processing systems, vol. 33, pp. 17 022–17 033, 2020

  48. [48]

    Dif- fwave: A versatile diffusion model for audio synthesis,

    Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “Dif- fwave: A versatile diffusion model for audio synthesis,”arXiv preprint arXiv:2009.09761, 2020

  49. [49]

    Long short-term memory,

    S. Hochreiter and J. Schmidhuber, “Long short-term memory,”Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997

  50. [50]

    Bidirectional recurrent neural net- works,

    M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural net- works,”IEEE Transactions on Signal Processing, vol. 45, no. 11, pp. 2673–2681, 1997

  51. [51]

    Librispeech: an asr corpus based on public domain audio books,

    V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2015, pp. 5206–5210

  52. [52]

    Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,

    H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,” in2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA). IEEE, 2017, pp. 1–5

  53. [53]

    Sdr–half-baked or well done?

    J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “Sdr–half-baked or well done?” inICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, pp. 626–630

  54. [54]

    Distance measures for speech processing,

    A. Gray and J. Markel, “Distance measures for speech processing,”IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 24, no. 5, pp. 380–391, 1976

  55. [55]

    S. R. Quackenbush, T. P. Barnwell, and M. A. Clements,Objective Measures of Speech Quality. Englewood Cliffs, NJ: Prentice Hall, 1988

  56. [56]

    An algorithm for intelligibility prediction of time–frequency weighted noisy speech,

    C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time–frequency weighted noisy speech,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 19, no. 7, pp. 2125–2136, 2011

  57. [57]

    Perceptual evaluation of speech quality (pesq): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,

    I.-T. Recommendation, “Perceptual evaluation of speech quality (pesq): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,”Rec. ITU-T P . 862, 2001

  58. [58]

    Decoupled weight decay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2017

  59. [59]

    Sgdr: Stochastic gradient descent with warm restarts,

    ——, “Sgdr: Stochastic gradient descent with warm restarts,”arXiv preprint arXiv:1608.03983, 2016