Pith. sign in

REVIEW 2 major objections 5 minor 41 references

Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications

T0 review · 2 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Playing modified loudspeaker signals carves a quiet zone around a voice assistant, and that alone improves speech recognition.

desk verdict A genuinely new loudspeaker-side spotformer that improves far-field ASR in tested conditions, but the robustness claim rests on a single reverberation time and the ASR gains are reported without error bars. read the letter →

arxiv 2501.08104 v1 pith:ILGJHCBN submitted 2025-01-14 eess.AS cs.SD

classification eess.AScs.SD
keywords loudspeakerbeamformingspotformingvoicedrivenapplicationsspeechrecognitionsoundzonesperceptualdistortionworderrorrate
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a voice-driven application can be made more robust to loud playback not by adding echo cancellation or extra sensors, but by changing what the loudspeakers play. Its loudspeaker spotformer modifies the playback signal so that a low-acoustic-energy region forms around the device running automatic speech recognition, while keeping the sound at the listener position close to the original in a perceptual sense. If the claim holds, any system with fixed loudspeakers and a known listener region, such as a home cinema or car audio, can improve voice-command accuracy simply by processing the playback stream. The experiments support this: in simulations and in a real room, word error rate and word information lost improve consistently with the spotformer on, and a control parameter trades quiet-zone depth against perceived quality.

What carries the argument

The central object is the loudspeaker spotformer (LSp), an adaptation of microphone spotforming: instead of filtering microphone signals to select a region of interest, it filters loudspeaker playback signals to create a region of low energy around the VDA. It is built from three components: a spatial covariance matrix $R_M(\omega)$ obtained by integrating free-field transfer vectors over a Gaussian-torus probability density $p_M$; a perceptual distortion measure $D(\hat{s},\hat{\epsilon})=\|P_s\hat{\epsilon}\|_2^2$ based on tonal masking that limits how far each control point's signal may deviate from the reference; and a convex optimization that minimizes the expected quiet-zone energy subject to the distortion constraints, with a frequency weighting $\alpha(\omega_k)$ that protects the speech band below about 100 Hz.

What would settle it

Run the identical loudspeaker and listener geometry in rooms with increasing reverberation time, and measure both the acoustic energy reduction at the VDA microphone array and the ASR word error rate with LSp on versus off; if the energy reduction shrinks or the WER gap closes as $T_{60}$ grows, the direct-path-only assumption is the limiting factor.

Watch

Extended reading notes

Core claim

The central claim is that placing the VDA inside a quiet zone created by playback-side beamforming is sufficient to improve ASR when loudspeaker audio is the dominant interferer. More precisely, the loudspeaker spotformer minimizes the expected acoustic energy in a torus-shaped region $M$ that models the microphone array's possible positions, using free-field direct-path transfer functions of the form $\hat{h}(x_s,x_r,\omega)=e^{-j\omega\|x_s-x_r\|_2/c}/(4\pi\|x_s-x_r\|_2)$, and it constrains the deviation from the unmodified signal at control points around the listener by a tonal-masking distortion bound. The optimization is solved frame by frame in the frequency domain, and robustness to position errors comes from spatially averaging the covariance matrix rather than from estimating the true room transfer function. Measured energy reductions reach at least 7 dB for the largest allowed distortion, the listener MOS stays around 4.4 out of 5, and both WER and WIL improve across all tested SIRs, including in a real room with $T_{60}\approx 220$ ms.

Load-bearing premise

The load-bearing premise is that a free-field, direct-path-only acoustic model is enough to design a quiet zone that survives real reverberation, so if reflections are stronger than those tested the quiet zone can refill and the ASR gain disappear.

Editorial extensions

If this is right

  • When loudspeaker playback is the dominant interferer, the modified playback signal by itself is enough to improve ASR; no echo cancellation or extra microphones are required.
  • The distortion parameter $d$ gives a predictable trade-off: raising it deepens the quiet zone (more energy reduction at the VDA) while lowering objective audio quality only slightly.
  • The improvement persists when the VDA also applies microphone beamforming, including MVDR and a microphone spotformer, so the playback-side and microphone-side methods stack.
  • The method's robustness transfers from simulation to a real room with similar reverberation time, suggesting it can be deployed without per-room transfer-function estimation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One extension the paper leaves implicit is tracking: recomputing the spotformer as the listener or the VDA moves, which today is limited by the cost of solving the convex problem each frame.
  • Because the distortion constraint is applied at control points around a single listener, a binaural or head-tracked listener may notice spatial-image changes; a multi-zone variant could preserve interaural cues.
  • In a car or living room with a known geometry, the same algorithm could be integrated with acoustic echo cancellation when playback signals are shared, potentially eliminating the residual interferer completely.
  • The free-field assumption suggests a testable scaling law: quiet-zone depth should degrade as reverberation time grows, so measuring energy reduction at several $T_{60}$ values would map the method's operating range.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The manuscript proposes a loudspeaker spotformer (LSp) that modifies playback signals from multiple loudspeakers so that a low-acoustic-energy region is created around the microphone array of a voice-driven application (VDA), while bounding perceptually weighted distortion at the listener's location. The algorithm is formulated as a convex optimization problem (Eq. 12) using free-field direct-path transfer functions (Eq. 1), a spatial covariance matrix (Eqs. 5-6) with a torus-shaped uncertainty distribution (Eq. 7), and a masking-based distortion constraint. Experiments in simulation and in a real room with T60 approximately 220 ms show that LSp reduces the received energy at the VDA and improves Whisper-based word error rate and word information lost on average across the tested conditions. The paper includes a link to the implementation code.

Significance. The idea of using loudspeaker beamforming to create a quiet zone around a voice-driven application without requiring acoustic echo cancellation or additional sensors is novel and practically appealing. The optimization formulation is clean, and the energy-reduction results are consistent between simulation and the real room, with code provided for reproducibility. However, the central ASR claim rests on average curves without statistical support, and the robustness claim is validated under a single reverberation condition; these issues must be addressed before the contribution is fully established.

major comments (2)
  1. [Sec. II (Eq. 1) and Sec. IV-B] The optimization uses free-field, direct-path-only transfer functions (Eq. 1), and the robustness claim rests on the spatial averaging in R_M (Eq. 6) and the distortion constraint at the control points. The spatial averaging integrates only over microphone positions, not over room responses, and the distortion constraint acts at the listener control points, not at the VDA region. The method is validated only at T60 approximately 220 ms in one simulated and one real room. Since Sec. II explicitly states that the algorithm relies on its robustness, the absence of experiments at higher reverberation (e.g., T60 > 500 ms) or an analytical bound on when the direct-path model remains valid leaves the central robustness claim unsupported. Please add multi-condition experiments or an analysis of the critical distance and its effect on the achievable null depth.
  2. [Sec. IV-C and Fig. 4] The ASR results are presented as averaged WER/WIL curves without error bars, confidence intervals, or significance tests. The caption states that outliers were removed, but no removal criterion is given, making the reported improvements (which appear modest at mid-to-high SIR) impossible to evaluate. Additionally, the distortion parameter d=5 was chosen by informal listening by the authors, although Sec. III-B states that d=1 is calibrated to the just-noticeable difference; this calibration should be justified or replaced by a systematic listening test. Please provide per-condition variance, specify the outlier-removal rule, and test the statistical significance of the LSp-on versus LSp-off differences.
minor comments (5)
  1. [Sec. IV-A] The sentence 'The number of modelled control points is P = 9placed in space sampled from a normal distribution...' is missing a space and should be reworded.
  2. [Sec. IV-A] The notation '3σr = 3σz = 0.095 m' and later 'with a standard deviation of 3σ = 0.2 m' is ambiguous; please state explicitly whether σ or 3σ is meant in each case.
  3. [Sec. IV-A] The sentence beginning 'the regularisation parameter α(ωk) = 0 in (12) equals 0 if |ωk| ≤ 200π rad/s...' is grammatically garbled and should be rewritten for clarity.
  4. [Fig. 4 caption] The caption should specify the SIR values used on the x-axis and the number of test utterances per condition; the statement 'outliers were removed' needs a precise definition.
  5. [Sec. IV-B] The MOS results in Fig. 3a are reported as averages without indicating the spread across test signals and validation points; please include variance information or state that the spread is negligible.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the ASR claim is tested against an external ASR system and held-out voice commands, not derived from the fitted constants.

full rationale

The central claim that the loudspeaker spotformer (LSp) improves ASR performance is evaluated empirically with Whisper Medium [39] and LibriSpeech voice commands [40] at multiple SIRs, in both a simulated reverberant room and a real room. This outcome is not a restatement of the optimization objective: the optimization in (12) minimizes the spatially averaged VDA energy defined by (5)-(7) under a perceptual distortion constraint, whereas the reported WER/WIL values in Fig. 4 are produced by an external, untrained-for-this-task ASR system. The measured energy reduction in Fig. 3b is indeed closely tied to the optimized objective, but the paper does not present that energy reduction as a prediction; it is the design goal. The free-field, direct-path-only RTF in Eq. (1) is an explicit modeling simplification, not a hidden reuse of the conclusion, and the paper acknowledges this by relying on the LSp algorithm's robustness and by validating in a real room. Parameter d=5 is hand-set by listening, not fitted to minimize WER, so no fitted input is renamed as a prediction. The self-citations [16] and [18] supply building blocks (region-based spotforming and perceptual masking constraints), but the ASR evaluation does not reduce to those references, and no uniqueness theorem is imported from the authors' prior work. The main caveat, that the direct-path model may fail at higher reverberation, is a robustness and correctness risk rather than a circularity.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim rests mainly on domain assumptions about sound propagation, listener and microphone position uncertainty, and perceptual modeling. The method has no free parameters fitted to the ASR outcome; parameters d, sigma, alpha, and P_ref boosts are hand-set before evaluating WER/WIL.

free parameters (5)
  • Distortion budget d = 5 (ASR tests)
    Controls the maximum allowed perceptual distortion at listener control points. Calibrated conceptually to d=1 for just-noticeable distortion; d=5 is chosen by the authors' own listening in Sec. IV-C, so it is hand-tuned rather than derived.
  • Spatial uncertainty standard deviations sigma_r, sigma_z = sigma_r = sigma_z = 0.0317 m (3 sigma = 0.095 m)
    Define the torus-like region M around the microphone array in Eq. (7). Larger values increase robustness but reduce achievable energy reduction. Chosen from assumed measurement accuracy.
  • Low-frequency weighting alpha(omega_k) = 0 for f <= 100 Hz, 1 otherwise
    User-defined weighting in the objective of Eq. (12); preserves speech-relevant low frequencies at the expense of less energy reduction there.
  • Masking-matrix boost for low frequencies = 100x the maximum inverse masking curve
    Applied to P_ref elements below 100 Hz, per [30], to keep that band unchanged; an ad hoc constant affecting the distortion constraint.
  • Control-point sampling spread = 9 points, 3 sigma = 0.2 m
    Listener control points sampled around x_u; this parameterizes the assumed listener position uncertainty in the distortion constraints.
assumptions (5)
  • domain assumption Free-field direct-path RTF model (Eq. 1)
    All room transfer functions are modeled as direct path only, ignoring reflections and reverberation during optimization. Robustness to reverberation is expected to come from spatial averaging in R_M and is only tested for T60 about 220 ms.
  • domain assumption Spatial probability model for microphone region (Eq. 7)
    The VDA microphone positions are assumed to follow a truncated Gaussian torus distribution with chosen mu_r, mu_z, sigma_r, sigma_z. This is the robustness mechanism, but it is a modeling assumption.
  • domain assumption Perceptual masking distortion model from van de Par et al. [19]
    The constraint uses a tonal-masking-based measure to predict audibility of distortion. It is objective and computational, but the paper does not validate it with subjective listening tests.
  • domain assumption A priori knowledge of loudspeaker, user, and VDA locations
    The algorithm requires the locations to lie in known or estimated regions, stated in Sec. I as reasonable for cars and living rooms.
  • domain assumption Single ASR engine, Whisper Medium, is representative
    ASR performance is measured with one recognizer; improvements may not transfer to other recognizers, especially those trained differently.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications." pith.science (2026). https://pith.science/paper/ILGJHCBN

@misc{pith2026250108104,
  author       = {Pith},
  title        = {Pith review of: Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ILGJHCBN}},
  note         = {Machine review of arXiv:2501.08104}
}
read the original abstract

In this paper we propose a robust loudspeaker beamforming algorithm which is used to enhance the performance of voice driven applications in scenarios where the loudspeakers introduce the majority of the noise, e.g. when music is playing loudly. The loudspeaker beamformer modifies the loudspeaker playback signals to create a low-acoustic-energy region around the device that implements automatic speech recognition for a voice driven application (VDA). The algorithm utilises a distortion measure based on human auditory perception to limit the distortion perceived by human listeners. Simulations and real-world experiments show that the proposed loudspeaker beamformer improves the speech recognition performance in all tested scenarios. Moreover, the algorithm allows to further reduce the acoustic energy around the VDA device at the expense of reduced objective audio quality at the listener's location.

Figures

Figures reproduced from arXiv: 2501.08104 by the authors.

Figure 1
Figure 1. The loudspeaker spotformer (LSp) setup. In (a), a topview schematic [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. A schematic view of the setup used in both simulations and real-world () [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The results of the objective speech quality metric (including zoom-in) [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: The results for word error rate (WER, upper row) and word [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

41 extracted references · 39 canonical work pages

  1. [1]

    Robust Speech Recognition Via Anchor Word Representations,

    B. King, I. Chen, Y . Vaizman, Y . Liu, R. Maas, S. H. K. Parthasarathi and B. Hoffmeister, “Robust Speech Recognition Via Anchor Word Representations,” Proc. Interspeech 2017 , Stockholm, Sweden, 2017, pp. 2471-2575

  2. [2]

    An experiment of dual-LTE MPTCP with In- Car V oice Assistant,

    V . A. Vu and M. Akselrod, “An experiment of dual-LTE MPTCP with In- Car V oice Assistant,”2021 IEEE 93rd Vehicular Technology Conference (VTC2021-Spring), Helsinki, Finland, 2021, pp. 1-5

  3. [3]

    V oice assistants in private households: a conceptual framework for future research in an interdisciplinary field,

    B. Minder, P. Wolf, M. Baldauf and S. Verma, “V oice assistants in private households: a conceptual framework for future research in an interdisciplinary field,” Humanit Soc Sci Commun 10, 173, 2023

  4. [4]

    Far-Field Automatic Speech Recognition,

    R. Haeb-Umbach, J. Heymann, L. Drude, S. Watanabe, M. Delcroix and T. Nakatani, “Far-Field Automatic Speech Recognition,” in Proceedings of the IEEE , vol. 109, no. 2, pp. 124-148, Feb. 2021

  5. [5]

    Whisper-AT: Noise- Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,

    Y . Gong, S. Khurana, L. Karlinsky and J. Glass, “Whisper-AT: Noise- Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,” Proc. Interspeech 2023 , 2023, pp. 2798-2802

  6. [6]

    Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation,

    A. Rouditchenko et al., “Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation,” arXiv:2406.10082 [eess.AS], June 2024

  7. [7]

    A Con- solidated Perspective on Multimicrophone Speech Enhancement and Source Separation,

    S. Gannot, E. Vincent, S. Markovich-Golan and A. Ozerov, “A Con- solidated Perspective on Multimicrophone Speech Enhancement and Source Separation,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, no. 4, pp. 692-730, April 2017

  8. [8]

    A Review of the State of the Art and Future Challenges of Deep Learning-Based Beamforming,

    H. A. Kassir, Z. D. Zaharis, P. I. Lazaridis, N. V . Kantartzis, T. V . Yioultsis and T. D. Xenos, “A Review of the State of the Art and Future Challenges of Deep Learning-Based Beamforming,” in IEEE Access , vol. 10, pp. 80869-80882, 2022

Show all 41 references
  1. [9]

    SpatialNet: Extensively Learning Spatial In- formation for Multichannel Joint Speech Separation, Denoising and Dereverberation,

    C. Quan and X. Li, “SpatialNet: Extensively Learning Spatial In- formation for Multichannel Joint Speech Separation, Denoising and Dereverberation,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 1310-1323, 2024

  2. [10]

    The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And Results,

    H. Chen et al., “The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And Results,” 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, Singapore, 2022, pp. 9266-9270

  3. [11]

    A proposed method to improve the WER of an ASR system in the noisy reverberant room

    M.E. Sadeghi, H. Sheikhzadeh and M. J. Emadi, “A proposed method to improve the WER of an ASR system in the noisy reverberant room”, in Journal of the Franklin Institute , vol. 361 (1), pp. 99-109, 2024

  4. [12]

    Soft Speech, Loud World: Bone Conduction Microphones Enhance V oice Assistant Interaction,

    C. Manzanillo, R. Chettiar, R. Soroushmojdehi, L. Ying, J. Dong and M. Mohsenvand, “Soft Speech, Loud World: Bone Conduction Microphones Enhance V oice Assistant Interaction,” 2024 IEEE International Confer- ence on Consumer Electronics (ICCE) , Las Vegas, NV , USA, 2024, pp. 1-5

  5. [13]

    mmMIC: Multi-modal Speech Recognition based on mmWave Radar,

    L. Fan, L. Xie, X. Lu, Y . Li, C. Wang and S. Lu, “mmMIC: Multi-modal Speech Recognition based on mmWave Radar,” IEEE INFOCOM 2023 - IEEE Conference on Computer Communications , New York City, NY , USA, 2023, pp. 1-10

  6. [14]

    Interface for Barge-in Free Spoken Dialogue System Based on Sound Field Reproduction and Microphone Array,

    S. Miyabe, Y . Hinamoto, H. Saruwatari, K. Shikano and Y . Tatekura, “Interface for Barge-in Free Spoken Dialogue System Based on Sound Field Reproduction and Microphone Array,” EURASIP J. Adv. Signal Process. 2007, 057470

  7. [15]

    Spotforming: Spatial Filtering With Distributed Arrays for Position-Selective Sound Acquisition,

    M. Taseska and E. A. P. Habets, “Spotforming: Spatial Filtering With Distributed Arrays for Position-Selective Sound Acquisition,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 7, pp. 1291-1304, July 2016

  8. [16]

    A robust region-based near- field beamformer,

    J. Martinez, N. Gaubitch and W. B. Kleijn, “A robust region-based near- field beamformer,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , South Brisbane, QLD, Aus- tralia, 2015, pp. 2494-2498

  9. [17]

    Signal-Adaptive and Perceptually Optimized Sound Zones With Variable Span Trade-Off Filters,

    T. Lee, J. K. Nielsen and M. G. Christensen, “Signal-Adaptive and Perceptually Optimized Sound Zones With Variable Span Trade-Off Filters,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 2412-2426, 2020

  10. [18]

    Block-Based Perceptually Adaptive Sound Zones with Reproduction Error Constraints,

    N. de Koeijer, M. B. Møller, J. Martinez, P. Mart ´ınez-Nuevo and R. C. Hendriks, “Block-Based Perceptually Adaptive Sound Zones with Reproduction Error Constraints,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing

  11. [19]

    A Perceptual Model for Sinusoidal Audio Coding Based on Spectral Integration

    S. van de Par, A. Kohlrausch, R. Heusdens, J. Jensen and S.H. Jensen, “A Perceptual Model for Sinusoidal Audio Coding Based on Spectral Integration”. EURASIP J. Adv. Signal Process . 2005, 317529 (2005)

  12. [20]

    Code related to “Loudspeaker Beamforming to Enhance Speech Recognition Performance of V oice Driven Applica- tions

    D. de Groot, 2025, “Code related to “Loudspeaker Beamforming to Enhance Speech Recognition Performance of V oice Driven Applica- tions””, 4TU.ResearchData. [Online]. Available: https://doi.org/10.4121/ 36b9065e-278e-40ee-b359-6cd734561f86

  13. [21]

    Time-delay estimation for TOA-based localization of multiple sensors,

    R. Heusdens and N. Gaubitch, “Time-delay estimation for TOA-based localization of multiple sensors,” 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Florence, Italy, 2014, pp. 609-613

  14. [22]

    Ahrens, Analytic Methods of Sound Field Synthesis , Springer, Berlin, Ger., Jan

    J. Ahrens, Analytic Methods of Sound Field Synthesis , Springer, Berlin, Ger., Jan. 2012

  15. [23]

    M.H. Hayes. Discrete-Time Random Processes. In: Statistical digital signal processing and modeling . John Wiley & Sons, 1996

  16. [24]

    Horn and C

    R. Horn and C. Johnson, Matrix Analysis (2nd Ed.) . Cambridge: Cam- bridge University Press, 2012

  17. [25]

    Brandstein and D

    M. Brandstein and D. B. Ward, Microphone Arrays: Signal Processing Techniques and Applications, Springer, Berlin, June 2001

  18. [26]

    N. L. Johnson, S. Kotz and N. Balakrishnan, Continuous Univariate Distributions, Volume 1 (2nd Ed.) , Wiley, 1994

  19. [27]

    Boyd and L

    S. Boyd and L. Vandenberghe, Convex Optimisation, Cambridge Uni- versity Press, 2009

  20. [28]

    Grant and S

    M. Grant and S. Boyd. CVX: Matlab software for disciplined convex programming, version 2.2. https://cvxr.com/cvx, 2014

  21. [29]

    Deng and D

    L. Deng and D. O’Shaughnessy, Speech Processing: A Dynamic and Optimization-Oriented Approach (1st Ed.) , CRC Press, 2003

  22. [30]

    Increasing Loudness in Audio Signals: A Perceptually Motivated Approach to Preserve Audio Quality,

    A. Jeannerot, N. de Koeijer, P. Mart ´ınez-Nuevo, M. B. Møller, J. Dyreby and P. Prandoni, “Increasing Loudness in Audio Signals: A Perceptually Motivated Approach to Preserve Audio Quality,” 2022 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICAS...

  23. [31]

    Smith, Spectral Audio Signal Processing

    J.O. Smith, Spectral Audio Signal Processing. http://ccrma.stanford.edu/ ∼jos/sasp/, online book, 2011 edition, accessed Apr. 2024

  24. [32]

    W. H. Press, B. P. Flannery, A. S. Teukolsky, W. T. Vetterling,Numerical Recipes in C : The Art of Scientific Computing , Cambridge University Press, Oct. 1992

  25. [33]

    Image method for efficiently simulating small room acoustics,

    J. B. Allen and D. A. Berkley, “Image method for efficiently simulating small room acoustics,” J. Acoust. Soc. Am. , vol. 65, no. 4, pp. 943–950, Apr. 1979

  26. [34]

    Room impulse response generator

    E. A. P. Habets, “Room impulse response generator”, Sept. 2010. GitHub Repository, https://github.com/ehabets/RIR-Generator, accessed May 2023

  27. [35]

    ViSQOLAudio: An objective audio quality metric for low bitrate codecs,

    A. Hines, E. Gillen, D. Kelly, J. Skoglund, A. Kokaram and N. Harte. “ViSQOLAudio: An objective audio quality metric for low bitrate codecs,” in J. Acoust. Soc. Am. vol. 137, no. 6, pp. EL449–EL455, 2015

  28. [36]

    ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric,

    M. Chinen, F. S. C. Lim, J. Skoglund, N. Gureev, F. O’Gorman and A. Hines, “ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric,” 2020 Twelfth International Conference on Quality of Multimedia Experience (QoMEX), Athlone, Ireland, 2020, pp. 1-6

  29. [37]

    ITU- T P.800, International Telecommunications Union, Geneva, Switzerland, 1996

    Methods for subjective determination of transmission quality, Rec. ITU- T P.800, International Telecommunications Union, Geneva, Switzerland, 1996

  30. [38]

    From WER and RIL to MER and WIL: improved evaluation measures for connected speech recognition,

    A. C. Morris, V . Maier and P. Green, “From WER and RIL to MER and WIL: improved evaluation measures for connected speech recognition,” Proc. Interspeech 2004 , 2004, pp. 2765-2768. 5

  31. [39]

    Robust Speech Recognition via Large-Scale Weak Super- vision,

    A. Radford, J.W. Kim, T. Xu, G. Brockman, C. McLeavey and I. Sutskever, “Robust Speech Recognition via Large-Scale Weak Super- vision,” arXiv:2212.04356 [eess.AS], Dec. 2022

  32. [40]

    Librispeech: An ASR corpus based on public domain audio books,

    V . Panayotov, G. Chen, D. Povey and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, QLD, Australia, 2015, pp. 5206-5210

  33. [41]

    Detecting Inaudible V oice Commands via Acoustic Attenuation by Multi-channel Micro- phones,

    X. Ji, G. Zhang, X. Li, G. Qu, X. Cheng and W. Xu, “Detecting Inaudible V oice Commands via Acoustic Attenuation by Multi-channel Micro- phones,” in IEEE Transactions on Dependable and Secure Computing , 2024. 6

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.