Pith. sign in

REVIEW 3 major objections 4 minor 38 references

Accelerating Audio Research with Robotic Dummy Heads

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A 3D-printed robotic dummy head records repeatable binaural audio while moving, bringing realistic motion into objective audio evaluation.

desk verdict Useful open-source robotic dummy head with mostly sound validation, but the quiet-motor claim needs an in-situ ear-microphone measurement before the headline contribution is fully established. read the letter →

arxiv 2505.04548 v1 pith:3TZYFJGX submitted 2025-05-07 eess.AS cs.HCcs.ROcs.SD

classification eess.AScs.HCcs.ROcs.SD
keywords roboticdummyheadbinauralaudiohead-relatedtransferfunctionspatially-dynamicrecordingsrepeatableexperimentsmotornoiseMVDRbeamforming3D-printedacoustics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that a low-cost, 3D-printed dummy head on a quiet turntable can move, talk, and listen during audio recordings without contaminating them with motor noise. This would close the gap between stationary acoustic mannequins, which are realistic but immobile, and mobile robots, which have been too noisy for clean recordings. The authors validate acoustic realism by comparing the head's HRTF, interaural level differences, and mouth radiation pattern with a KEMAR mannequin, and they show that artificially mixed recordings closely match physically mixed ones even while the talker rotates. Eight repeated recordings agree closely, supporting the claim that dynamic recordings are repeatable and suitable for objective evaluation of audio algorithms.

What carries the argument

The load-bearing mechanism is the quiet turntable: a gear-free, direct-drive stepper motor driven by a specialized control algorithm and operated below 0.4 revolutions per second, where motor harmonics are relatively spectrally white, with a 3D-printed structure that dampens vibration and no cooling fan. This lets the head rotate during recording without audible contamination. The other half is the 3D-printed dummy head, whose measured HRTF, interaural level and time differences, and mouth radiation pattern are qualitatively similar to a KEMAR, giving lifelike spatial cues at low cost and low mass. Together the two parts convert a binaural mannequin from a stationary measurement instrument into a repeatable moving source and listener.

What would settle it

Place a measurement microphone at or inside the ear canal of the rotating head in a quiet room and compare spectra with the motor still and rotating at 0.2 to 0.4 rev/s; if motor harmonics exceed the stationary noise floor at speech frequencies, the claim that recordings are uncontaminated and repeatable fails for those speeds.

Watch

Extended reading notes

Core claim

The central claim is that spatially-dynamic audio recordings can be made repeatable and physically realistic by using a quiet, robotically rotated acoustic dummy head. The device combines a 3D-printed head with two in-ear microphones and a mouth loudspeaker, mounted on a direct-drive stepper motor turntable whose motor noise stays at or below whisper level near conversational distances. Benchmark experiments show that the normalized mean-squared error between artificially mixed and physically mixed waveforms remains low at all tested rotation speeds, and repeated recordings are highly consistent. The authors conclude that these robot-enabled recordings are repeatable and suited for objective evaluation, and they demonstrate the utility by applying a motion-robust minimum-variance distortionless-response (MVDR) beamformer to a moving talker, finding that the presence of motion, not its rate, causes a severe drop in high-frequency SNR gain.

Load-bearing premise

The paper's motor-noise evidence comes from a microphone 1.0 m away, so the load-bearing assumption is that the ear microphones, mounted close to the turntable, also record clean audio while the head is moving.

Editorial extensions

If this is right

  • Researchers can automate the labor-intensive data collection for binaural and spatial audio experiments by scripting head rotations instead of repositioning loudspeakers or people.
  • Separate noise-only and target recordings from a moving source can be scaled and summed to synthesize arbitrary SNR conditions without extra recording passes, with error comparable to ambient noise.
  • Dynamic scenarios involving a rotating talker become objectively evaluable: repeated recordings with identical motion show high repeatability, which human actors cannot provide.
  • Because the design uses standard 3D printing and a low-cost stepper motor, laboratories can build the device and reproduce dynamic binaural experiments without a KEMAR or an anechoic robot facility.
  • The beamforming result indicates that algorithm evaluation must include motion as a condition, since even slow rotation changes performance in ways stationary tests miss.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If self-noise at the ear microphones proves as low as the far-field measurements suggest, the same platform could generate large labeled datasets for sound localization, head-pose estimation, and cocktail-party separation with physically moving sources rather than simulated room impulse responses.
  • An immediate test the paper does not report is placing a probe microphone at or inside the ear position during rotation; this would directly confirm the binaural recordings are free of motor contamination at the speeds recommended for experiments.
  • The preliminary beamformer result suggests motion itself may disrupt relative transfer function tracking; varying the rotation trajectory or the adaptation time constant would reveal whether the high-frequency drop is fundamental or an artifact of the 200 ms forgetting factor.
  • The one-axis turntable could be extended to a second rotation axis or a mobile base, letting researchers approximate natural head and body motion while keeping the same repeatable recording protocol.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents a low-cost, open-source robotic dummy head: a 3D-printed acoustic mannequin with binaural ear microphones and a mouth loudspeaker, mounted on a stepper-motor turntable. The authors argue that the device combines the acoustic realism of conventional mannequins with precise, quiet mobility, enabling automated spatially-stationary experiments and repeatable spatially-dynamic recordings. Validation includes HRTF comparisons against a KEMAR mannequin (qualitative magnitude plots, quantitative ILD and ITD), motor-noise measurements at 1 m distance, repeatability tests over eight repetitions, a comparison of artificially mixed versus physically mixed recordings, and a demonstration of a motion-robust binaural MVDR beamformer. The design files are provided as open source.

Significance. If the central claim is sustained, the device would be a useful research tool for repeatable dynamic binaural audio experiments and for generating large labeled spatial-audio datasets. The paper's strengths include its explicit external validation against KEMAR (with a reported ITD RMSE of 67.9 microseconds), the use of eight repeated recordings to demonstrate repeatability, the comparison of artificial versus physical mixing, and the release of open-source design files. These features make the claims independently checkable and are appropriate for a tools-focused audio research paper.

major comments (3)
  1. [Section 2.2, Figs. 5-7] The 'quiet motor' claim is load-bearing for the entire paper, but it is validated only with a microphone placed 1.0 m from the device, not at the binaural ear-microphone positions used in the actual experiments. The ear microphones are mounted in the printed head close to the turntable motor, so structure-borne and near-field airborne motor noise could be substantially higher at those positions than at 1 m. The paper should report an in-situ motor-on versus motor-off measurement at the ear microphones, ideally in terms of the resulting SNR or noise spectrum relative to the target speech level, because this is the quantity that determines whether the device is acoustically unobtrusive for objective evaluation.
  2. [Section 3, Figs. 9-10] The repeatability and artificial/physical mixing experiments cannot detect deterministic motor noise, so they do not by themselves establish that the motor is quiet. In Fig. 9, if the target recording contains motor noise, that same noise appears in both the artificially mixed and physically mixed signals, so high agreement is expected regardless of the noise amplitude. In Fig. 10, repeated target recordings from the same stepper motor driven by the same control signal will contain nearly identical motor noise, and comparing each repetition to the sample-wise average will show high repeatability even if the motor noise is substantial. These results demonstrate that the motor noise is repeatable, not that it is absent or inaudible. A separate in-situ noise measurement is needed to support the quiet-motor claim.
  3. [Section 4.3, Fig. 11] The beamforming comparison between static and moving talkers could be confounded by motor noise, because any motor noise radiated during motion is folded into the target or mixture signals rather than modeled as an independent interferer. If motor noise is present at the ear microphones during the moving condition, the observed high-frequency SNR-gain drop could be partly attributable to that noise rather than to the effect of motion on RTF estimation. The authors should either provide the in-situ motor-noise measurement recommended above, or analyze the beamforming result with motor-off reference recordings, before interpreting the 'surprising result' as a motion-related phenomenon.
minor comments (4)
  1. [Fig. 10 caption] The caption for Fig. 10 appears to be copied from Fig. 8 and describes a beamforming setup, whereas the text and the figure itself describe repeatability of repeated target speech recordings. The caption should be corrected to describe what is actually plotted.
  2. [Section 2.1, Fig. 3 text] The text reports a 'root mean-squared error (MSE) of 67.9 µs'; since the quantity is in microseconds, this should be called RMSE rather than MSE, or the abbreviation should be defined consistently.
  3. [Fig. 2 and Fig. 3] The HRTF comparison is mostly qualitative; adding a quantitative frequency-dependent magnitude error metric (for example, average absolute dB difference per frequency band between the printed head and KEMAR) would strengthen the acoustic-realism claim.
  4. [Section 4.2, Eq. (5)] The forgetting factor α is said to correspond to a time constant τ, but the explicit relationship (such as α = exp(-1/(τ·fs)) or the equivalent per-frame formula) is not given. Please state this relationship so the parameter setting is reproducible.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: all load-bearing validation is against external benchmarks (KEMAR, SPL-calibrated measurements) rather than fitted or self-referential quantities.

full rationale

The paper's central claims—acoustic realism of the printed head, quietness of the motor, and repeatability of robot-enabled recordings—are each validated against external references rather than against the paper's own fitted parameters. The HRTF and ILD comparisons in Figs. 2–3 are benchmarked against a calibrated GRAS 45BC KEMAR; the motor-noise measurements in Figs. 5–7 are calibrated with an SPL meter and compared with external references such as conversational speech and a whisper; and the repeatability benchmarks in Figs. 9–10 compare physically mixed and artificially mixed waveforms and repeated recordings. No equation in the paper defines an output quantity in terms of the quantity it is meant to predict, and no fitted parameter is repackaged as a prediction. The self-citations to prior work on 3D-printed head simulators [24]–[26] are contextual: the current paper independently measures the printed head's HRTF, ILD, ITD, and mouth-simulator radiation pattern against KEMAR equipment, so the cited prior work is not load-bearing. The beamforming experiment in Section 4 is presented as a demonstration of the device's utility, not as a derived theoretical result, and its outcome is an observation rather than a prediction forced by construction. The skeptical concern about motor noise not being measured directly at the ear microphones is an evidentiary gap or correctness risk, not a circularity; the paper never equates the 1-m measurement with the ear-position condition or defines the quiet-motor claim in terms of the benchmark outcomes. Therefore the derivation chain is self-contained, and no circular step can be exhibited.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

No fundamentally new physical entities are introduced. The only free parameters are operational choices in the demo and hardware. The central validation is against an external KEMAR reference and measured sound levels.

free parameters (2)
  • Motor speed limit = 0.4 rev/s
    Chosen by hand as the maximum speed at which motor noise is 'relatively spectrally white' (Sec. 2.2). It affects the dynamic experiments but is not fitted to the evaluation metrics.
  • Beamformer forgetting factor tau = 200 ms
    Set to a single value for the MVDR+CW beamformer demo (Sec. 4.2); a standard algorithm parameter, not optimized against the reported SNR gain.
assumptions (3)
  • domain assumption KEMAR mannequin is the accepted acoustic reference for dummy-head realism.
    The paper compares HRTF, ILD, and ITD of the printed head to KEMAR measurements (Figs. 2-3) and interprets similarity as evidence of acoustic realism.
  • domain assumption Spectral subtraction-based background noise removal is valid for the motor noise measurements.
    In Fig. 5, background noise 'significant below 500 Hz' is removed by spectral subtraction, assuming stationarity and independence from motor noise.
  • domain assumption Artificial mixing of separately recorded noise and target signals accurately models a natural mixture when no equipment moves between recordings.
    Used in Sec. 3 to create mixtures for SNR calculations; standard practice in speech enhancement evaluation but an assumption nonetheless.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Accelerating Audio Research with Robotic Dummy Heads." pith.science (2026). https://pith.science/paper/3TZYFJGX

@misc{pith2026250504548,
  author       = {Pith},
  title        = {Pith review of: Accelerating Audio Research with Robotic Dummy Heads},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3TZYFJGX}},
  note         = {Machine review of arXiv:2505.04548}
}
read the original abstract

This work introduces a robotic dummy head that fuses the acoustic realism of conventional audiological mannequins with the mobility of robots. The proposed device is capable of moving, talking, and listening as people do, and can be used to automate spatially-stationary audio experiments, thus accelerating the pace of audio research. Critically, the device may also be used as a moving sound source in dynamic experiments, due to its quiet motor. This feature differentiates our work from previous robotic acoustic research platforms. Validation that the robot enables high quality audio data collection is provided through various experiments and acoustic measurements. These experiments also demonstrate how the robot might be used to study adaptive binaural beamforming. Design files are provided as open-source to stimulate novel audio research.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

38 extracted references · 36 canonical work pages

  1. [9]

    Mecha- tronic Generation of Datasets for Acoustics Research,

    Austin Lu, Ethaniel Moore, Arya Nallanthighall, Kanad Sarkar, Manan Mittal, Ryan M Corey, Paris Smaragdis, and Andrew Singer, “Mecha- tronic Generation of Datasets for Acoustics Research,” in 2022 International Workshop on Acoustic Signal Enhancement (IWAENC) , 2022, pp. 1–5

  2. [25]

    3d-printed acoustic head simulators that talk and move,

    Xinran Yue, Arya Nallanthighall, Manan Mittal, Austin Lu, Kanad Sarkar, Ryan M Corey, Paris Smaragdis, and Andrew C Singer, “3d-printed acoustic head simulators that talk and move,” The Journal of the Acoustical Society of America , vol. 153, no. 3 supplement, pp. A38–A38, 2023

  3. [1]

    The PASCAL CHiME speech separation and recognition challenge,

    Jon Barker, Emmanuel Vincent, Ning Ma, Heidi Christensen, and Phil Green, “The PASCAL CHiME speech separation and recognition challenge,” Computer Speech & Language , vol. 27, no. 3, pp. 621– 633, 2013

  4. [2]

    Evaluation of objective quality measures for speech enhancement,

    Yi Hu and Philipos C. Loizou, “Evaluation of objective quality measures for speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 16, no. 1, pp. 229–238, 2008

  5. [3]

    Performance measurement in blind audio source separation,

    E. Vincent, R. Gribonval, and C. Fevotte, “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1462–1469, 2006

  6. [4]

    Low latency two stage beamforming with distributed microphone arrays using a planewave decomposition,

    Manan Mittal, Ryan M. Corey, Yongjie Zhuang, and Andrew C. Singer, “Low latency two stage beamforming with distributed microphone arrays using a planewave decomposition,” in 2024 18th International Workshop on Acoustic Signal Enhancement (IWAENC) , 2024, pp. 180–184

  7. [5]

    HRFT Measurements of a KEMAR Dummy-head Microphone,

    Bill Gardner, Keith Martin, et al., “HRFT Measurements of a KEMAR Dummy-head Microphone,” 1994

  8. [6]

    Audio signal processing in the 21st century: The important outcomes of the past 25 years,

    Ga¨el Richard, Paris Smaragdis, Sharon Gannot, Patrick A Naylor, Shoji Makino, Walter Kellermann, and Akihiko Sugiyama, “Audio signal processing in the 21st century: The important outcomes of the past 25 years,” IEEE Signal Processing Magazine , vol. 40, no. 5, pp. 12–26, 2023

Show all 38 references
  1. [7]

    Micbots: Collecting large realistic datasets for speech and audio research using mobile robots,

    Jonathan Le Roux, Emmanuel Vincent, John R. Hershey, and Daniel P.W. Ellis, “Micbots: Collecting large realistic datasets for speech and audio research using mobile robots,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 5635–5639

  2. [8]

    MIRACLE—a microphone array impulse response dataset for acoustic learning,

    Adam Kujawski, Art JR Pelling, and Ennes Sarradj, “MIRACLE—a microphone array impulse response dataset for acoustic learning,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2024, no. 1, pp. 32, 2024

  3. [10]

    Reverberant sound localization with a robot head based on direct-path relative transfer function,

    Xiaofei Li, Laurent Girin, Fabien Badeig, and Radu Horaud, “Reverberant sound localization with a robot head based on direct-path relative transfer function,” in 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2016, pp. 2819–2826

  4. [11]

    The cocktail party robot: Sound source separation and localisation with an active binaural head,

    Antoine Deleforge and Radu Horaud, “The cocktail party robot: Sound source separation and localisation with an active binaural head,” in Proceedings of the seventh annual ACM/IEEE international conference on Human-Robot Interaction , 2012, pp. 431–438

  5. [12]

    Acoustic space learning for sound-source separation and localization on binaural manifolds,

    Antoine Deleforge, Florence Forbes, and Radu Horaud, “Acoustic space learning for sound-source separation and localization on binaural manifolds,” International Journal of Neural Systems , vol. 25, no. 01, 2015, Art no. 1440003

  6. [13]

    Automating acoustic signal processing experiments and audio machine learning datasets using robots,

    Austin Lu, “Automating acoustic signal processing experiments and audio machine learning datasets using robots,” M.S. thesis, University of Illinois at Urbana-Champaign, 2024

  7. [14]

    Comparison of binaural RTF-vector- based direction of arrival estimation methods exploiting an external microphone,

    Daniel Fejgin and Simon Doclo, “Comparison of binaural RTF-vector- based direction of arrival estimation methods exploiting an external microphone,” in 2021 29th European Signal Processing Conference (EUSIPCO). IEEE, 2021, pp. 241–245

  8. [15]

    The fifth ’CHiME’ speech separation and recognition challenge: dataset, task and baselines,

    Jon Barker, Shinji Watanabe, Emmanuel Vincent, and Jan Trmal, “The fifth ’CHiME’ speech separation and recognition challenge: dataset, task and baselines,” arXiv preprint arXiv:1803.10609 , 2018

  9. [16]

    Adaptive crosstalk cancellation and spatialization for dynamic group conversation enhancement using mobile and wearable devices,

    Ryan M Corey, Manan Mittal, Kanad Sarkar, and Andrew C Singer, “Adaptive crosstalk cancellation and spatialization for dynamic group conversation enhancement using mobile and wearable devices,” in 2022 International Workshop on Acoustic Signal Enhancement (IWAENC) . IEEE, 2022...

  10. [17]

    The LOCATA challenge data corpus for acoustic source localization and tracking,

    Heinrich W L ¨ollmann, Christine Evers, Alexander Schmidt, Heinrich Mellmann, Hendrik Barfuss, Patrick A Naylor, and Walter Kellermann, “The LOCATA challenge data corpus for acoustic source localization and tracking,” in 2018 IEEE 10th Sensor array and multichannel signal proc...

  11. [18]

    The LOCATA challenge: Acoustic source localization and tracking,

    Christine Evers, Heinrich W L ¨ollmann, Heinrich Mellmann, Alexander Schmidt, Hendrik Barfuss, Patrick A Naylor, and Walter Kellermann, “The LOCATA challenge: Acoustic source localization and tracking,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28,...

  12. [19]

    Ego- noise reduction using a motor data-guided multichannel dictionary,

    Alexander Schmidt, Antoine Deleforge, and Walter Kellermann, “Ego- noise reduction using a motor data-guided multichannel dictionary,” in 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2016, pp. 1281–1286

  13. [20]

    Gear noise and the sideband phenomenon,

    AK Dale, “Gear noise and the sideband phenomenon,” ASME Paper, vol. 84, 1987

  14. [21]

    Emmanuel Vincent, Tuomas Virtanen, and Sharon Gannot, Audio source separation and speech enhancement , John Wiley & Sons, 2018

  15. [22]

    The second ‘chime’ speech separation and recognition challenge: Datasets, tasks and baselines,

    Emmanuel Vincent, Jon Barker, Shinji Watanabe, Jonathan Le Roux, Francesco Nesta, and Marco Matassoni, “The second ‘chime’ speech separation and recognition challenge: Datasets, tasks and baselines,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Process...

  16. [23]

    The signal separation evaluation campaign (2007– 2010): Achievements and remaining challenges,

    Emmanuel Vincent, Shoko Araki, Fabian Theis, Guido Nolte, Pau Bofill, Hiroshi Sawada, Alexey Ozerov, Vikrham Gowreesunker, Dominik Lutter, and Ngoc QK Duong, “The signal separation evaluation campaign (2007– 2010): Achievements and remaining challenges,” Signal Processing, vol...

  17. [24]

    Acoustic impulse responses for wearable audio devices,

    Ryan M. Corey, Naoki Tsuda, and Andrew C Singer, “Acoustic impulse responses for wearable audio devices,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, pp. 216–220

  18. [26]

    Comparison of the acoustic effects of face masks on speech,

    Ryan M Corey, Uriah Jones, and Andrew C Singer, “Comparison of the acoustic effects of face masks on speech,” The Hearing Journal , vol. 74, no. 1, pp. 36–38, 2021

  19. [27]

    Directiv- ity of artificial and human speech,

    Teemu Halkosaari, Markus Vaalgamaa, and Matti Karjalainen, “Directiv- ity of artificial and human speech,” Journal of the Audio Engineering Society, vol. 53, no. 7/8, pp. 620–631, 2005

  20. [28]

    Horizontal directivity of low-and high-frequency energy in speech and singing,

    Brian B Monson, Eric J Hunter, and Brad H Story, “Horizontal directivity of low-and high-frequency energy in speech and singing,” The Journal of the Acoustical Society of America , vol. 132, no. 1, pp. 433–441, 2012

  21. [29]

    Simultaneous measurement of impulse response and distortion with a swept-sine technique,

    Angelo Farina, “Simultaneous measurement of impulse response and distortion with a swept-sine technique,” Preprints-Audio Engineering Society, 2000

  22. [30]

    A multi- loudspeaker binaural room impulse response dataset with high-resolution translational and rotational head coordinates in a listening room,

    Yue Qiao, Ryan Miguel Gonzales, and Edgar Choueiri, “A multi- loudspeaker binaural room impulse response dataset with high-resolution translational and rotational head coordinates in a listening room,” Frontiers in Signal Processing , vol. 4, 2024

  23. [31]

    A hybrid framework for ego noise cancellation of a robot,

    G¨okhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Yuji Hasegawa, Hiroshi Tsujino, and Jun-ichi Imura, “A hybrid framework for ego noise cancellation of a robot,” in 2010 IEEE International Conference on Robotics and Automation . IEEE, 2010, pp. 3623–3628

  24. [32]

    Noise Reduction in Spur Gear Systems,

    Aurelio Liguori, Enrico Armentani, Alcide Bertocco, Andrea Formato, Arcangelo Pellegrino, and Francesco Villecco, “Noise Reduction in Spur Gear Systems,” Entropy, vol. 22, no. 11, 2020

  25. [33]

    An Efficient Stepper Motor Audio Noise Filter,

    Fitzgerald J Archibald, “An Efficient Stepper Motor Audio Noise Filter,” Texas Instruments White Paper , 2008

  26. [34]

    Enhancement of remotely controlled laboratory for Active Noise Control and acoustic experiments,

    I. Khan, M. ˙Zmuda, P. Konopka, I. Gustavsson, and L. H ˚akansson, “Enhancement of remotely controlled laboratory for Active Noise Control and acoustic experiments,” in 2014 11th International Conference on Remote Engineering and Virtual Instrumentation (REV) , 2014, pp. 285– 290

  27. [35]

    A comparison of spectra of loud and whispered speech,

    Igor V N ´abelek and Sumalai Maroonroge, “A comparison of spectra of loud and whispered speech,” The Journal of the Acoustical Society of America, vol. 75, no. S1, pp. S83–S83, 1984

  28. [36]

    RTF-steered binaural MVDR beamforming incorporating an external microphone for dynamic acoustic scenarios,

    Nico G ¨oßling and Simon Doclo, “RTF-steered binaural MVDR beamforming incorporating an external microphone for dynamic acoustic scenarios,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 416–420

  29. [37]

    A consolidated perspective on multimicrophone speech enhancement and source separation,

    Sharon Gannot, Emmanuel Vincent, Shmulik Markovich-Golan, and Alexey Ozerov, “A consolidated perspective on multimicrophone speech enhancement and source separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 4, pp. 692–730, 2017

  30. [38]

    Performance analysis of the covariance subtraction method for relative transfer function estimation and comparison to the covariance whitening method,

    Shmulik Markovich-Golan and Sharon Gannot, “Performance analysis of the covariance subtraction method for relative transfer function estimation and comparison to the covariance whitening method,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing ...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.