Pith. sign in

REVIEW 5 major objections 7 minor 57 references

MaskClip: Detachable Clip-on Piezoelectric Sensing of Mask Surface Vibrations for Real-time Noise-Robust Speech Input

T0 review · 5 major / 7 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read A detachable piezoelectric clip on a face mask can capture the wearer's speech through mask-surface vibrations, keeping character error at 6.1% in noisy conditions—far below a conventional pin microphone's 19.7%.

desk verdict A genuinely useful incremental hardware paper—mask-surface piezo clip with a plausible noise-robustness story—whose main numbers come from a manikin and therefore need human confirmation before the broader claims land. read the letter →

arxiv 2505.02180 v1 pith:WNVNUPTU submitted 2025-05-04 cs.SD cs.ARcs.HCeess.AS

classification cs.SDcs.ARcs.HCeess.AS
keywords noise-suppressivemicrophonespeechenhancementpiezoelectricsensingwearabledevicevoiceuserinterfacefacemaskvibrationwhisper
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that the wearer's voice can be captured cleanly by sensing vibrations on the outside of a face mask, using a detachable clip with a piezoelectric element, so that ambient noise is physically rejected at the sensor rather than removed later by software. If this holds, medical staff, cleanroom workers, and others who must wear masks could get reliable hands-free voice input in noisy settings without heavy signal processing or skin-contact devices. The reported numbers support the claim: character error rates of 5.1% in quiet and 6.1% in noisy conditions, versus 9.4% and 19.7% for a conventional pin microphone under the same test setup. The authors also report higher subjective audio quality scores from 102 listeners in a MUSHRA test. The load-bearing premise is that a head-and-torso simulator with a mouth simulator behaves like a real person wearing a mask; the paper itself notes that walking and head movements remain unverified.

What carries the argument

The key machinery is the detachable stainless-steel clip with an inward-facing piezoelectric element: the element converts mechanical deformation of the mask surface into voltage, and the clip's rigidity and vibration transmission allow the sensor to receive the signal even when it is not touching the face. The clip is attached to a small circuit with a preamplifier, a 24-bit ADC, and an ESP32 microcontroller that streams audio wirelessly, so no on-device denoising is needed. The authors' systematic comparison of clip material, sensor orientation, and position establishes that this combination maximizes pick-up of speech-induced vibration while keeping the device lightweight and detachable for hygiene.

What would settle it

Record MaskClip on human speakers in a noisy room with two interfering talkers at 30 cm, using the same Whisper transcription pipeline; if the character error rate is substantially above 6.1% or degrades with head movement, the central transferability claim fails.

Watch

Extended reading notes

Core claim

The central discovery is that a piezoelectric sensor clipped to the outer surface of a face mask can act as a vibration microphone for the wearer's speech, using the mask itself as the transmission medium. Speech sets the mask surface vibrating through the face and jaw; ambient noise, arriving as airborne pressure fluctuations, does not produce comparable vibration of the mask, so the sensor preferentially picks up the wearer's voice. A systematic parameter sweep on a head-and-torso simulator found that an inward-facing sensor on a stainless steel clip placed about 10 mm from the mask's left edge gives the strongest signal—roughly 70 dB—and still 45–50 dB at 30–60 mm positions where skin contact is absent. In speech-recognition tests with Whisper-Large-V3, MaskClip achieved a character error rate of 6.1% with two interfering talkers 30 cm away plus background noise, compared with 19.7% for a pin microphone; software denoising applied to the pin microphone produced worse results (43.1% with a denoiser, 26.3% with a separation model). The paper interprets this as evidence that hardware-based capture at the mask surface is a practical alternative to computation-heavy speech enhancement.

Load-bearing premise

The measured noise-robustness was obtained with a head-and-torso simulator and mouth simulator, not with human wearers, so the claim depends on the vibration coupling between a real face, a real mask, and the clip being similar enough that the 6.1% character error rate transfers to real people.

Editorial extensions

If this is right

  • Voice-controlled equipment in operating rooms could work hands-free without software denoising and without breaking sterile technique.
  • The detachable design lets masks be disposed of or cleaned separately, addressing a hygiene barrier in mask-integrated microphones.
  • The approach extends to whispered speech, where MaskClip remained competitive with a standard microphone in subjective listening tests.
  • Because the sensor reads vibrations rather than air pressure, motion artifact and friction noise that plague throat and NAM microphones may be reduced, although dynamic conditions are not yet verified.
  • The minimal rise in error from quiet to noisy conditions suggests the mask-vibration channel is largely insensitive to acoustic interference, pointing toward stable voice-input performance across changing noise levels.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension not tested in the paper is applying the same clip principle to other protective equipment—hard hats, goggles, respirators—whose surfaces vibrate with speech, potentially generalizing the noise-robustness claim beyond face masks.
  • The roughly 1% character-error degradation from quiet to noisy conditions hints that a small array of piezoelectric clips on the mask could enable spatial separation of multiple talkers, an extension the paper mentions but does not test.
  • The paper's observation that software denoisers worsened pin-microphone recognition in this setup may reflect a mismatch between those models' training conditions and the close-range, non-stationary interference used here; this is an editorial reading, not a claim the paper makes.
  • If vibration coupling transfers from the head-and-torso simulator to real wearers, comparing MaskClip against a bone-conduction microphone on the same noisy-mask task would clarify which vibration pathway—mask surface or tissue-borne—carries the speech signal most faithfully.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper proposes MaskClip, a detachable clip-on piezoelectric sensor that attaches to a face mask and records the wearer's speech by sensing mask-surface vibrations. The authors argue that this hardware approach inherently suppresses ambient noise because environmental sound does not sufficiently vibrate the mask. They describe the device's implementation, a systematic parameter sweep on a head-and-torso simulator (HATS) to choose clip material, sensor orientation, and position, and two evaluations: an ASR-based CER comparison on HATS recordings (MaskClip 5.1% quiet / 6.1% noisy vs. Pin-mic 9.4% / 19.7%) and a MUSHRA subjective listening test with 102 participants. The paper claims practical value for medical, cleanroom, and industrial settings and lists human-subject and dynamic-movement evaluation as future work.

Significance. If the reported HATS results transfer to human wearers, the contribution is a low-cost, hygiene-preserving, hardware-only noise-suppression mechanism for mask-based voice input, with a plausible physical principle and a controlled evaluation setup. The use of a standardized HATS, publicly available speech/noise corpora, and a subjective MUSHRA study with a reasonably large participant pool are strengths. The key unresolved question is whether the vibration-coupling path on a real human face, with soft tissue and articulation, is similar enough to the HATS manikin to support the stated conclusions; the authors themselves acknowledge this gap. The paper is therefore a promising demonstration of a principle, but its central generalization claim is not yet supported by the evidence presented.

major comments (5)
  1. [Section 4.3 / Section 5.2] The central noise-robustness claim (CER 5.1% quiet / 6.1% noisy vs. Pin-mic 9.4% / 19.7%) rests entirely on recordings made with a SAMAR4700M HATS and mouth simulator, not on any human wearer. The physical mechanism depends on how speech-induced vibration propagates from a soft, articulating human face through the mask to the clip; the paper itself states that performance under dynamic usage conditions such as walking and head movements is unverified (Section 5.2) and lists human-subject evaluation only as future work (Section 6). Because the stated real-world applications are human-facing, the reported CER values cannot yet be claimed to transfer to people. Please add at least a static human-subject recording experiment in quiet and noise, or substantially reframe the conclusion as a HATS-specific validation.
  2. [Section 4.3 / Fig. 7] The headline CER values are reported as single point estimates with no confidence intervals, error bars, or significance tests. Given that the core comparison (MaskClip 5.1/6.1 vs. Pin-mic 9.4/19.7) is the paper's primary quantitative evidence, please report per-utterance variability (e.g., box plots, bootstrap confidence intervals, or repeated trials) so that the difference can be assessed statistically.
  3. [Section 4.3] The software baselines Denoiser and Sepformer applied to the Pin-mic signal yield CER 43.1% and 26.3%, respectively, which is substantially worse than the unprocessed Pin-mic baseline (19.7%). This is surprising because these models are designed to improve speech quality and recognition, and the large degradation suggests a mismatch in how they were applied (for example, model training domain, sampling rate, or input conditioning) rather than a fair comparison. Please specify the exact model versions, preprocessing, and processing chain used for these baselines, and explain why they degrade performance; without this, the claim that MaskClip 'outperforms' state-of-the-art software methods is uninterpretable.
  4. [Section 3.2] The hardware description is internally inconsistent: the text names the Seeed ESP32 (S3 Xiao) as the microcontroller and then states that 'the wireless transmission is handled by the nRF52840's integrated Bluetooth 5.0 module.' Please clarify which microcontroller is actually used, or correct the description. This matters because the paper's contribution includes the practical low-power wireless streaming device.
  5. [Section 3.3 / Section 3.4] The optimal configuration (stainless steel clip, inward-facing sensor, 10 mm position) was chosen from a parameter sweep conducted on the same HATS apparatus that was then used for the main evaluation. This introduces a risk of overfitting the sensor configuration to the specific manikin, mask mounting, and mouth simulator. The paper should either justify the transferability of this configuration independently or test it on a second setup and on human subjects to rule out that the choice is an artifact of the HATS rig.
minor comments (7)
  1. [Section 4.3] The sentence 'Both microphones were positioned 2cm from the mask's center on the opposite side' is unclear, since MaskClip is attached to the mask surface; please specify which microphones were at 2 cm and what 'opposite side' means.
  2. [Section 4.4.2] The whispered-speech comparison mixes 'standard microphone' and 'noise-free standard microphone' conditions; please define all MUSHRA conditions consistently in both the text and Fig. 8.
  3. [Section 4.4] The text reports values such as '73.0±9.8' and refers to 'standard errors' in the figure caption, but does not state whether the ± values are standard deviations or standard errors; please specify explicitly.
  4. [Section 4.3] There is a typo in 'The MaskCIP system' — should be 'MaskClip'.
  5. [Section 3.2] The device weight is described as 'approximately 20g'; please provide a measured value if available, since the claimed lightweight design is part of the contribution.
  6. [Figure 5] The y-axis in Fig. 5 has no visible axis label; the caption mentions decibels, but the axis itself should be labeled (e.g., 'Signal level (dB)').
  7. [References] The LibriMix dataset is cited via reference [55], which is actually a paper on acoustic sensing ('Sensing to Hear'); LibriMix is normally attributed to Cosentino et al. Please verify and correct the citation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the CER and MUSHRA results are independent measurements, and the sensor-configuration sweep is an engineering choice rather than a fitted prediction.

full rationale

The paper's central claims are empirical rather than derived from a model, so there is no derivation chain that could collapse into its inputs. Section 4.3 reports CER values (MaskClip 5.1% quiet / 6.1% noisy; Pin-mic 9.4% / 19.7%) obtained by running Whisper-Large-V3 on recorded audio, and Section 4.4 reports independent MUSHRA ratings from 102 participants; neither outcome was used to fit any parameter. The final sensor configuration was selected in Section 3.4 from a systematic sweep using signal-level measurements (e.g., approximately 70 dB at 10 mm), not from the later CER or MUSHRA outcome, so the evaluation metrics are not forced by construction. Self-citations to the authors' prior WhisperMask and SilentMask work appear only as related-work context and do not carry a load-bearing premise or invoke any uniqueness claim. The paper explicitly acknowledges the HATS-to-real-wearer and dynamic-motion generalization question in Section 5.2, which is an external-validity limitation rather than circularity. Overall, the evaluation is self-contained and the reported comparisons stand as independent measurements.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The paper is an empirical system paper rather than a derivation, so the ledger is dominated by domain assumptions about vibration coupling, the fidelity of the HATS setup, and the validity of the evaluation metrics. The only hand-tuned quantity that affects the central result is the sensor mounting configuration chosen from a parameter sweep.

free parameters (1)
  • Sensor mounting configuration = stainless steel clip, inward-facing sensor, 10 mm from left edge
    Selected from a parameter sweep in Section 3.3 to maximize measured signal level on the HATS rig; the final CER evaluation uses this configuration, so any overfitting to the test setup transfers to all reported results.
assumptions (5)
  • domain assumption Mask surface vibrations carry enough speech information for automatic speech recognition
    Section 3.1 asserts voice vibrations transmit through the mask and clip; this is the physical basis of the device and is only demonstrated empirically, not derived.
  • domain assumption Ambient acoustic noise does not excite the mask and clip enough to corrupt the piezo signal
    Section 3.1 states the mechanism predominantly captures sounds that induce mask movement while filtering out ambient noise; the paper tests this in an anechoic chamber with calibrated speakers, but not in general environments.
  • domain assumption The HATS mouth simulator reproduces human speech-induced mask vibration coupling
    All acoustic and CER measurements use a SAMAR4700M HATS (Sections 3.3 and 4.2) rather than human wearers; generalization to real faces and masks is assumed.
  • domain assumption Whisper-Large-V3 CER is a valid metric for comparing speech-capture quality
    Section 4.3 uses Whisper-Large-V3 as the recognizer and CER as the objective metric; this is a reasonable engineering choice but adds ASR-specific error sources.
  • domain assumption Online MUSHRA listening results with participants' own headphones reflect audio quality differences
    Section 4.4 describes an unmonitored web-based MUSHRA test; headphone and listening environment variability are not controlled.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MaskClip: Detachable Clip-on Piezoelectric Sensing of Mask Surface Vibrations for Real-time Noise-Robust Speech Input." pith.science (2026). https://pith.science/paper/WNVNUPTU

@misc{pith2026250502180,
  author       = {Pith},
  title        = {Pith review of: MaskClip: Detachable Clip-on Piezoelectric Sensing of Mask Surface Vibrations for Real-time Noise-Robust Speech Input},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WNVNUPTU}},
  note         = {Machine review of arXiv:2505.02180}
}
read the original abstract

Masks are essential in medical settings and during infectious outbreaks but significantly impair speech communication, especially in environments with background noise. Existing solutions often require substantial computational resources or compromise hygiene and comfort. We propose a novel sensing approach that captures only the wearer's voice by detecting mask surface vibrations using a piezoelectric sensor. Our developed device, MaskClip, employs a stainless steel clip with an optimally positioned piezoelectric sensor to selectively capture speech vibrations while inherently filtering out ambient noise. Evaluation experiments demonstrated superior performance with a low Character Error Rate of 6.1\% in noisy environments compared to conventional microphones. Subjective evaluations by 102 participants also showed high satisfaction scores. This approach shows promise for applications in settings where clear voice communication must be maintained while wearing protective equipment, such as medical facilities, cleanrooms, and industrial environments.

Figures

Figures reproduced from arXiv: 2505.02180 by the authors.

Figure 1
Figure 1. Overview of the MaskClip system architecture and operational principles. The left illustration shows the device [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Comparison of speech recognition performance between MaskClip and a unidirectional microphone. The upper [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Detailed hardware implementation of the MaskClip device. The left image shows the physical device implementation, [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Overview of experimental conditions used for systematic evaluation of MaskClip’s sensor configuration parameters. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: shows the results of testing different sensor configuration parameters for the MaskClip device. The graph compares [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Comprehensive overview of the speech recognition evaluation process. (a) Comparison between software-based and [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Final results of speech recognition evaluation. The graph compares Character Error Rates (CER) under various condi [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Subjective audio quality evaluation results for normal speech (left) and whispered speech (right) across different speech [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Illustration of MaskClip’s potential applications across different environments. Left: A medical professional using [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

57 extracted references · 34 canonical work pages

  1. [1]

    Al-Karawi

    Khamis A. Al-Karawi. 2024. Face mask effects on speaker verification perfor- mance in the presence of noise. Multimedia Tools and Applications 83, 2 (2024), 4811–4824. https://doi.org/10.1007/s11042-023-15824-w

  2. [2]

    B. B. Bauer. 1962. A Century of Microphones. Proceedings of the IRE 50, 5 (1962), 719–729. https://doi.org/10.1109/JRPROC.1962.288106

  3. [3]

    Christopher Beach, Nazmul Karim, and Alexander J. Casson. 2019. A Graphene- Based Sleep Mask for Comfortable Wearable Eye Tracking. In 2019 41st Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). 6693–6696. https://doi.org/10.1109/EMBC.2019.8857198 MaskClip: Detachable Clip-on Piezoelectric Sensing of Mask Surf...

  4. [4]

    Ishan Chatterjee, Maruchi Kim, Vivek Jayaram, Shyamnath Gollakota, Ira Kemel- macher, Shwetak Patel, and Steven M. Seitz. 2022. ClearBuds: Wireless Binaural Earbuds for Learning-Based Speech Enhancement. In Proceedings of the 20th Annual International Conference on Mobile Systems, Applications and Services (Portland, Oregon) (MobiSys ’22). Association for...

  5. [5]

    Han Ding, Yizhan Wang, Hao Li, Cui Zhao, Ge Wang, Wei Xi, and Jizhong Zhao

  6. [7]

    Di Duan, Yongliang Chen, Weitao Xu, and Tianxing Li. 2024. EarSE: Bringing Robust Speech Enhancement to COTS Headphones. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 7, 4, Article 158 (Jan. 2024), 33 pages. https://doi. org/10.1145/3631447

  7. [8]

    Alexandre Défossez, Gabriel Synnaeve, and Yossi Adi. 2020. Real Time Speech Enhancement in the Waveform Domain. In Proc. Interspeech 2020 . 3291–3295. https://doi.org/10.21437/Interspeech.2020-2409

  8. [9]

    Engin Erzin. 2009. Improving Throat Microphone Speech Recognition by Joint Analysis of Throat and Acoustic Microphone Recordings. IEEE Transactions on Audio, Speech, and Language Processing 17, 7 (2009), 1316–1324. https://doi.org/ 10.1109/TASL.2009.2016733

Show all 57 references
  1. [10]

    Masaaki Fukumoto. 2018. SilentVoice: Unnoticeable Voice Input by Ingressive Speech. In Proceedings of the 31st Annual ACM Symposium on User Interface Soft- ware and Technology (Berlin, Germany) (UIST ’18). Association for Computing Ma- chinery, New York, NY, USA, 237–246. http...

  2. [11]

    Gonzalez, Lam A

    Jose A. Gonzalez, Lam A. Cheah, Angel M. Gomez, Phil D. Green, James M. Gilbert, Stephen R. Ell, Roger K. Moore, and Ed Holdsworth. 2017. Direct Speech Reconstruction From Articulatory Sensor Data by Machine Learning. IEEE/ACM Transactions on Audio, Speech, and Language Proces...

  3. [12]

    Zengrong Guo and Rong-Hao Liang. 2023. TexonMask: Facial Expression Recog- nition Using Textile Electrodes on Commodity Facemasks. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Ger- many) (CHI ’23). Association for Computing Machiner...

  4. [13]

    Tatsuya Hirahara, Makoto Otani, Shota Shimizu, Tomoki Toda, Keigo Nakamura, Yoshitaka Nakajima, and Kiyohiro Shikano. 2010. Silent-speech enhancement using body-conducted vocal-tract resonance signals. Speech Communication 52, 4 (2010), 301–313. https://doi.org/10.1016/j.speco...

  5. [14]

    Hirotaka Hiraki, Shusuke Kanazawa, Takahiro Miura, Manabu Yoshida, Masaaki Mochimaru, and Jun Rekimoto. 2023. External noise reduction using Whisper- Mask, a mask-type wearable microphone. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems (...

  6. [15]

    Hirotaka Hiraki, Shusuke Kanazawa, Takahiro Miura, Manabu Yoshida, Masaaki Mochimaru, and Jun Rekimoto. 2024. WhisperMask: a noise suppressive mask- type microphone for whisper speech. In Proceedings of the Augmented Humans International Conference 2024 (Melbourne, VIC, Austra...

  7. [16]

    Hirotaka Hiraki and Jun Rekimoto. 2021. SilentMask: Mask-Type Silent Speech Interface with Measurement of Mouth Movement. In Augmented Humans Confer- ence 2021 (Rovaniemi, Finland) (AHs’21). Association for Computing Machinery, New York, NY, USA, 86–90. https://doi.org/10.1145...

  8. [17]

    Yuchen Hu, Chen Chen, Chao-Han Huck Yang, Ruizhe Li, Chao Zhang, Pin-Yu Chen, and Engsiong Chng. 2024. Large Language Models are Efficient Learners of Noise-Robust Speech Recognition. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austr...

  9. [18]

    Robert Ingalls. 1987. Throat microphone. The Journal of the Acoustical Society of America 81, 3 (03 1987), 809–809. https: //doi.org/10.1121/1.394659 arXiv:https://pubs.aip.org/asa/jasa/article- pdf/81/3/809/12095011/809_1_online.pdf

  10. [19]

    Junki Kawaguchi and Mitsuharu Matsumoto. 2022. Noise Reduction Combining a General Microphone and a Throat Microphone. Sensors 22, 12 (2022). https: //doi.org/10.3390/s2212s4473

  11. [20]

    Prerna Khanna, Tanmay Srivastava, Shijia Pan, Shubham Jain, and Phuc Nguyen

  12. [21]

    Ryoga Kumazaki and Akifumi Inoue. 2020. Development and Evaluation of a Mask-Type Display Transforming the Wearer’s Impression. InProceedings of 31st Australian Conference on Human-Computer-Interaction (Fremantle, WA, Australia) (OzCHI ’19). Association for Computing Machinery...

  13. [22]

    Yusuke Kunimi, Masa Ogata, Hirotaka Hiraki, Motoshi Itagaki, Shusuke Kanazawa, and Masaaki Mochimaru. 2022. E-MASK: A Mask-Shaped Inter- face for Silent Speech Interaction with Flexible Strain Sensors. In Augmented Humans 2022 (Kashiwa, Chiba, Japan) (AHs 2022). Association fo...

  14. [24]

    Siddique Latif, Junaid Qadir, Adnan Qayyum, Muhammad Usama, and Shahzad Younis. 2021. Speech Technology for Healthcare: Opportunities, Challenges, and State of the Art. IEEE Reviews in Biomedical Engineering 14 (2021), 342–356. https://doi.org/10.1109/RBME.2020.3006860

  15. [25]

    Hyein Lee, Yoonji Kim, and Andrea Bianchi. 2020. MAScreen: Augmenting Speech with Visual Cues of Lip Motions, Facial Expressions, and Text Using a Wearable Display. In SIGGRAPH Asia 2020 Emerging Technologies (Virtual Event, Republic of Korea) (SA ’20). Association for Computi...

  16. [26]

    Te-Won Lee. 1998. Independent Component Analysis. Springer US, Boston, MA, 27–66. https://doi.org/10.1007/978-1-4757-2851-4_2

  17. [27]

    Boon Pang Lim. 2010. Computational differences between whispered and non- whispered speech. PhD Thesis UIUC

  18. [28]

    Yi Luo and Nima Mesgarani. 2019. Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation. IEEE/ACM Trans- actions on Audio, Speech, and Language Processing 27, 8 (2019), 1256–1266. https://doi.org/10.1109/TASLP.2019.2915167

  19. [29]

    Andri Mirzal. 2017. NMF versus ICA for blind source separation.Advances in Data Analysis and Classification 11, 1 (2017), 25–48. https://doi.org/10.1007/s11634- 014-0192-4

  20. [30]

    Nakajima, H

    Y. Nakajima, H. Kashioka, K. Shikano, and N. Campbell. 2003. Non-audible murmur recognition input interface using stethoscopic microphone attached to the skin. In 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP ’03). ,...

  21. [31]

    Hye Yeon Nam, Iyleah Hernandez, and Brendan Harmon. 2020. Unmasked. In Adjunct Publication of the 33rd Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’20 Adjunct). Association for Computing Machinery, New York, NY, USA, 111–113. https...

  22. [32]

    Muhammed Zahid Ozturk, Chenshu Wu, Beibei Wang, Min Wu, and K. J. Ray Liu. 2023. RadioSES: mmWave-Based Audioradio Speech Enhancement and Separation System. IEEE/ACM Trans. Audio, Speech and Lang. Proc. 31 (March 2023), 1333–1347. https://doi.org/10.1109/TASLP.2023.3250846

  23. [33]

    Paliwal and A

    K. Paliwal and A. Basu. 1987. A speech enhancement method based on Kalman filtering. In ICASSP ’87. IEEE International Conference on Acoustics, Speech, and Signal Processing, Vol. 12. 177–180. https://doi.org/10.1109/ICASSP.1987.1169756

  24. [34]

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In Proceedings of the 40th International Conference on Machine Learning (Honolulu, Hawaii, USA) (ICML’23). JMLR.org, Ar...

  25. [35]

    Saeidi, T

    R. Saeidi, T. Niemi, H. Karppelin, J. Pohjalainen, T. Kinnunen, and P. Alku. 2015. Speaker recognition for speech under face cover. In Proc. Interspeech 2015 . 1012–

  26. [36]

    Mose Sakashita, Keisuke Kawahara, Amy Koike, Kenta Suzuki, Ippei Suzuki, and Yoichi Ochiai. 2016. Yadori: Mask-Type User Interface for Manipulation of Puppets. In ACM SIGGRAPH 2016 Emerging Technologies (Anaheim, California) (SIGGRAPH ’16). Association for Computing Machinery,...

  27. [37]

    Philipp Schilk, Niccolò Polvani, Andrea Ronco, Milos Cernak, and Michele Magno. 2023. In-Ear-Voice: Towards Milli-Watt Audio Enhancement With Bone- Conduction Microphones for In-Ear Sensing Platforms. In Proceedings of the 8th ACM/IEEE Conference on Internet of Things Design a...

  28. [38]

    Antonia Schulte, Rodrigo Suarez-Ibarrola, Daniel Wegen, Philippe-Fabian Pohlmann, Elina Petersen, and Arkadiusz Miernik. 2020. Automatic speech recognition in the operating room – An essential contemporary tool or a re- dundant gadget? A survey evaluation among physicians in f...

  29. [39]

    Shota Shimizu, Makoto Otani, and Tatsuya Hirahara. 2009. Frequency charac- teristics of several non-audible murmur (NAM) microphones. Acoustical Science and Technology 30, 2 (2009), 139–142. https://doi.org/10.1250/ast.30.139

  30. [40]

    Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, and Jianyuan Zhong. 2021. Attention Is All You Need In Speech Separation. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 21–25. https://doi.org/10.1109/ICASSP...

  31. [41]

    Ke Sun and Xinyu Zhang. 2021. UltraSE: single-channel speech enhancement using ultrasound. In Proceedings of the 27th Annual International Conference on Mobile Computing and Networking (New Orleans, Louisiana) (MobiCom ’21). Association for Computing Machinery, New York, NY, U...

  32. [42]

    Yutaro Suzuki, Kodai Sekimori, Yuki Yamato, Yusuke Yamasaki, Buntarou Shizuki, and Shin Takahashi. 2020. A Mouth Gesture Interface Featuring a Mutual- Capacitance Sensor Embedded in a Surgical Mask. In Human-Computer Interac- tion. Multimodal and Natural Interaction , Masaaki ...

  33. [43]

    Thibodeau, Rachel B

    Linda M. Thibodeau, Rachel B. Thibodeau-Nielsen, Chi Mai Q. Tran, and Ryan T. S. Jacob. 2021. Communicating During COVID-19: The Effect of Transparent Masks for Speech Recognition in Noise. Ear and Hearing 42, 4 (Jul 2021), 772–781. https://doi.org/10.1097/AUD.0000000000001065

  34. [44]

    Vishal Varun Tipparaju, Di Wang, Jingjing Yu, Fang Chen, Francis Tsow, Erica Forzani, Nongjian Tao, and Xiaojun Xian. 2020. Respiration pattern recognition by wearable mask device. Biosensors and Bioelectronics 169 (2020), 112590. https: //doi.org/10.1016/j.bios.2020.112590

  35. [45]

    Vishal Varun Tipparaju, Xiaojun Xian, Devon Bridgeman, Di Wang, Francis Tsow, Erica Forzani, and Nongjian Tao. 2020. Reliable Breathing Tracking With Wearable Mask Device. IEEE Sensors Journal 20, 10 (2020), 5510–5518. https: //doi.org/10.1109/JSEN.2020.2969635

  36. [46]

    Toscano and Caroline M

    Joseph C. Toscano and Caroline M. Toscano. 2021. Effects of face masks on speech recognition in multi-talker babble noise. PLOS ONE 16, 2 (Feb 2021), e0246842. https://doi.org/10.1371/journal.pone.0246842

  37. [47]

    Bandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Look Once to Hear: Target Speech Hearing with Noisy Examples. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association f...

  38. [48]

    Gopakumar

    Amritha Vijayan, Bipil Mary Mathai, Karthik Valsalan, Riyanka Raji Johnson, Lani Rachel Mathew, and K. Gopakumar. 2017. Throat microphone speech recog- nition using mfcc. In 2017 International Conference on Networks & Advances in Computational Technologies (NetACT). 392–395. h...

  39. [49]

    Mou Wang, Junqi Chen, Xiao-Lei Zhang, and Susanto Rahardja. 2022. End-to- End Multi-Modal Speech Recognition on an Air and Bone Conducted Speech Corpus. IEEE/ACM Trans. Audio, Speech and Lang. Proc. 31 (Nov. 2022), 513–524. https://doi.org/10.1109/TASLP.2022.3224305

  40. [50]

    Gordon Wichern, Joe Antognini, Michael Flynn, Licheng Richard Zhu, Emmett McQuinn, Dwight Crow, Ethan Manilow, and Jonathan Le Roux. 2019. WHAM!: Extending Speech Separation to Noisy Environments. In Proc. Interspeech. 1368–

  41. [51]

    Takumi Yamamoto, Katsutoshi Masai, Anusha Withana, and Yuta Sugiura. 2023. Masktrap: Designing and Identifying Gestures to Transform Mask Strap into an Input Interface. In Proceedings of the 28th International Conference on Intelligent User Interfaces (Sydney, NSW, Australia)(...

  42. [52]

    NAKAJIMA Yoshitaka, KASHIOKA Hideki, CAMPBELL Nick, and SHIKANO Kiyohiro. 2005. Non-Audible Murmur (NAM) Recognition.IEICE TRANSACTIONS on Information and Systems E89-D, 1 (2005)

  43. [53]

    Petr Zelinka and Milan Sigmund. 2010. Towards reliable speech recognition in operating room noise environment. In 20th International Conference Radioelek- tronika 2010. 1–4. https://doi.org/10.1109/RADIOELEK.2010.5478597

  44. [54]

    Jun Zhang, Jingyue Wu, Yiyi Qiu, Aiguo Song, Weifeng Li, Xin Li, and Yecheng Liu. 2023. Intelligent speech technologies for transcription, disease diagnosis, and medical equipment interactive control in smart hospitals: A review.Computers in Biology and Medicine 153 (2023), 10...

  45. [55]

    Qian Zhang, Dong Wang, Run Zhao, Yinggang Yu, and Junjie Shen. 2021. Sensing to Hear: Speech Enhancement for Mobile Devices Using Acoustic Signals. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 5, 3, Article 137 (Sept. 2021), 30 pages

  46. [1016]

    https://doi.org/10.21437/Interspeech.2015-275

  47. [1372]

    https://doi.org/10.21437/Interspeech.2019-2821

  48. [2021]

    In Proceedings of the 22nd International Workshop on Mobile Computing Systems and Applications (Virtual, United Kingdom) (HotMobile ’21)

    JawSense: Recognizing Unvoiced Sound Using a Low-Cost Ear-Worn System. In Proceedings of the 22nd International Workshop on Mobile Computing Systems and Applications (Virtual, United Kingdom) (HotMobile ’21). Association for Computing Machinery, New York, NY, USA, 44–49. https...

  49. [2022]

    UltraSpeech: Speech Enhancement by Interaction between Ultrasound and Speech. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 6, 3, Article 111 (Sept. 2022), 25 pages. https://doi.org/10.1145/3550303

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.