Pith. sign in

REVIEW 3 major objections 5 minor 70 references

Aurchestra claims the first real-time system that lets users independently adjust the volume of up to five overlapping sound classes on resource-constrained hearables.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 19:54 UTC pith:7D75LW4K

load-bearing objection A genuinely useful hearable system with one overclaimed headline: 5-target robustness is only demonstrated synthetically, while in-the-wild tests stop at 2 targets. the 3 major comments →

arxiv 2603.00395 v3 pith:7D75LW4K submitted 2026-02-28 cs.SD cs.LGeess.AS

Fine-grained Soundscape Control for Augmented Hearing

classification cs.SD cs.LGeess.AS
keywords augmented hearingsoundscape controlmulti-output sound extractionreal-time audio processinghearablessound event detectiondeep learning for audio
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper sets out to prove that a hearable can treat a busy acoustic scene as a set of adjustable tracks, not a single mixture. Aurchestra combines a multi-output extraction network that produces one separated audio stream per user-selected class (up to five) with a dynamic interface that shows only the sound classes currently present. The authors show the network runs in real time on 6 ms streaming chunks, beats a prior single-target system in signal quality while using less than half the parameters, and improves listening and interaction in user studies. If correct, this moves hearables beyond binary noise cancellation toward programmable, personalized soundscapes.

Core claim

At the paper's core is the claim that a small neural network can output multiple independent target-sound streams in real time, conditioned on a user's selection. The extraction network processes causal STFT chunks through dual-path time-frequency modeling blocks, with FiLM layers injecting a multi-hot encoding of chosen classes; a dynamic mapping assigns each selected class to one of O=5 output streams in alphabetical order, so the network need not compute all 20 possible classes. On a test of 20 sound classes, the model achieves 11.99 dB SNRi (improvement over the mixture) at 0.5M parameters, compared with 7.29 dB for the prior single-target approach at 1.2M parameters, and it maintains st

What carries the argument

The key object is the multi-output extraction network: a causal STFT-domain dual-path model with B blocks that alternately model frequency and time, a temporal stage of unidirectional LSTMs (or MLP-Mixer blocks on one platform), and FiLM conditioning that injects the user's multi-hot class selection into each block. The crucial design is the dynamic output mapping: the network emits only O=5 streams and learns to assign each selected class to the stream matching its alphabetical order within the active set, avoiding the need for 20 fixed output heads or permutation-invariant training. This keeps the model around 0.5M parameters and preserves separation for up to five targets. A dual-window S

Load-bearing premise

That synthetic mixtures of isolated sound events with spatial filtering are representative of real dense acoustic scenes; the in-the-wild evaluation only contained one or two target sounds, so the 'up to five overlapping' headline is currently verified only in simulation.

What would settle it

Run Aurchestra on a real street or construction site where five or more target classes (speech, traffic, birds, alarm, siren) genuinely overlap, and compare separation metrics or listening-test scores against the synthetic test; if performance drops substantially, the five-target claim rests on the synthetic distribution rather than real-world acoustics.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • With the multi-stream output, hearables can apply per-class effects (volume, EQ, pitch) in real time, opening the door to hearing aids that emphasize alarms and de-emphasize traffic.
  • The system runs on compact boards typically used in hearing aids and earbuds, suggesting this level of control is feasible for battery-powered wearables.
  • The dynamic interface cuts sound-selection time by 67.9% in the study, indicating context-aware menus substantially reduce interaction overhead for users.
  • Because the extraction and detection models are trained on a 20-class taxonomy, the system could be extended to user-configured sound categories with additional training data.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The five-target capability is demonstrated on synthetic mixtures and real scenes with one or two targets; a dense real-world scene with five concurrent classes is the natural next test, and the outcome could either confirm or narrow the claim.
  • The alphabetical-order output mapping is a pragmatic way to avoid permutation training, but it may become a bottleneck if classes are added or if users choose different subsets across time; a learned assignment could generalize better.
  • The system's class-level streams are a foundation for speaker-level selection: combining scene-level separation with talker identification could create a hearing aid that isolates both the sound class and the specific person.
  • The 20-class fixed taxonomy is a proof-of-concept; open-set or hierarchical classification would be needed for everyday use where novel sounds appear.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces Aurchestra, a hearable system for fine-grained soundscape control. It proposes a multi-output sound extraction network conditioned on a multi-hot class encoding, with output streams assigned alphabetically among selected classes; hardware-tailored variants for Orange Pi, Raspberry Pi, and GreenWaves GAP9; and a dynamic interface driven by a fine-tuned AST-based sound event detector. Evaluations include synthetic Scaper/CIPIC benchmarks, hardware latency/power measurements, in-the-wild listening tests, and interface timing/usability studies. The abstract claims this is the first system to provide fine-grained, real-time soundscape control on resource-constrained hearables, with robust performance for up to 5 overlapping target sounds.

Significance. If the results hold, Aurchestra would be a meaningful advance over single-target semantic hearing: it produces separate per-class streams with independent gains, runs in real time on embedded platforms, and reduces interface overhead through automatic class surfacing. Strengths include the dynamic output-to-class mapping, hardware-specific architecture variants with measured latency/power, the comparison against Waveformer, and a real-world pilot with user studies. However, the central five-target capability is validated only on synthetic data, and several key behavioral claims lack inferential statistics. The paper provides an audio demo URL, though no code release statement is included.

major comments (3)
  1. [Abstract; §3.3.1; §4.3; Table 2] The headline claim of 'robust performance for upto 5 overlapping target sounds' is supported only by Table 2, which evaluates on synthetic Scaper/CIPIC mixtures generated on-the-fly with the same pipeline used for training (§3.3.1: 1–5 target classes, 1–2 interfering classes, 5–15 dB target SNR, 0–10 dB interferer SNR, fixed 3–5 s event durations, CIPIC HRTFs). The in-the-wild evaluation in §4.3 states that 'Each recording contained 1-2 target sounds,' and reports only subjective MOS scores; no reference-based objective metrics are given for any real recording. The Limitations section (§5) also does not flag this gap. This is a correctness risk for the paper's most distinctive claim, not merely a missing ablation. Please either provide real-world recordings with 3–5 simultaneous target classes and objective evaluation, or explicitly restrict the five-target claim to synthetic conditions
  2. [§4.3; §4.4.1] The key behavioral claims—background-noise suppression +1.54 points, overall listening experience +0.95 points, target-clarity parity, and a 67.9% reduction in selection time—are reported as point estimates with no confidence intervals, effect sizes, or significance tests. The n=17 listening study and n=7 interface study are small, and it is unclear whether ratings are averaged per participant or per clip, or whether participants were treated as random effects. Since 'substantial improvements' and 'significant reduction' appear in the abstract and Section 1, the absence of inferential statistics is load-bearing. Please report paired tests with participant-level analysis, or soften the claims to descriptive observations.
  3. [§3.3.1 vs §3.4.2] There is an internal inconsistency in the training distribution for the SED model. §3.3.1 states that training mixtures contained 1–5 target classes, while §3.4.2 states that the AST fine-tuning data used 1–3 target classes. Figure 4 and Table 4 evaluate the SED model on up to 5 simultaneous sources. The reported 93.2% accuracy at 5 sources therefore cannot be attributed unambiguously to the described training procedure. Please clarify which target-class counts were used for SED fine-tuning and whether the 5-source test condition was seen in training.
minor comments (5)
  1. [Abstract; §1; §3.1.3] Typos and wording: 'upto' should be 'up to' (Abstract and §1); 'has two key benefits to having a fixed' in §3.1.3 is ungrammatical.
  2. [Footnote 1] 'with PC chair approval' is unclear; presumably 'IRB approval' or 'per chair approval' was intended.
  3. [§4.1.3; Table 2] The statement that performance 'declines with 4 or more' targets is not consistently reflected in the 5-output rows (e.g., SI-SNRi is 9.87 at 4 targets and 9.84 at 5 targets). Specify which output configuration is being discussed and support the trend with paired tests or rephrase.
  4. [Figures 7 and 8] No measure of variability is shown for the MOS ratings or per-class clarity scores. Please add error bars, per-participant scatter, or confidence intervals.
  5. [Data availability] No code or dataset release statement is included. The audio demo URL is helpful, but a clear availability statement would aid reproducibility.

Circularity Check

0 steps flagged

No significant circularity: central results are empirical benchmarks on held-out synthetic/real data; self-citations are implementation references, not load-bearing.

full rationale

Walking the paper's derivation chain, the central capabilities are established empirically rather than by construction. The multi-output extraction network is trained on on-the-fly Scaper/CIPIC binaural mixtures and evaluated on held-out synthetic mixtures from the same generation procedure (Tables 1 and 2); no fitted parameter is later renamed as a prediction. The mapping from multi-hot target selection to output streams is a deterministic alphabetical ordering given in §3.1.3, and the network's ability to realize that mapping is measured on held-out data, so the mapping is not defined in terms of the result. The SED/dynamic-interface component is fine-tuned on synthetic mixtures and benchmarked against YAMNet and pretrained AST on a held-out set with reported accuracy/precision/recall/F1, so its improvement is an empirical result rather than an imported self-citation. The 11.99 dB vs 7.29 dB SNRi comparison in Table 1 is against the external Waveformer baseline with reported parameter counts, supporting the efficiency claim independently of the authors' prior work. Self-citations to Semantic Hearing [55], NeuralAids [26], and TF-MLPNet [25] appear as baseline systems, low-latency STFT implementation details, and architectural components; none is invoked as a uniqueness theorem or as the sole evidence for the headline 'up to 5 overlapping target sounds' capability. The gap between the synthetic 5-target evaluation and the in-the-wild 1-2 target MOS study is a real external-validity concern, but it is a correctness risk, not circularity: the simulation is not identical to the claimed result by construction, and no equation reduces the claim to its training input. No specific circular step meeting the quoted-reduction standard was found.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 0 invented entities

No new physical entities are postulated. The central claims rest on training-domain realism and hand-chosen system parameters; the ledger reflects that. The class taxonomy, output-stream count, window length, and mixture SNR ranges are all choices that influence every reported number.

free parameters (6)
  • Output stream count O = 5
    Chosen in §3.1.2/§4.1.3 as the balance between flexibility and performance; Table 2 shows large swings with O=1/5/20 (11.99→9.56→3.51 dB SNRi for single-source mixtures).
  • FiLM placement = all TF blocks
    Selected from ablation in Table 3; improves Orange Pi SNRi from 11.76 to 12.26 dB. The paper does not specify a validation split for this selection.
  • Per-platform architecture dimensions = Orange Pi D=32,H=64,B=6; Raspberry Pi D=16,H=64,B=3; NeuralAids D=32,H=32,B=6
    Hand-tuned hyperparameters in §3.2 that determine the accuracy/latency tradeoffs; no automated search is reported.
  • SED analysis window and threshold = 5 s window; F1-maximizing threshold on validation
    §3.4.3/§4.2: window length is a latency/accuracy compromise; the classification threshold is selected on validation and directly affects which classes the interface displays.
  • Synthetic mixture SNR ranges = targets 5–15 dB, interferers 0–10 dB
    §3.3.1: defines the training/test distribution; all reported extraction numbers are conditioned on these ranges.
  • Target class taxonomy = 20 AudioSet classes + 141 interferers
    §3.3.1: manually selected classes; the set defines task difficulty and the dynamic interface's vocabulary.
axioms (6)
  • domain assumption Linear binaural additivity x(t)=Σs_i(t)+n(t) and per-class volume mixing at the output.
    §3.1.1: each class is reduced to a stream and mixed with [0.5,0.5] and v_i; real room reverberation and crosstalk are approximated only by HRTF convolution.
  • domain assumption Scaper-synthesized mixtures with HRTFs adequately represent real-world auditory scenes.
    §3.3.1; this is the weakest premise because the in-the-wild evaluation never exceeds two target sounds.
  • standard math Causal dual-window STFT achieves 10 ms algorithmic latency with near-perfect reconstruction.
    §3.1.4, citing [59] and [26]; needed for the real-time claim and assumes the modified synthesis window does not introduce network-visible artifacts.
  • domain assumption AudioSet pre-trained AST representations transfer to the 20-class fine-tuning task.
    §3.4.2: no comparison to training AST from scratch; the fine-tuning approach assumes the pretrained encoder is a good feature extractor for dense overlapping scenes.
  • domain assumption CIPIC HRTFs generalize across unseen wearers.
    §3.3.1: training uses a fixed 43-subject HRTF set; the system is later worn by new participants, so cross-subject binaural transfer is assumed.
  • domain assumption Separation metrics (SNRi, SI-SNRi) and MOS scales capture the claimed soundscape-control benefit.
    Used throughout §4; there is no objective blind test of whether per-class volume changes match the user's intent, only subjective ratings.

pith-pipeline@v1.3.0-alltime-deepseek · 19887 in / 15798 out tokens · 154311 ms · 2026-08-02T19:54:56.348732+00:00 · methodology

0 comments
read the original abstract

Hearables are becoming ubiquitous, yet their sound controls remain blunt: users can either enable global noise suppression or focus on a single target sound. Real-world acoustic scenes, however, contain many simultaneous sources that users may want to adjust independently. We introduce Aurchestra, the first system to provide fine-grained, real-time soundscape control on resource-constrained hearables. Our system has two key components: (1) a dynamic interface that surfaces only active sound classes and (2) a real-time, on-device multi-output extraction network that generates separate streams for each selected class, achieving robust performance for upto 5 overlapping target sounds, and letting users mix their environment by customizing per-class volumes, much like an audio engineer mixes tracks. We optimize the model architecture for multiple compute-limited platforms and demonstrate real-time performance on 6 ms streaming audio chunks. Across real-world environments in previously unseen indoor and outdoor scenarios, our system enables expressive per-class sound control and achieves substantial improvements in target-class enhancement and interference suppression. Our results show that the world need not be heard as a single, undifferentiated stream: with Aurchestra, the soundscape becomes truly programmable.

Figures

Figures reproduced from arXiv: 2603.00395 by Aseem Gauri, Malek Itani, Seunghyun Oh, Shyamnath Gollakota.

Figure 1
Figure 1. Figure 1: Aurchestra transforms the auditory world into a programmable studio. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Multi-output sound extraction architecture. 3.1.4 Network Architecture. The network in [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: (a) CDF plots of the runtime of the three proposed networks on their respective deployment platforms. (b) Run￾time and power consumption of running the NeuralAids model at different GAP9 clock frequencies. multi-hot encoding into the network. We compare three placement strategies: (1) applying FiLM only to the first block, (2) applying FiLM to all blocks, and (3) applying FiLM to all blocks except the firs… view at source ↗
Figure 4
Figure 4. Figure 4: Sound Event Detection performance comparison on 5-second audio segments. [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Runtime CDF of the fine-tuned AST model across different iPhone platforms. All platforms complete inference faster than the 5-second audio segment duration. and iPhone 12 Pro have median runtimes of 2.9s and 3.3s, re￾spectively. Since all devices process 5s audio segments within 5s, the SED model runs in real time on all tested devices. 4.3 In-The-Wild Evaluation To evaluate our system’s real-world perform… view at source ↗
Figure 6
Figure 6. Figure 6: In-the-wild scenarios. The wearer and sound sources were free to move, and head rotation was uncontrolled [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: User listening study results comparing Au [PITH_FULL_IMAGE:figures/full_fig_p011_7.png] view at source ↗
Figure 9
Figure 9. Figure 9: Response times for sound selection. Our dy￾namic interface reduces selection time by 67.9% compared to the static interface by displaying only detected sounds rather than all 20 categories. Error bars show standard deviation. sounds they heard previously or expect to encounter soon. Users may want to configure their preferences in advance, e.g., choosing to focus on “speech” before entering a meeting room … view at source ↗
Figure 11
Figure 11. Figure 11: User preferences survey results (N=7). (a) Preferred device. (b) Expected response time. (c) Sound selection method. (d) Volume adjustment method. (e) Number of simultaneous sounds. (f) Usage situations. A User Preferences Survey The participants in our user study also participated in a survey to understand their preferences for sound filtering applications. The participants were allowed to pick multiple … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

70 extracted references · 5 canonical work pages

  1. [1]

    Algazi, R.O

    V.R. Algazi, R.O. Duda, D.M. Thompson, and C. Avendano. 2001. The CIPIC HRTF database. 99-102 pages. doi:10.1109/ASPAA.2001.969552 12

  2. [2]

    Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. 2023. AudioLM: A Language Modeling Approach to Audio Generation.IEEE/ACM Trans. Audio, Speech and Lang. Proc.31 (June 2023), 2523–2533. doi:10.1109/ TASLP.2023.3288409

  3. [3]

    Justin Chan, Nada Ali, Ali Najafi, Anna Meehan, Lisa Mancl, Emily Gal- lagher, Randall Bly, and Shyamnath Gollakota. 2022. An off-the-shelf otoacoustic-emission probe for hearing screening via a smartphone. Nature Biomedical Engineering6 (10 2022), 1–11. doi:10.1038/s41551- 022-00947-6

  4. [4]

    Mancl, Emily Gal- lagher, Randall Bly, Shwetak Patel, and Shyamnath Gollakota

    Justin Chan, Antonio Glenn, Malek Itani, Lisa R. Mancl, Emily Gal- lagher, Randall Bly, Shwetak Patel, and Shyamnath Gollakota. 2023. Wireless Earbuds for Low-Cost Hearing Screening. InProceedings of the 21st Annual International Conference on Mobile Systems, Applications and Services(Helsinki, Finland)(MobiSys ’23). Association for Com- puting Machinery,...

  5. [5]

    Ruei-Che Chang, Chia-Sheng Hung, Bing-Yu Chen, Dhruv Jain, and An- hong Guo. 2024. SoundShift: Exploring Sound Manipulations for Acces- sible Mixed-Reality Awareness. InProceedings of the 2024 ACM Design- ing Interactive Systems Conference(Copenhagen, Denmark)(DIS ’24). Association for Computing Machinery, New York, NY, USA, 116–132. doi:10.1145/3643834.3661556

  6. [6]

    Ishan Chatterjee, Maruchi Kim, Vivek Jayaram, Shyamnath Gollakota, Ira Kemelmacher, Shwetak Patel, and Steven M Seitz. 2022. ClearBuds: wireless binaural earbuds for learning-based speech enhancement. In MobiSys

  7. [7]

    Tao Chen, Xiaoran Fan, Yongjie Yang, and Longfei Shangguan. 2023. Towards Remote Auscultation with Commodity Earphones. InProceed- ings of the 20th ACM Conference on Embedded Networked Sensor Systems (Boston, Massachusetts)(SenSys ’22). Association for Computing Ma- chinery, New York, NY, USA, 853–854. doi:10.1145/3560905.3568084

  8. [8]

    Tuochao Chen, Malek Itani, Sefik Eskimez, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Hearable devices with sound bubbles. Nature Electronics(2024)

  9. [9]

    Tuochao Chen, D Shin, Hakan Erdogan, and Sinan Hersek. 2025. Sound- Sculpt: Direction and Semantics Driven Ambisonic Target Sound Ex- traction. InInterspeech 2025. 943–947. doi:10.21437/Interspeech.2025- 1379

  10. [10]

    Tao Chen, Yongjie Yang, Xiaoran Fan, Xiuzhen Guo, Jie Xiong, and Longfei Shangguan. 2024. Exploring the Feasibility of Remote Car- diac Auscultation Using Earphones. InProceedings of the 30th Annual International Conference on Mobile Computing and Networking(Wash- ington D.C., DC, USA)(ACM MobiCom ’24). Association for Computing Machinery, New York, NY, U...

  11. [11]

    K. M. de Paiva Vianna, M. R. Alves Cardoso, and R. M. Rodrigues. 2015. Noise pollution and annoyance: an urban soundscapes study.Noise & Health17, 76 (May–Jun 2015), 125–133. doi:10.4103/1463-1741.155833

  12. [12]

    Marc Delcroix, Jorge Bennasar Vázquez, Tsubasa Ochiai, Keisuke Ki- noshita, Yasunori Ohishi, and Shoko Araki. 2022. SoundBeam: Tar- get sound extraction conditioned on sound-class labels and enroll- ment clues for increased performance and continuous learning.arXiv preprint arXiv:2204.03895(2022)

  13. [13]

    Xiaoran Fan, Longfei Shangguan, Siddharth Rupavatharam, Yanyong Zhang, Jie Xiong, Yunfei Ma, and Richard Howard. 2021. HeadFi: bringing intelligence to all headphones. InProceedings of the 27th Annual International Conference on Mobile Computing and Networking (New Orleans, Louisiana)(MobiCom ’21). Association for Computing Machinery, New York, NY, USA, 1...

  14. [14]

    Xiaoran Fan and Trausti Thormundsson. 2023. Design Earable Sensing Systems: Perspectives and Lessons Learned from Industry. InAdjunct Proceedings of the 2023 ACM International Joint Conference on Pervasive and Ubiquitous Computing and the 2023 ACM International Symposium on Wearable Computers

  15. [15]

    Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra. 2022. FSD50K: An Open Dataset of Human-Labeled Sound Events. arXiv:2010.00475 [cs.SD]

  16. [16]

    Ruohan Gao and Kristen Grauman. 2019. Co-separating sounds of visual objects. InProceedings of the IEEE/CVF International Conference on Computer Vision

  17. [17]

    Gemmeke, Daniel P

    Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter

  18. [18]

    Beat Gfeller, Dominik Roblek, and Marco Tagliasacchi. 2021. One-shot conditional audio filtering of arbitrary sounds. InICASSP. IEEE

  19. [19]

    Arushi Goel, Sreyan Ghosh, Jaehyeon Kim, Sonal Kumar, Zhifeng Kong, Sang gil Lee, Chao-Han Huck Yang, Ramani Duraiswami, Dinesh Manocha, Rafael Valle, and Bryan Catanzaro. 2025. Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models. arXiv:2507.08128 [cs.SD] https://arxiv.org/abs/2507.08128

  20. [20]

    Yuan Gong, Yu-An Chung, and James Glass. 2021. AST: Audio Spec- trogram Transformer. InInterspeech 2021. 571–575. doi:10.21437/ Interspeech.2021-698

  21. [21]

    Jiarui Hai, Helin Wang, Dongchao Yang, Karan Thakkar, Najim Dehak, and Mounya Elhilali. 2024. DPM-TSE: A Diffusion Probabilistic Model for Target Sound Extraction. InICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1196–

  22. [22]

    Justin han, Sharat Raju, Rajalakshmi Nandakumar, Randall Bly, and Shyamnath Gollakota. 2019. Detecting middle ear fluid using smart- phones.Science Translational Medicine11 (05 2019), eaav1102. doi:10. 1126/scitranslmed.aav1102

  23. [23]

    Guilin Hu, Malek Itani, Tuochao Chen, and Shyamnath Gollakota. 2025. Proactive Hearing Assistants that Isolate Egocentric Conversations. InProceedings of the 2025 Conference on Empirical Methods in Natu- ral Language Processing. Association for Computational Linguistics, Suzhou, China, 25377–25394. doi:10.18653/v1/2025.emnlp-main.1289

  24. [24]

    Jeremy Zhengqi Huang, Jaylin Herskovitz, Liang-Yuan Wu, Cecily Morrison, and Dhruv Jain. 2025. Weaving Sound Information to Sup- port Real-Time Sensemaking of Auditory Environments: Co-Designing with a DHH User(CHI ’25)

  25. [25]

    Malek Itani, Tuochao Chen, and Shyamnath Gollakota. 2025. TF- MLPNet: Tiny Real-Time Neural Speech Separation. InClarity Chal- lenge, InterSpeech

  26. [26]

    Malek Itani, Tuochao Chen, Arun Raghavan, Gavriel Kohlberg, and Shyamnath Gollakota. 2025. Wireless Hearables With Programmable Speech AI Accelerators(ACM MOBICOM ’25). ACM

  27. [27]

    Malek Itani, Ashton Graves, Sefik Emre Eskimez, and Shyamnath Gollakota. 2025. Neural Speech Extraction with Human Feedback. In Interspeech 2025. 4998–5002. doi:10.21437/Interspeech.2025-214

  28. [28]

    Froehlich

    Dhruv Jain, Kelly Mack, Akli Amrous, Matt Wright, Steven Goodman, Leah Findlater, and Jon E. Froehlich. 2020. HomeSound: An Iterative Field Deployment of an In-Home Sound Awareness System for Deaf or Hard of Hearing Users. InACM CHI

  29. [29]

    Dhruv Jain, Hung Ngo, Pratyush Patel, Steven Goodman, Leah Find- later, and Jon Froehlich. 2020. SoundWatch: Exploring Smartwatch- Based Deep Learning Approaches to Support Sound Awareness for Deaf and Hard of Hearing Users. InACM SIGACCESS ASSETS

  30. [30]

    Yincheng Jin, Yang Gao, Xiaotao Guo, Jun Wen, Zhengxiong Li, and Zhanpeng Jin. 2022. EarHealth: an earphone-based acoustic otoscope for detection of multiple ear diseases in daily life(MobiSys ’22). ACM. 13

  31. [31]

    Fahim Kawsar, Chulhong Min, Akhil Mathur, and Alessandro Monta- nari. 2018. Earables for Personal-Scale Behavior Analytics.IEEE Perva- sive Computing17, 3 (2018), 83–89. doi:10.1109/MPRV.2018.03367740

  32. [32]

    Kevin Kilgour, Beat Gfeller, Qingqing Huang, Aren Jansen, Scott Wis- dom, and Marco Tagliasacchi. 2022. Text-Driven Separation of Arbi- trary Sounds.arXiv preprint arXiv:2204.05738(2022)

  33. [33]

    Gierad Laput, Karan Ahuja, Mayank Goel, and Chris Harrison. 2018. Ubicoustics: Plug-and-Play Acoustic Activity Recognition. InACM UIST

  34. [34]

    Xubo Liu, Haohe Liu, Qiuqiang Kong, Xinhao Mei, Jinzheng Zhao, Qiushi Huang, Mark D Plumbley, and Wenwu Wang. 2022. Separate What You Describe: Language-Queried Audio Source Separation.arXiv preprint arXiv:2203.15147(2022)

  35. [35]

    Lane, Tanzeem Choudhury, and An- drew T

    Hong Lu, Wei Pan, Nicholas D. Lane, Tanzeem Choudhury, and An- drew T. Campbell. 2009. SoundSense: Scalable Sound Sensing for People-Centric Applications on Mobile Phones. InACM MobiSys

  36. [36]

    Vimal Mollyn, Karan Ahuja, Dhruv Verma, Chris Harrison, and Mayank Goel. 2022. SAMoSA: Sensing Activities with Motion and Subsampled Audio.IMWUT(2022)

  37. [37]

    Vimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison, and Karan Ahuja. 2023. IMUPoser: Full-Body Pose Estimation Using IMUs in Phones, Watches, and Earbuds. InCHI(Hamburg, Germany)(CHI ’23). ACM

  38. [38]

    Alessandro Montanari, Ashok Thangarajan, Khaldoon Al-Naimi, An- drea Ferlini, Yang Liu, Ananta Narayanan Balaji, and Fahim Kawsar

  39. [39]

    2020.Noise files for the DISCO dataset

    Furnon Nicolas. 2020.Noise files for the DISCO dataset. https://github. com/nfurnon/disco

  40. [40]

    Tsubasa Ochiai, Marc Delcroix, Yuma Koizumi, Hiroaki Ito, Keisuke Kinoshita, and Shoko Araki. 2020. Listen to What You Want: Neural Network-based Universal Sound Selector.arXiv e- prints, Article arXiv:2006.05712 (2020), arXiv:2006.05712 pages. arXiv:2006.05712 [eess.AS]

  41. [41]

    Yuki Okamoto, Shota Horiguchi, Masaaki Yamamoto, Keisuke Imoto, and Yohei Kawaguchi. 2022. Environmental Sound Extraction Using Onomatopoeic Words. InICASSP. IEEE

  42. [42]

    Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron Courville. 2018. FiLM: visual reasoning with a general con- ditioning layer. InProceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Arti- ficial Intelligence Conference and Eighth AAAI Symposium on Educa- tional Advances i...

  43. [43]

    Darius Petermann and Minje Kim. 2022. Spain-Net: Spatially-Informed Stereophonic Music Source Separation. InICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 106–110. doi:10.1109/ICASSP43922.2022.9746277

  44. [44]

    Karol J. Piczak. 2015. ESC: Dataset for Environmental Sound Classifi- cation. InACM Multimedia

  45. [45]

    Manoj Plakal and Daniel P. W. Ellis. 2020. YAMNet: A Pre-trained Audio Event Classifier. https://github.com/tensorflow/models/tree/ master/research/audioset/yamnet. TensorFlow Model Garden

  46. [46]

    Jay Prakash, Zhijian Yang, Yu-Lin Wei, Haitham Hassanieh, and Romit Roy Choudhury. 2020. EarSense: earphones as a teeth activity sensor. InProceedings of the 26th Annual International Conference on Mobile Computing and Networking(London, United Kingdom)(Mobi- Com ’20). Association for Computing Machinery, New York, NY, USA, Article 40, 13 pages. doi:10.11...

  47. [47]

    Adam Pullin, Jake Stuchbury-Wass, Mathias Ciliberto, Kayla-Jade Butkow, Philipp Lepold, Tobias Röddiger, and Cecilia Mascolo. 2025. Ear-ECG Denoising Using Heart Sounds and the Extended Kalman Filter. InIEEE-EMBS International Conference on Body Sensor Networks

  48. [48]

    2017.MUSDB18 - a corpus for music separation

    Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner. 2017.MUSDB18 - a corpus for music separation

  49. [49]

    Tobias Röddiger, Tobias King, Dylan Ray Roodt, Christopher Clarke, and Michael Beigl. 2023. OpenEarable: Open Hardware Earable Sensing Platform(UbiComp/ISWC ’22 Adjunct)

  50. [50]

    Justin Salamon, Duncan MacConnell, Mark Cartwright, Peter Li, and Juan Pablo Bello. 2017. Scaper: A library for soundscape synthesis and augmentation. InW ASPAA. doi:10.1109/WASPAA.2017.8170052

  51. [51]

    Christian J Steinmetz and Joshua D Reiss. 2020. auraloss: Audio focused loss functions in PyTorch. InDigital music research network one-day workshop (DMRN+ 15)

  52. [52]

    Jake Stuchbury-Wass, Andrea Ferlini, and Cecilia Mascolo. 2023. Mul- timodal Attention Networks for Human Activity Recognition From Earable Devices(UbiComp/ISWC ’22 Adjunct). ACM

  53. [53]

    Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiao- hua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. 2021. MLP-Mixer: An all-MLP Architecture for Vision. InNeurips

  54. [54]

    Bandhav Veluri, Justin Chan, Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2023. Real-Time Target Sound Extraction. InICASSP

  55. [55]

    Bandhav Veluri, Malek Itani, Justin Chan, Takuya Yoshioka, and Shyam- nath Gollakota. 2023. Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables. InACM UIST

  56. [56]

    Bandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Look Once to Hear: Target Speech Hearing with Noisy Examples. InACM CHI

  57. [57]

    Keigo Wakayama, Tomoko Kawase, Takafumi Moriya, Marc Delcroix, Hiroshi Sato, Tsubasa Ochiai, Masahiro Yasuda, and Shoko Araki. 2025. Real-time TSE demonstration via SoundBeam with KD. InInterspeech

  58. [58]

    Helin Wang, Jiarui Hai, Yen-Ju Lu, Karan Thakkar, Mounya Elhilali, and Najim Dehak. 2025. SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer. InICASSP. 1–5

  59. [59]

    Zhong-Qiu Wang, Gordon Wichern, Shinji Watanabe, and Jonathan Le Roux. 2022. STFT-domain neural speech enhancement with very low algorithmic latency.Trans. on Audio, Speech, and Language Processing (2022)

  60. [60]

    Xudong Xu, Bo Dai, and Dahua Lin. 2019. Recursive visual sound separation using minus-plus net. InIEEE/CVF ICCV

  61. [61]

    Dongchao Yang, Jinchuan Tian, Xuejiao Tan, Rongjie Huang, Songxi- ang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian, Xixin Wu, Zhou Zhao, and Helen Meng. 2023. UniAudio: An Audio Founda- tion Model Toward Universal Audio Generation.ArXivabs/2310.00704 (2023). https://api.semanticscholar.org/CorpusID:263334347

  62. [62]

    Bryan, and Minje Kim

    Haici Yang, Shivani Firodiya, Nicholas J. Bryan, and Minje Kim. 2022. Don’t Separate, Learn To Remix: End-To-End Neural Remixing With Joint Optimization. InICASSP. 116–120. doi:10.1109/ICASSP43922. 2022.9746077

  63. [63]

    Qiang Yang, Yang Liu, Jake Stuchbury-Wass, Mathias Ciliberto, Tobias Röddiger, Kayla-Jade Butkow, Adam Pullin, Emeli Panariti, Dong Ma, and Cecilia Mascolo. 2025. HearForce: Force Estimation for Manual Toothbrushing with Earables. (2025). doi:10.17863/CAM.122079

  64. [64]

    Koji Yatani and Khai N. Truong. 2012. BodyScope: A Wearable Acoustic Sensor for Activity Recognition. InUbiComp

  65. [65]

    Plumbley, and Wenwu Wang

    Yi Yuan, Xubo Liu, Haohe Liu, Mark D. Plumbley, and Wenwu Wang

  66. [70]

    InICASSP

    FlowSep: Language-Queried Sound Separation with Rectified Flow Matching. InICASSP. 1–5. 14 Figure 11: User preferences survey results (N=7).(a) Preferred device. (b) Expected response time. (c) Sound selection method. (d) Volume adjustment method. (e) Number of simultaneous sounds. (f) Usage situations. A User Preferences Survey The participants in our us...

  67. [1200]

    doi:10.1109/ICASSP48485.2024.10447219

  68. [2017]

    InIEEE ICASSP

    Audio Set: An ontology and human-labeled dataset for audio events. InIEEE ICASSP

  69. [2024]

    arXiv:2410.04775 [cs.ET] https://arxiv.org/abs/2410.04775

    OmniBuds: A Sensory Earable Platform for Advanced Bio- Sensing and On-Device Machine Learning. arXiv:2410.04775 [cs.ET] https://arxiv.org/abs/2410.04775

  70. [2025]

    https://openreview.net/forum?id=eQE5YiQexy