REVIEW 3 major objections 5 minor 70 references
Aurchestra claims the first real-time system that lets users independently adjust the volume of up to five overlapping sound classes on resource-constrained hearables.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 19:54 UTC pith:7D75LW4K
load-bearing objection A genuinely useful hearable system with one overclaimed headline: 5-target robustness is only demonstrated synthetically, while in-the-wild tests stop at 2 targets. the 3 major comments →
Fine-grained Soundscape Control for Augmented Hearing
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
At the paper's core is the claim that a small neural network can output multiple independent target-sound streams in real time, conditioned on a user's selection. The extraction network processes causal STFT chunks through dual-path time-frequency modeling blocks, with FiLM layers injecting a multi-hot encoding of chosen classes; a dynamic mapping assigns each selected class to one of O=5 output streams in alphabetical order, so the network need not compute all 20 possible classes. On a test of 20 sound classes, the model achieves 11.99 dB SNRi (improvement over the mixture) at 0.5M parameters, compared with 7.29 dB for the prior single-target approach at 1.2M parameters, and it maintains st
What carries the argument
The key object is the multi-output extraction network: a causal STFT-domain dual-path model with B blocks that alternately model frequency and time, a temporal stage of unidirectional LSTMs (or MLP-Mixer blocks on one platform), and FiLM conditioning that injects the user's multi-hot class selection into each block. The crucial design is the dynamic output mapping: the network emits only O=5 streams and learns to assign each selected class to the stream matching its alphabetical order within the active set, avoiding the need for 20 fixed output heads or permutation-invariant training. This keeps the model around 0.5M parameters and preserves separation for up to five targets. A dual-window S
Load-bearing premise
That synthetic mixtures of isolated sound events with spatial filtering are representative of real dense acoustic scenes; the in-the-wild evaluation only contained one or two target sounds, so the 'up to five overlapping' headline is currently verified only in simulation.
What would settle it
Run Aurchestra on a real street or construction site where five or more target classes (speech, traffic, birds, alarm, siren) genuinely overlap, and compare separation metrics or listening-test scores against the synthetic test; if performance drops substantially, the five-target claim rests on the synthetic distribution rather than real-world acoustics.
If this is right
- With the multi-stream output, hearables can apply per-class effects (volume, EQ, pitch) in real time, opening the door to hearing aids that emphasize alarms and de-emphasize traffic.
- The system runs on compact boards typically used in hearing aids and earbuds, suggesting this level of control is feasible for battery-powered wearables.
- The dynamic interface cuts sound-selection time by 67.9% in the study, indicating context-aware menus substantially reduce interaction overhead for users.
- Because the extraction and detection models are trained on a 20-class taxonomy, the system could be extended to user-configured sound categories with additional training data.
Where Pith is reading between the lines
- The five-target capability is demonstrated on synthetic mixtures and real scenes with one or two targets; a dense real-world scene with five concurrent classes is the natural next test, and the outcome could either confirm or narrow the claim.
- The alphabetical-order output mapping is a pragmatic way to avoid permutation training, but it may become a bottleneck if classes are added or if users choose different subsets across time; a learned assignment could generalize better.
- The system's class-level streams are a foundation for speaker-level selection: combining scene-level separation with talker identification could create a hearing aid that isolates both the sound class and the specific person.
- The 20-class fixed taxonomy is a proof-of-concept; open-set or hierarchical classification would be needed for everyday use where novel sounds appear.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Aurchestra, a hearable system for fine-grained soundscape control. It proposes a multi-output sound extraction network conditioned on a multi-hot class encoding, with output streams assigned alphabetically among selected classes; hardware-tailored variants for Orange Pi, Raspberry Pi, and GreenWaves GAP9; and a dynamic interface driven by a fine-tuned AST-based sound event detector. Evaluations include synthetic Scaper/CIPIC benchmarks, hardware latency/power measurements, in-the-wild listening tests, and interface timing/usability studies. The abstract claims this is the first system to provide fine-grained, real-time soundscape control on resource-constrained hearables, with robust performance for up to 5 overlapping target sounds.
Significance. If the results hold, Aurchestra would be a meaningful advance over single-target semantic hearing: it produces separate per-class streams with independent gains, runs in real time on embedded platforms, and reduces interface overhead through automatic class surfacing. Strengths include the dynamic output-to-class mapping, hardware-specific architecture variants with measured latency/power, the comparison against Waveformer, and a real-world pilot with user studies. However, the central five-target capability is validated only on synthetic data, and several key behavioral claims lack inferential statistics. The paper provides an audio demo URL, though no code release statement is included.
major comments (3)
- [Abstract; §3.3.1; §4.3; Table 2] The headline claim of 'robust performance for upto 5 overlapping target sounds' is supported only by Table 2, which evaluates on synthetic Scaper/CIPIC mixtures generated on-the-fly with the same pipeline used for training (§3.3.1: 1–5 target classes, 1–2 interfering classes, 5–15 dB target SNR, 0–10 dB interferer SNR, fixed 3–5 s event durations, CIPIC HRTFs). The in-the-wild evaluation in §4.3 states that 'Each recording contained 1-2 target sounds,' and reports only subjective MOS scores; no reference-based objective metrics are given for any real recording. The Limitations section (§5) also does not flag this gap. This is a correctness risk for the paper's most distinctive claim, not merely a missing ablation. Please either provide real-world recordings with 3–5 simultaneous target classes and objective evaluation, or explicitly restrict the five-target claim to synthetic conditions
- [§4.3; §4.4.1] The key behavioral claims—background-noise suppression +1.54 points, overall listening experience +0.95 points, target-clarity parity, and a 67.9% reduction in selection time—are reported as point estimates with no confidence intervals, effect sizes, or significance tests. The n=17 listening study and n=7 interface study are small, and it is unclear whether ratings are averaged per participant or per clip, or whether participants were treated as random effects. Since 'substantial improvements' and 'significant reduction' appear in the abstract and Section 1, the absence of inferential statistics is load-bearing. Please report paired tests with participant-level analysis, or soften the claims to descriptive observations.
- [§3.3.1 vs §3.4.2] There is an internal inconsistency in the training distribution for the SED model. §3.3.1 states that training mixtures contained 1–5 target classes, while §3.4.2 states that the AST fine-tuning data used 1–3 target classes. Figure 4 and Table 4 evaluate the SED model on up to 5 simultaneous sources. The reported 93.2% accuracy at 5 sources therefore cannot be attributed unambiguously to the described training procedure. Please clarify which target-class counts were used for SED fine-tuning and whether the 5-source test condition was seen in training.
minor comments (5)
- [Abstract; §1; §3.1.3] Typos and wording: 'upto' should be 'up to' (Abstract and §1); 'has two key benefits to having a fixed' in §3.1.3 is ungrammatical.
- [Footnote 1] 'with PC chair approval' is unclear; presumably 'IRB approval' or 'per chair approval' was intended.
- [§4.1.3; Table 2] The statement that performance 'declines with 4 or more' targets is not consistently reflected in the 5-output rows (e.g., SI-SNRi is 9.87 at 4 targets and 9.84 at 5 targets). Specify which output configuration is being discussed and support the trend with paired tests or rephrase.
- [Figures 7 and 8] No measure of variability is shown for the MOS ratings or per-class clarity scores. Please add error bars, per-participant scatter, or confidence intervals.
- [Data availability] No code or dataset release statement is included. The audio demo URL is helpful, but a clear availability statement would aid reproducibility.
Circularity Check
No significant circularity: central results are empirical benchmarks on held-out synthetic/real data; self-citations are implementation references, not load-bearing.
full rationale
Walking the paper's derivation chain, the central capabilities are established empirically rather than by construction. The multi-output extraction network is trained on on-the-fly Scaper/CIPIC binaural mixtures and evaluated on held-out synthetic mixtures from the same generation procedure (Tables 1 and 2); no fitted parameter is later renamed as a prediction. The mapping from multi-hot target selection to output streams is a deterministic alphabetical ordering given in §3.1.3, and the network's ability to realize that mapping is measured on held-out data, so the mapping is not defined in terms of the result. The SED/dynamic-interface component is fine-tuned on synthetic mixtures and benchmarked against YAMNet and pretrained AST on a held-out set with reported accuracy/precision/recall/F1, so its improvement is an empirical result rather than an imported self-citation. The 11.99 dB vs 7.29 dB SNRi comparison in Table 1 is against the external Waveformer baseline with reported parameter counts, supporting the efficiency claim independently of the authors' prior work. Self-citations to Semantic Hearing [55], NeuralAids [26], and TF-MLPNet [25] appear as baseline systems, low-latency STFT implementation details, and architectural components; none is invoked as a uniqueness theorem or as the sole evidence for the headline 'up to 5 overlapping target sounds' capability. The gap between the synthetic 5-target evaluation and the in-the-wild 1-2 target MOS study is a real external-validity concern, but it is a correctness risk, not circularity: the simulation is not identical to the claimed result by construction, and no equation reduces the claim to its training input. No specific circular step meeting the quoted-reduction standard was found.
Axiom & Free-Parameter Ledger
free parameters (6)
- Output stream count O =
5
- FiLM placement =
all TF blocks
- Per-platform architecture dimensions =
Orange Pi D=32,H=64,B=6; Raspberry Pi D=16,H=64,B=3; NeuralAids D=32,H=32,B=6
- SED analysis window and threshold =
5 s window; F1-maximizing threshold on validation
- Synthetic mixture SNR ranges =
targets 5–15 dB, interferers 0–10 dB
- Target class taxonomy =
20 AudioSet classes + 141 interferers
axioms (6)
- domain assumption Linear binaural additivity x(t)=Σs_i(t)+n(t) and per-class volume mixing at the output.
- domain assumption Scaper-synthesized mixtures with HRTFs adequately represent real-world auditory scenes.
- standard math Causal dual-window STFT achieves 10 ms algorithmic latency with near-perfect reconstruction.
- domain assumption AudioSet pre-trained AST representations transfer to the 20-class fine-tuning task.
- domain assumption CIPIC HRTFs generalize across unseen wearers.
- domain assumption Separation metrics (SNRi, SI-SNRi) and MOS scales capture the claimed soundscape-control benefit.
read the original abstract
Hearables are becoming ubiquitous, yet their sound controls remain blunt: users can either enable global noise suppression or focus on a single target sound. Real-world acoustic scenes, however, contain many simultaneous sources that users may want to adjust independently. We introduce Aurchestra, the first system to provide fine-grained, real-time soundscape control on resource-constrained hearables. Our system has two key components: (1) a dynamic interface that surfaces only active sound classes and (2) a real-time, on-device multi-output extraction network that generates separate streams for each selected class, achieving robust performance for upto 5 overlapping target sounds, and letting users mix their environment by customizing per-class volumes, much like an audio engineer mixes tracks. We optimize the model architecture for multiple compute-limited platforms and demonstrate real-time performance on 6 ms streaming audio chunks. Across real-world environments in previously unseen indoor and outdoor scenarios, our system enables expressive per-class sound control and achieves substantial improvements in target-class enhancement and interference suppression. Our results show that the world need not be heard as a single, undifferentiated stream: with Aurchestra, the soundscape becomes truly programmable.
Figures
Reference graph
Works this paper leans on
-
[1]
V.R. Algazi, R.O. Duda, D.M. Thompson, and C. Avendano. 2001. The CIPIC HRTF database. 99-102 pages. doi:10.1109/ASPAA.2001.969552 12
arXiv 2001
-
[2]
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. 2023. AudioLM: A Language Modeling Approach to Audio Generation.IEEE/ACM Trans. Audio, Speech and Lang. Proc.31 (June 2023), 2523–2533. doi:10.1109/ TASLP.2023.3288409
arXiv 2023
-
[3]
Justin Chan, Nada Ali, Ali Najafi, Anna Meehan, Lisa Mancl, Emily Gal- lagher, Randall Bly, and Shyamnath Gollakota. 2022. An off-the-shelf otoacoustic-emission probe for hearing screening via a smartphone. Nature Biomedical Engineering6 (10 2022), 1–11. doi:10.1038/s41551- 022-00947-6
doi:10.1038/s41551- 2022
-
[4]
Mancl, Emily Gal- lagher, Randall Bly, Shwetak Patel, and Shyamnath Gollakota
Justin Chan, Antonio Glenn, Malek Itani, Lisa R. Mancl, Emily Gal- lagher, Randall Bly, Shwetak Patel, and Shyamnath Gollakota. 2023. Wireless Earbuds for Low-Cost Hearing Screening. InProceedings of the 21st Annual International Conference on Mobile Systems, Applications and Services(Helsinki, Finland)(MobiSys ’23). Association for Com- puting Machinery,...
-
[5]
Ruei-Che Chang, Chia-Sheng Hung, Bing-Yu Chen, Dhruv Jain, and An- hong Guo. 2024. SoundShift: Exploring Sound Manipulations for Acces- sible Mixed-Reality Awareness. InProceedings of the 2024 ACM Design- ing Interactive Systems Conference(Copenhagen, Denmark)(DIS ’24). Association for Computing Machinery, New York, NY, USA, 116–132. doi:10.1145/3643834.3661556
arXiv 2024
-
[6]
Ishan Chatterjee, Maruchi Kim, Vivek Jayaram, Shyamnath Gollakota, Ira Kemelmacher, Shwetak Patel, and Steven M Seitz. 2022. ClearBuds: wireless binaural earbuds for learning-based speech enhancement. In MobiSys
2022
-
[7]
Tao Chen, Xiaoran Fan, Yongjie Yang, and Longfei Shangguan. 2023. Towards Remote Auscultation with Commodity Earphones. InProceed- ings of the 20th ACM Conference on Embedded Networked Sensor Systems (Boston, Massachusetts)(SenSys ’22). Association for Computing Ma- chinery, New York, NY, USA, 853–854. doi:10.1145/3560905.3568084
arXiv 2023
-
[8]
Tuochao Chen, Malek Itani, Sefik Eskimez, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Hearable devices with sound bubbles. Nature Electronics(2024)
2024
-
[9]
Tuochao Chen, D Shin, Hakan Erdogan, and Sinan Hersek. 2025. Sound- Sculpt: Direction and Semantics Driven Ambisonic Target Sound Ex- traction. InInterspeech 2025. 943–947. doi:10.21437/Interspeech.2025- 1379
-
[10]
Tao Chen, Yongjie Yang, Xiaoran Fan, Xiuzhen Guo, Jie Xiong, and Longfei Shangguan. 2024. Exploring the Feasibility of Remote Car- diac Auscultation Using Earphones. InProceedings of the 30th Annual International Conference on Mobile Computing and Networking(Wash- ington D.C., DC, USA)(ACM MobiCom ’24). Association for Computing Machinery, New York, NY, U...
arXiv 2024
-
[11]
K. M. de Paiva Vianna, M. R. Alves Cardoso, and R. M. Rodrigues. 2015. Noise pollution and annoyance: an urban soundscapes study.Noise & Health17, 76 (May–Jun 2015), 125–133. doi:10.4103/1463-1741.155833
arXiv 2015
-
[12]
Marc Delcroix, Jorge Bennasar Vázquez, Tsubasa Ochiai, Keisuke Ki- noshita, Yasunori Ohishi, and Shoko Araki. 2022. SoundBeam: Tar- get sound extraction conditioned on sound-class labels and enroll- ment clues for increased performance and continuous learning.arXiv preprint arXiv:2204.03895(2022)
Pith/arXiv arXiv 2022
-
[13]
Xiaoran Fan, Longfei Shangguan, Siddharth Rupavatharam, Yanyong Zhang, Jie Xiong, Yunfei Ma, and Richard Howard. 2021. HeadFi: bringing intelligence to all headphones. InProceedings of the 27th Annual International Conference on Mobile Computing and Networking (New Orleans, Louisiana)(MobiCom ’21). Association for Computing Machinery, New York, NY, USA, 1...
arXiv 2021
-
[14]
Xiaoran Fan and Trausti Thormundsson. 2023. Design Earable Sensing Systems: Perspectives and Lessons Learned from Industry. InAdjunct Proceedings of the 2023 ACM International Joint Conference on Pervasive and Ubiquitous Computing and the 2023 ACM International Symposium on Wearable Computers
2023
-
[15]
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra. 2022. FSD50K: An Open Dataset of Human-Labeled Sound Events. arXiv:2010.00475 [cs.SD]
Pith/arXiv arXiv 2022
-
[16]
Ruohan Gao and Kristen Grauman. 2019. Co-separating sounds of visual objects. InProceedings of the IEEE/CVF International Conference on Computer Vision
2019
-
[17]
Gemmeke, Daniel P
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter
-
[18]
Beat Gfeller, Dominik Roblek, and Marco Tagliasacchi. 2021. One-shot conditional audio filtering of arbitrary sounds. InICASSP. IEEE
2021
-
[19]
Arushi Goel, Sreyan Ghosh, Jaehyeon Kim, Sonal Kumar, Zhifeng Kong, Sang gil Lee, Chao-Han Huck Yang, Ramani Duraiswami, Dinesh Manocha, Rafael Valle, and Bryan Catanzaro. 2025. Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models. arXiv:2507.08128 [cs.SD] https://arxiv.org/abs/2507.08128
Pith/arXiv arXiv 2025
-
[20]
Yuan Gong, Yu-An Chung, and James Glass. 2021. AST: Audio Spec- trogram Transformer. InInterspeech 2021. 571–575. doi:10.21437/ Interspeech.2021-698
2021
-
[21]
Jiarui Hai, Helin Wang, Dongchao Yang, Karan Thakkar, Najim Dehak, and Mounya Elhilali. 2024. DPM-TSE: A Diffusion Probabilistic Model for Target Sound Extraction. InICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1196–
2024
-
[22]
Justin han, Sharat Raju, Rajalakshmi Nandakumar, Randall Bly, and Shyamnath Gollakota. 2019. Detecting middle ear fluid using smart- phones.Science Translational Medicine11 (05 2019), eaav1102. doi:10. 1126/scitranslmed.aav1102
2019
-
[23]
Guilin Hu, Malek Itani, Tuochao Chen, and Shyamnath Gollakota. 2025. Proactive Hearing Assistants that Isolate Egocentric Conversations. InProceedings of the 2025 Conference on Empirical Methods in Natu- ral Language Processing. Association for Computational Linguistics, Suzhou, China, 25377–25394. doi:10.18653/v1/2025.emnlp-main.1289
-
[24]
Jeremy Zhengqi Huang, Jaylin Herskovitz, Liang-Yuan Wu, Cecily Morrison, and Dhruv Jain. 2025. Weaving Sound Information to Sup- port Real-Time Sensemaking of Auditory Environments: Co-Designing with a DHH User(CHI ’25)
2025
-
[25]
Malek Itani, Tuochao Chen, and Shyamnath Gollakota. 2025. TF- MLPNet: Tiny Real-Time Neural Speech Separation. InClarity Chal- lenge, InterSpeech
2025
-
[26]
Malek Itani, Tuochao Chen, Arun Raghavan, Gavriel Kohlberg, and Shyamnath Gollakota. 2025. Wireless Hearables With Programmable Speech AI Accelerators(ACM MOBICOM ’25). ACM
2025
-
[27]
Malek Itani, Ashton Graves, Sefik Emre Eskimez, and Shyamnath Gollakota. 2025. Neural Speech Extraction with Human Feedback. In Interspeech 2025. 4998–5002. doi:10.21437/Interspeech.2025-214
-
[28]
Froehlich
Dhruv Jain, Kelly Mack, Akli Amrous, Matt Wright, Steven Goodman, Leah Findlater, and Jon E. Froehlich. 2020. HomeSound: An Iterative Field Deployment of an In-Home Sound Awareness System for Deaf or Hard of Hearing Users. InACM CHI
2020
-
[29]
Dhruv Jain, Hung Ngo, Pratyush Patel, Steven Goodman, Leah Find- later, and Jon Froehlich. 2020. SoundWatch: Exploring Smartwatch- Based Deep Learning Approaches to Support Sound Awareness for Deaf and Hard of Hearing Users. InACM SIGACCESS ASSETS
2020
-
[30]
Yincheng Jin, Yang Gao, Xiaotao Guo, Jun Wen, Zhengxiong Li, and Zhanpeng Jin. 2022. EarHealth: an earphone-based acoustic otoscope for detection of multiple ear diseases in daily life(MobiSys ’22). ACM. 13
2022
-
[31]
Fahim Kawsar, Chulhong Min, Akhil Mathur, and Alessandro Monta- nari. 2018. Earables for Personal-Scale Behavior Analytics.IEEE Perva- sive Computing17, 3 (2018), 83–89. doi:10.1109/MPRV.2018.03367740
arXiv 2018
-
[32]
Kevin Kilgour, Beat Gfeller, Qingqing Huang, Aren Jansen, Scott Wis- dom, and Marco Tagliasacchi. 2022. Text-Driven Separation of Arbi- trary Sounds.arXiv preprint arXiv:2204.05738(2022)
Pith/arXiv arXiv 2022
-
[33]
Gierad Laput, Karan Ahuja, Mayank Goel, and Chris Harrison. 2018. Ubicoustics: Plug-and-Play Acoustic Activity Recognition. InACM UIST
2018
-
[34]
Xubo Liu, Haohe Liu, Qiuqiang Kong, Xinhao Mei, Jinzheng Zhao, Qiushi Huang, Mark D Plumbley, and Wenwu Wang. 2022. Separate What You Describe: Language-Queried Audio Source Separation.arXiv preprint arXiv:2203.15147(2022)
Pith/arXiv arXiv 2022
-
[35]
Lane, Tanzeem Choudhury, and An- drew T
Hong Lu, Wei Pan, Nicholas D. Lane, Tanzeem Choudhury, and An- drew T. Campbell. 2009. SoundSense: Scalable Sound Sensing for People-Centric Applications on Mobile Phones. InACM MobiSys
2009
-
[36]
Vimal Mollyn, Karan Ahuja, Dhruv Verma, Chris Harrison, and Mayank Goel. 2022. SAMoSA: Sensing Activities with Motion and Subsampled Audio.IMWUT(2022)
2022
-
[37]
Vimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison, and Karan Ahuja. 2023. IMUPoser: Full-Body Pose Estimation Using IMUs in Phones, Watches, and Earbuds. InCHI(Hamburg, Germany)(CHI ’23). ACM
2023
-
[38]
Alessandro Montanari, Ashok Thangarajan, Khaldoon Al-Naimi, An- drea Ferlini, Yang Liu, Ananta Narayanan Balaji, and Fahim Kawsar
-
[39]
2020.Noise files for the DISCO dataset
Furnon Nicolas. 2020.Noise files for the DISCO dataset. https://github. com/nfurnon/disco
2020
-
[40]
Tsubasa Ochiai, Marc Delcroix, Yuma Koizumi, Hiroaki Ito, Keisuke Kinoshita, and Shoko Araki. 2020. Listen to What You Want: Neural Network-based Universal Sound Selector.arXiv e- prints, Article arXiv:2006.05712 (2020), arXiv:2006.05712 pages. arXiv:2006.05712 [eess.AS]
Pith/arXiv arXiv 2020
-
[41]
Yuki Okamoto, Shota Horiguchi, Masaaki Yamamoto, Keisuke Imoto, and Yohei Kawaguchi. 2022. Environmental Sound Extraction Using Onomatopoeic Words. InICASSP. IEEE
2022
-
[42]
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron Courville. 2018. FiLM: visual reasoning with a general con- ditioning layer. InProceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Arti- ficial Intelligence Conference and Eighth AAAI Symposium on Educa- tional Advances i...
2018
-
[43]
Darius Petermann and Minje Kim. 2022. Spain-Net: Spatially-Informed Stereophonic Music Source Separation. InICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 106–110. doi:10.1109/ICASSP43922.2022.9746277
arXiv 2022
-
[44]
Karol J. Piczak. 2015. ESC: Dataset for Environmental Sound Classifi- cation. InACM Multimedia
2015
-
[45]
Manoj Plakal and Daniel P. W. Ellis. 2020. YAMNet: A Pre-trained Audio Event Classifier. https://github.com/tensorflow/models/tree/ master/research/audioset/yamnet. TensorFlow Model Garden
2020
-
[46]
Jay Prakash, Zhijian Yang, Yu-Lin Wei, Haitham Hassanieh, and Romit Roy Choudhury. 2020. EarSense: earphones as a teeth activity sensor. InProceedings of the 26th Annual International Conference on Mobile Computing and Networking(London, United Kingdom)(Mobi- Com ’20). Association for Computing Machinery, New York, NY, USA, Article 40, 13 pages. doi:10.11...
arXiv 2020
-
[47]
Adam Pullin, Jake Stuchbury-Wass, Mathias Ciliberto, Kayla-Jade Butkow, Philipp Lepold, Tobias Röddiger, and Cecilia Mascolo. 2025. Ear-ECG Denoising Using Heart Sounds and the Extended Kalman Filter. InIEEE-EMBS International Conference on Body Sensor Networks
2025
-
[48]
2017.MUSDB18 - a corpus for music separation
Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner. 2017.MUSDB18 - a corpus for music separation
2017
-
[49]
Tobias Röddiger, Tobias King, Dylan Ray Roodt, Christopher Clarke, and Michael Beigl. 2023. OpenEarable: Open Hardware Earable Sensing Platform(UbiComp/ISWC ’22 Adjunct)
2023
-
[50]
Justin Salamon, Duncan MacConnell, Mark Cartwright, Peter Li, and Juan Pablo Bello. 2017. Scaper: A library for soundscape synthesis and augmentation. InW ASPAA. doi:10.1109/WASPAA.2017.8170052
arXiv 2017
-
[51]
Christian J Steinmetz and Joshua D Reiss. 2020. auraloss: Audio focused loss functions in PyTorch. InDigital music research network one-day workshop (DMRN+ 15)
2020
-
[52]
Jake Stuchbury-Wass, Andrea Ferlini, and Cecilia Mascolo. 2023. Mul- timodal Attention Networks for Human Activity Recognition From Earable Devices(UbiComp/ISWC ’22 Adjunct). ACM
2023
-
[53]
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiao- hua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. 2021. MLP-Mixer: An all-MLP Architecture for Vision. InNeurips
2021
-
[54]
Bandhav Veluri, Justin Chan, Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2023. Real-Time Target Sound Extraction. InICASSP
2023
-
[55]
Bandhav Veluri, Malek Itani, Justin Chan, Takuya Yoshioka, and Shyam- nath Gollakota. 2023. Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables. InACM UIST
2023
-
[56]
Bandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Look Once to Hear: Target Speech Hearing with Noisy Examples. InACM CHI
2024
-
[57]
Keigo Wakayama, Tomoko Kawase, Takafumi Moriya, Marc Delcroix, Hiroshi Sato, Tsubasa Ochiai, Masahiro Yasuda, and Shoko Araki. 2025. Real-time TSE demonstration via SoundBeam with KD. InInterspeech
2025
-
[58]
Helin Wang, Jiarui Hai, Yen-Ju Lu, Karan Thakkar, Mounya Elhilali, and Najim Dehak. 2025. SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer. InICASSP. 1–5
2025
-
[59]
Zhong-Qiu Wang, Gordon Wichern, Shinji Watanabe, and Jonathan Le Roux. 2022. STFT-domain neural speech enhancement with very low algorithmic latency.Trans. on Audio, Speech, and Language Processing (2022)
2022
-
[60]
Xudong Xu, Bo Dai, and Dahua Lin. 2019. Recursive visual sound separation using minus-plus net. InIEEE/CVF ICCV
2019
-
[61]
Dongchao Yang, Jinchuan Tian, Xuejiao Tan, Rongjie Huang, Songxi- ang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian, Xixin Wu, Zhou Zhao, and Helen Meng. 2023. UniAudio: An Audio Founda- tion Model Toward Universal Audio Generation.ArXivabs/2310.00704 (2023). https://api.semanticscholar.org/CorpusID:263334347
Pith/arXiv arXiv 2023
-
[62]
Haici Yang, Shivani Firodiya, Nicholas J. Bryan, and Minje Kim. 2022. Don’t Separate, Learn To Remix: End-To-End Neural Remixing With Joint Optimization. InICASSP. 116–120. doi:10.1109/ICASSP43922. 2022.9746077
arXiv 2022
-
[63]
Qiang Yang, Yang Liu, Jake Stuchbury-Wass, Mathias Ciliberto, Tobias Röddiger, Kayla-Jade Butkow, Adam Pullin, Emeli Panariti, Dong Ma, and Cecilia Mascolo. 2025. HearForce: Force Estimation for Manual Toothbrushing with Earables. (2025). doi:10.17863/CAM.122079
-
[64]
Koji Yatani and Khai N. Truong. 2012. BodyScope: A Wearable Acoustic Sensor for Activity Recognition. InUbiComp
2012
-
[65]
Plumbley, and Wenwu Wang
Yi Yuan, Xubo Liu, Haohe Liu, Mark D. Plumbley, and Wenwu Wang
-
[70]
InICASSP
FlowSep: Language-Queried Sound Separation with Rectified Flow Matching. InICASSP. 1–5. 14 Figure 11: User preferences survey results (N=7).(a) Preferred device. (b) Expected response time. (c) Sound selection method. (d) Volume adjustment method. (e) Number of simultaneous sounds. (f) Usage situations. A User Preferences Survey The participants in our us...
-
[1200]
doi:10.1109/ICASSP48485.2024.10447219
arXiv 2024
-
[2017]
InIEEE ICASSP
Audio Set: An ontology and human-labeled dataset for audio events. InIEEE ICASSP
-
[2024]
arXiv:2410.04775 [cs.ET] https://arxiv.org/abs/2410.04775
OmniBuds: A Sensory Earable Platform for Advanced Bio- Sensing and On-Device Machine Learning. arXiv:2410.04775 [cs.ET] https://arxiv.org/abs/2410.04775
-
[2025]
https://openreview.net/forum?id=eQE5YiQexy
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.