Pith. sign in

REVIEW 3 major objections 5 minor 87 references

Spatial Speech Translation: Translating Across Space With Binaural Hearables

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read A binaural hearable pipeline translates multiple concurrent speakers in real time while preserving each speaker's direction and voice characteristics, and it generalizes from synthetic training to unseen real-world environments.

desk verdict A solid proof-of-concept system paper that integrates separation, streaming expressive translation, and binaural rendering for hearables; the real-world evaluation is genuine evidence, but the synthetic-to-real BRIR transfer is the main open risk. read the letter →

arxiv 2504.18715 v1 pith:BEXIYJ42 submitted 2025-04-25 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords spatialspeechtranslationbinauralhearablesjointlocalizationandsourceseparationsimultaneousexpressiverenderingHRTFgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that hearables can translate several speakers in the wearer's environment at once, while keeping each speaker's direction and voice in the translated audio. It introduces a complete real-time pipeline: binaural source separation and localization, simultaneous expressive speech translation, and binaural rendering. On real-world unseen environments, it reports ASR-BLEU up to 22.01 despite interference, and user-perceived speaker similarity rising from 1.81 to 3.45. The key claim is that models trained entirely on synthetic binaural mixtures generalize to real rooms, outdoor spaces, and unseen wearers without hardware-specific training data.

What carries the argument

The load-bearing component is a search-based joint localization and separation network. The 360-degree space is divided into 36 angular regions; for each region, the binaural input is time-shifted by the interaural time difference corresponding to that angle and fed to a streaming TF-GridNet that is trained to output the separated source if a speaker is present at that angle and silence otherwise. Interaural phase and level differences are concatenated with the spectrogram as features, and false duplicates from multipath are removed by clustering separation outputs by segment-wise similarity. Around this core, the translation module is a simultaneous speech-to-text model (StreamSpeech-style with Conformer encoder and CTC-guided READ/WRITE policy), followed by a text-to-unit model and an expressivity-preserving vocoder conditioned on an expressive embedding extracted from the source speech; the whole translation model is fine-tuned on the separation model's imperfect outputs to become robust to residual interference. Finally, binaural rendering convolves translated monaural speech with a generic HRTF at the estimated angle for ITD and applies an ILD compensation scale computed from the separated source.

What would settle it

Collect binaural recordings of two concurrent French speakers in a room with a wearer whose head size differs substantially from the 18 cm average assumed in training, or place speakers closer than 0.75 m, then run the released pipeline and measure localization precision and recall plus ASR-BLEU; if precision or recall collapses well below the reported indoor values (97% and 98%) or ASR-BLEU drops to the no-separation baseline, the claimed synthetic-to-real generalization is falsified.

Watch

Extended reading notes

Core claim

The paper introduces spatial speech translation, a concept and system that takes a binaural mixture from microphones at the two ears, identifies how many speakers are present and from which angles, separates each voice, translates them simultaneously into the wearer's language while preserving prosody and vocal identity, and renders the translated speech binaurally so it appears to come from the original speaker's direction. The central empirical claim is that this works in real time on Apple M2 silicon and generalizes to unseen real-world environments and wearers: in six indoor and four outdoor venues, the joint localization and separation step reaches 97% precision and recall indoors and 92% and 94% outdoors with a median angle error of 6.8 degrees; fine-tuning the translation model on separation outputs raises ASR-BLEU from 18.06 to 22.07, and the full expressive system raises perceived speaker similarity from 1.81 to 3.45 while keeping median perceived direction error at 16.7 degrees versus 15.0 degrees for the original speech. The paper also reports that generic-HRTF rendering with ILD compensation brings interaural time difference error to 72.3 microseconds and interaural level difference error to 0.16 dB.

Load-bearing premise

The whole system rests on the claim that a separation model trained only on synthetic binaural mixtures, built from 77 room and head configurations, will work on real heads, real rooms, and real distances without any recordings made with the actual hardware; if that synthetic-to-real bridge fails, the spatial translation benefit disappears.

Editorial extensions

If this is right

  • Multi-speaker environments become translatable in real time; existing speech translators that assume a single speaker fail under interference.
  • The synthetic training recipe removes the need to collect data with each hearable device, room, language, or wearer; new language pairs can be added by generating new synthetic mixtures.
  • Listeners can follow who is speaking in a conversation because direction, prosody, and voice identity survive translation.
  • Because the pipeline runs on Apple M2 silicon with a real-time factor below one, it is deployable on commodity AR and wearable hardware today.
  • Latency can be traded against accuracy via chunk size (one to four seconds), giving tunable behavior for casual versus high-stakes use.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the synthetic-to-real bridge is as robust as reported, the same separation-trained-on-BRIR recipe should extend to other wearable geometries (for example, earbuds with smaller microphone spacing) and to more languages without any new real-world data, but only if the time-difference-of-arrival shift model is recalibrated to the new inter-microphone distance.
  • The fine-tune-on-upstream-distortion trick is a general design principle: cascaded speech systems (speech-to-text plus machine translation, diarization plus translation) should train their downstream model on the actual output distribution of their front end, not only on clean corpora.
  • The rendering delay-compensation scheme (apply current spatial cues to delayed translated audio) suggests a broader principle for streaming augmented-reality audio: spatial metadata can be decoupled from audio content latency; this could be tested with moving speakers by measuring whether direction perception remains accurate while a speaker walks during the translation delay.
  • Because the system outputs text as an intermediate, a natural extension is spatially anchored transcripts on an AR display; the paper mentions this direction but does not evaluate it.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces 'spatial speech translation', a hearable pipeline that localizes and separates concurrent speakers from binaural microphone input, translates each separated stream in real time with expressive speech-to-speech translation, and renders the translated speech binaurally at each speaker's original direction. The separation/localization model is trained on synthetic mixtures built from CoVoST2 speech, WHAM! noise, and 77 BRIR configurations from four public datasets, and the translation model is fine-tuned on the imperfect separation outputs. Evaluation includes synthetic benchmarks, real-world indoor/outdoor recordings with a Sony WH-1000XM4 prototype, listening studies with 29 participants, latency and noise-cancellation preference studies, and objective metrics such as ASR-BLEU, VSim, localization precision/recall, and delta-ITD/delta-ILD. The central claim is that this is the first real-time binaural hearable system that translates multiple speakers while preserving spatial cues and speaker voice characteristics, and that the synthetic training recipe generalizes to unseen wearers and environments without hardware-specific training data.

Significance. If the claims hold, this is a significant contribution to hearable and speech-translation research. The paper is the first to integrate spatial perception into end-to-end speech translation, and it provides a decomposable, reproducible pipeline with code and dataset release. The strongest evidence is the real-world user study (10 environments, 29 participants), the localization precision/recall results, the delta-ITD/delta-ILD improvements, and the demonstration that fine-tuning on separation outputs improves ASR-BLEU on both synthetic and real-world data. The synthetic-to-real transfer strategy is an appealing practical contribution, and the runtime analysis shows the pipeline is plausible for on-device use. However, the generalization claim rests on a narrow real-world validation, and several headline numbers are reported without measures of variability, which tempers the strength of the conclusions.

major comments (3)
  1. [§3.1.3, §4, §5] The separation and localization model is trained exclusively on synthetic mixtures generated by convolving CoVoST2 monaural speech with 77 BRIR configurations from CIPIC, RRBRIR, ASH, and CATTRIR, plus WHAM! noise, but the deployed hardware places SP15C microphones on the outside of the Sony WH-1000XM4 earcups. The real-world evaluation uses only loudspeaker reproductions of CoVoST2 test clips, not human talkers, and distances of 0.75–2.5 m. This is the load-bearing bridge for the paper's claim that the system generalizes to unseen wearers and environments without hardware-specific training data, yet §7 does not flag the microphone-placement mismatch, near-field effects, or loudspeaker-only speech sources as open risks. Because separation errors propagate directly to translation and rendering, I would like either additional experiments with human talkers and varied mic placements or an explicit, detailed scope discussion in the limitations section.
  2. [§5.1, Table 2, Fig. 6] No confidence intervals, standard deviations, or significance tests are reported for any of the headline subjective or objective numbers, including semantic consistency (3.35 vs. 1.15), speaker similarity (1.81 vs. 3.45), ASR-BLEU differences (18.06 vs. 22.07), localization precision/recall, or perceived angular error medians. The claim that the rendered English speech has 'similar' localization error to the original French speech is based on a median comparison without any measure of spread or a paired test. The manuscript should report per-participant/per-sample variability and appropriate statistical tests for the human ratings and for the metric comparisons that support the main claims.
  3. [§5.2.1, Fig. 11] The localization results are reported as means over participants in indoor and outdoor groups, but the precision/recall values appear to be per-participant binary decisions; no details are given on how partial detections (e.g., two speakers but one false positive) are counted, or how the 90th-percentile AoA error is computed across mixtures with different numbers of sources. Since the clustering false-positive elimination is a key algorithmic component, a precise definition of the evaluation protocol and error aggregation would strengthen the reproducibility of the localization claims.
minor comments (5)
  1. [§3.1.3] There are typos: 'diferent' and 'binural' should be 'different' and 'binaural'.
  2. [§6.2] The list of four model configurations labels both item (3) and item (4) as 'Finetuned S2T with Expressive T2S'; one of them should be 'Finetuned S2T with non-Expressive T2S' to match Table 4.
  3. [Fig. 12 caption] The caption says 'we compute the ΔITD and ΔITD between each input binaural French speech chunk and the rendered English speech chunk'; the second metric should be ΔILD.
  4. [Abstract, §1, Table 2] The abstract reports BLEU 'up to 22.01' while the introduction and Table 2 report 22.07; the numbers should be made consistent and the model configuration (fine-tuned non-expressive vs. fine-tuned expressive) should be stated in the abstract.
  5. [§5.1.1 and §5.1.2] The listening survey and spatial perception study report only mean values; adding the number of ratings per condition and error bars in Figs. 6 and 7 would make the results much easier to interpret.

Circularity Check

1 steps flagged · score 4.0 of 10

Rendering ILD metric is by construction (output ILD is set equal to input ILD); central translation and user-study claims are independent.

  1. self definitional [Section 3.3 (ILD compensation equation); Section 6.3 / Table 6.]
    "we first compute its ILD from the separated binaural source speech as ILD_i = ||y_i^r||_1 / ||y_i^l||_1. We then scale the translated signal based on the binaural HRTF response and ILD_i as: [o_i^l, o_i^r] = [o_i * h^l(theta=theta_i, phi=0), (ILD_i / (||h^l||_1 / ||h^r||_1)) o_i * h^r(theta=theta_i, phi=0)]"

    The right channel is scaled by ILD_i normalized by the HRTF channel norms, so the output ILD is forced to equal ILD_i by the formula itself. Table 6 then reports a ΔILD of 0.16 dB for this method and credits it with preserving spatial cues. That is not an independent prediction: it is the same quantity that was inserted into the output being measured on the output. The nontrivial residual is only how accurately the separated ILD_i matches the original source ILD, plus the ITD contribution from the generic HRTF. The central BLEU and user-study results do not depend on this tautology.

full rationale

The joint separation/localization and translation chain is self-contained and evaluated against external data. Separation is trained on synthetic mixtures from CoVoST2, WHAM!, and four public BRIR datasets and tested on held-out real-world recordings; localization precision/recall and AoA errors are not fitted to the reported numbers. The S2T module uses external StreamSpeech and SeamlessExpressive components, with fine-tuning on separation outputs evaluated on real-world recordings and standard test sets. Existing self-citations to search-based separation and TF-GridNet are architectural borrowings, not load-bearing uniqueness claims. The only reduction-by-construction element is the ILD compensation benchmark: the rendering equation explicitly transfers the input ILD to the output, so the reported ΔILD improvement over the generic-HRTF baseline is a definitional identity rather than an empirical discovery. This is a secondary validation metric; the paper's central spatial-translation claim is independently supported by human localization tests and BLEU results.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on a stack of external datasets and pretrained models; the only hand-set parameters are the search grid granularity, a power threshold, and a loss weight. No new physical entities are introduced.

free parameters (3)
  • N=36 angular search regions = 36
    Divides 360 degrees into 10-degree bins (Section 3.1.2); chosen by the authors and sets the localization resolution of the search-based separation.
  • Power threshold for source detection = 1e-2
    Windowed power threshold of 1e-2 over 0.75 s windows used to discard angular regions without a speaker (Section 3.1.2); affects precision and recall and is chosen heuristically.
  • Multi-resolution loss weight lambda = 0.1
    Weighting between L1 and multi-resolution spectrogram loss in the separation training objective (Section 3.1.3).
assumptions (4)
  • domain assumption Far-field sources and fixed microphone separation d = 18 cm for TDoA alignment
    Equation in Section 3.1.1 uses TDoA(theta) = d sin(theta)/c with d = 18 cm and c = 340 m/s; near-field sources or atypical head sizes would violate the alignment and degrade separation, though synthetic HRTF augmentation partially covers variation.
  • domain assumption Synthetic BRIR mixtures from CIPIC, RRBRIR, ASH, and CATTRIR are representative of real-world reverberation and HRTF variability
    Section 3.1.3 builds 90,000 training mixtures from four public BRIR datasets and claims generalization to unseen wearers and multipath environments; if this transfer fails, the separation module fails.
  • domain assumption Generic CIPIC HRTF is adequate for ITD rendering
    Section 3.3 uses CIPIC HRTF for both channels and only compensates ILD; evaluation shows small delta ITD, but personalized HRTFs are not used.
  • domain assumption ASR-BLEU on the rendered speech is a valid proxy for translation quality
    Tables 2, 4, and 5 score translation by first transcribing the synthesized speech with an ASR model; ASR errors on expressive or reverberant speech are conflated with translation errors.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Spatial Speech Translation: Translating Across Space With Binaural Hearables." pith.science (2026). https://pith.science/paper/BEXIYJ42

@misc{pith2026250418715,
  author       = {Pith},
  title        = {Pith review of: Spatial Speech Translation: Translating Across Space With Binaural Hearables},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BEXIYJ42}},
  note         = {Machine review of arXiv:2504.18715}
}
read the original abstract

Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce spatial speech translation, a novel concept for hearables that translate speakers in the wearer's environment, while maintaining the direction and unique voice characteristics of each speaker in the binaural output. To achieve this, we tackle several technical challenges spanning blind source separation, localization, real-time expressive translation, and binaural rendering to preserve the speaker directions in the translated audio, while achieving real-time inference on the Apple M2 silicon. Our proof-of-concept evaluation with a prototype binaural headset shows that, unlike existing models, which fail in the presence of interference, we achieve a BLEU score of up to 22.01 when translating between languages, despite strong interference from other speakers in the environment. User studies further confirm the system's effectiveness in spatially rendering the translated speech in previously unseen real-world reverberant environments. Taking a step back, this work marks the first step towards integrating spatial perception into speech translation.

Figures

Figures reproduced from arXiv: 2504.18715 by the authors.

Figure 1
Figure 1. "Spatial speech translation" is an intelligent hearable system that translates speakers in the wearer’s auditory space, [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of spatial speech translation. The input to our pipeline is a binaural noisy speech mixture in the source [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Spatial cues extraction and binaural rendering. (A) shows search-based joint localization and separation. We divide [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Simultaneous and expressive speech-to-speech translation. (1) In simultaneous speech to text (S2T) translation, a [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Real-world evaluation settings. A-J show ten different unseen multipath environments tested in our real-world [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Subjective evaluation of semantic consistency and speaker similarity. The left figure shows the mean opinion score [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 7
Figure 7. Figure 7: Real-world evaluation for user-perceived spatial [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: Subjective latency and noise cancellation evaluation. (A) shows the preference for different translation latencies [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Contextualization of the VSim metric. For each [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Tradeoff between ASR-BLEU and AL. We ran our [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Real-world evaluation for joint separation and [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 12
Figure 12. Figure 12: Binaural rendering with motion. (A) AoA error, (B) [PITH_FULL_IMAGE:figures/full_fig_p017_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

87 extracted references · 63 canonical work pages

  1. [1]

    Alex Agranovich, Eliya Nachmani, Oleg Rybakov, Yifan Ding, Ye Jia, Nadav Bar, Heiga Zen, and Michelle Tadmor Ramanovich. 2024. SimulTron: On-Device Simultaneous Speech to Speech Translation. arXiv:2406.02133 [eess.AS] https: //arxiv.org/abs/2406.02133

  2. [2]

    Dmitry Alexandrovsky, Susanne Putze, Michael Bonfert, Sebastian Höffner, Pitt Michelmann, Dirk Wenig, Rainer Malaka, and Jan David Smeddinck. 2020. Unmet Needs and Opportunities for Mobile Translation AI. CHI (2020)

  3. [3]

    Algazi, R.O

    V.R. Algazi, R.O. Duda, D.M. Thompson, and C. Avendano. 2001. The CIPIC HRTF database. , 99-102 pages. https://doi.org/10.1109/ASPAA.2001.969552

  4. [4]

    Apple. 2024. Listen with Personalized Spatial Audio for AirPods and Beats. https://support.apple.com/en-us/102596

  5. [5]

    Shoko Araki, Hiroshi Sawada, and Shoji Makino. 2007. Blind Speech Separation in a Meeting Situation with Maximum SNR Beamformers. In ICASSP

  6. [6]

    Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey, Chung-Cheng Chiu, Semih Yavuz, Ruoming Pang, Wei Li, and Colin Raffel. 2019. Monotonic Infinite Lookback Attention for Simultaneous Machine Translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Korhonen, David Traum, and Lluís Màrquez (Eds.)

  7. [7]

    Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick Von Platen, Yatharth Saraf, Juan Pino, et al

  8. [8]

    Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. CoRR abs/1409.0473 (2014). https://api.semanticscholar.org/CorpusID:11212020

Show all 87 references
  1. [9]

    Loïc Barrault, Yu-An Chung, Mariano Meglioli, David Dale, Ning Dong, Paul- Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoff- man, Christopher Klaiber, Pengwei Li, Daniel Licht, Jean Maillard, Alice Rakotoari- son, Kaushik Sadagopan, Guillaume Wenzek, Et...

  2. [10]

    Google Blog. 2024. Translate with Google Pixel Buds. https://support.google. com/googlepixelbuds/answer/7573100?hl=en

  3. [11]

    Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. 2023. AudioLM: a Language Modeling Approach to Audio Generation. arXiv:2209.03143 [cs.SD]

  4. [12]

    Alessandro Carlini, Camille Bordeau, and Maxime Ambard. 2024. Auditory localization: a comprehensive practical review. Frontiers Psycholo. (2024)

  5. [13]

    Ishan Chatterjee, Maruchi Kim, Vivek Jayaram, Shyamnath Gollakota, Ira Kemel- macher, Shwetak Patel, and Steven M. Seitz. 2022. ClearBuds: wireless binaural earbuds for learning-based speech enhancement. InProceedings of the 20th Annual International Conference on Mobile Syste...

  6. [14]

    Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei. 2022. WavLM: Large-Scale Self-Supervised Pre-...

  7. [15]

    Tuochao Chen, Malek Itani, Sefik Eskimez, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Hearable devices with sound bubbles. Nature Electronics 7 (11 2024), 1047–1058. https://doi.org/10.1038/s41928-024-01276-z

  8. [16]

    Tuochao Chen, Qirui Wang, Bohan Wu, Malek Itani, Sefik Emre Eskimez, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Target conversation extraction: Source separation using turn-taking dynamics. In InterSpeech

  9. [17]

    Shanbo Cheng, Zhichao Huang, Tom Ko, Hang Li, Ningxin Peng, Lu Xu, and Qini Zhang. 2024. Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent. arXiv:2407.21646 [cs.CL] https://arxiv.org/ abs/2407.21646

  10. [18]

    Kyunghyun Cho and Masha Esipova. 2016. Can neural machine translation do simultaneous translation? arXiv:1606.02012 [cs.CL]

  11. [19]

    Seamless Communication, Loïc Barrault, Yu-An Chung, Mariano Coria Megli- oli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, John Hoffman, Min-Jae Hwang, Hiro- fumi Inaguma, Christopher Klaiber, Ilia Kulikov, Pengwe...

  12. [20]

    Samuele Cornell, Zhong-Qiu Wang, Yoshiki Masuyama, Shinji Watanabe, Manuel Pariente, and Nobutaka Ono. 2023. Multi-Channel Target Speaker Extraction CHI ’25, April 26-May 1, 2025, Yokohama, Japan Chen, Wang, He and Gollakota with Refinement: The WavLab Submission to the Second...

  13. [21]

    Fahim Dalvi, Nadir Durrani, Hassan Sajjad, and Stephan Vogel. 2018. Incremental Decoding and Training Methods for Simultaneous Translation in Neural Machine Translation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Li...

  14. [22]

    Qianqian Dong, Zhiying Huang, Qiao Tian, Chen Xu, Tom Ko, Yunlong Zhao, Siyuan Feng, Tang Li, Kexin Wang, Xuxin Cheng, Fengpeng Yue, Ye Bai, Xi Chen, Lu Lu, Zejun Ma, Yuping Wang, Mingxuan Wang, and Yuxuan Wang. 2023. PolyVoice: Language Models for Speech to Speech Translation...

  15. [23]

    Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, and Holger Schwenk. 2023. SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations. In Proceedings of the 61st Annual Mee...

  16. [24]

    Maha Elbayad, Laurent Besacier, and Jakob Verbeek. 2020. Efficient Wait-k Models for Simultaneous Machine Translation. In InterSpeech

  17. [25]

    Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang, Xiaofei Wang, Zhuo Chen, and Xuedong Huang. 2022. Personalized speech enhancement: new models and Comprehensive evaluation. In ICASSP

  18. [26]

    Qingkai Fang, Zhengrui Ma, Yan Zhou, Min Zhang, and Yang Feng. 2024. CTC- based Non-autoregressive Textless Speech-to-Speech Translation. In Findings of the Association for Computational Linguistics: ACL 2024

  19. [27]

    Ritwik Giri, Shrikant Venkataramani, Jean-Marc Valin, Umut Isik, and Arvindh Krishnaswamy. 2021. Personalized PercepNet: Real-time, Low-complexity Target Voice Separation and Enhancement. In InterSpeech

  20. [28]

    Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher

  21. [29]

    Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang

  22. [30]

    Cong Han, Yi Luo, and Nima Mesgarani. 2020. Real-time binaural speech separa- tion with preserved spatial cues. In ICASSP). IEEE

  23. [31]

    IoSR-Surrey. 2016. IoSR-surrey/realroombrirs: Binaural impulse responses cap- tured in real rooms. https://github.com/IoSR-Surrey/RealRoomBRIRs

  24. [32]

    IoSR-Surrey. 2023. Simulated Room Impulse Responses. https://iosr.uk/software/ index.php

  25. [33]

    Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2023. Creating speech zones with self-distributing acoustic swarms. Nature Communi- cations 14 (09 2023). https://doi.org/10.1038/s41467-023-40869-8

  26. [34]

    Vivek Jayaram, Ira Kemelmacher-Shlizerman, and Steven M. Seitz. 2023. HRTF Estimation in the Wild. In UIST. ACM

  27. [35]

    Teerapat Jenrungrot, Vivek Jayaram, Steve Seitz, and Ira Kemelmacher- Shlizerman. 2020. The Cone of Silence: Speech Separation by Localization. In Advances in Neural Information Processing Systems

  28. [36]

    Chunyu Kit and Tak Ming Wong. 2008. Comparative Evaluation of Online Machine Translation Systems with Legal Texts. Law Library Journal 100 (2008), 299–321. https://api.semanticscholar.org/CorpusID:13481458

  29. [37]

    Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in neural information processing systems 33 (2020), 17022–17033

  30. [38]

    Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, and Wei-Ning Hsu. 2024. Voicebox: text-guided multilingual universal speech generation at scale. In Proceedings of the 37th International Conf...

  31. [39]

    Marie Lebert. 2022. A short history of translation through the ages. https://www.iapti.org/iaptiarticle/a-short-history-of-translation-through-the- ages-marie-lebert-2/

  32. [40]

    Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, and Wei-Ning Hsu. 2022. Direct Speech-to-Speech Translation With Discrete Units. In Proceedings of the 60th Annual Meeting of the Association for Co...

  33. [41]

    Tong Lei, Zhongshu Hou, Yuxiang Hu, Wanyu Yang, Tianchi Sun, Xiaobin Rong, Dahan Wang, Kai Chen, and Jing Lu. 2023. A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing Aids. In ICASSP. 1–2

  34. [42]

    Levinson

    Stephen C. Levinson. 2016. Turn-taking in Human Communication – Origins and Implications for Language Processing. Trends in Cog. Sci. (2016)

  35. [43]

    Liebling, Katherine Heller, Samantha Robertson, and Wesley Hanwen Deng

    Daniel J. Liebling, Katherine Heller, Samantha Robertson, and Wesley Hanwen Deng. 2022. Opportunities for Human-centered Evaluation of Machine Transla- tion Systems. In NAACL-HLT

  36. [44]

    Mingbo Ma, Liang Huang, Hao Xiong, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Xing Li, Hua Wu, and Haifeng Wang. 2019. STACL: Simultaneous Translation with Implicit Anticipation and Controllable Latency using Prefix-to-Prefix Framework....

  37. [45]

    Xutai Ma, Juan Pino, James Cross, Liezl Puzon, and Jiatao Gu. 2020. Monotonic Multihead Attention. In ICLR

  38. [46]

    Evgeny Matusov, Stephan Kanthak, and Hermann Ney. 2005. On the Integration of Speech Recognition and Statistical Machine Translation. 3177–3180. https: //doi.org/10.21437/Interspeech.2005-726

  39. [47]

    Tobias May, Steven Van De Par, and Armin Kohlrausch. 2010. A probabilistic model for robust localization based on a binaural auditory front-end. IEEE Transactions on audio, speech, and language processing 19, 1 (2010), 1–13

  40. [48]

    Mymanu. 2024. Mymanu Click S. https://mymanu.com/products/mymanu-clik-s

  41. [49]

    H. Ney. 1999. Speech translation: coupling of recognition and translation. In ICASSP99, Vol. 1

  42. [50]

    NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia-Gonzalez, Prang...

  43. [51]

    NPR. 2023. Finding your place in the galaxy with the help of Star Trek. https: //www.npr.org/2023/10/14/1205714903/star-trek

  44. [52]

    All Things Considered NPR. 1998. Babelfish, a Translator Inspired by ’The Hitchhiker’s Guide’. https://www.npr.org/1998/02/12/1036190/babelfish-a- translator-inspired-by-the-hitchhikers-guide

  45. [53]

    Lucas Nunes Vieira, Minako O’Hagan, and Carol O’Sullivan. 2020. Understanding the societal impacts of machine translation: a critical review of the literature on medical and legal use cases. Information Communication and Society (06 2020). https://doi.org/10.1080/1369118X.2020.1776370

  46. [54]

    Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, and Ann Lee. 2022. Enhanced direct speech-to-speech translation using self-supervised pre-training and data augmentation. In InterSpeech

  47. [55]

    Santhosh Kumar, Adam Lopez, Damianos G

    Matt Post, G. Santhosh Kumar, Adam Lopez, Damianos G. Karakos, Chris Callison- Burch, and Sanjeev Khudanpur. 2013. Improved speech-to-text translation with the Fisher and Callhome Spanish-English speech translation corpus. In Interna- tional Workshop on Spoken Language Translation

  48. [56]

    Liu, Ron J

    Colin Raffel, Minh-Thang Luong, Peter J. Liu, Ron J. Weiss, and Douglas Eck

  49. [57]

    Yi Ren, Jinglin Liu, Xu Tan, Chen Zhang, Tao Qin, Zhou Zhao, and Tie-Yan Liu

  50. [58]

    Dario Rethage, Jordi Pons, and Xavier Serra. 2018. A Wavenet for Speech De- noising. In ICASSP

  51. [59]

    Liebling, Michal Lahav, Katherine Heller, Mark Díaz, Samy Bengio, and Niloufar Salehi (Eds.)

    Samantha Robertson, Wesley Deng, Timnit Gebru, Margaret Mitchell, Daniel J. Liebling, Michal Lahav, Katherine Heller, Mark Díaz, Samy Bengio, and Niloufar Salehi (Eds.). 2021. Three Directions for the Design of Human-Centered Machine Translation

  52. [60]

    In Proceedings of the 34th International Conference on Machine Learning - Volume 70 (Sydney, NSW, Australia)(ICML’17)

    Online and linear-time attention by enforcing monotonic alignments. In Proceedings of the 34th International Conference on Machine Learning - Volume 70 (Sydney, NSW, Australia)(ICML’17). JMLR.org, 2837–2846

  53. [61]

    SDK. 2023. Steam Audio. https://valvesoftware.github.io/steam-audio/

  54. [62]

    In Annual Meeting of the Association for Computational Linguistics

    SimulSpeech: End-to-End Simultaneous Speech to Text Translation. In Annual Meeting of the Association for Computational Linguistics

  55. [63]

    Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian. 2023. NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers. arXiv:2304.09116 [eess.AS] https://arxiv.org/abs/2304.09116

  56. [64]

    Matthias Sperber and Matthias Paulik. 2020. Speech Translation and the End-to- End Promise: Taking Stock of Where We Are. InAnnual Meeting of the Association for Computational Linguistics

  57. [65]

    Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, Hannah Muckenhirn, Dirk Padfield, James Qin, Danny Rozenberg, Tara Sainath, Johan Schalkwyk, Matt Sharif...

  58. [66]

    Timekettle. [n. d.]. Timekettle WT2 Edge/W3 Real-time Translator Earbuds, 2-way simultaneous interpretation. https://www.timekettle.co/products/wt2- edge-online-voice-language-translator-earbuds Spatial Speech Translation: Translating Across Space With Binaural Hearables CHI ’...

  59. [67]

    ShanonPearce. 2022. Shanonpearce/ash-listening-set: A dataset of filters for head- phone correction and binaural synthesis of spatial audio systems on headphones. https://github.com/ShanonPearce/ASH-Listening-Set/tree/main

  60. [68]

    Bandhav Veluri, Justin Chan, Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2023. Real-Time Target Sound Extraction. In ICASSP. 1–5

  61. [69]

    Bandhav Veluri, Malek Itani, Justin Chan, Takuya Yoshioka, and Shyamnath Gollakota. 2023. Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (...

  62. [70]

    Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, and Jianyuan Zhong. 2021. Attention Is All You Need In Speech Separation. In ICASSP 2021

  63. [71]

    Anran Wang, Maruchi Kim, Hao Zhang, and Shyamnath Gollakota. 2022. Hybrid Neural Networks for On-device Directional Hearing. AAAI (2022)

  64. [72]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems , Vol. 30

  65. [73]

    Peidong Wang, Eric Sun, Jian Xue, Yu Wu, Long Zhou, Yashesh Gaur, Shujie Liu, and Jinyu Li. 2022. LAMASSU: A Streaming Language-Agnostic Multilin- gual Speech Recognition and Translation Model Using Neural Transducers. In Interspeech. https://api.semanticscholar.org/CorpusID:258968116

  66. [74]

    Waverly. 2024. Waverly labs Earbuds. https://www.waverlylabs.com/

  67. [75]

    Bandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Look Once to Hear: Target Speech Hearing with Noisy Examples. In CHI (Honolulu, HI, USA) (CHI ’24). ACM, Article 37, 16 pages

  68. [76]

    Zhongweiyang Xu and Romit Roy Choudhury. 2022. Learning to Separate Voices by Spatial Regions. ICML (2022)

  69. [77]

    Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei. 2023. Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers. arXiv:2301.02111 [cs.CL]

  70. [78]

    Changhan Wang Jiatao Gu Juan Pino Xutai Ma, Mohammad Javad Dousti. 2020. Simuleval: An evaluation toolkit for simultaneous translation. In Proceedings of the EMNLP

  71. [79]

    Mu Yang, Naoyuki Kanda, Xiaofei Wang, Junkun Chen, Peidong Wang, Jian Xue, Jinyu Li, and Takuya Yoshioka. 2024. Diarist: Streaming Speech Translation with Speaker Diarization. In ICASSP. 10866–10870

  72. [80]

    Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith St...

  73. [81]

    Shaolei Zhang and Yang Feng. 2021. Universal Simultaneous Machine Translation with Mixture-of-Experts Wait-k Policy. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics

  74. [82]

    Jian Xue, Peidong Wang, Jinyu Li, Matt Post, and Yashesh Gaur. 2022. Large- Scale Streaming End-to-End Speech Translation with Neural Transducers. In Interspeech

  75. [85]

    Shaolei Zhang, Qingkai Fang, Shoutao Guo, Zhengrui Ma, Min Zhang, and Yang Feng. 2024. StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

  76. [87]

    Kateřina Žmolíková, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, and Tomohiro Nakatani. 2017. Speaker-Aware Neural Network Based Beamformer for Speaker Extraction in Speech Mixtures. InProc. Interspeech 2017

  77. [2017]

    Non-autoregressive neural machine translation. (2017)

  78. [2020]

    In InterSpeech

    Conformer: Convolution-augmented Transformer for Speech Recognition. In InterSpeech

  79. [2022]

    In InterSpeech

    XLS-R: Self-supervised cross-lingual speech representation learning at scale. In InterSpeech

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.