Pith. sign in

REVIEW 5 major objections 10 minor 23 references

AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition

T0 review · 5 major / 10 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read AISHELL-5, an open-source multi-channel Mandarin corpus of over 100 hours of free conversation recorded inside a moving electric car, shows that today's best ASR systems still make character errors above 16% even after fine-tuning.

desk verdict A genuinely useful open dataset with a real annotation-quality gap; the benchmark claims are conditional on an audit, but the resource deserves referee time. read the letter →

arxiv 2505.23036 v1 pith:UYEKE6J2 submitted 2025-05-29 cs.SD eess.AS

classification cs.SDeess.AS
keywords AISHELL-5in-carspeechrecognitionmulti-channelaudiomulti-speakerdiarizationMandarinASRsourceseparationreal-worldnoisedatasetopen-sourcebenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces AISHELL-5, which it presents as the first open-source in-car multi-channel multi-speaker Mandarin speech dataset. The recordings comprise more than 100 hours of free conversation inside an electric vehicle under more than 60 designed driving scenarios, captured by four far-field microphones mounted on the doors plus near-field headset microphones that serve as clean reference signals; a separate 40-hour set of real road and cabin noise supports simulated mixtures. Alongside the corpus, the authors release a reproducible baseline whose frontend suppresses echo and separates speakers before recognition. On the two evaluation tracks, the best fine-tuned system reaches a 16.65% character error rate when speaker timestamps are given, but only 47.18% when the system must do diarization itself, showing that in-car noise and overlapping speech remain unsolved. If the dataset is sound, it gives the field a public benchmark tightly matched to a commercially important scenario.

What carries the argument

The load-bearing object is the dataset design: a hybrid electric car with four far-field microphones placed above the door handles and high-fidelity near-field headset microphones worn by each of two to four speakers, recorded across more than 60 scenarios that vary speed, windows, sunroof, air conditioning, stereo, and day/night condition. The near-field channels provide clean reference speech used to build time-aligned annotations, while the four far-field channels supply the actual multi-channel task. The baseline system's mechanism is a two-stage frontend: acoustic echo cancellation followed by blind source separation (including an Independent Vector Analysis stage and a learned spatial separation model), then voice-activity-detection segmentation and an ASR module; track two receives this frontend, while track one is evaluated directly on oracle-segmented speech. This pairing makes it possible to attribute errors separately to recognition and to diarization or frontend processing.

What would settle it

Independently re-annotate a random sample of the released near-field sessions (for example, three hours across the train and test splits) and measure agreement on speaker identity and segment boundaries; if disagreement exceeds a small preset percentage, the ground truth and all baseline scores drawn from it are unreliable, regardless of whether the corpus exists.

Watch

Extended reading notes

Core claim

The paper's central claim is that a sizable, real-recorded, open corpus can finally make the in-car scenario a measurable research problem, and that current models are far from solving it. The authors record natural multi-speaker conversation in a real hybrid car with four door-mounted far-field microphones and per-speaker near-field references, then annotate sessions in a portable text-grid format with speaker identities, segment timestamps, and transcriptions. They show that large pretrained speech recognizers degrade sharply on this data, and that a specifically tuned frontend plus fine-tuning yields the best published numbers on the corpus: 16.65% CER on the track with oracle timestamps and 47.18% cpCER on the full diarization-plus-recognition track. The contribution is therefore both the corpus and the reproducible measurement of how far in-car ASR still has to go.

Load-bearing premise

The load-bearing premise is that the TextGrid annotations, built from near-field reference audio, contain correct speaker identities, segment boundaries, and transcriptions; if those labels are systematically wrong, every reported error rate and the whole benchmark's usefulness collapses.

Editorial extensions

If this is right

  • Future in-car ASR systems can now be compared on one public corpus with two clearly defined tracks, making progress measurable.
  • The 40-hour noise collection lets researchers synthesize in-car mixtures at controlled SNRs and train separation or enhancement systems without new vehicle recordings.
  • The large gap between 16.65% CER on oracle-segmented speech and 47.18% cpCER on the full pipeline shows that speaker diarization and frontend processing, not raw recognition, are currently the dominant source of error.
  • Fine-tuning an open-source recognizer on the release improves its in-car CER from 20.16% to 16.65%, demonstrating that domain adaptation matters more than raw model scale on this benchmark.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the release includes near-field audio only in the training set, a natural next step is to use the two test sets to isolate diarization error: a reader could feed oracle timestamps to the evaluation pipeline for Eval2 and compare against the published cpCER.
  • The fixed four-door microphone layout ties the corpus to one vehicle geometry; extending the protocol to other cars or seating layouts would show which conclusions are specific to this geometry and which transfer.
  • Using the 40 hours of real noise together with the near-field speech would let researchers build synthetic mixtures at arbitrary overlaps; such simulations would be a cheap test of whether the challenges exposed here can be overcome with data augmentation alone.
  • The claim of being 'the first' open-source corpus of this kind could be checked by a literature sweep; if an earlier public in-car multi-channel multi-speaker dataset exists, the novelty claim would narrow to the scale and scenario coverage.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 10 minor

Summary. This paper describes AISHELL-5, an open-source in-car Mandarin speech dataset: over 100 hours of multi-channel audio captured by four door-mounted far-field microphones and near-field headset microphones, recorded in an electric vehicle across more than 60 driving scenarios, plus 40 hours of real environmental noise. The authors adopt the ICMC-ASR challenge's Eval1 (oracle-timestamp CER) and Eval2 (diarization and recognition cpCER) protocols, provide a baseline with AEC, IVA, and SpatialNet frontends followed by several ASR models, and report best results of 16.65% CER on Eval1 and 47.18% cpCER on Eval2. The paper claims that this is the first open-source corpus of its kind and that it demonstrates the difficulty of current ASR models in in-car multi-speaker conditions.

Significance. If the corpus is released as described, it fills a genuine gap: there is no large public in-car multi-channel multi-speaker Mandarin corpus, and the 40-hour real noise set plus the open baseline code would support both frontend and recognition research. The diversity of the recording scenarios, the use of real driving conditions, and the planned public release of data and code are clear strengths. However, the value of the benchmark depends entirely on the TextGrid reference labels, whose quality is not demonstrated, and the current evaluation design prevents external users from checking the Eval labels; these issues must be resolved before the dataset can serve as a trustworthy benchmark.

major comments (5)
  1. [Section 2; Table 2; Introduction] Annotation quality is the load-bearing premise of the benchmark, but no evidence for it is provided. Section 2 states only that each session's TextGrid contains timestamps, speaker IDs, gender, and transcription text; there is no annotation protocol, inter-annotator agreement metric, manual audit rate, or post-correction validation. The Introduction says data issues including "audio truncation and mismatched transcription" were fixed, but the fixing procedure and its verification are not described. Because Table 3 reports CER and cpCER computed against these TextGrids, systematic errors in boundaries, speaker assignment, or transcripts propagate directly into every evaluation result and into future use of the corpus as ground truth. Table 2 shows that near-field audio is available only for Train, so external users cannot independently verify the Eval labels; please provide a label-quality audit or release a human-verified subset of the Eval data.
  2. [Section 3.2; Abstract] The reproducibility claim is not fully supported by the manuscript's own text. Section 3.2 states that "we only share the decoding results in the baseline system without providing the related code" for at least the Zipformer/Icefall part of the ASR module, while the abstract promises an "open-access, reproducible baseline system." Please state explicitly which components are released (frontend, Wenet training, Zipformer training, decoding), provide the missing code, or revise the reproducibility claim to match what is actually distributed.
  3. [Abstract; Section 3.3; Table 3] The abstract's claim that the recognition module "accurately transcribes the content of each individual speaker" is contradicted by the reported results: the best Eval1 CER is 16.65% and the best Eval2 cpCER is 47.18% (Table 3). These numbers demonstrate a challenging benchmark, not accurate transcription. The abstract and conclusion should be reworded to describe the baseline as an open, reproducible reference system whose results quantify the difficulty of the task.
  4. [Section 3.3; Table 3] The fine-tuning setup is reported inconsistently: Table 3 lists the Paraformer-Finetuned model as trained for 10 epochs, while Section 3.3 says "After fine-tuning for 20 epochs with the same baseline training data." This must be reconciled. In addition, Table 3's header should label the numeric columns explicitly (Eval1 with AEC+IVA, Eval2 with AEC+IVA, Eval2 with SpatialNet), since the current header is ambiguous.
  5. [Section 3.1; Eq. (2)] Equation (2) is not a correct description of IVA and is dimensionally inconsistent: A is defined in Eq. (1) as an M by N mixing matrix, so det(A) is not defined, and A(y(t)) is not a meaningful expression. The IVA objective involves frequency-wise demixing matrices and source models. Because the paper advertises a reproducible frontend, either correct the equation or replace it with a precise citation to the actual implementation.
minor comments (10)
  1. [Abstract] The abstract contains the typo "AISHLL-5" for "AISHELL-5."
  2. [Section 1] The phrase "the recognition progress" should be "the recognition process," and "some researches" should be "some researchers" or "some studies."
  3. [Section 2; Table 1] Table 1 appears twice, and the two copies are inconsistent: the first shows speed bins D-F as 0-40 km/h and omits J-L, while the second shows D-F as 0-60 km/h and includes J-L at 80-120 km/h.
  4. [Section 2; Table 2] Section 2 says 260 participants were recorded, but Table 2 lists 147+6+6+6=165 speakers across Train, Dev, Eval1, and Eval2; please explain how the remaining participants are used or excluded.
  5. [Section 3.1] The text states M=4 and N=2 in AISHELL-5, but Section 2 says 2-4 speakers are recorded per session; clarify how N is set for sessions with three or four speakers.
  6. [Section 3.1; Eq. (1)] Equation (1) and the surrounding text use s(t) for both the source signals and the mixed audio; the second occurrence should be y(t).
  7. [Section 3.3; Table 3] Table 3 reports only point estimates, and Eval1 and Eval2 contain only 18 sessions and 6 speakers each (Table 2); small differences such as 20.16 versus 16.65 on Eval1 should be interpreted with caution. Please add confidence intervals, bootstrap estimates, or an explicit caveat.
  8. [Section 3.3] The statement that E-Branchformer achieved the best Eval1 performance at 26.05% is only true among the four models trained from scratch, since Paraformer's 20.16% without fine-tuning is lower; please clarify the comparison scope.
  9. [Footnotes; References] The dataset URL in footnote 2 contains a space ("aishell 5") and is not a valid URL; also, no data license is stated. Provide a working link or DOI and a license.
  10. [References] Reference [12] contains a stray space in "V oice" and should be "Voice."

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the paper is a dataset release and benchmark, not a derivation; the only self-referential elements are the provenance of the ICMC-ASR-derived data and baseline, which are not load-bearing.

full rationale

The paper's central claim is that AISHELL-5 is the first open-source in-car multi-channel multi-speaker Mandarin ASR dataset. This is an availability/contribution claim, not a derived result, and it is not established by fitting a parameter to data and then predicting that same data. The reported Eval1 CER and Eval2 cpCER values in Table 3 are measurements obtained by running ASR systems on the released test sets against the provided TextGrid labels; they are not predictions derived from the labels by construction. The paper contains no equation that reduces to its inputs. The main self-referential aspect is provenance: the paper states that 'we fixed all the data-related issues of the ICMC-ASR challenge dataset, including audio truncation and mismatched transcription, and now officially open-source it, called AISHELL-5', and the baseline is 'based on the ICMC-ASR baseline'. Both of these are explicit provenance statements, and the 'first open-source' claim does not rest on the authors' prior work as proof. The evaluation protocol is adopted from the authors' own ICMC-ASR challenge, but adopting a protocol is not circular. The unmeasured annotation quality of the TextGrids is a serious validity and correctness risk, but it is not a circularity: the benchmark scores inherit the labels empirically, not by definition. Overall, the paper is not circular; any concerns are about annotation verification, reproducibility of the Zipformer portion, or the strength of the 'first' claim, all of which are outside the circularity definition used here.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim is a dataset release, not a derivation, so no fitted free parameters are core to the claim; baseline hyperparameters, such as learning rate 0.002 and 100 epochs, are stated but are not used to establish the dataset's existence. The load-bearing unproven premises are annotation quality, representativeness of the recording setup, and the standard-Mandarin speaker pool. No new physical or theoretical entities are introduced.

assumptions (3)
  • domain assumption TextGrid timestamps and transcripts derived from near-field headset audio are accurate enough to serve as ground truth.
    Section 2 describes scripts provided in TextGrid format but reports no inter-annotator agreement, manual audit rate, or label-quality metric.
  • domain assumption Four door-mounted far-field microphones plus near-field headsets capture the intended in-car acoustic scenes.
    Section 2 specifies the hardware but provides no calibration, channel verification, or acoustic validation.
  • domain assumption Participants with 'no notable accents' are representative of the target Mandarin in-car ASR population.
    Section 2 states 260 participants with no notable accents; dialectal diversity is not addressed and is not acknowledged as a limitation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition." pith.science (2026). https://pith.science/paper/UYEKE6J2

@misc{pith2026250523036,
  author       = {Pith},
  title        = {Pith review of: AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UYEKE6J2}},
  note         = {Machine review of arXiv:2505.23036}
}
read the original abstract

This paper delineates AISHELL-5, the first open-source in-car multi-channel multi-speaker Mandarin automatic speech recognition (ASR) dataset. AISHLL-5 includes two parts: (1) over 100 hours of multi-channel speech data recorded in an electric vehicle across more than 60 real driving scenarios. This audio data consists of four far-field speech signals captured by microphones located on each car door, as well as near-field signals obtained from high-fidelity headset microphones worn by each speaker. (2) a collection of 40 hours of real-world environmental noise recordings, which supports the in-car speech data simulation. Moreover, we also provide an open-access, reproducible baseline system based on this dataset. This system features a speech frontend model that employs speech source separation to extract each speaker's clean speech from the far-field signals, along with a speech recognition module that accurately transcribes the content of each individual speaker. Experimental results demonstrate the challenges faced by various mainstream ASR models when evaluated on the AISHELL-5. We firmly believe the AISHELL-5 dataset will significantly advance the research on ASR systems under complex driving scenarios by establishing the first publicly available in-car ASR benchmark.

Figures

Figures reproduced from arXiv: 2505.23036 by the authors.

Figure 1
Figure 1. Structure of our baseline system, including train process and inference process subsets of AISHELL-5 is provided in Table2 [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

23 extracted references · 16 canonical work pages

  1. [1]

    Introduction Unlike common automatic speech recognition (ASR) applica- tions in the home or smart assistant scenarios, in-car ASR sys- tems encounter a range of unique challenges. These systems must contend with various internal and external noise sources, including wind, engine sound, tire noise, car stereos, and nearby vehicles, all of which contribute ...

  2. [2]

    Additionally, each speaker wears a high-fidelity micro- phone to collect near-field audio for data annotation

    Dataset The AISHELL-5 dataset is recorded inside a hybrid electric car, with a far-field microphone placed above the door handles of all four doors to capture far-field audio from different areas of the car. Additionally, each speaker wears a high-fidelity micro- phone to collect near-field audio for data annotation. A total of 260 participants are involv...

  3. [3]

    The baseline system con- sists of two primary sub-modules: speech frontend processing and automatic speech recognition (ASR)

    Baseline We develop a multi-channel in-car speech transcription system3 based on the ICMC-ASR baseline 4. The baseline system con- sists of two primary sub-modules: speech frontend processing and automatic speech recognition (ASR). We provide the data preprocessing pipeline in the baseline system, and after process- ing the data, we train each sub-module ...

  4. [4]

    AISHELL-5 is suitable for speech sep- aration, speech enhancement, noise reduction, and automatic speech recognition (ASR) tasks targeting in-car scenarios

    Conclusions This paper introduces AISHELL-5, currently the largest open- source multi-channel, multi-speaker free-talk in-car speech dataset, tailored for two key issues in current in-car speech pro- cessing techniques: the complex acoustic environment within the car and the frequent occurrence of overlapping speech from driver and passengers. AISHELL-5 i...

  5. [5]

    AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diariza- tion in conference scenario,

    Y . Fu, L. Cheng, S. Lv, Y . Jv, Y . Kong, Z. Chen, Y . Hu, L. Xie, J. Wu, H. Buet al., “AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diariza- tion in conference scenario,”arXiv preprint arXiv:2104.03603, 2021

  6. [6]

    M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,

    F. Yu, S. Zhang, Y . Fu, L. Xie, S. Zheng, Z. Du, W. Huang, P. Guo, Z. Yan, B. Ma, X. Xu, and H. Bu, “M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,” inICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022

  7. [7]

    Librimix: An open-source dataset for generalizable speech separation,

    J. Cosentino, M. Pariente, S. Cornell, A. Deleforge, and E. Vin- cent, “Librimix: An open-source dataset for generalizable speech separation,”arXiv preprint arXiv:2005.11262, 2020

  8. [8]

    Lib- rispeech: An asr corpus based on public domain audio books,

    V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- rispeech: An asr corpus based on public domain audio books,” in2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2015, pp. 5206–5210

Show all 23 references
  1. [9]

    Wham!: Extend- ing speech separation to noisy environments,

    G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. L. Roux, “Wham!: Extend- ing speech separation to noisy environments,”arXiv preprint arXiv:1907.01160, 2019

  2. [10]

    Realistic multi- microphone data simulation for distant speech recognition,

    M. Ravanelli, P. Svaizer, and M. Omologo, “Realistic multi- microphone data simulation for distant speech recognition,”arXiv preprint arXiv:1711.09470, 2017

  3. [11]

    Effective noise-aware data simulation for domain- adaptive speech enhancement leveraging dynamic stochastic per- turbation,

    C.-C. Wang, L.-W. Chen, H.-S. Lee, B. Chen, and H.- M. Wang, “Effective noise-aware data simulation for domain- adaptive speech enhancement leveraging dynamic stochastic per- turbation,” in2024 IEEE Spoken Language Technology Workshop (SLT), 2024

  4. [12]

    The iscslp 2022 intelligent cock- pit speech recognition challenge (icsrc): Dataset, tracks, baseline and results,

    A. Zhang, F. Yu, K. Huang, L. Xie, L. Wang, E. S. Chng, H. Bu, B. Zhang, W. Chen, and X. Xu, “The iscslp 2022 intelligent cock- pit speech recognition challenge (icsrc): Dataset, tracks, baseline and results,” in2022 13th International Symposium on Chinese Spoken Language Proc...

  5. [13]

    ICMC-ASR: the ICASSP 2024 in-car multi- channel automatic speech recognition challenge,

    H. Wang, P. Guo, Y . Li, A. Zhang, J. Sun, L. Xie, W. Chen, P. Zhou, H. Bu, X. Xu, B. Zhang, Z. Chen, J. Wu, L. Wang, E. S. Chng, and S. Li, “ICMC-ASR: the ICASSP 2024 in-car multi- channel automatic speech recognition challenge,” inIEEE Inter- national Conference on Acoustics...

  6. [14]

    Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,

    Z. Gao, S. Zhang, I. McLoughlin, and Z. Yan, “Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,” inINTERSPEECH, 2022

  7. [15]

    Robust speech recognition via large-scale weak su- pervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak su- pervision,” 2022

  8. [16]

    Funaudiollm: V oice understanding and genera- tion foundation models for natural interaction between humans and llms,

    K. An, Q. Chen, C. Deng, Z. Du, C. Gao, Z. Gao, Y . Gu, T. He, H. Hu, K. Hu, S. Ji, Y . Li, Z. Li, H. Lu, H. Luo, X. Lv, B. Ma, Z. Ma, C. Ni, C. Song, J. Shi, X. Shi, H. Wang, W. Wang, Y . Wang, Z. Xiao, Z. Yan, Y . Yang, B. Zhang, Q. Zhang, S. Zhang, N. Zhao, and S. Zheng, “F...

  9. [17]

    Qwen2-audio technical report,

    Y . Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y . Leng, Y . Lv, J. He, J. Lin, C. Zhou, and J. Zhou, “Qwen2-audio technical report,” 2024. [Online]. Available: https://arxiv.org/abs/2407.10759

  10. [18]

    Independent vector analysis: An extension of ICA to multivariate components,

    T. Kim, T. Eltoft, and T. Lee, “Independent vector analysis: An extension of ICA to multivariate components,” inIndependent Component Analysis and Blind Signal Separation, 6th Interna- tional Conference, ICA 2006, Charleston, SC, USA, March 5-8, 2006, Proceedings, J. P. Rosca,...

  11. [19]

    Speech-transformer: A no- recurrence sequence-to-sequence model for speech recognition,

    L. Dong, S. Xu, and B. Xu, “Speech-transformer: A no- recurrence sequence-to-sequence model for speech recognition,” in2018 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2018, Calgary, AB, Canada, April 15- 20, 2018. IEEE, 2018

  12. [20]

    Conformer: Convolution-augmented transformer for speech recognition,

    A. Gulati, J. Qin, C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in 21st Annual Conference of the International Speech Communi- cation Association, Interspeech 2020, ...

  13. [21]

    E-branchformer: Branchformer with enhanced merging for speech recognition,

    K. Kim, F. Wu, Y . Peng, J. Pan, P. Sridhar, K. J. Han, and S. Watanabe, “E-branchformer: Branchformer with enhanced merging for speech recognition,” inIEEE Spoken Language Technology Workshop, SLT 2022, Doha, Qatar, January 9- 12, 2023. IEEE, 2022, pp. 84–91. [Online]. Availa...

  14. [22]

    Zipformer: A faster and better encoder for automatic speech recognition,

    Z. Yao, L. Guo, X. Yang, W. Kang, F. Kuang, Y . Yang, Z. Jin, L. Lin, and D. Povey, “Zipformer: A faster and better encoder for automatic speech recognition,” inThe Twelfth International Conference on Learning Representations, 2023

  15. [23]

    Wenet: Production oriented stream- ing and non-streaming end-to-end speech recognition toolkit,

    Z. Yao, D. Wu, X. Wang, B. Zhang, F. Yu, C. Yang, Z. Peng, X. Chen, L. Xie, and X. Lei, “Wenet: Production oriented stream- ing and non-streaming end-to-end speech recognition toolkit,” in 22nd Annual Conference of the International Speech Communi- cation Association, Interspe...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.