REVIEW 2 major objections 5 minor 54 references
Multi-model ASR consensus recovers usable utterance timestamps from long, noisy child-speech recordings and turns CHILDES into training data that cuts out-of-domain child WER by up to 19.5%.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 00:47 UTC pith:XCV47KUP
load-bearing objection Solid engineering release: multi-model consensus turns noisy CHILDES into a usable 283 h child-ASR set with real held-out gains; proxy-only timestamp eval is the main soft spot, not a load-bearing flaw. the 2 major comments →
CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
BEACON recovers reliable utterance onsets and offsets for long-form child speech by aligning multiple off-the-shelf ASR word streams to a trusted CHAT transcript and taking a consensus vote over the resulting candidate spans; the resulting curated CHILDES clips, once quality-filtered, supply transferable supervision that lowers out-of-domain child ASR error.
What carries the argument
BEACON (Boundary Estimation via Alignment CONsensus): multi-model inference of word-level timestamps, per-model alignment of those timestamps to the reference utterance sequence via search-window determination, candidate generation, and monotonic DP path selection, followed by majority-consensus fusion (or abstention) of the per-model time intervals.
Load-bearing premise
The claim that lower word-error rate of a fixed external ASR on the cut clips is a faithful proxy for timestamp accuracy, without any direct measurement of onset or offset error against human ground-truth times.
What would settle it
Manually time-align a representative sample of CHILDES utterances to gold onset/offset labels and check whether BEACON's consensus boundaries show lower absolute timing error than the original CHILDES timestamps, BatchAlign2, and FASA on that same sample.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces BEACON, a corpus-agnostic pipeline that recovers utterance-level timestamps for long-form audio paired with a trusted transcript by (i) decoding with four diverse off-the-shelf ASR systems, (ii) aligning each model’s word stream to the reference via chunked edit-distance search and monotonic DP, and (iii) fusing per-model spans by majority consensus (VERIFIED / RECOVERED / UNRESOLVED). Applied to English CHILDES, the authors release a 413-hour general-purpose set that preserves raw CHAT annotations and a stricter 283-hour ASR-training subset with verbatim normalization and insertion/deletion filtering. Timestamp quality is assessed via a bounded-WER proxy on clips cut by each method (Table II); downstream usefulness is shown by fine-tuning Whisper, Parakeet, and Canary on the ASR subset and evaluating on four non-overlapping child benchmarks (RSR, MyST, OGI Kids, CMU Kids), reporting average relative WER reductions of 19.5%, 3.6%, and 5.6% respectively versus zero-shot (Table III).
Significance. Child ASR remains data-scarce; a large, publicly released, speaker-attributed CHILDES-derived resource with corrected timestamps is of clear practical value. The multi-model consensus design is a sensible, reusable recipe for any long-form corpus with trusted text but unreliable times, and the authors ship code and curated data. The downstream experiment is well controlled: three model families, a high-precision low-yield FASA baseline, and four held-out benchmarks that do not overlap training. If the released clips remain usable supervised pairs—as Table III indicates—the work supplies a concrete training resource and a general timestamp-curation method that other labs can apply. Strengths to credit explicitly: public data/code release, multi-family OOD evaluation, and an independent proxy ASR (Granite) outside the ensemble.
major comments (2)
- [Section V-A, Table II] Section V-A and Table II: the claim that recovered timestamps are “more accurate” rests solely on a bounded-WER proxy from a fixed external ASR. Lower proxy WER is consistent with better audio–text agreement but does not measure absolute onset/offset error against human-annotated times; any systematic bias shared by the proxy and the ensemble could inflate apparent quality. This is load-bearing for the accuracy narrative of BEACON itself (as distinct from the downstream usefulness claim in V-B). The manuscript should either (a) report a small human re-annotation sample with absolute boundary error, or (b) substantially strengthen the caveats so that Table II is framed only as a relative ranking of clip usability, not as absolute timestamp accuracy.
- [Section III-B/C, Section VI] Section III-B/C and Section VI: free parameters (L_chunk, τ_wer, ε/θ/κ, P_miss, p_tok, w_jmp, insertion/deletion threshold 0.25, etc.) are set by light manual inspection with no ablation or sensitivity analysis. The Limitations section acknowledges this, but the operating point is still presented as the default recipe. Because yield, RECOVERED/UNRESOLVED rates, and the ASR-subset size depend on these choices, a minimal sensitivity study (or at least reporting status-label counts and yield under a few nearby settings) is needed to support the claim that BEACON is a robust, transferable curation pipeline rather than a single tuned run on CHILDES.
minor comments (5)
- [Abstract, Table I] Abstract and Table I: the ASR subset is variously “283-hour” / “282.6 h”; keep a single rounded figure consistently.
- [Figure 1] Figure 1 is dense; the three-stage flow (inference → per-model alignment → ensemble voting) would be clearer with a short caption walk-through of one example utterance (e.g., “want to play”).
- [Section III-C, Eq. (6)] Eq. (6): the agreement rule mixes absolute endpoint tolerance and IoU-with-center; a one-sentence intuition for why both clauses are needed would help readers implement the consensus.
- [Table III] Table III: report absolute hours retained after the ins/del filter for FASA vs. BEACON side-by-side in the table header so yield differences are immediately visible when reading ΔWER.
- [Section II] Related Work: briefly note how BEACON differs from ROVER-style hypothesis combination (Fiscus 1997), which operates on word sequences rather than time-boundary votes, to avoid conflating the two ensemble traditions.
Circularity Check
No significant circularity: empirical curation pipeline evaluated on held-out benchmarks with an external proxy ASR; no claim reduces to its inputs by construction.
full rationale
BEACON is an engineering pipeline (multi-ASR word streams → per-model edit-distance/DP alignment → consensus voting) applied to trusted CHAT transcripts; it does not claim a first-principles derivation whose output is forced by its inputs. Timestamp quality is scored with a fixed external ASR (Granite) never used in the ensemble (Section V-A). Downstream usefulness is measured by fine-tuning three off-the-shelf models and evaluating WER on four non-overlapping child benchmarks (RSR, MyST, OGI Kids, CMU Kids; Table III). Relative WER reductions are experimental outcomes, not fitted constants renamed as predictions. Hyperparameters are set by light inspection and acknowledged as unablated (Section VI); that is a limitation, not circularity. Mild reuse of Parakeet/Canary both as ensemble members and as fine-tuning bases, and of Qwen3-ASR for the post-hoc ins/del filter, is ordinary transfer/filtering practice and does not make the reported ΔWER true by definition. No self-definitional equation, no uniqueness theorem imported from the authors, and no load-bearing self-citation chain was found. Score 0 is therefore appropriate.
Axiom & Free-Parameter Ledger
free parameters (5)
- L_chunk =
100
- τ_wer =
0.75
- ε, θ, κ (ensemble agreement) =
0.5 s / 0.9 / 0.25
- P_miss, p_tok, w_jmp, etc. =
see footnote 6
- insertion/deletion filter 0.25 =
0.25
axioms (3)
- domain assumption Word streams from diverse ASR systems are sufficiently independent that majority consensus reduces model-specific boundary error.
- domain assumption Speaker-turn order in the reference transcript is approximately monotonic in the decoded word stream (soft turn-order prior).
- ad hoc to paper Lower bounded WER of a fixed external ASR on cut clips is a valid relative ranking of timestamp quality.
invented entities (1)
-
BEACON consensus status labels (VERIFIED / RECOVERED / UNRESOLVED)
no independent evidence
read the original abstract
CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and language development. However, utterance-level timestamps provided in this corpus are often noisy, incomplete, or misaligned with the audio. As a result, utterances cannot always be reliably localized within long recordings, which limits the direct use of these data for training and evaluating speech models. In this work, we propose BEACON (Boundary Estimation via Alignment CONsensus), an ensemble timestamp-curation framework that refines utterance-level timestamps by aggregating knowledge from multiple off-the-shelf ASR models. Specifically, each model's word-level timestamp predictions are first aligned to provided human transcripts, and the final utterance time boundaries are determined by a consensus voting strategy. The framework is corpus-agnostic and applies to any long-form recording paired with a trusted transcript whose timestamps are unreliable or missing, offering a general recipe for timestamp curation. Leveraging this pipeline, we curate and release a 413-hour general-purpose child-speech dataset with corrected utterance-level timestamps, together with a 283-hour quality-controlled subset for ASR training. Fine-tuning on this subset yields up to an average 19.5% relative WER reduction on four out-of-domain child-speech benchmarks.
Figures
Reference graph
Works this paper leans on
-
[1]
Benchmarking children’s asr with supervised and self-supervised speech foundation models,
R. Fan, N. B. Shankar, and A. Alwan, “Benchmarking children’s asr with supervised and self-supervised speech foundation models,”arXiv preprint arXiv:2406.10507, 2024
Pith/arXiv arXiv 2024
-
[2]
Automatic speech recognition (asr) systems for children: A systematic literature review,
V . Bhardwaj, M. T. Ben Othman, V . Kukreja, Y . Belkhier, M. Bajaj, B. S. Goud, A. U. Rehman, M. Shafiq, and H. Hamam, “Automatic speech recognition (asr) systems for children: A systematic literature review,” Applied Sciences, vol. 12, no. 9, p. 4419, 2022
2022
-
[3]
X. Shi, X. Wang, Z. Guo, Y . Wang, P. Zhang, X. Zhang, Z. Guo, H. Hao, Y . Xi, B. Yang, J. Xu, J. Zhou, and J. Lin, “Qwen3-asr technical report,”arXiv preprint arXiv:2601.21337, 2026. [Online]. Available: https://arxiv.org/abs/2601.21337
Pith/arXiv arXiv 2026
-
[4]
M. Sekoyan, N. R. Koluguri, N. Tadevosyan, P. Zelasko, T. Bartley, N. Karpov, J. Balam, and B. Ginsburg, “Canary-1b-v2 & parakeet-tdt- 0.6b-v3: Efficient and high-performance models for multilingual asr and ast,”arXiv preprint arXiv:2509.14128, 2025. [Online]. Available: https://arxiv.org/abs/2509.14128
arXiv 2025
-
[5]
Robust speech recognition via large-scale weak supervi- sion,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervi- sion,” inInternational conference on machine learning. PMLR, 2023, pp. 28 492–28 518
2023
-
[6]
Exploring the effect of differences in the acoustic correlates of adults’ and children’s speech in the context of automatic speech recognition,
S. Ghai and R. Sinha, “Exploring the effect of differences in the acoustic correlates of adults’ and children’s speech in the context of automatic speech recognition,”EURASIP Journal on Audio, Speech, and Music Processing, vol. 2010, no. 1, p. 318785, 2010
2010
-
[7]
A review of asr technologies for children’s speech,
M. Gerosa, D. Giuliani, S. Narayanan, and A. Potamianos, “A review of asr technologies for children’s speech,” inProceedings of the 2nd Workshop on Child, Computer and Interaction, 2009, pp. 1–8
2009
-
[8]
Transfer learning from adult to chil- dren for speech recognition: Evaluation, analysis and recommendations,
P. G. Shivakumar and P. Georgiou, “Transfer learning from adult to chil- dren for speech recognition: Evaluation, analysis and recommendations,” Computer speech & language, vol. 63, p. 101077, 2020
2020
-
[9]
Speech production variability in fricatives of children and adults: Results of functional data analysis,
L. L. Koenig, J. C. Lucero, and E. Perlman, “Speech production variability in fricatives of children and adults: Results of functional data analysis,”The Journal of the Acoustical Society of America, vol. 124, no. 5, pp. 3158–3170, 2008
2008
-
[10]
Stop consonant voicing and intraoral pressure contours in women and children,
L. L. Koenig and J. C. Lucero, “Stop consonant voicing and intraoral pressure contours in women and children,”The Journal of the Acoustical Society of America, vol. 123, no. 2, pp. 1077–1088, 2008
2008
-
[11]
Analysis of children’s speech: Duration, pitch and formants,
S. Lee, A. Potamianos, and S. Narayanan, “Analysis of children’s speech: Duration, pitch and formants,” inProc. Eurospeech 1997, 1997, pp. 473–476
1997
-
[12]
Acoustics of children’s speech: Developmental changes of tem- poral and spectral parameters,
——, “Acoustics of children’s speech: Developmental changes of tem- poral and spectral parameters,”The Journal of the Acoustical Society of America, vol. 105, no. 3, pp. 1455–1468, 1999
1999
-
[13]
V owel acoustic space development in children: A synthesis of acoustic and anatomic data,
H. K. V orperian and R. D. Kent, “V owel acoustic space development in children: A synthesis of acoustic and anatomic data,”Journal of Speech, Language, and Hearing Research, vol. 50, no. 6, pp. 1510–1545, 2007
2007
-
[14]
Ticl: Text- embedding knn for speech in-context learning unlocks speech recogni- tion abilities of large multimodal models,
H. Zheng, Y . Yegorova, and M. Hasegawa-Johnson, “Ticl: Text- embedding knn for speech in-context learning unlocks speech recogni- tion abilities of large multimodal models,” inICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 17 912–17 916
2026
-
[15]
Ticl+: A case study on speech in-context learning for children’s speech recognition,
——, “Ticl+: A case study on speech in-context learning for children’s speech recognition,”arXiv preprint arXiv:2512.18263, 2025
arXiv 2025
-
[16]
Sicl-at: An- other way to adapt auditory llm to low-resource task,
H. Zheng, S. Wang, Z. Jin, and M. Hasegawa-Johnson, “Sicl-at: An- other way to adapt auditory llm to low-resource task,”arXiv preprint arXiv:2601.18904, 2026
Pith/arXiv arXiv 2026
-
[17]
Fsa- grpo: Teaching auditory llms to use few-shot demonstrations,
H. Zheng, S. Wang, X. Fan, Z. Jin, and M. Hasegawa-Johnson, “Fsa- grpo: Teaching auditory llms to use few-shot demonstrations,”arXiv preprint arXiv:2606.02615, 2026
Pith/arXiv arXiv 2026
-
[18]
Benchmarking training paradigms, dataset composition, and model scaling for child asr in espnet,
A. Ying, N. B. Shankar, C.-J. Lin, M. Shi, P. Wang, H.-j. Shim, S. Arora, A. Alwan, S. Watanabeet al., “Benchmarking training paradigms, dataset composition, and model scaling for child asr in espnet,”arXiv preprint arXiv:2508.16576, 2025
Pith/arXiv arXiv 2025
-
[19]
Age-aware adapter tuning for children’s speech recognition,
J. Li, “Age-aware adapter tuning for children’s speech recognition,” arXiv preprint arXiv:2606.05440, 2026
Pith/arXiv arXiv 2026
-
[20]
Adapta- tion of whisper models to child speech recognition,
R. Jain, A. Barcovschi, M. Yiwere, P. Corcoran, and H. Cucu, “Adapta- tion of whisper models to child speech recognition,” inProc. Interspeech 2023, 2023, pp. 5242–5246
2023
-
[21]
Kid- whisper: Towards bridging the performance gap in automatic speech recognition for children vs. adults,
A. A. Attia, J. Liu, W. Ai, D. Demszky, and C. Espy-Wilson, “Kid- whisper: Towards bridging the performance gap in automatic speech recognition for children vs. adults,” inProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, vol. 7, no. 1, 2024, pp. 74–80
2024
-
[22]
Sparsely shared lora on whisper for child speech recognition,
W. Liu, Y . Qin, Z. Peng, and T. Lee, “Sparsely shared lora on whisper for child speech recognition,” inICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 11 751–11 755
2024
-
[23]
Mind the shift: Using delta ssl embeddings to enhance child asr,
Z. Wang, N. B. Shankar, K. Zhang, Z. Wang, and A. Alwan, “Mind the shift: Using delta ssl embeddings to enhance child asr,”arXiv preprint arXiv:2601.20142, 2026
arXiv 2026
-
[24]
R. Fan and A. Alwan, “Draft: A novel framework to reduce domain shifting in self-supervised learning and its application to children’s asr,” arXiv preprint arXiv:2206.07931, 2022
Pith/arXiv arXiv 2022
-
[25]
Towards better domain adaptation for self-supervised models: A case study of child asr,
R. Fan, Y . Zhu, J. Wang, and A. Alwan, “Towards better domain adaptation for self-supervised models: A case study of child asr,”IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1242–1252, 2022
2022
-
[26]
A wav2vec2-based experimental study on self-supervised learning methods to improve child speech recognition,
R. Jain, A. Barcovschi, M. Y . Yiwere, D. Bigioi, P. Corcoran, and H. Cucu, “A wav2vec2-based experimental study on self-supervised learning methods to improve child speech recognition,”IEEE Access, vol. 11, pp. 46 938–46 948, 2023
2023
-
[27]
MacWhinney,The CHILDES Project: Tools for Analyzing Talk, 3rd ed
B. MacWhinney,The CHILDES Project: Tools for Analyzing Talk, 3rd ed. Mahwah, NJ: Lawrence Erlbaum Associates, 2000
2000
-
[28]
Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,” inProc. INTERSPEECH 2017, 2017, pp. 498–502
2017
-
[29]
Scaling speech tech- nology to 1,000+ languages,
V . Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, A. Elkahky, Z. Ni, A. Vyas, M. Fazel-Zarandi, A. Baevski, Y . Adi, X. Zhang, W.-N. Hsu, A. Conneau, and M. Auli, “Scaling speech tech- nology to 1,000+ languages,”Journal of Machine Learning Research, vol. 25, no. 97, pp. 1–52, 2024, arXiv:2305.13516
Pith/arXiv arXiv 2024
-
[30]
A recursive algorithm for the forced alignment of very long audio segments,
P. J. Moreno, C. Joerg, J.-M. Van Thong, and O. Glickman, “A recursive algorithm for the forced alignment of very long audio segments,” in Proc. 5th International Conference on Spoken Language Processing (ICSLP), 1998
1998
-
[31]
CTC- Segmentation of large corpora for german end-to-end speech recogni- tion,
L. K ¨urzinger, D. Winkelbauer, L. Li, T. Watzel, and G. Rigoll, “CTC- Segmentation of large corpora for german end-to-end speech recogni- tion,” inSpeech and Computer (SPECOM 2020), ser. Lecture Notes in Computer Science, vol. 12335. Springer, 2020, pp. 267–278
2020
-
[32]
ALISA: An automatic lightly supervised speech segmentation and alignment tool,
A. Stan, Y . Mamiya, J. Yamagishi, P. Bell, O. Watts, R. A. J. Clark, and S. King, “ALISA: An automatic lightly supervised speech segmentation and alignment tool,”Computer Speech & Language, vol. 35, pp. 116– 133, 2016
2016
-
[33]
A linear memory CTC-based algorithm for text-to-voice alignment of very long audio recordings,
G. Doras, Y . Teytaut, and A. Roebel, “A linear memory CTC-based algorithm for text-to-voice alignment of very long audio recordings,” Applied Sciences, vol. 13, no. 3, p. 1854, 2023
2023
-
[34]
Whisperx: Time- accurate speech transcription of long-form audio,
M. Bain, J. Huh, T. Han, and A. Zisserman, “Whisperx: Time- accurate speech transcription of long-form audio,” inProceedings of INTERSPEECH, 2023, pp. 4489–4493. [Online]. Available: https://arxiv.org/abs/2303.00747
Pith/arXiv arXiv 2023
-
[35]
D. Liu and J. Xiong, “Fasa: a flexible and automatic speech aligner for extracting high-quality aligned children speech data,”arXiv preprint arXiv:2406.17926, 2024
Pith/arXiv arXiv 2024
-
[36]
Kids: a database of children’s speech,
M. S. Eskenazi, “Kids: a database of children’s speech,” Ph.D. disser- tation, Acoustical Society of America, 1996
1996
-
[37]
The ogi kids’ speech corpus and recognizers,
K. Shobaki, J.-P. Hosom, and R. Cole, “The ogi kids’ speech corpus and recognizers,” inProc. of ICSLP. Citeseer, 2000, pp. 564–567
2000
-
[38]
Multi-task learning for speech attribute detection of children’s speech
M. Shahin, B. Ahmed, and J. Epps, “Multi-task learning for speech attribute detection of children’s speech.”
-
[39]
Tidigits,
R. G. Leonard and G. Doddington, “Tidigits,” 1993
1993
-
[40]
Diagnostic accuracy of sentence recall and past tense measures for identifying children’s language impairments,
S. M. Redmond, A. C. Ash, T. T. Christopulos, and T. Pfaff, “Diagnostic accuracy of sentence recall and past tense measures for identifying children’s language impairments,”Journal of Speech, Language, and Hearing Research, vol. 62, no. 7, pp. 2438–2454, 2019
2019
-
[41]
My science tutor (myst)–a large corpus of children’s conversational speech,
S. Pradhan, R. Cole, and W. Ward, “My science tutor (myst)–a large corpus of children’s conversational speech,” inProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), 2024, pp. 12 040– 12 045
2024
-
[42]
Children’s speech recognition with application to interactive books and tutors,
A. Hagen, B. Pellom, and R. Cole, “Children’s speech recognition with application to interactive books and tutors,” in2003 IEEE Workshop on Automatic Speech Recognition and Understanding (IEEE Cat. No. 03EX721). IEEE, 2003, pp. 186–191
2003
-
[43]
Word-minimality, epenthesis and coda licensing in the early acquisition of english,
K. Demuth, J. Culbertson, and J. Alter, “Word-minimality, epenthesis and coda licensing in the early acquisition of english,”Language and speech, vol. 49, no. 2, pp. 137–173, 2006
2006
-
[44]
The pf-star british english childrens speech corpus,
M. Russell, “The pf-star british english childrens speech corpus,”The Speech Ark Limited, 2006
2006
-
[45]
The pf star children’s speech corpus,
A. Batliner, M. Blomberg, S. D’Arcy, D. Elenius, D. Giuliani, M. Gerosa, C. Hacker, M. Russell, S. Steidl, and M. Wong, “The pf star children’s speech corpus,” 2005
2005
-
[46]
Tball data collection: the making of a young children’s speech corpus,
A. Kazemzadeh, H. You, M. Iseli, B. Jones, X. Cui, M. Heritage, P. Price, E. Anderson, S. Narayanan, and A. Alwan, “Tball data collection: the making of a young children’s speech corpus,” inProc. Interspeech 2005, 2005, pp. 1581–1584
2005
-
[47]
The kaldi speech recognition toolkit,
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y . Qian, P. Schwarzet al., “The kaldi speech recognition toolkit,” inIEEE 2011 workshop on automatic speech recognition and understanding. IEEE Signal Processing Society, 2011
2011
-
[48]
Automation of Language Sample Analysis,
H. Liu, B. MacWhinney, D. Fromm, and A. Lanzi, “Automation of Language Sample Analysis,”Journal of Speech, Language, and Hearing Research, vol. 66, no. 7, pp. 2421–2433, Jul. 2023
2023
-
[49]
Sailalign: Robust long speech-text alignment,
A. Katsamanis, M. P. Black, P. G. Georgiou, L. Goldstein, and S. Narayanan, “Sailalign: Robust long speech-text alignment,” inWork- shop on New Tools and Methods for Very-Large-Scale Phonetics Re- search (VLSRP), 2011
2011
-
[50]
Automatic long audio alignment and confidence scoring for conversational arabic speech,
M. Elmahdy, M. Hasegawa-Johnson, and E. Mustafawi, “Automatic long audio alignment and confidence scoring for conversational arabic speech,” inProceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14). European Language Resources Association (ELRA), 2014, pp. 3062–3066
2014
-
[51]
A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (rover),
J. G. Fiscus, “A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (rover),” inIEEE Workshop on Automatic Speech Recognition and Understanding (ASRU). IEEE, 1997, pp. 347–354
1997
-
[52]
Ensemble methods in machine learning,
T. G. Dietterich, “Ensemble methods in machine learning,” inInterna- tional workshop on multiple classifier systems. Springer, 2000, pp. 1–15
2000
-
[53]
Efficient sequence transduction by jointly predicting tokens and durations,
H. Xu, F. Jia, S. Majumdar, H. Huang, S. Watanabe, and B. Ginsburg, “Efficient sequence transduction by jointly predicting tokens and durations,” inProceedings of the 40th International Conference on Machine Learning. PMLR, 2023, pp. 38 462–38 484. [Online]. Available: https://arxiv.org/abs/2304.06795
Pith/arXiv arXiv 2023
-
[54]
Granite-speech: Open-source speech-aware llms with strong english asr capabilities,
G. Saon, A. Dekel, A. Brooks, T. Nagano, A. Daniels, A. Satt, A. Mittal, B. Kingsbury, D. Haws, E. Moraiset al., “Granite-speech: Open-source speech-aware llms with strong english asr capabilities,”arXiv preprint arXiv:2505.08699, 2025
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.