Pith. sign in

REVIEW 2 major objections 5 minor 54 references

Multi-model ASR consensus recovers usable utterance timestamps from long, noisy child-speech recordings and turns CHILDES into training data that cuts out-of-domain child WER by up to 19.5%.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 00:47 UTC pith:XCV47KUP

load-bearing objection Solid engineering release: multi-model consensus turns noisy CHILDES into a usable 283 h child-ASR set with real held-out gains; proxy-only timestamp eval is the main soft spot, not a load-bearing flaw. the 2 major comments →

arxiv 2607.03670 v1 pith:XCV47KUP submitted 2026-07-04 eess.AS eess.SP

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

classification eess.AS eess.SP
keywords child speechautomatic speech recognitiontimestamp curationensemble alignmentCHILDESlong-form audioforced alignmentdataset release
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Long naturalistic child-adult recordings in CHILDES come with trusted human transcripts but often with noisy, missing, or misaligned utterance timestamps, so the audio cannot be cut into clean clips for speech-model training. The authors introduce BEACON, a fully automatic pipeline that runs several diverse off-the-shelf ASR systems, aligns each system's word-level timestamps back to the trusted transcript, and fuses the resulting candidate boundaries by majority consensus (or abstains when models disagree). Applied to English CHILDES, the pipeline yields a 413-hour general-purpose release that keeps the original raw CHAT annotations and a stricter 283-hour ASR-ready subset with verbatim-normalized transcripts and extra audio-text agreement filters. Fine-tuning three different ASR families on the ASR subset improves word-error rates on four held-out child-speech benchmarks that never appear in the training data, with the largest average relative reduction reaching 19.5%. The same recipe is presented as corpus-agnostic: any long recording paired with a trusted transcript can be timestamp-curated without new human annotation.

Core claim

BEACON recovers reliable utterance onsets and offsets for long-form child speech by aligning multiple off-the-shelf ASR word streams to a trusted CHAT transcript and taking a consensus vote over the resulting candidate spans; the resulting curated CHILDES clips, once quality-filtered, supply transferable supervision that lowers out-of-domain child ASR error.

What carries the argument

BEACON (Boundary Estimation via Alignment CONsensus): multi-model inference of word-level timestamps, per-model alignment of those timestamps to the reference utterance sequence via search-window determination, candidate generation, and monotonic DP path selection, followed by majority-consensus fusion (or abstention) of the per-model time intervals.

Load-bearing premise

The claim that lower word-error rate of a fixed external ASR on the cut clips is a faithful proxy for timestamp accuracy, without any direct measurement of onset or offset error against human ground-truth times.

What would settle it

Manually time-align a representative sample of CHILDES utterances to gold onset/offset labels and check whether BEACON's consensus boundaries show lower absolute timing error than the original CHILDES timestamps, BatchAlign2, and FASA on that same sample.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces BEACON, a corpus-agnostic pipeline that recovers utterance-level timestamps for long-form audio paired with a trusted transcript by (i) decoding with four diverse off-the-shelf ASR systems, (ii) aligning each model’s word stream to the reference via chunked edit-distance search and monotonic DP, and (iii) fusing per-model spans by majority consensus (VERIFIED / RECOVERED / UNRESOLVED). Applied to English CHILDES, the authors release a 413-hour general-purpose set that preserves raw CHAT annotations and a stricter 283-hour ASR-training subset with verbatim normalization and insertion/deletion filtering. Timestamp quality is assessed via a bounded-WER proxy on clips cut by each method (Table II); downstream usefulness is shown by fine-tuning Whisper, Parakeet, and Canary on the ASR subset and evaluating on four non-overlapping child benchmarks (RSR, MyST, OGI Kids, CMU Kids), reporting average relative WER reductions of 19.5%, 3.6%, and 5.6% respectively versus zero-shot (Table III).

Significance. Child ASR remains data-scarce; a large, publicly released, speaker-attributed CHILDES-derived resource with corrected timestamps is of clear practical value. The multi-model consensus design is a sensible, reusable recipe for any long-form corpus with trusted text but unreliable times, and the authors ship code and curated data. The downstream experiment is well controlled: three model families, a high-precision low-yield FASA baseline, and four held-out benchmarks that do not overlap training. If the released clips remain usable supervised pairs—as Table III indicates—the work supplies a concrete training resource and a general timestamp-curation method that other labs can apply. Strengths to credit explicitly: public data/code release, multi-family OOD evaluation, and an independent proxy ASR (Granite) outside the ensemble.

major comments (2)
  1. [Section V-A, Table II] Section V-A and Table II: the claim that recovered timestamps are “more accurate” rests solely on a bounded-WER proxy from a fixed external ASR. Lower proxy WER is consistent with better audio–text agreement but does not measure absolute onset/offset error against human-annotated times; any systematic bias shared by the proxy and the ensemble could inflate apparent quality. This is load-bearing for the accuracy narrative of BEACON itself (as distinct from the downstream usefulness claim in V-B). The manuscript should either (a) report a small human re-annotation sample with absolute boundary error, or (b) substantially strengthen the caveats so that Table II is framed only as a relative ranking of clip usability, not as absolute timestamp accuracy.
  2. [Section III-B/C, Section VI] Section III-B/C and Section VI: free parameters (L_chunk, τ_wer, ε/θ/κ, P_miss, p_tok, w_jmp, insertion/deletion threshold 0.25, etc.) are set by light manual inspection with no ablation or sensitivity analysis. The Limitations section acknowledges this, but the operating point is still presented as the default recipe. Because yield, RECOVERED/UNRESOLVED rates, and the ASR-subset size depend on these choices, a minimal sensitivity study (or at least reporting status-label counts and yield under a few nearby settings) is needed to support the claim that BEACON is a robust, transferable curation pipeline rather than a single tuned run on CHILDES.
minor comments (5)
  1. [Abstract, Table I] Abstract and Table I: the ASR subset is variously “283-hour” / “282.6 h”; keep a single rounded figure consistently.
  2. [Figure 1] Figure 1 is dense; the three-stage flow (inference → per-model alignment → ensemble voting) would be clearer with a short caption walk-through of one example utterance (e.g., “want to play”).
  3. [Section III-C, Eq. (6)] Eq. (6): the agreement rule mixes absolute endpoint tolerance and IoU-with-center; a one-sentence intuition for why both clauses are needed would help readers implement the consensus.
  4. [Table III] Table III: report absolute hours retained after the ins/del filter for FASA vs. BEACON side-by-side in the table header so yield differences are immediately visible when reading ΔWER.
  5. [Section II] Related Work: briefly note how BEACON differs from ROVER-style hypothesis combination (Fiscus 1997), which operates on word sequences rather than time-boundary votes, to avoid conflating the two ensemble traditions.

Circularity Check

0 steps flagged

No significant circularity: empirical curation pipeline evaluated on held-out benchmarks with an external proxy ASR; no claim reduces to its inputs by construction.

full rationale

BEACON is an engineering pipeline (multi-ASR word streams → per-model edit-distance/DP alignment → consensus voting) applied to trusted CHAT transcripts; it does not claim a first-principles derivation whose output is forced by its inputs. Timestamp quality is scored with a fixed external ASR (Granite) never used in the ensemble (Section V-A). Downstream usefulness is measured by fine-tuning three off-the-shelf models and evaluating WER on four non-overlapping child benchmarks (RSR, MyST, OGI Kids, CMU Kids; Table III). Relative WER reductions are experimental outcomes, not fitted constants renamed as predictions. Hyperparameters are set by light inspection and acknowledged as unablated (Section VI); that is a limitation, not circularity. Mild reuse of Parakeet/Canary both as ensemble members and as fine-tuning bases, and of Qwen3-ASR for the post-hoc ins/del filter, is ordinary transfer/filtering practice and does not make the reported ΔWER true by definition. No self-definitional equation, no uniqueness theorem imported from the authors, and no load-bearing self-citation chain was found. Score 0 is therefore appropriate.

Axiom & Free-Parameter Ledger

5 free parameters · 3 axioms · 1 invented entities

The central claim rests on a large set of hand-chosen alignment and voting thresholds plus domain assumptions that ASR word streams roughly preserve turn order and that majority agreement implies correct boundaries. No new physical entities are postulated; the free parameters are purely algorithmic.

free parameters (5)
  • L_chunk = 100
    Chunk size (100 tokens) for coarse search-window determination; chosen empirically, never ablated.
  • τ_wer = 0.75
    WER threshold (0.75) that admits a span as a candidate; controls precision/recall of local matches.
  • ε, θ, κ (ensemble agreement) = 0.5 s / 0.9 / 0.25
    Endpoint tolerance 0.5 s, IoU 0.9, duration scale 0.25 that decide whether two model spans 'agree'; set by light manual inspection.
  • P_miss, p_tok, w_jmp, etc. = see footnote 6
    Full set of DP path costs (P_miss=1.1, p_tok=0.02, w_jmp=0.0188, …) listed in footnote 6; all empirical.
  • insertion/deletion filter 0.25 = 0.25
    ASR-training subset drops clips whose ins or del rate exceeds 0.25; threshold chosen without sensitivity study.
axioms (3)
  • domain assumption Word streams from diverse ASR systems are sufficiently independent that majority consensus reduces model-specific boundary error.
    Invoked throughout §III-C and the design of the four-model ensemble; never proved, only motivated by architectural diversity.
  • domain assumption Speaker-turn order in the reference transcript is approximately monotonic in the decoded word stream (soft turn-order prior).
    Used to justify the jump-penalty term (Eq. 3) and the Viterbi path; acknowledged as approximate in §III-B.
  • ad hoc to paper Lower bounded WER of a fixed external ASR on cut clips is a valid relative ranking of timestamp quality.
    Core evaluation premise of §V-A; no absolute boundary ground truth is available.
invented entities (1)
  • BEACON consensus status labels (VERIFIED / RECOVERED / UNRESOLVED) no independent evidence
    purpose: Categorize each utterance according to whether the original timestamp is corroborated, a new majority span is adopted, or no reliable majority exists.
    Purely definitional bookkeeping; no independent physical existence claimed.

pith-pipeline@v1.1.0-grok45 · 19714 in / 2826 out tokens · 24252 ms · 2026-07-12T00:47:33.234876+00:00 · methodology

0 comments
read the original abstract

CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and language development. However, utterance-level timestamps provided in this corpus are often noisy, incomplete, or misaligned with the audio. As a result, utterances cannot always be reliably localized within long recordings, which limits the direct use of these data for training and evaluating speech models. In this work, we propose BEACON (Boundary Estimation via Alignment CONsensus), an ensemble timestamp-curation framework that refines utterance-level timestamps by aggregating knowledge from multiple off-the-shelf ASR models. Specifically, each model's word-level timestamp predictions are first aligned to provided human transcripts, and the final utterance time boundaries are determined by a consensus voting strategy. The framework is corpus-agnostic and applies to any long-form recording paired with a trusted transcript whose timestamps are unreliable or missing, offering a general recipe for timestamp curation. Leveraging this pipeline, we curate and release a 413-hour general-purpose child-speech dataset with corrected utterance-level timestamps, together with a 283-hour quality-controlled subset for ASR training. Fine-tuning on this subset yields up to an average 19.5% relative WER reduction on four out-of-domain child-speech benchmarks.

Figures

Figures reproduced from arXiv: 2607.03670 by Brian Kingsbury, Dancheng Liu, Haolong Zheng, Jinjun Xiong, Mark A. Hasegawa-Johnson, Samuel Thomas, Vishal Sunder, Xinyu Liang, Yuanzhuo Hu, Zhizheng Wu.

Figure 1
Figure 1. Figure 1: Overview of the BEACON pipeline from numerous independent corpora; however, its long-form recordings often come with noisy timestamps, making them difficult to use for ASR training and evaluation. Also, to the best of our knowledge, many corpora are private and not publicly accessible [42]–[46]. All of these make our large￾scale, high-quality labeled and publicly released dataset more valuable. Curating CH… view at source ↗
Figure 2
Figure 2. Figure 2: Ensemble voting for one utterance. C is the largest block of mutually￾agreeing votes and V the number of valid votes. C. Step 3: Ensemble voting Step 2 yields, for each utterance, up to one span per model; these are the votes. The ensemble decides each utterance’s timestamp by a majority consensus over those spans, or abstains when the models do not agree. For utterance ui , each model m that produced a ma… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

54 extracted references · 12 linked inside Pith

  1. [1]

    Benchmarking children’s asr with supervised and self-supervised speech foundation models,

    R. Fan, N. B. Shankar, and A. Alwan, “Benchmarking children’s asr with supervised and self-supervised speech foundation models,”arXiv preprint arXiv:2406.10507, 2024

  2. [2]

    Automatic speech recognition (asr) systems for children: A systematic literature review,

    V . Bhardwaj, M. T. Ben Othman, V . Kukreja, Y . Belkhier, M. Bajaj, B. S. Goud, A. U. Rehman, M. Shafiq, and H. Hamam, “Automatic speech recognition (asr) systems for children: A systematic literature review,” Applied Sciences, vol. 12, no. 9, p. 4419, 2022

  3. [3]

    Qwen3-asr technical report,

    X. Shi, X. Wang, Z. Guo, Y . Wang, P. Zhang, X. Zhang, Z. Guo, H. Hao, Y . Xi, B. Yang, J. Xu, J. Zhou, and J. Lin, “Qwen3-asr technical report,”arXiv preprint arXiv:2601.21337, 2026. [Online]. Available: https://arxiv.org/abs/2601.21337

  4. [4]

    Canary-1b-v2 & parakeet-tdt- 0.6b-v3: Efficient and high-performance models for multilingual asr and ast,

    M. Sekoyan, N. R. Koluguri, N. Tadevosyan, P. Zelasko, T. Bartley, N. Karpov, J. Balam, and B. Ginsburg, “Canary-1b-v2 & parakeet-tdt- 0.6b-v3: Efficient and high-performance models for multilingual asr and ast,”arXiv preprint arXiv:2509.14128, 2025. [Online]. Available: https://arxiv.org/abs/2509.14128

  5. [5]

    Robust speech recognition via large-scale weak supervi- sion,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervi- sion,” inInternational conference on machine learning. PMLR, 2023, pp. 28 492–28 518

  6. [6]

    Exploring the effect of differences in the acoustic correlates of adults’ and children’s speech in the context of automatic speech recognition,

    S. Ghai and R. Sinha, “Exploring the effect of differences in the acoustic correlates of adults’ and children’s speech in the context of automatic speech recognition,”EURASIP Journal on Audio, Speech, and Music Processing, vol. 2010, no. 1, p. 318785, 2010

  7. [7]

    A review of asr technologies for children’s speech,

    M. Gerosa, D. Giuliani, S. Narayanan, and A. Potamianos, “A review of asr technologies for children’s speech,” inProceedings of the 2nd Workshop on Child, Computer and Interaction, 2009, pp. 1–8

  8. [8]

    Transfer learning from adult to chil- dren for speech recognition: Evaluation, analysis and recommendations,

    P. G. Shivakumar and P. Georgiou, “Transfer learning from adult to chil- dren for speech recognition: Evaluation, analysis and recommendations,” Computer speech & language, vol. 63, p. 101077, 2020

  9. [9]

    Speech production variability in fricatives of children and adults: Results of functional data analysis,

    L. L. Koenig, J. C. Lucero, and E. Perlman, “Speech production variability in fricatives of children and adults: Results of functional data analysis,”The Journal of the Acoustical Society of America, vol. 124, no. 5, pp. 3158–3170, 2008

  10. [10]

    Stop consonant voicing and intraoral pressure contours in women and children,

    L. L. Koenig and J. C. Lucero, “Stop consonant voicing and intraoral pressure contours in women and children,”The Journal of the Acoustical Society of America, vol. 123, no. 2, pp. 1077–1088, 2008

  11. [11]

    Analysis of children’s speech: Duration, pitch and formants,

    S. Lee, A. Potamianos, and S. Narayanan, “Analysis of children’s speech: Duration, pitch and formants,” inProc. Eurospeech 1997, 1997, pp. 473–476

  12. [12]

    Acoustics of children’s speech: Developmental changes of tem- poral and spectral parameters,

    ——, “Acoustics of children’s speech: Developmental changes of tem- poral and spectral parameters,”The Journal of the Acoustical Society of America, vol. 105, no. 3, pp. 1455–1468, 1999

  13. [13]

    V owel acoustic space development in children: A synthesis of acoustic and anatomic data,

    H. K. V orperian and R. D. Kent, “V owel acoustic space development in children: A synthesis of acoustic and anatomic data,”Journal of Speech, Language, and Hearing Research, vol. 50, no. 6, pp. 1510–1545, 2007

  14. [14]

    Ticl: Text- embedding knn for speech in-context learning unlocks speech recogni- tion abilities of large multimodal models,

    H. Zheng, Y . Yegorova, and M. Hasegawa-Johnson, “Ticl: Text- embedding knn for speech in-context learning unlocks speech recogni- tion abilities of large multimodal models,” inICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 17 912–17 916

  15. [15]

    Ticl+: A case study on speech in-context learning for children’s speech recognition,

    ——, “Ticl+: A case study on speech in-context learning for children’s speech recognition,”arXiv preprint arXiv:2512.18263, 2025

  16. [16]

    Sicl-at: An- other way to adapt auditory llm to low-resource task,

    H. Zheng, S. Wang, Z. Jin, and M. Hasegawa-Johnson, “Sicl-at: An- other way to adapt auditory llm to low-resource task,”arXiv preprint arXiv:2601.18904, 2026

  17. [17]

    Fsa- grpo: Teaching auditory llms to use few-shot demonstrations,

    H. Zheng, S. Wang, X. Fan, Z. Jin, and M. Hasegawa-Johnson, “Fsa- grpo: Teaching auditory llms to use few-shot demonstrations,”arXiv preprint arXiv:2606.02615, 2026

  18. [18]

    Benchmarking training paradigms, dataset composition, and model scaling for child asr in espnet,

    A. Ying, N. B. Shankar, C.-J. Lin, M. Shi, P. Wang, H.-j. Shim, S. Arora, A. Alwan, S. Watanabeet al., “Benchmarking training paradigms, dataset composition, and model scaling for child asr in espnet,”arXiv preprint arXiv:2508.16576, 2025

  19. [19]

    Age-aware adapter tuning for children’s speech recognition,

    J. Li, “Age-aware adapter tuning for children’s speech recognition,” arXiv preprint arXiv:2606.05440, 2026

  20. [20]

    Adapta- tion of whisper models to child speech recognition,

    R. Jain, A. Barcovschi, M. Yiwere, P. Corcoran, and H. Cucu, “Adapta- tion of whisper models to child speech recognition,” inProc. Interspeech 2023, 2023, pp. 5242–5246

  21. [21]

    Kid- whisper: Towards bridging the performance gap in automatic speech recognition for children vs. adults,

    A. A. Attia, J. Liu, W. Ai, D. Demszky, and C. Espy-Wilson, “Kid- whisper: Towards bridging the performance gap in automatic speech recognition for children vs. adults,” inProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, vol. 7, no. 1, 2024, pp. 74–80

  22. [22]

    Sparsely shared lora on whisper for child speech recognition,

    W. Liu, Y . Qin, Z. Peng, and T. Lee, “Sparsely shared lora on whisper for child speech recognition,” inICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 11 751–11 755

  23. [23]

    Mind the shift: Using delta ssl embeddings to enhance child asr,

    Z. Wang, N. B. Shankar, K. Zhang, Z. Wang, and A. Alwan, “Mind the shift: Using delta ssl embeddings to enhance child asr,”arXiv preprint arXiv:2601.20142, 2026

  24. [24]

    Draft: A novel framework to reduce domain shifting in self-supervised learning and its application to children’s asr,

    R. Fan and A. Alwan, “Draft: A novel framework to reduce domain shifting in self-supervised learning and its application to children’s asr,” arXiv preprint arXiv:2206.07931, 2022

  25. [25]

    Towards better domain adaptation for self-supervised models: A case study of child asr,

    R. Fan, Y . Zhu, J. Wang, and A. Alwan, “Towards better domain adaptation for self-supervised models: A case study of child asr,”IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1242–1252, 2022

  26. [26]

    A wav2vec2-based experimental study on self-supervised learning methods to improve child speech recognition,

    R. Jain, A. Barcovschi, M. Y . Yiwere, D. Bigioi, P. Corcoran, and H. Cucu, “A wav2vec2-based experimental study on self-supervised learning methods to improve child speech recognition,”IEEE Access, vol. 11, pp. 46 938–46 948, 2023

  27. [27]

    MacWhinney,The CHILDES Project: Tools for Analyzing Talk, 3rd ed

    B. MacWhinney,The CHILDES Project: Tools for Analyzing Talk, 3rd ed. Mahwah, NJ: Lawrence Erlbaum Associates, 2000

  28. [28]

    Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,

    M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,” inProc. INTERSPEECH 2017, 2017, pp. 498–502

  29. [29]

    Scaling speech tech- nology to 1,000+ languages,

    V . Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, A. Elkahky, Z. Ni, A. Vyas, M. Fazel-Zarandi, A. Baevski, Y . Adi, X. Zhang, W.-N. Hsu, A. Conneau, and M. Auli, “Scaling speech tech- nology to 1,000+ languages,”Journal of Machine Learning Research, vol. 25, no. 97, pp. 1–52, 2024, arXiv:2305.13516

  30. [30]

    A recursive algorithm for the forced alignment of very long audio segments,

    P. J. Moreno, C. Joerg, J.-M. Van Thong, and O. Glickman, “A recursive algorithm for the forced alignment of very long audio segments,” in Proc. 5th International Conference on Spoken Language Processing (ICSLP), 1998

  31. [31]

    CTC- Segmentation of large corpora for german end-to-end speech recogni- tion,

    L. K ¨urzinger, D. Winkelbauer, L. Li, T. Watzel, and G. Rigoll, “CTC- Segmentation of large corpora for german end-to-end speech recogni- tion,” inSpeech and Computer (SPECOM 2020), ser. Lecture Notes in Computer Science, vol. 12335. Springer, 2020, pp. 267–278

  32. [32]

    ALISA: An automatic lightly supervised speech segmentation and alignment tool,

    A. Stan, Y . Mamiya, J. Yamagishi, P. Bell, O. Watts, R. A. J. Clark, and S. King, “ALISA: An automatic lightly supervised speech segmentation and alignment tool,”Computer Speech & Language, vol. 35, pp. 116– 133, 2016

  33. [33]

    A linear memory CTC-based algorithm for text-to-voice alignment of very long audio recordings,

    G. Doras, Y . Teytaut, and A. Roebel, “A linear memory CTC-based algorithm for text-to-voice alignment of very long audio recordings,” Applied Sciences, vol. 13, no. 3, p. 1854, 2023

  34. [34]

    Whisperx: Time- accurate speech transcription of long-form audio,

    M. Bain, J. Huh, T. Han, and A. Zisserman, “Whisperx: Time- accurate speech transcription of long-form audio,” inProceedings of INTERSPEECH, 2023, pp. 4489–4493. [Online]. Available: https://arxiv.org/abs/2303.00747

  35. [35]

    Fasa: a flexible and automatic speech aligner for extracting high-quality aligned children speech data,

    D. Liu and J. Xiong, “Fasa: a flexible and automatic speech aligner for extracting high-quality aligned children speech data,”arXiv preprint arXiv:2406.17926, 2024

  36. [36]

    Kids: a database of children’s speech,

    M. S. Eskenazi, “Kids: a database of children’s speech,” Ph.D. disser- tation, Acoustical Society of America, 1996

  37. [37]

    The ogi kids’ speech corpus and recognizers,

    K. Shobaki, J.-P. Hosom, and R. Cole, “The ogi kids’ speech corpus and recognizers,” inProc. of ICSLP. Citeseer, 2000, pp. 564–567

  38. [38]

    Multi-task learning for speech attribute detection of children’s speech

    M. Shahin, B. Ahmed, and J. Epps, “Multi-task learning for speech attribute detection of children’s speech.”

  39. [39]

    Tidigits,

    R. G. Leonard and G. Doddington, “Tidigits,” 1993

  40. [40]

    Diagnostic accuracy of sentence recall and past tense measures for identifying children’s language impairments,

    S. M. Redmond, A. C. Ash, T. T. Christopulos, and T. Pfaff, “Diagnostic accuracy of sentence recall and past tense measures for identifying children’s language impairments,”Journal of Speech, Language, and Hearing Research, vol. 62, no. 7, pp. 2438–2454, 2019

  41. [41]

    My science tutor (myst)–a large corpus of children’s conversational speech,

    S. Pradhan, R. Cole, and W. Ward, “My science tutor (myst)–a large corpus of children’s conversational speech,” inProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), 2024, pp. 12 040– 12 045

  42. [42]

    Children’s speech recognition with application to interactive books and tutors,

    A. Hagen, B. Pellom, and R. Cole, “Children’s speech recognition with application to interactive books and tutors,” in2003 IEEE Workshop on Automatic Speech Recognition and Understanding (IEEE Cat. No. 03EX721). IEEE, 2003, pp. 186–191

  43. [43]

    Word-minimality, epenthesis and coda licensing in the early acquisition of english,

    K. Demuth, J. Culbertson, and J. Alter, “Word-minimality, epenthesis and coda licensing in the early acquisition of english,”Language and speech, vol. 49, no. 2, pp. 137–173, 2006

  44. [44]

    The pf-star british english childrens speech corpus,

    M. Russell, “The pf-star british english childrens speech corpus,”The Speech Ark Limited, 2006

  45. [45]

    The pf star children’s speech corpus,

    A. Batliner, M. Blomberg, S. D’Arcy, D. Elenius, D. Giuliani, M. Gerosa, C. Hacker, M. Russell, S. Steidl, and M. Wong, “The pf star children’s speech corpus,” 2005

  46. [46]

    Tball data collection: the making of a young children’s speech corpus,

    A. Kazemzadeh, H. You, M. Iseli, B. Jones, X. Cui, M. Heritage, P. Price, E. Anderson, S. Narayanan, and A. Alwan, “Tball data collection: the making of a young children’s speech corpus,” inProc. Interspeech 2005, 2005, pp. 1581–1584

  47. [47]

    The kaldi speech recognition toolkit,

    D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y . Qian, P. Schwarzet al., “The kaldi speech recognition toolkit,” inIEEE 2011 workshop on automatic speech recognition and understanding. IEEE Signal Processing Society, 2011

  48. [48]

    Automation of Language Sample Analysis,

    H. Liu, B. MacWhinney, D. Fromm, and A. Lanzi, “Automation of Language Sample Analysis,”Journal of Speech, Language, and Hearing Research, vol. 66, no. 7, pp. 2421–2433, Jul. 2023

  49. [49]

    Sailalign: Robust long speech-text alignment,

    A. Katsamanis, M. P. Black, P. G. Georgiou, L. Goldstein, and S. Narayanan, “Sailalign: Robust long speech-text alignment,” inWork- shop on New Tools and Methods for Very-Large-Scale Phonetics Re- search (VLSRP), 2011

  50. [50]

    Automatic long audio alignment and confidence scoring for conversational arabic speech,

    M. Elmahdy, M. Hasegawa-Johnson, and E. Mustafawi, “Automatic long audio alignment and confidence scoring for conversational arabic speech,” inProceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14). European Language Resources Association (ELRA), 2014, pp. 3062–3066

  51. [51]

    A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (rover),

    J. G. Fiscus, “A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (rover),” inIEEE Workshop on Automatic Speech Recognition and Understanding (ASRU). IEEE, 1997, pp. 347–354

  52. [52]

    Ensemble methods in machine learning,

    T. G. Dietterich, “Ensemble methods in machine learning,” inInterna- tional workshop on multiple classifier systems. Springer, 2000, pp. 1–15

  53. [53]

    Efficient sequence transduction by jointly predicting tokens and durations,

    H. Xu, F. Jia, S. Majumdar, H. Huang, S. Watanabe, and B. Ginsburg, “Efficient sequence transduction by jointly predicting tokens and durations,” inProceedings of the 40th International Conference on Machine Learning. PMLR, 2023, pp. 38 462–38 484. [Online]. Available: https://arxiv.org/abs/2304.06795

  54. [54]

    Granite-speech: Open-source speech-aware llms with strong english asr capabilities,

    G. Saon, A. Dekel, A. Brooks, T. Nagano, A. Daniels, A. Satt, A. Mittal, B. Kingsbury, D. Haws, E. Moraiset al., “Granite-speech: Open-source speech-aware llms with strong english asr capabilities,”arXiv preprint arXiv:2505.08699, 2025