Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet

As of 21 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2508.16576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16576 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:16:31.419659Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:16:31.102623Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:16:31.542417Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2eafadc7-12b5-49dc-bbec-8791a94cbc41 · outbound

This paper cites Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:16:31.551433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.102623Z digest=sha256:1be9b40edb231e4faae8f5ace0c68232c55ed3acd10eaea36575e4db849bb6ae

Observation 74fa6b2a-e06d-405b-88c5-1978a31ba304 · outbound

This paper cites fine-tuning), front-end representations, representation types (continuous vs.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet fine-tuning), front-end representations, representation types (continuous vs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.594125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.108835Z digest=sha256:f3ab31a9491197979a82d92e1ace8389f35feb46fecc76b866d3fc270f40da75

Observation 0994f503-58d8-4df7-a075-5efcf7235c8a · outbound

This paper cites Child speech corpora The experiments are conducted on three child speech corpora summarized in Table 1.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Child speech corpora The experiments are conducted on three child speech corpora summarized in Table 1

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.571858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.115680Z digest=sha256:a8c6ff535ddfbe5c62a592ed0e3388609ccc354dbe92a96a33fa4d4defd7796c

Observation 6445c31f-81b7-4e2a-8b84-eeb5c398b3ae · outbound

This paper cites Flat-start training vs.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Flat-start training vs

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T17:16:32.543442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.121205Z digest=sha256:945a9cf4527ae622f981010f2fb4d56ab5d02629c83d5cd4897d215874330549

Observation 472239de-2e27-4c49-913c-fd43565cd75c · outbound

This paper cites While fine-tuned models generally perform best, flat-start models help mitigate biases in SSL representations, which are predominantly trained on adult speech.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet While fine-tuned models generally perform best, flat-start models help mitigate biases in SSL representations, which are predominantly trained on adult speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.517526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.127321Z digest=sha256:e65bbc282c869196111e983f86d0adb5376f94480d1ac1814bf7383790db3fdf

Observation 5dfb216a-0a63-428b-a632-db515e58a32a · outbound

This paper cites Additional support was provided by FWO-SBO grant S004923N: NEFL, KU Leuven C24M/22/025, and FWO grant V401325N.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Additional support was provided by FWO-SBO grant S004923N: NEFL, KU Leuven C24M/22/025, and FWO grant V401325N

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.487710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.132633Z digest=sha256:c0f1e5602672cd89ccaee08c3c97d60a6748ecc07d9d774c09ba45d5fdd875b3

Observation 3ca39a5d-ac42-416b-bfb5-6d082c70b7a1 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T17:16:31.138386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:16:31.138386Z digest=sha256:815abccce691daf24ab8bab74a06f861ff7cdf2041d40338d8381b231a94cbac

Observation 1c608275-75bb-4a9b-b0da-6fd801d3376b · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.434279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.143121Z digest=sha256:9c4eaa376db41a45e4c6cc58f5e3d5f3083a4241c00a5cf3460da4339b8bdc12

Observation 20aec226-25a8-42af-8994-e4e19a5732d3 · outbound

This paper cites Wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.405360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.148533Z digest=sha256:c76709bbca86fa990d94e14804da8b47fa67c857d9a8ed3448e9e7a552d1c70f

Observation df379917-991e-4840-9d00-441e59491b8a · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Robust speech recognition via large-scale weak supervision,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.368367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.154323Z digest=sha256:1036cb4c2b34367fdcfb6c531b6c6f4286543403a0ced8f220824c8622a4b8bb

Observation fad0f4ce-27fe-4651-be4f-58991921b337 · outbound

This paper cites Owsm v3.1: Better and faster open whisper-style speech models based on e-branchformer,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Owsm v3.1: Better and faster open whisper-style speech models based on e-branchformer,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.347480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.160946Z digest=sha256:fafd18ff2530d40440891abd59d7275c3a90391c13b59b1241530939b6ccd72a

Observation fd70485c-8065-4542-b9d2-7cde717a2f52 · outbound

This paper cites Less is more: Accurate speech recognition & translation without web-scale data,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Less is more: Accurate speech recognition & translation without web-scale data,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.320551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.167905Z digest=sha256:e78fdf355911688421482e944773150240fee759d54288298ad7ddde29205635

Observation 55526a59-4dda-47a1-b8a1-b5289fe3fc4e · outbound

This paper cites Acoustics of children’s speech: Developmental changes of temporal and spectral parameters,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Acoustics of children’s speech: Developmental changes of temporal and spectral parameters,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.302029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.174142Z digest=sha256:47fc53b0722174235ada0378bb8a29e0a4fd499059968fb35d51fa79681317e6

Observation 2cba373b-c6f9-41d9-a4f4-66bdc9a0b23e · outbound

This paper cites On the difficulties of automatic speech recognition for kindergarten-aged children,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet On the difficulties of automatic speech recognition for kindergarten-aged children,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.282771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.182289Z digest=sha256:677696033d2023014c6145cf3f4b0053d899a34c4dd5c64d891694b52340d98f

Observation 721ca13b-b691-4858-bcae-c30ae3bae5dc · outbound

This paper cites Challenges remain in building ASR for sponta- neous preschool children speech in naturalistic educational envi- ronments,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Challenges remain in building ASR for sponta- neous preschool children speech in naturalistic educational envi- ronments,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.262963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.190094Z digest=sha256:4d6098e80dbad5131e19ce064c99dd9d2f9e95f2f948af810c0c0fe8c9e3079c

Observation 9b42ae76-4c44-480f-9976-8758c7177332 · outbound

This paper cites V ocal tract length perturbation (vtlp) improves speech recognition,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet V ocal tract length perturbation (vtlp) improves speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.241891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.196402Z digest=sha256:990ed34d9e4e526f4531df128a67b3f25e31d482a09707e427e7fb3225284e8d

Observation c17c774b-058e-4795-aabb-e42dd1a2eee0 · outbound

This paper cites Prosodic adaptations to pitch perturbation in run- ning speech,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Prosodic adaptations to pitch perturbation in run- ning speech,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.223689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.206230Z digest=sha256:e9f99e223e74992eb9b01631b1292e8f80c0c2dc4e421e55d346a078168fb219

Observation 6ad89029-9f22-40f4-a7a9-f5face5f7dee · outbound

This paper cites V oice Conversion Based Data Aug- mentation to Improve Children’s Speech Recognition in Limited Data Scenario,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet V oice Conversion Based Data Aug- mentation to Improve Children’s Speech Recognition in Limited Data Scenario,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.199171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.213002Z digest=sha256:ff906a97ee84ef943c1fb6ae7139f490626e930c9e0e7751908b96ebe30b5a76

Observation d5cdc2ba-57ac-4600-b839-6d8372153ad2 · outbound

This paper cites Improved children’s automatic speech recognition combining adapters and synthetic data augmentation,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Improved children’s automatic speech recognition combining adapters and synthetic data augmentation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.176800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.217767Z digest=sha256:10f53b6fd0eccc0ac878b28692576213dba18345bd7627f65d057ad9bf6c6b1e

Observation 7d53c064-1cad-4b49-aa08-d4c1440aa027 · outbound

This paper cites Transfer learning from adult to children for speech recognition: Evaluation, analysis and recommendations,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Transfer learning from adult to children for speech recognition: Evaluation, analysis and recommendations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.150832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.222384Z digest=sha256:d688852968fe4c93f43394c839ce88894e9c7ff499e7f90df275f640c6bd10c8

Observation 594556bd-3869-4098-9d78-6f4021ce2dc6 · outbound

This paper cites Benchmarking children’s asr with supervised and self-supervised speech foundation models,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Benchmarking children’s asr with supervised and self-supervised speech foundation models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.129392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.227688Z digest=sha256:33ec0a0e4aa23ba78297382333d47a5f1472147f35a6c48b6ac1783e55aa6e7a

Observation 494033fe-5ea3-4c63-8eed-37988ad64a49 · outbound

This paper cites Kid-whisper: Towards bridging the per- formance gap in automatic speech recognition for children vs. adults,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Kid-whisper: Towards bridging the per- formance gap in automatic speech recognition for children vs. adults,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.105117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.233899Z digest=sha256:bfd0ec55b25c9fc28ecf9eb8a95a089a77db104cb2755ff5e1872217d0aff2d5

Observation 1cca7780-710b-4d60-8ce2-f568aaeffc6c · outbound

This paper cites Ml-superb 2.0: Benchmarking multilingual speech models across modeling constraints, languages, and datasets,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Ml-superb 2.0: Benchmarking multilingual speech models across modeling constraints, languages, and datasets,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.085269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.238912Z digest=sha256:4ff4b1dd48572b0a6833279a5c26b62240e1a19754578af2b407f3300754fa73

Observation 3dbe124c-4e8c-42d6-89e8-63a6dad6700f · outbound

This paper cites On the evaluation of speech foundation models for spoken language understanding,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet On the evaluation of speech foundation models for spoken language understanding,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.062965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.243601Z digest=sha256:25ecf293fba5cba5f0a3b59a2924c283326dc718bc87da0ab7685a059c249221

Observation 773600a9-fd3d-4dbc-b007-063dc4c5e629 · outbound

This paper cites Speech self-supervised representations bench- marking: a case for larger probing heads,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Speech self-supervised representations bench- marking: a case for larger probing heads,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.043459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.250018Z digest=sha256:d9fdb9c1a9d88d915aeb1acbfc08bc7d2960fd79630964b22819664d265f6550

Observation 2773c104-cbbb-4c8e-9750-de3cafca9114 · outbound

This paper cites Analysis of self-supervised speech models on chil- dren’s speech and infant vocalizations,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Analysis of self-supervised speech models on chil- dren’s speech and infant vocalizations,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:32.018584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.258012Z digest=sha256:72370c980193506ac0286afc45d0208cabe9c51714c05b300d9ee0258d34cf3a

Observation 71112622-77f6-48da-ab32-6ecdba487a96 · outbound

This paper cites Towards better domain adaptation for self- supervised models: A case study of child asr,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Towards better domain adaptation for self- supervised models: A case study of child asr,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.990930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.263848Z digest=sha256:1fe83520498219d380c62c5e0da2c8c161ba83441052c40b6b6e05e726d4772d

Observation babb52d0-be38-4d49-9bd7-1f6dd553f87b · outbound

This paper cites Towards universal speech discrete tokens: A case study for ASR and TTS,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Towards universal speech discrete tokens: A case study for ASR and TTS,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.968788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.270445Z digest=sha256:a8c41c1af2f72260030a9279e663db09c07c92a204165b9a39e498baa9b02856

Observation b2ef6f22-7db0-4e07-be71-8765b62ee2b2 · outbound

This paper cites Exploration of efficient end-to-end ASR using discretized input from self-supervised learning,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Exploration of efficient end-to-end ASR using discretized input from self-supervised learning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.946700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.276584Z digest=sha256:7bfe0ba408ae100c242dbc765f55c4732da240e958df511da91541dda3618d36

Observation 59c68ade-296d-41f4-b6f5-d309c9953c2a · outbound

This paper cites A wav2vec2-based experimental study on self- supervised learning methods to improve child speech recogni- tion,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet A wav2vec2-based experimental study on self- supervised learning methods to improve child speech recogni- tion,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.924405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.282942Z digest=sha256:adfb8d1f7f0f28d04cee50652854bcfb29f0fd7dbdc189140e93771eadac39ec

Observation 814d4849-537e-4311-b9f5-ad33a6eda4a0 · outbound

This paper cites Children’s speaker verification in low and zero resource conditions,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Children’s speaker verification in low and zero resource conditions,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.904823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.292512Z digest=sha256:82bf71c962842ba5c843bae9db110c9d1b38c35297bc1c2246d8bd9b9f0a8701

Observation 9afe58e0-cc83-44d8-97f7-1f37964f22d5 · outbound

This paper cites Childaugment: Data augmentation methods for zero-resource children’s speaker verification,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Childaugment: Data augmentation methods for zero-resource children’s speaker verification,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.874804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.299807Z digest=sha256:b295e1512168f7bcd54680ae1f282e655d69c0f34b37f4e13c0aac85f2c0b5c7

Observation d68023b3-f0d8-4629-a5f3-126e7027ff36 · outbound

This paper cites Effective preservation of higher-frequency contents in the context of short utterance based children’s speaker verification system,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Effective preservation of higher-frequency contents in the context of short utterance based children’s speaker verification system,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.853967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.309249Z digest=sha256:4a4599002237a82466d6b48ac1d21c7c17406d6bc2595c4003d27656f4b4bb9e

Observation cc235edc-9c73-4c81-b207-3b67131c1576 · outbound

This paper cites OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T17:16:31.314733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:16:31.314733Z digest=sha256:af3716fe763366c78c1e6e12c971c3356466834c44bfe49f7b08d6891dee875c

Observation e468ad48-0e82-4cb5-af13-8ef59eae4165 · outbound

This paper cites Espnet: End-to-end speech processing toolkit,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Espnet: End-to-end speech processing toolkit,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.830861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.320366Z digest=sha256:a9d35afb81f3c34a60a23b104affb03e27f1206c456037c4a5afc9d033cf1f4b

Observation 038c33fa-bf1b-4ecc-a9c3-8b8dd4df821d · outbound

This paper cites Reproducing whisper-style training using an open- source toolkit and publicly available data,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Reproducing whisper-style training using an open- source toolkit and publicly available data,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.814241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.325641Z digest=sha256:570450108bfa810775942c2c5ae358751cfe75b5ae593d51d3eece508f25d229

Observation f3a2d75d-323b-4ece-b301-ee4ddc39c05e · outbound

This paper cites Towards robust speech representation learning for thousands of languages,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Towards robust speech representation learning for thousands of languages,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.794888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.332728Z digest=sha256:4e5f418f8bd6e5708d829bc246cdca62b6c4f2401e230f0ccdbf0d80e5ff5e79

Observation fdd13e0a-4fac-41f3-9289-22fe6f801df7 · outbound

This paper cites E-branchformer: Branchformer with enhanced merging for speech recognition,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet E-branchformer: Branchformer with enhanced merging for speech recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.770338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.339183Z digest=sha256:7958b6327e23c779116daf4f4891fab157e463ade3597581f453c7482c6194d7

Observation fef85494-3100-4877-bc37-33968fc7e3c4 · outbound

This paper cites A comparative study on transformer vs rnn in speech applications,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet A comparative study on transformer vs rnn in speech applications,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.751586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.343888Z digest=sha256:14add744f0a02eb2068942b92d0127e88a2b5ec74bdcb373ce815cbcb8e4ddb8

Observation f4f9e820-4f9d-41b2-bcd4-c69025630516 · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Sequence Transduction with Recurrent Neural Networks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T17:16:31.352621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:16:31.352621Z digest=sha256:def071b94e12f42f087551898d4deb70045c7e3afca1ecbddd5b698100eb5627

Observation a504f094-8293-4cc0-9994-55f1f8bdb81f · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.729544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.360651Z digest=sha256:85dd777df15aa42df998bda8a794d591d900842c31334413577d9acb0b80a480

Observation 2a22ccac-1f5b-4b56-9a34-a33a7698790c · outbound

This paper cites OWSM-CTC: An open encoder-only speech foun- dation model for speech recognition, translation, and language identification,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet OWSM-CTC: An open encoder-only speech foun- dation model for speech recognition, translation, and language identification,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.706898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.368276Z digest=sha256:5bf18600dc7b5a1b5ab41bdb7ee8e297a3dcc316db916ff2f7631d2bd801b312

Observation c7b93c42-67f5-47c6-9344-4757ecb539f5 · outbound

This paper cites Pushing the limits of raw waveform speaker recognition,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Pushing the limits of raw waveform speaker recognition,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.687511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.378744Z digest=sha256:96eff0cbda4186667592bac74ba86f7fe0e83dfb4f750324eae1ed0443c82189

Observation b7adcdf7-cbfd-48eb-9ea0-cf29b959cac9 · outbound

This paper cites Espnet-spk: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Espnet-spk: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.666598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.388748Z digest=sha256:f1d2f59f6cdf0e92f8ee01b505a79391c248c5c929c78018fb237afb76b627a1

Observation 6a918204-7ae3-4e31-8952-583a8af7f7de · outbound

This paper cites My science tutor (MyST)–a large corpus of children‘s conversational speech,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet My science tutor (MyST)–a large corpus of children‘s conversational speech,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.644503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.394808Z digest=sha256:b5fdf5394219f75b6632e71390e7d714f3d8503ad33eb8ea6d31e0ae1538fc75

Observation 7e1175c4-7c7b-4688-ab3e-64f4d69acd9a · outbound

This paper cites The ogi kids’ speech corpus and recognizers,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet The ogi kids’ speech corpus and recognizers,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.625739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.401161Z digest=sha256:e52ce052dadeb99ce65861db28efea8613d878ee49f3bea06148839973a0b997

Observation ad1048c3-6eca-4cb5-80ea-0d2c0ed7bdc3 · outbound

This paper cites The cmu kids corpus,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet The cmu kids corpus,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.605906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.410934Z digest=sha256:2ddfc92b1684b9eb0b4869ecccc9581395c52ee9b14b21849da57e8d665ea077

Observation ccb227d1-072e-48b1-97f8-8f76f5dd4358 · outbound

This paper cites Superb: Speech processing universal perfor- mance benchmark,.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Superb: Speech processing universal perfor- mance benchmark,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:16:31.582622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.419659Z digest=sha256:0d0bf95df05220625f52903a346c3b337ae5808bc7cbe19267acb2e5f5804905

Pith citing papers

Observation 2eafadc7-12b5-49dc-bbec-8791a94cbc41 · inbound

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet cites this paper.

Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:16:31.551433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T17:16:31.102623Z digest=sha256:1be9b40edb231e4faae8f5ace0c68232c55ed3acd10eaea36575e4db849bb6ae

Observation cf1645f8-3434-461a-8948-80afd9fc8bd5 · inbound

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling cites this paper.

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T00:48:21.716770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:48:21.716770Z digest=sha256:10e766d2869d3b93a47b663fce3c420302ba604b634ca1c2d3cdda09ce2895e4