Pith. sign in

REVIEW 1 cited by

An open-source voice type classifier for child-centered daylong recordings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.12656 v3 pith:ZQTIZZ2R submitted 2020-05-26 eess.AS

classification eess.AS
keywords producedspeechadultchild-centeredchildrenlanguagerecordingsaudio
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Spontaneous conversations in real-world settings such as those found in child-centered recordings have been shown to be amongst the most challenging audio files to process. Nevertheless, building speech processing models handling such a wide variety of conditions would be particularly useful for language acquisition studies in which researchers are interested in the quantity and quality of the speech that children hear and produce, as well as for early diagnosis and measuring effects of remediation. In this paper, we present our approach to designing an open-source neural network to classify audio segments into vocalizations produced by the child wearing the recording device, vocalizations produced by other children, adult male speech, and adult female speech. To this end, we gathered diverse child-centered corpora which sums up to a total of 260 hours of recordings and covers 10 languages. Our model can be used as input for downstream tasks such as estimating the number of words produced by adult speakers, or the number of linguistic units produced by children. Our architecture combines SincNet filters with a stack of recurrent layers and outperforms by a large margin the state-of-the-art system, the Language ENvironment Analysis (LENA) that has been used in numerous child language studies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Employing self-supervised learning models for cross-linguistic child speech maturity classification

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 25+ language child-vocalization dataset, SpeechMaturity, improves self-supervised speech model classification of cry, laugh, and speech maturity, reaching 74.2% unweighted average recall.

Pith tools