Pith. sign in

REVIEW 1 cited by

Catplayinginthesnow: Impact of Prior Segmentation on a Model of Visually Grounded Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.08387 v2 pith:OHWPIBSL submitted 2020-06-15 cs.CL

classification cs.CL
keywords boundaryinformationmodelsegmentsunitsgroundedhigh-levelinvestigate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The language acquisition literature shows that children do not build their lexicon by segmenting the spoken input into phonemes and then building up words from them, but rather adopt a top-down approach and start by segmenting word-like units and then break them down into smaller units. This suggests that the ideal way of learning a language is by starting from full semantic units. In this paper, we investigate if this is also the case for a neural model of Visually Grounded Speech trained on a speech-image retrieval task. We evaluated how well such a network is able to learn a reliable speech-to-image mapping when provided with phone, syllable, or word boundary information. We present a simple way to introduce such information into an RNN-based model and investigate which type of boundary is the most efficient. We also explore at which level of the network's architecture such information should be introduced so as to maximise its performances. Finally, we show that using multiple boundary types at once in a hierarchical structure, by which low-level segments are used to recompose high-level segments, is beneficial and yields better results than using low-level or high-level segments in isolation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Never Reset Again: A Mathematical Framework for Continual Inference in Recurrent Neural Networks

    cs.LG 2024-12 conditional novelty 5.0 of 10

    A dual cross-entropy and KL-divergence loss lets recurrent networks maintain stable accuracy over very long streams without hidden-state resets, matching and sometimes slightly beating periodic reset baselines.

Pith tools