Pith. sign in

REVIEW 1 cited by

Large-scale representation learning from visually grounded untranscribed speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.08782 v1 pith:ZXKVDULX submitted 2019-09-19 cs.CV cs.CLcs.SDeess.AS

classification cs.CVcs.CLcs.SDeess.AS
keywords audiocaptionsgroundedimageslearninglossmodelsobtain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Systems that can associate images with their spoken audio captions are an important step towards visually grounded language learning. We describe a scalable method to automatically generate diverse audio for image captioning datasets. This supports pretraining deep networks for encoding both audio and images, which we do via a dual encoder that learns to align latent representations from both modalities. We show that a masked margin softmax loss for such models is superior to the standard triplet loss. We fine-tune these models on the Flickr8k Audio Captions Corpus and obtain state-of-the-art results---improving recall in the top 10 from 29.6% to 49.5%. We also obtain human ratings on retrieval outputs to better assess the impact of incidentally matching image-caption pairs that were not associated in the data, finding that automatic evaluation substantially underestimates the quality of the retrieved results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Text embedding models can be great data engineers

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Text embeddings of raw, text-serialized time series, compressed by a supervised variational information bottleneck, can match or beat hand-engineered pipelines on some classification benchmarks.

Pith tools