Pith. sign in

REVIEW

Auto-captions on GIF: A Large-scale Video-sentence Dataset for Vision-language Pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.02375 v1 pith:JBRCEG2B submitted 2020-07-05 cs.CV cs.CL

classification cs.CVcs.CL
keywords datasetvideoauto-captionspre-trainingvideo-sentencecaptioningdownstreamencoder-decoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we present Auto-captions on GIF, which is a new large-scale pre-training dataset for generic video understanding. All video-sentence pairs are created by automatically extracting and filtering video caption annotations from billions of web pages. Auto-captions on GIF dataset can be utilized to pre-train the generic feature representation or encoder-decoder structure for video captioning, and other downstream tasks (e.g., sentence localization in videos, video question answering, etc.) as well. We present a detailed analysis of Auto-captions on GIF dataset in comparison to existing video-sentence datasets. We also provide an evaluation of a Transformer-based encoder-decoder structure for vision-language pre-training, which is further adapted to video captioning downstream task and yields the compelling generalizability on MSR-VTT. The dataset is available at \url{http://www.auto-video-captions.top/2020/dataset}.

Discussion (0). Sign in to comment.

Pith tools