Pith. sign in

REVIEW 2 cited by

VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.03204 v1 pith:6SZV2WS7 submitted 2023-05-04 cs.CV cs.CL

classification cs.CVcs.CL
keywords videomodelpre-trainingvideo-to-textansweringcaptioningdatageneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a new two-stage pre-training framework for video-to-text generation tasks such as video captioning and video question answering: A generative encoder-decoder model is first jointly pre-trained on massive image-text data to learn fundamental vision-language concepts, and then adapted to video data in an intermediate video-text pre-training stage to learn video-specific skills such as spatio-temporal reasoning. As a result, our VideoOFA model achieves new state-of-the-art performance on four Video Captioning benchmarks, beating prior art by an average of 9.7 points in CIDEr score. It also outperforms existing models on two open-ended Video Question Answering datasets, showcasing its generalization capability as a universal video-to-text model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

    cs.MM 2024-12 conditional novelty 6.0 of 10

    SyncFlow jointly generates temporally aligned 16 FPS video and 48kHz audio from text using a dual-diffusion-transformer with modality adaptors, and reports better audio-video alignment than cascaded and contrastive baselines.

  2. Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A factorized autoregressive decoder, shared across video segments with cross-segment masking, produces denser, more localized captions online while saving about 20 percent compute versus a global decoder.

Pith tools