Pith. sign in

REVIEW 4 cited by

VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.00741 v2 pith:M7WRTU3M submitted 2024-10-01 cs.CL cs.CVcs.MM

classification cs.CLcs.CVcs.MM
keywords descriptionlongclipunderstandingvideocapabilitylong-descriptionpre-training
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive Language-Image Pre-training (CLIP) has been widely studied and applied in numerous applications. However, the emphasis on brief summary texts during pre-training prevents CLIP from understanding long descriptions. This issue is particularly acute regarding videos given that videos often contain abundant detailed contents. In this paper, we propose the VideoCLIP-XL (eXtra Length) model, which aims to unleash the long-description understanding capability of video CLIP models. Firstly, we establish an automatic data collection system and gather a large-scale VILD pre-training dataset with VIdeo and Long-Description pairs. Then, we propose Text-similarity-guided Primary Component Matching (TPCM) to better learn the distribution of feature space while expanding the long description capability. We also introduce two new tasks namely Detail-aware Description Ranking (DDR) and Hallucination-aware Description Ranking (HDR) for further understanding improvement. Finally, we construct a Long Video Description Ranking (LVDR) benchmark for evaluating the long-description capability more comprehensively. Extensive experimental results on widely-used text-video retrieval benchmarks with both short and long descriptions and our LVDR benchmark can fully demonstrate the effectiveness of our method.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A 1,680-question video benchmark shows leading multimodal models lag humans by ~15 points on visual knowledge, and a See-Think-Answer RL-trained model narrows the gap.

  2. Audio-Sync Video Generation with Multi-Stream Temporal Control

    cs.CV 2025-06 reject novelty 6.0 of 10

    MTV splits audio into speech, effects, and music to separately drive lip sync, event timing, and visual mood in video generation, trained on a new 392K-clip dataset.

  3. PanoWan: Lifting Diffusion Video Generation Models to 360{\deg} with Latitude/Longitude-aware Mechanisms

    cs.CV 2025-05 conditional novelty 6.0 of 10

    PanoWan adapts the Wan 2.1 text-to-video model to generate seamless 360-degree videos by remapping initial noise, rotating the latent grid during denoising, and padding the latent before VAE decoding, trained on a new...

  4. ContentV: Efficient Training of Video Generation Models with Limited Compute

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An 8B text-to-video model built by adapting Stable Diffusion 3.5 with a 3D video autoencoder reaches near-leading VBench scores after four weeks of NPU training.

Pith tools