Pith. sign in

REVIEW 1 cited by

Unleash the Potential of CLIP for Video Highlight Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01745 v1 pith:6DGZ2E2Z submitted 2024-04-02 cs.CV cs.AI

classification cs.CVcs.AI
keywords detectionhighlightknowledgemultimodalvideomodelstaskachieved
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal and large language models (LLMs) have revolutionized the utilization of open-world knowledge, unlocking novel potentials across various tasks and applications. Among these domains, the video domain has notably benefited from their capabilities. In this paper, we present Highlight-CLIP (HL-CLIP), a method designed to excel in the video highlight detection task by leveraging the pre-trained knowledge embedded in multimodal models. By simply fine-tuning the multimodal encoder in combination with our innovative saliency pooling technique, we have achieved the state-of-the-art performance in the highlight detection task, the QVHighlight Benchmark, to the best of our knowledge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sparse-Dense Side-Tuner for efficient Video Temporal Grounding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SDST is a parameter-efficient, anchor-free side-tuning architecture for video temporal grounding that matches or beats state-of-the-art methods with about 73% fewer trainable parameters.

Pith tools