Pith. sign in

REVIEW 3 cited by

Multi-granularity Correspondence Learning from Long-term Noisy Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.16702 v1 pith:3U457QVM submitted 2024-01-30 cs.CV

classification cs.CV
keywords nortonclip-captionlearningmisalignmenttemporalvideoaddressclips
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing video-language studies mainly focus on learning short video clips, leaving long-term temporal dependencies rarely explored due to over-high computational cost of modeling long videos. To address this issue, one feasible solution is learning the correspondence between video clips and captions, which however inevitably encounters the multi-granularity noisy correspondence (MNC) problem. To be specific, MNC refers to the clip-caption misalignment (coarse-grained) and frame-word misalignment (fine-grained), hindering temporal learning and video understanding. In this paper, we propose NOise Robust Temporal Optimal traNsport (Norton) that addresses MNC in a unified optimal transport (OT) framework. In brief, Norton employs video-paragraph and clip-caption contrastive losses to capture long-term dependencies based on OT. To address coarse-grained misalignment in video-paragraph contrast, Norton filters out the irrelevant clips and captions through an alignable prompt bucket and realigns asynchronous clip-caption pairs based on transport distance. To address the fine-grained misalignment, Norton incorporates a soft-maximum operator to identify crucial words and key frames. Additionally, Norton exploits the potential faulty negative samples in clip-caption contrast by rectifying the alignment target with OT assignment to ensure precise temporal modeling. Extensive experiments on video retrieval, videoQA, and action segmentation verify the effectiveness of our method. Code is available at https://lin-yijie.github.io/projects/Norton.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A weakly supervised optimal-transport method infers frame-phrase alignments in human motion from clip-level text, evaluated on only 30 test pairs.

  2. From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    GRACE pairs motion-weighted video tokens with AI-refined emotion text tokens using optimal transport, reporting new UAR and WAR records on DFEW, FERV39k, and MAFW.

  3. Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ARL detects ambiguous text-video pairs using uncertainty and similarity, then trains retrieval models with ambiguity-aware contrastive and triplet losses, achieving state-of-the-art on TVR and ActivityNet Captions.

Pith tools