Pith. sign in

REVIEW 1 cited by

ProTA: Probabilistic Token Aggregation for Text-Video Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.12216 v2 pith:LH3P4JH5 submitted 2024-04-18 cs.CV

classification cs.CV
keywords probabilisticaggregationcross-modalproposeprotatokencontentrepresentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the model aligning these asymmetric video-text pairs has a high risk of retrieving many false positive results. In this paper, we propose Probabilistic Token Aggregation (ProTA) to handle cross-modal interaction with content asymmetry. Specifically, we propose dual partial-related aggregation to disentangle and re-aggregate token representations in both low-dimension and high-dimension spaces. We propose token-based probabilistic alignment to generate token-level probabilistic representation and maintain the feature representation diversity. In addition, an adaptive contrastive loss is proposed to learn compact cross-modal distribution space. Based on extensive experiments, ProTA achieves significant improvements on MSR-VTT (50.9%), LSMDC (25.8%), and DiDeMo (47.2%).

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval

    cs.CV 2025-06 conditional novelty 4.0 of 10

    MamFusion adds a Mamba module and two temporal cross-attention modules to GMMFormer, achieving slight SumR improvements on three partially relevant video retrieval benchmarks.

Pith tools