Pith. sign in

REVIEW 5 cited by

Disentangled Representation Learning for Text-Video Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.07111 v1 pith:VOKQ2PVG submitted 2022-03-14 cs.CV

classification cs.CV
keywords interactionrepresentationdisentangledsequentialdifferentfunctionhierarchicalinputs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cross-modality interaction is a critical component in Text-Video Retrieval (TVR), yet there has been little examination of how different influencing factors for computing interaction affect performance. This paper first studies the interaction paradigm in depth, where we find that its computation can be split into two terms, the interaction contents at different granularity and the matching function to distinguish pairs with the same semantics. We also observe that the single-vector representation and implicit intensive function substantially hinder the optimization. Based on these findings, we propose a disentangled framework to capture a sequential and hierarchical representation. Firstly, considering the natural sequential structure in both text and video inputs, a Weighted Token-wise Interaction (WTI) module is performed to decouple the content and adaptively exploit the pair-wise correlations. This interaction can form a better disentangled manifold for sequential inputs. Secondly, we introduce a Channel DeCorrelation Regularization (CDCR) to minimize the redundancy between the components of the compared vectors, which facilitate learning a hierarchical representation. We demonstrate the effectiveness of the disentangled representation on various benchmarks, e.g., surpassing CLIP4Clip largely by +2.9%, +3.1%, +7.9%, +2.3%, +2.8% and +6.5% R@1 on the MSR-VTT, MSVD, VATEX, LSMDC, AcitivityNet, and DiDeMo, respectively.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A parameter-efficient video-text retrieval method that trains only 0.56M parameters on top of frozen CLIP and achieves 50.5% R@1 on MSRVTT.

  2. PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval

    cs.IR 2026-08 conditional novelty 5.0 of 10

    PHA-Net inserts shared prototype tokens into a three-level text-video alignment model and reports higher aggregate retrieval scores than the HBI baseline on four benchmarks, though several gains are small and unverified.

  3. MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval

    cs.CV 2025-10 conditional novelty 5.0 of 10

    MSAM introduces two drone-video/text datasets and a CLIP-based multi-semantic pooling model that reports 0.6–3.8 point R@1 gains over earlier video-text retrieval methods.

  4. Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ARL detects ambiguous text-video pairs using uncertainty and similarity, then trains retrieval models with ambiguity-aware contrastive and triplet losses, achieving state-of-the-art on TVR and ActivityNet Captions.

  5. MemVerse: Multimodal Memory for Lifelong Learning Agents

    cs.AI 2025-12 reject novelty 4.0 of 10

    MemVerse reports large gains on multimodal benchmarks by adding a hierarchical knowledge-graph memory plus fine-tuned parametric recall, but its strongest video-retrieval result uses ground-truth caption-video pairs i...

Pith tools