Pith. sign in

REVIEW 2 cited by

Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.12218 v1 pith:HNQSTB6R submitted 2023-05-20 cs.CV cs.AIcs.IR

classification cs.CVcs.AIcs.IR
keywords conceptsalignmentconceptualizationdisentangledset-to-setdicosafailleverage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local details or are computationally expensive. What's worse, they fail to leverage the heterogeneous concepts in data. In this paper, we propose the Disentangled Conceptualization and Set-to-set Alignment (DiCoSA) to simulate the conceptualizing and reasoning process of human beings. For disentangled conceptualization, we divide the coarse feature into multiple latent factors related to semantic concepts. For set-to-set alignment, where a set of visual concepts correspond to a set of textual concepts, we propose an adaptive pooling method to aggregate semantic concepts to address the partial matching. In particular, since we encode concepts independently in only a few dimensions, DiCoSA is superior at efficiency and granularity, ensuring fine-grained interactions using a similar computational complexity as coarse-grained alignment. Extensive experiments on five datasets, including MSR-VTT, LSMDC, MSVD, ActivityNet, and DiDeMo, demonstrate that our method outperforms the existing state-of-the-art methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ARL detects ambiguous text-video pairs using uncertainty and similarity, then trains retrieval models with ambiguity-aware contrastive and triplet losses, achieving state-of-the-art on TVR and ActivityNet Captions.

  2. Uneven Event Modeling for Partially Relevant Video Retrieval

    cs.CV 2025-06 conditional novelty 5.0 of 10

    UEM retrieves partially relevant videos by adaptively segmenting frames into uneven events and refining the best-matching event with text-conditioned attention.

Pith tools