Pith. sign in

REVIEW 1 cited by

Progressive Feature Mining and External Knowledge-Assisted Text-Pedestrian Image Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.11994 v1 pith:OSR2A4H5 submitted 2023-08-23 cs.CV

classification cs.CV
keywords featuresimageretrievaldiversityexternalfeaturemethodmining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-Pedestrian Image Retrieval aims to use the text describing pedestrian appearance to retrieve the corresponding pedestrian image. This task involves not only modality discrepancy, but also the challenge of the textual diversity of pedestrians with the same identity. At present, although existing research progress has been made in text-pedestrian image retrieval, these methods do not comprehensively consider the above-mentioned problems. Considering these, this paper proposes a progressive feature mining and external knowledge-assisted feature purification method. Specifically, we use a progressive mining mode to enable the model to mine discriminative features from neglected information, thereby avoiding the loss of discriminative information and improving the expression ability of features. In addition, to further reduce the negative impact of modal discrepancy and text diversity on cross-modal matching, we propose to use other sample knowledge of the same modality, i.e., external knowledge to enhance identity-consistent features and weaken identity-inconsistent features. This process purifies features and alleviates the interference caused by textual diversity and negative sample correlation features of the same modal. Extensive experiments on three challenging datasets demonstrate the effectiveness and superiority of the proposed method, and the retrieval performance even surpasses that of the large-scale model-based method on large-scale datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    QUAG, a query-centric audio-visual cognition network, reports state-of-the-art moment retrieval and segmentation on HIREST and competitive video summarization on TVSum.

Pith tools