REVIEW 15 cited by
Semantically Self-Aligned Network for Text-to-Image Part-aware Person Re-identification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Text-to-image person re-identification (ReID) aims to search for images containing a person of interest using textual descriptions. However, due to the significant modality gap and the large intra-class variance in textual descriptions, text-to-image ReID remains a challenging problem. Accordingly, in this paper, we propose a Semantically Self-Aligned Network (SSAN) to handle the above problems. First, we propose a novel method that automatically extracts semantically aligned part-level features from the two modalities. Second, we design a multi-view non-local network that captures the relationships between body parts, thereby establishing better correspondences between body parts and noun phrases. Third, we introduce a Compound Ranking (CR) loss that makes use of textual descriptions for other images of the same identity to provide extra supervision, thereby effectively reducing the intra-class variance in textual features. Finally, to expedite future research in text-to-image ReID, we build a new database named ICFG-PEDES. Extensive experiments demonstrate that SSAN outperforms state-of-the-art approaches by significant margins. Both the new ICFG-PEDES database and the SSAN code are available at https://github.com/zifyloo/SSAN.
Forward citations
Cited by 15 Pith papers
-
Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching
A dual-granularity identity association mechanism with dynamic confidence weighting improves weakly supervised text-to-person matching, achieving 73.06% Rank-1 on CUHK-PEDES.
-
Rethinking Text-Based Image Retrieval in Specific Domain
A new multi-match surveillance TBIR benchmark and a fine-tuning framework with cross-modal soft labels and intra-modal distillation improve mAP@20 by 7.8 points over standard contrastive tuning.
-
LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search
LightAIR combines a frozen action-word codebook, null-space appearance projection, and a Riemannian-style gradient rectification to set new state-of-the-art results on text-based person anomaly search and four TIPR be...
-
Achieving Text-based Person Retrieval with Any Granularity
A multi-grained dataset, evaluation benchmark, and CMAM model are proposed for text-based person retrieval across five levels of query granularity, with CMAM outperforming existing methods on the new benchmark.
-
Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval
GA-DMS with the WebPerson dataset sets new state-of-the-art Rank-1 accuracy on CUHK-PEDES, ICFG-PEDES, and RSTPReid.
-
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
Distilling three pattern-specific alignments from a large human-centric teacher yields a 5M-parameter student that approaches teacher-level generalization on many downstream tasks.
-
Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification
A test-time MLLM interaction module and a text reorganization augmentation improve text-to-image person re-identification ranking across four benchmarks.
-
Knowing Where to Focus: Attention-Guided Alignment for Text-based Person Search
Attention-Guided Masking and Text Enrichment improve text-based person retrieval, achieving 78.36% Rank-1 on CUHK-PEDES, 67.31% on ICFG-PEDES, and 67.4% on RSTPReid.
-
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
Introduces the Pedestrian Anomaly Behavior (PAB) benchmark and a Cross-Modal Pose-aware model for retrieving pedestrians from text descriptions of normal or anomalous actions, reporting 84.93% R@1.
-
Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification
A unified taxonomy and survey of single- to multi-modal person ReID, plus a Transformer-based VI-ReID baseline that is solid but not state-of-the-art.
-
Adapting Vision-Language Models Without Labels: A Comprehensive Survey
A survey that organizes unsupervised vision-language model adaptation by unlabeled-data availability into four paradigms: data-free transfer, domain transfer, episodic test-time, and online test-time adaptation.
-
CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval
CAMeL combines stylized synthetic tasks, a hard-negative memory queue, and dual-speed meta-updates to improve text-based person retrieval after fine-tuning.
-
Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval
GCKD combines graph-based cross-domain propagation with momentum knowledge distillation to improve text-to-image person retrieval without target-domain annotations.
-
Improving Text-based Person Search via Part-level Cross-modal Correspondence
A shared-token encoder-decoder plus a commonality-based margin ranking loss achieves state-of-the-art text-to-image person retrieval on three benchmarks.
-
Enhancing Visual Representation for Text-based Person Searching
VFE-TPS adds text-guided masked image modeling and identity-supervised feature calibration to CLIP and reports state-of-the-art Rank-1 accuracy on three text-based person search benchmarks, though not above a cited Ra...
Discussion (0). Continue with ORCID to comment.