Using synthetic formal languages and a new discriminative evaluation metric, the paper shows that fine-tuning outperforms in-context learning on in-distribution language generalization but both perform equally on out-of-distribution generalization across 18 LLMs.
Semantically self-aligned network for text-to- image part-aware person re-identification
9 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Empirical study of a fully synthetic data generation pipeline for text-based person retrieval that tests its use as a replacement or augmentation for real data across scenarios.
InterPartAbility adds an open-vocabulary patch-phrase interaction module and a perturbation-based interpretability protocol to TI-ReID, claiming SOTA explainability scores with competitive retrieval accuracy on three benchmarks.
UATTA adapts pre-trained text-image models at test time without labels by using disagreement in bidirectional retrieval rankings to estimate and mitigate uncertainty for improved person search.
SC-LMKB uses LLM-generated data with cross-domain fusion to cut hallucinations and delivers up to 72.6% gains on cross-modality retrieval tasks over standard semantic communication.
CRST improves ultra-low-resolution text-to-image person retrieval by 5.7% Rank-1 and 5.3% mAP on average across three datasets while stabilizing mixed-resolution galleries.
A decoupled two-stage training pipeline with a single vision encoder enables joint image-to-image and text-to-image person re-identification by avoiding cross-task interference, with I2I pre-training and textual supervision shown to benefit both tasks.
ROGLE introduces automated pseudo region-sentence pairs via RSM and multi-granular learning to boost fine-grained alignment in text-based person search, plus the P-VLG benchmark with over 100k annotated regions.
A multi-view semantic reformulation and feature compensation method using LLMs and VLMs improves text-to-image person retrieval accuracy without training and reaches SOTA on three datasets.
citing papers explorer
-
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
Using synthetic formal languages and a new discriminative evaluation metric, the paper shows that fine-tuning outperforms in-context learning on in-distribution language generalization but both perform equally on out-of-distribution generalization across 18 LLMs.
-
An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval
Empirical study of a fully synthetic data generation pipeline for text-based person retrieval that tests its use as a replacement or augmentation for real data across scenarios.
-
InterPartAbility: Phrase-Region Grounding for Interpretable Text-to-Image Person Re-Identification
InterPartAbility adds an open-vocabulary patch-phrase interaction module and a perturbation-based interpretability protocol to TI-ReID, claiming SOTA explainability scores with competitive retrieval accuracy on three benchmarks.
-
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
UATTA adapts pre-trained text-image models at test time without labels by using disagreement in bidirectional retrieval rankings to estimate and mitigate uncertainty for improved person search.
-
Semantic Communication with an LLM-enabled Knowledge Base
SC-LMKB uses LLM-generated data with cross-domain fusion to cut hallucinations and delivers up to 72.6% gains on cross-modality retrieval tasks over standard semantic communication.
-
Cross-Resolution Semantic Transfer for Robust Text-to-Image Retrieval in Low-Resolution Surveillance
CRST improves ultra-low-resolution text-to-image person retrieval by 5.7% Rank-1 and 5.3% mAP on average across three datasets while stabilizing mixed-resolution galleries.
-
Towards Resolving Optimization Conflicts Between Image- and Text-Based Person Re-Identification
A decoupled two-stage training pipeline with a single vision encoder enables joint image-to-image and text-to-image person re-identification by avoiding cross-task interference, with I2I pre-training and textual supervision shown to benefit both tasks.
-
ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search
ROGLE introduces automated pseudo region-sentence pairs via RSM and multi-granular learning to boost fine-grained alignment in text-based person search, plus the P-VLG benchmark with over 100k annotated regions.
-
Towards Robust Text-to-Image Person Retrieval: Multi-View Reformulation for Semantic Compensation
A multi-view semantic reformulation and feature compensation method using LLMs and VLMs improves text-to-image person retrieval accuracy without training and reaches SOTA on three datasets.