TIMAM combines an adversarial modality discriminator, norm-softmax identification losses, and a cross-modal projection matching loss with a BERT plus bidirectional LSTM text encoder, achieving state-of-the-art text-to-image retrieval on CUHK-PEDES, Flickr30K, CUB, and Flowers.
Improving text-based person search by spatial matching and adaptive threshold
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2019 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Adversarial Representation Learning for Text-to-Image Matching
TIMAM combines an adversarial modality discriminator, norm-softmax identification losses, and a cross-modal projection matching loss with a BERT plus bidirectional LSTM text encoder, achieving state-of-the-art text-to-image retrieval on CUHK-PEDES, Flickr30K, CUB, and Flowers.