TIMAM combines an adversarial modality discriminator, norm-softmax identification losses, and a cross-modal projection matching loss with a BERT plus bidirectional LSTM text encoder, achieving state-of-the-art text-to-image retrieval on CUHK-PEDES, Flickr30K, CUB, and Flowers.
Evaluation of output embeddings for fine- grained image classification
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Adversarial Representation Learning for Text-to-Image Matching
TIMAM combines an adversarial modality discriminator, norm-softmax identification losses, and a cross-modal projection matching loss with a BERT plus bidirectional LSTM text encoder, achieving state-of-the-art text-to-image retrieval on CUHK-PEDES, Flickr30K, CUB, and Flowers.