Pith. sign in

REVIEW 1 cited by

Image-Text Multi-Modal Representation Learning by Adversarial Backpropagation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1612.08354 v1 pith:UAWP3UNZ submitted 2016-12-26 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords multi-modalimage-textinformationlearningfeaturepairadversarialbackpropagation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present novel method for image-text multi-modal representation learning. In our knowledge, this work is the first approach of applying adversarial learning concept to multi-modal learning and not exploiting image-text pair information to learn multi-modal feature. We only use category information in contrast with most previous methods using image-text pair information for multi-modal embedding. In this paper, we show that multi-modal feature can be achieved without image-text pair information and our method makes more similar distribution with image and text in multi-modal feature space than other methods which use image-text pair information. And we show our multi-modal feature has universal semantic information, even though it was trained for category prediction. Our model is end-to-end backpropagation, intuitive and easily extended to other multi-modal learning work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Cross Modal Systems Leverage Semantic Relationships?

    cs.CV 2019-09 reject novelty 4.0 of 10

    The authors introduce SemanticMap, a cosine-similarity based evaluation metric for cross-modal retrieval, and a single-stream network that encodes text as images, but the metric can be trivially gamed by collapsing em...

Pith tools