REVIEW 2 cited by
Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We study how to generate captions that are not only accurate in describing an image but also discriminative across different images. The problem is both fundamental and interesting, as most machine-generated captions, despite phenomenal research progresses in the past several years, are expressed in a very monotonic and featureless format. While such captions are normally accurate, they often lack important characteristics in human languages - distinctiveness for each caption and diversity for different images. To address this problem, we propose a novel conditional generative adversarial network for generating diverse captions across images. Instead of estimating the quality of a caption solely on one image, the proposed comparative adversarial learning framework better assesses the quality of captions by comparing a set of captions within the image-caption joint space. By contrasting with human-written captions and image-mismatched captions, the caption generator effectively exploits the inherent characteristics of human languages, and generates more discriminative captions. We show that our proposed network is capable of producing accurate and diverse captions across images.
Forward citations
Cited by 2 Pith papers
-
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings
An unpaired image captioning method that aligns image features to a visually structured sentence embedding space, using a robust min-distance loss and concept-conditioned adversarial training, achieves state-of-the-ar...
-
Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning
Seq-CVAE learns a latent variable for every word position in an image caption, guided by a backward language model, and produces more diverse yet accurate captions than previous approaches.
Discussion (0). Continue with ORCID to comment.