Pith. sign in

REVIEW 1 cited by

JECL: Joint Embedding and Cluster Learning for Image-Text Pairs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.01860 v3 pith:IGVFJYB6 submitted 2019-01-04 cs.LG stat.ML

classification cs.LGstat.ML
keywords jeclclusterimage-captionpairsassignmentsclusteringdatadistribution
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose JECL, a method for clustering image-caption pairs by training parallel encoders with regularized clustering and alignment objectives, simultaneously learning both representations and cluster assignments. These image-caption pairs arise frequently in high-value applications where structured training data is expensive to produce, but free-text descriptions are common. JECL trains by minimizing the Kullback-Leibler divergence between the distribution of the images and text to that of a combined joint target distribution and optimizing the Jensen-Shannon divergence between the soft cluster assignments of the images and text. Regularizers are also applied to JECL to prevent trivial solutions. Experiments show that JECL outperforms both single-view and multi-view methods on large benchmark image-caption datasets, and is remarkably robust to missing captions and varying data sizes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Delineating Knowledge Domains in the Scientific Literature Using Visual Information

    cs.DL 2019-08 conditional novelty 6.0 of 10

    Visual fingerprints of figures in arXiv papers yield inter-field distances that correlate with citation and text distances, and neural-network diagram counts rose before citation counts for key deep learning papers.

Pith tools