Pith. sign in

REVIEW 2 cited by

Learning to generalize to new compositions in image understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1608.07639 v1 pith:WCJMRLKK submitted 2016-08-27 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords structuredgeneralizationrepresentationsimageslearningbeencapturecombinations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recurrent neural networks have recently been used for learning to describe images using natural language. However, it has been observed that these models generalize poorly to scenes that were not observed during training, possibly depending too strongly on the statistics of the text in the training data. Here we propose to describe images using short structured representations, aiming to capture the crux of a description. These structured representations allow us to tease-out and evaluate separately two types of generalization: standard generalization to new images with similar scenes, and generalization to new combinations of known entities. We compare two learning approaches on the MS-COCO dataset: a state-of-the-art recurrent network based on an LSTM (Show, Attend and Tell), and a simple structured prediction model on top of a deep network. We find that the structured model generalizes to new compositions substantially better than the LSTM, ~7 times the accuracy of predicting structured representations. By providing a concrete method to quantify generalization for unseen combinations, we argue that structured representations and compositional splits are a useful benchmark for image captioning, and advocate compositional models that capture linguistic and visual structure.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning

    cs.CV 2025-06 reject novelty 5.0 of 10

    EVA reports SOTA on MIT-States, UT-Zappos, and C-GQA in closed- and open-world CZSL by combining MoE adapters with semantic variant alignment, but missing backbone-matched baselines weaken the claim.

  2. Learning Clustering-based Prototypes for Compositional Zero-shot Learning

    cs.CV 2025-02 conditional novelty 5.0 of 10

    ClusPro improves compositional zero-shot learning by representing each primitive with multiple online-clustered prototypes and adding prototype-anchored contrastive and decorrelation losses.

Pith tools