Pith. sign in

REVIEW 1 cited by

Elucidating image-to-set prediction: An analysis of models, losses and datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.05709 v2 pith:3N6OXDBX submitted 2019-04-11 cs.CV

classification cs.CV
keywords image-to-setpredictiondatasetsmodelsanalysisbenchmarkimagenamely
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we identify an important reproducibility challenge in the image-to-set prediction literature that impedes proper comparisons among published methods, namely, researchers use different evaluation protocols to assess their contributions. To alleviate this issue, we introduce an image-to-set prediction benchmark suite built on top of five public datasets of increasing task complexity that are suitable for multi-label classification (VOC, COCO, NUS-WIDE, ADE20k and Recipe1M). Using the benchmark, we provide an in-depth analysis where we study the key components of current models, namely the choice of the image representation backbone as well as the set predictor design. Our results show that (1) exploiting better image representation backbones leads to higher performance boosts than enhancing set predictors, and (2) modeling both the label co-occurrences and ordering has a slight positive impact in terms of performance, whereas explicit cardinality prediction only helps when training on complex datasets, such as Recipe1M. To facilitate future image-to-set prediction research, we make the code, best models and dataset splits publicly available at: https://github.com/facebookresearch/image-to-set.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Token-based Object Detection with Video

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A token-based autoregressive video detector represents objects as discrete-token 3D tracklets and achieves 91.14 mAP on UA-DETRAC, but its improvement over the static baseline mostly reflects redundant sliding windows...

Pith tools