Pith. sign in

REVIEW 8 cited by

The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1811.00982 v2 pith:O6V2D5XX submitted 2018-11-02 cs.CV

The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale

classification cs.CV
keywords imagesdetectionvisualobjectrelationshipannotationsimageopen
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present Open Images V4, a dataset of 9.2M images with unified annotations for image classification, object detection and visual relationship detection. The images have a Creative Commons Attribution license that allows to share and adapt the material, and they have been collected from Flickr without a predefined list of class names or tags, leading to natural class statistics and avoiding an initial design bias. Open Images V4 offers large scale across several dimensions: 30.1M image-level labels for 19.8k concepts, 15.4M bounding boxes for 600 object classes, and 375k visual relationship annotations involving 57 classes. For object detection in particular, we provide 15x more bounding boxes than the next largest datasets (15.4M boxes on 1.9M images). The images often show complex scenes with several objects (8 annotated objects per image on average). We annotated visual relationships between them, which support visual relationship detection, an emerging task that requires structured reasoning. We provide in-depth comprehensive statistics about the dataset, we validate the quality of the annotations, we study how the performance of several modern models evolves with increasing amounts of training data, and we demonstrate two applications made possible by having unified annotations of multiple types coexisting in the same images. We hope that the scale, quality, and variety of Open Images V4 will foster further research and innovation even beyond the areas of image classification, object detection, and visual relationship detection.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. High-Resolution Image Synthesis with Latent Diffusion Models

    cs.CV 2021-12 conditional novelty 7.0

    Latent diffusion models achieve state-of-the-art inpainting and competitive results on unconditional generation, scene synthesis, and super-resolution by performing the diffusion process in the latent space of pretrai...

  2. Optuna: A Next-generation Hyperparameter Optimization Framework

    cs.LG 2019-07 unverdicted novelty 7.0

    Optuna introduces a define-by-run hyperparameter optimization framework with efficient searching, pruning, and support for distributed to interactive use cases.

  3. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

    cs.AI 2026-07 conditional novelty 6.0

    SpatialCLI's Call-Learn-Internalize recipe lifts Qwen3-VL-8B from 29.3% to 84.6% on MindCube with spatial tools and 73.8% without tools by distilling successful tool trajectories into direct reasoning.

  4. TaCarla: A comprehensive benchmarking dataset for end-to-end autonomous driving

    cs.RO 2026-02 conditional novelty 6.0

    TaCarla releases 2.85M CARLA Leaderboard 2.0 frames with nuScenes-style sensors, multi-task annotations, planning baselines, and a text-based rarity score.

  5. Scaling Robot Learning with Semantically Imagined Experience

    cs.RO 2023-02 unverdicted novelty 6.0

    Augmenting robot datasets via diffusion-based semantic inpainting enables manipulation policies to solve unseen tasks with new objects and improves robustness to novel distractors.

  6. MHSA: A Lightweight Framework for Mitigating Hallucinations via Steered Attention in LVLMs

    cs.CV 2026-05 unverdicted novelty 5.0

    MHSA mitigates hallucinations in LVLMs by training an MLP to steer cross-modal attention, extending detection work to mitigation via attention replacement at inference.

  7. Data Selection for training Semantic Segmentation CNNs with cross-dataset weak supervision

    cs.CV 2019-07 unverdicted novelty 5.0

    Two data selection techniques (GMM visual similarity and bounding-box diversity) reduce required weakly labeled images by up to 100x on Open Images and 20x on Cityscapes while maintaining semantic segmentation performance.

  8. Obj-GloVe: Scene-Based Contextual Object Embedding

    cs.CV 2019-07 unverdicted novelty 5.0

    Obj-GloVe is a contextual embedding for visual objects derived from scene co-occurrences using the GloVe method, shown useful for object detection and text-to-image synthesis.