Pith. sign in

REVIEW 1 cited by

Tackling the Unannotated: Scene Graph Generation with Bias-Reduced Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.07832 v1 pith:OXP6MOMZ submitted 2020-08-18 cs.CV cs.LG

classification cs.CVcs.LG
keywords modelsscenevisualgraphcompetitivediscrepancyentitiesgeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Predicting a scene graph that captures visual entities and their interactions in an image has been considered a crucial step towards full scene comprehension. Recent scene graph generation (SGG) models have shown their capability of capturing the most frequent relations among visual entities. However, the state-of-the-art results are still far from satisfactory, e.g. models can obtain 31% in overall recall R@100, whereas the likewise important mean class-wise recall mR@100 is only around 8% on Visual Genome (VG). The discrepancy between R and mR results urges to shift the focus from pursuing a high R to a high mR with a still competitive R. We suspect that the observed discrepancy stems from both the annotation bias and sparse annotations in VG, in which many visual entity pairs are either not annotated at all or only with a single relation when multiple ones could be valid. To address this particular issue, we propose a novel SGG training scheme that capitalizes on self-learned knowledge. It involves two relation classifiers, one offering a less biased setting for the other to base on. The proposed scheme can be applied to most of the existing SGG models and is straightforward to implement. We observe significant relative improvements in mR (between +6.6% and +20.4%) and competitive or better R (between -2.4% and 0.3%) across all standard SGG tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing

    cs.CV 2025-01 conditional novelty 5.0 of 10

    G2 generates a location-free scene graph from an image and feeds it, with confidence-based token weighting, into an LLM to produce visual commonsense answers and explanations.

Pith tools