Pith. sign in

REVIEW 1 cited by

Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open World

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.13233 v2 pith:KRXMRW4O submitted 2023-03-23 cs.CV

classification cs.CV
keywords generationgraphpredicatescenecacaomodelspredicatescross-modal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scene Graph Generation (SGG) aims to extract <subject, predicate, object> relationships in images for vision understanding. Although recent works have made steady progress on SGG, they still suffer long-tail distribution issues that tail-predicates are more costly to train and hard to distinguish due to a small amount of annotated data compared to frequent predicates. Existing re-balancing strategies try to handle it via prior rules but are still confined to pre-defined conditions, which are not scalable for various models and datasets. In this paper, we propose a Cross-modal prediCate boosting (CaCao) framework, where a visually-prompted language model is learned to generate diverse fine-grained predicates in a low-resource way. The proposed CaCao can be applied in a plug-and-play fashion and automatically strengthen existing SGG to tackle the long-tailed problem. Based on that, we further introduce a novel Entangled cross-modal prompt approach for open-world predicate scene graph generation (Epic), where models can generalize to unseen predicates in a zero-shot manner. Comprehensive experiments on three benchmark datasets show that CaCao consistently boosts the performance of multiple scene graph generation models in a model-agnostic way. Moreover, our Epic achieves competitive performance on open-world predicate prediction. The data and code for this paper are publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A hierarchical prompt framework with entity clustering, LLM region descriptions, and VLM-based selection improves open-vocabulary scene graph generation on Visual Genome and Open Images v6.

Pith tools