Pith. sign in

Paper Citation Record · LEDGER

Open World Scene Graph Generation using Vision Language Models

As of 9 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 3 inbound Pith citation observations for arXiv:2506.08189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08189 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:22:12.194352Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:44:08.297552Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T20:42:58.205735Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c0fc7967-51b6-447b-b055-61b378c7e6ae · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Open World Scene Graph Generation using Vision Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.087934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.087934Z digest=sha256:d75d150fe1e4bfb6f5dff488817065d0aecd2946487d3448d7c3785cd03cf14d

Observation f25714bf-0374-43a8-a38c-dad42bafb371 · outbound

This paper cites GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives.

Open World Scene Graph Generation using Vision Language Models GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.091525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.091525Z digest=sha256:f25bab8b0958ff388b8af37b02a1f845fca45d61613bbb77ad07a2d29aaecf9b

Observation 58085406-ab58-40f7-89c5-380f19210270 · outbound

This paper cites Expanding scene graph boundaries: fully open-vocabulary scene graph generation via visual-concept alignment and retention.

Open World Scene Graph Generation using Vision Language Models Expanding scene graph boundaries: fully open-vocabulary scene graph generation via visual-concept alignment and retention

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.620025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.094288Z digest=sha256:78679dadef6635e12d024849e9b45476c2b02975d2761969342c6f19a4a812e2

Observation 12ae2883-1520-484f-a502-e06b47a57183 · outbound

This paper cites Reltr: Relation transformer for scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 45(9): 11169–11183, 2023.

Open World Scene Graph Generation using Vision Language Models Reltr: Relation transformer for scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 45(9): 11169–11183, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.612736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.096830Z digest=sha256:efcb3e6b75d1ad802239423267c76eac5ac20e6df566c87d38dce940a4644ea8

Observation 347bb022-99ea-4d76-b760-7e683b7f98cd · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Open World Scene Graph Generation using Vision Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.099201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.099201Z digest=sha256:ad1861c50abff88a49120fc3c39542d68ff45b56610fc00f74365195425f9fa0

Observation ab1b86b8-b2a4-407e-8684-09326248bc2c · outbound

This paper cites Prism-0: A predicate-rich scene graph genera- tion framework for zero-shot open-vocabulary tasks.arXiv preprint arXiv:2504.00844, 2025.

Open World Scene Graph Generation using Vision Language Models Prism-0: A predicate-rich scene graph genera- tion framework for zero-shot open-vocabulary tasks.arXiv preprint arXiv:2504.00844, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.101904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.101904Z digest=sha256:3464cc7f02b35d41152696181ee8699ee925d3506a9fbc261ec9b78493692c00

Observation 7e752d72-4886-4e04-9aa2-c3206c38b5b6 · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

Open World Scene Graph Generation using Vision Language Models SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.104420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.104420Z digest=sha256:f32285a0480d3b3bd268ff141a226810fc8ad3eedfae83a3be65b9bb6762757c

Observation e4064992-17d0-47ec-9142-2b387f3f5749 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Open World Scene Graph Generation using Vision Language Models Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.107012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.107012Z digest=sha256:4326cb4eea8f81b31a4f031631478e8a75b6ff74dea041fd941c91f9b69f153d

Observation 5d2effcd-34b2-470b-8d48-da7c88326306 · outbound

This paper cites To- wards open-vocabulary scene graph generation with prompt- 7 based finetuning.

Open World Scene Graph Generation using Vision Language Models To- wards open-vocabulary scene graph generation with prompt- 7 based finetuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.605387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.109479Z digest=sha256:bd72e3dec0e6b078720661a0dcd83d16fbba5a59ee5255d0a4e2dbf3e7dc37b2

Observation ebe5fd25-ee12-4f92-81b4-9fd2e32e8673 · outbound

This paper cites Scene Graph Reasoning for Visual Question Answering.

Open World Scene Graph Generation using Vision Language Models Scene Graph Reasoning for Visual Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.111648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.111648Z digest=sha256:7729488e9908d0487b13149e1b431a6a6b5442bbd5c527d3173aed862adfc1e8

Observation 39fd54ed-26af-4a45-a6aa-a5146ce61db2 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Open World Scene Graph Generation using Vision Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.114195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.114195Z digest=sha256:4955550dfb9e27be4ca82191d871a62e7df15fd6f7ccf14bf1515d2a188f5f10

Observation b4b5ce4d-ae68-41f6-aab5-5e5c72e934fe · outbound

This paper cites Enhancing scene graph generation with hierarchical relationships and commonsense knowledge.

Open World Scene Graph Generation using Vision Language Models Enhancing scene graph generation with hierarchical relationships and commonsense knowledge

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.594374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.116526Z digest=sha256:72a555d53b36235165c7c80b47b31ba9a1600a204dd703b4b4c116c073e82164

Observation 552fd6b8-9ebf-4403-a986-1561a10867d6 · outbound

This paper cites Image retrieval using scene graphs.

Open World Scene Graph Generation using Vision Language Models Image retrieval using scene graphs

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.587225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.118645Z digest=sha256:35028d2ab3fb3fe7cf1d79e457ad292a6fccd0bea1465aafcffc9d0a475e9b14

Observation 696ec004-e2bf-40ce-9e50-a1f600cdf44b · outbound

This paper cites Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency.

Open World Scene Graph Generation using Vision Language Models Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:22:12.250576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.120736Z digest=sha256:79b7310b9bd165aa63c1f7c216879c39dd9ce3c99b217a462c296b249a17b737

Observation e21e7d9e-d8eb-4955-8b94-c9698e1a00c8 · outbound

This paper cites Llm4sgg: large language models for weakly supervised scene graph generation.

Open World Scene Graph Generation using Vision Language Models Llm4sgg: large language models for weakly supervised scene graph generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.579047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.122900Z digest=sha256:2dbd94094f6750317b10dc5b03e07bb3c669de48eb8680148ef3ea22a91e0140

Observation c181adb0-57cc-41bd-bad9-4d7bad47c8ea · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

Open World Scene Graph Generation using Vision Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.571701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.125074Z digest=sha256:036bedb7e1c98d1eca64a2f3a08d4a46563b2e930593a26ce702ed82d0f03883

Observation 96942250-333e-4d51-94be-96c8e55e2d9c · outbound

This paper cites an unresolved cited work.

Open World Scene Graph Generation using Vision Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:22:12.563917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.126991Z digest=sha256:8fa1539e91f177f23bf84d1a045500d2e647891327f819ddcdfe1ce95996807d

Observation 234f1add-b521-4887-b86a-e8b758effe78 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Open World Scene Graph Generation using Vision Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.556567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.129134Z digest=sha256:a6a3cf97f44d1a8309cae44630d3bc054a581024d999412b717d4c7e1848627e

Observation 7060702d-f3f2-47d5-a9cf-5442e3405a86 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

Open World Scene Graph Generation using Vision Language Models Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.549301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.131213Z digest=sha256:20d5bb5d756f4892b2773e805ec5936bb40b9af581d16131571c7329e0e09dc7

Observation cc27b857-6ec1-4e1f-82c9-3807eda49d7a · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Open World Scene Graph Generation using Vision Language Models Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.133250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.133250Z digest=sha256:447c2bfb4e68f58abb82a6f6e0d353eb267560dde59c915fe17f2f26911447b9

Observation 5b0f3a94-2ee0-4e36-8300-118e09f4fc59 · outbound

This paper cites Sgtr: End-to- end scene graph generation with transformer.

Open World Scene Graph Generation using Vision Language Models Sgtr: End-to- end scene graph generation with transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.537311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.135454Z digest=sha256:07aba233d14e81242b727282f3989af93bd79fc6bf1c5650440db591a9b5d53f

Observation a643e086-5fad-41cd-8288-23c100320f15 · outbound

This paper cites From pixels to graphs: Open-vocabulary scene graph generation with vision-language models.

Open World Scene Graph Generation using Vision Language Models From pixels to graphs: Open-vocabulary scene graph generation with vision-language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.530043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.137587Z digest=sha256:23698492a9ade79f91ef171c2139ab0dd49db429a66b85888d1daddfeb9f6504

Observation 2d53b965-e8b6-454c-b67f-ab098aab04df · outbound

This paper cites Gps-net: Graph property sensing network for scene graph generation.

Open World Scene Graph Generation using Vision Language Models Gps-net: Graph property sensing network for scene graph generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.523285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.139572Z digest=sha256:821cf2f5ec0509ccec685be4b3004207b92ba5cb3d0be682ed4c26ab7d9af838

Observation e5deb73b-4c2c-4de7-bd78-7351c852f8e3 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Open World Scene Graph Generation using Vision Language Models Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.141548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.141548Z digest=sha256:0e9b1ac97a80702b5c2a76c3be6b188b1614409cc692ff28c419ea5322c95d1c

Observation ba7951ae-a5d1-4c26-98c6-f0aad3cb06fa · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Open World Scene Graph Generation using Vision Language Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.143903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.143903Z digest=sha256:34142c48aa95cb54ecef415fdf34d7e6da74a104b68ea5d36d395a4e9147cc1d

Observation d68dbdc3-7d4b-4cf5-95da-3c6769641e7d · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Open World Scene Graph Generation using Vision Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.146147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.146147Z digest=sha256:a01acd2b1737af49cb2a80853594edb94556ad3771ca802c99e89dde6ce12ed1

Observation 4f4b881e-e4f8-4a2b-9b7c-e647a1c58216 · outbound

This paper cites Relation-aware hierarchical prompt for open-vocabulary scene graph generation.

Open World Scene Graph Generation using Vision Language Models Relation-aware hierarchical prompt for open-vocabulary scene graph generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.505101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.148127Z digest=sha256:b7117031173adfa2a542d04b4671ee4c7754e414494b1b8cb16d63e4c722a8a9

Observation 6637806e-20e0-4cd8-9a6c-e7ef2fca36c3 · outbound

This paper cites Visual relationship detection with language priors.

Open World Scene Graph Generation using Vision Language Models Visual relationship detection with language priors

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.498561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.150609Z digest=sha256:efb190b5c440c54309c10c2734d56901c06f87102f3e2234e39019a0fbf96f0d

Observation 988bbf7c-5eca-4d04-8db9-ae985e5cb788 · outbound

This paper cites hello gpt-4.

Open World Scene Graph Generation using Vision Language Models hello gpt-4

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.491902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.152786Z digest=sha256:4766de2cd7490efafe76156699db973f64d25fd506f0975463c297dac08b8051

Observation 0eb9fa80-0d12-430a-a583-b5d2f23f67e6 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Open World Scene Graph Generation using Vision Language Models Learning transferable visual models from natural language supervi- sion

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.154832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.154832Z digest=sha256:f9f8544bd61f9feed26bdbc4bc90ff97ee4d8ef641670373486923658f41cee4

Observation 99885462-b79e-4011-8808-e3941a838617 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Open World Scene Graph Generation using Vision Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.156964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.156964Z digest=sha256:fa2a81b1690f915bbd0a1309304b2905141a668916dcee6873af235d1bfa7b49

Observation 812dc2d4-3b6a-4774-8498-ebbc3dfc3623 · outbound

This paper cites Learning to compose dynamic tree structures for visual contexts.

Open World Scene Graph Generation using Vision Language Models Learning to compose dynamic tree structures for visual contexts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.481150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.159122Z digest=sha256:8cd77af1e07e53aba265f75b76b57820d6bd62beda5cccab4a25961602f10877

Observation 503d2f66-2ac8-4fb3-8522-b3c0e83d734c · outbound

This paper cites Unbiased scene graph generation from bi- ased training.

Open World Scene Graph Generation using Vision Language Models Unbiased scene graph generation from bi- ased training

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.473950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.161392Z digest=sha256:77cb8d25b4be33a818c82b22f1f1460221f65e298292bab1032f20aeca65cc88

Observation 65b3abb5-eb71-4ee9-b3a4-86a26d7bf75b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Open World Scene Graph Generation using Vision Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.163539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.163539Z digest=sha256:617f71a319221e7cf371b02a57ed50a32230a4b59dfca4355f4b1708d17a7415

Observation dd5f0b44-2676-4109-b62b-4b17ab9b1d5e · outbound

This paper cites Graph-structured representations for visual question answer- ing.

Open World Scene Graph Generation using Vision Language Models Graph-structured representations for visual question answer- ing

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.466765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.166021Z digest=sha256:f0788c268ebfabdb145ce2518a91705400c4b46732104727902fc16b7b2ec66b

Observation f63bdcee-c9a4-4717-8cd2-4cd0d6011ecc · outbound

This paper cites Structured sparse r-cnn for direct scene graph generation.

Open World Scene Graph Generation using Vision Language Models Structured sparse r-cnn for direct scene graph generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.460126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.168519Z digest=sha256:ce384653a4619b2f9c7871b71fd35aadbafd23cf7409d38198274dec0697a9f1

Observation 95efe330-ac5c-42a4-88bf-5db75ab49aa4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Open World Scene Graph Generation using Vision Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.170659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.170659Z digest=sha256:a9c685281210341c25ea6713b3062e227e4431d9922bb66e5c979b0ab4ef027e

Observation de25351e-2698-4140-9294-4e3b9e37c319 · outbound

This paper cites Scene graph generation by iterative message passing.

Open World Scene Graph Generation using Vision Language Models Scene graph generation by iterative message passing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.453061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.173090Z digest=sha256:88d81079912578d6f02dd8a87ee76dc68dec781a128d462017411f06058a2915

Observation 61c721cb-7ad3-4e2d-a19c-af7d4c3059f5 · outbound

This paper cites Llava-spacesgg: Visual instruct tuning for open-vocabulary scene graph generation with enhanced spatial relations.

Open World Scene Graph Generation using Vision Language Models Llava-spacesgg: Visual instruct tuning for open-vocabulary scene graph generation with enhanced spatial relations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.446937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.175109Z digest=sha256:faa526931b49eebb3a7ff7d79df62292aebbd9431db8676ce7bed9203d610687

Observation 4c8bca80-1123-4596-8c8e-04df8d28fe3a · outbound

This paper cites Panoptic scene graph gen- eration.

Open World Scene Graph Generation using Vision Language Models Panoptic scene graph gen- eration

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.440066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.177220Z digest=sha256:4d328d4a201d3b9dddc9ed3d5cd9f0d99f7f06a0ba72276e7d0328cf6a8b5d91

Observation 9d37b875-eaef-4109-87f9-bdf8b9876244 · outbound

This paper cites Depth anything v2.Advances in Neural Information Processing Systems, 37: 21875–21911, 2024.

Open World Scene Graph Generation using Vision Language Models Depth anything v2.Advances in Neural Information Processing Systems, 37: 21875–21911, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.433926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.179475Z digest=sha256:ac224dec60819307c15709085a63c4249252db51c57f970f0dcd3cb48c77ff80

Observation 8fa9a057-99ad-4258-87da-e326a9954e9a · outbound

This paper cites Cross-modal rela- tionship inference for grounding referring expressions.

Open World Scene Graph Generation using Vision Language Models Cross-modal rela- tionship inference for grounding referring expressions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.426876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.181569Z digest=sha256:583a3e115775d26969f3da7a1cf0f798bc4c57ffc6dc9b9f06021ebed84f81cc

Observation ba671b8e-4c7b-42d5-91d1-bb93b9c4bb5c · outbound

This paper cites Visually-prompted language model for fine- grained scene graph generation in an open world.

Open World Scene Graph Generation using Vision Language Models Visually-prompted language model for fine- grained scene graph generation in an open world

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.418720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.183631Z digest=sha256:7e770352936c2b9e6a20e93667225ca89abfe8e0ef4c0d000d0e10e7ee27c493

Observation 0d6453c6-b1c5-4294-8718-c055ff3cd14a · outbound

This paper cites Open-vocabulary object detection using captions.

Open World Scene Graph Generation using Vision Language Models Open-vocabulary object detection using captions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.411258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.185895Z digest=sha256:a17a763b1f2338ad5e291db9ccbefad4eb16ddfff7807d9453f017bb14bb8f4f

Observation db98507f-16df-4147-ae74-ecbab09f577c · outbound

This paper cites Neural Motifs: Scene Graph Parsing with Global Context.

Open World Scene Graph Generation using Vision Language Models Neural Motifs: Scene Graph Parsing with Global Context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.187935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.187935Z digest=sha256:98a1e0a50eba641b69e0cb89903a167bd6f0eaf84e47a1fce35aab0cdb8717a9

Observation e0d9e044-a009-4de3-a517-797cc1b16b3a · outbound

This paper cites Graphical contrastive losses for scene graph parsing.

Open World Scene Graph Generation using Vision Language Models Graphical contrastive losses for scene graph parsing

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.403580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.190254Z digest=sha256:b13f549b79d026cf3c1715a6325ef107c17f9793c212c76ebc47b9486972e360

Observation eb752ca1-0f56-4403-8f17-e3ca98a7e91e · outbound

This paper cites Learning to generate language- supervised and open-vocabulary scene graph using pre-trained visual-semantic space.

Open World Scene Graph Generation using Vision Language Models Learning to generate language- supervised and open-vocabulary scene graph using pre-trained visual-semantic space

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.396290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.192405Z digest=sha256:91c86f2f99e90481da4c840249686d23c05323c8d9274baf7f673145654242b4

Observation d22aa359-343e-48d7-ad0c-99ab81d9d6c2 · outbound

This paper cites There is aXin the image.

Open World Scene Graph Generation using Vision Language Models There is aXin the image

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.389239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:22:12.194352Z digest=sha256:9db5ccacf135312855e956e0174c73a63067444bb175239b8ad44568d882ddcb

Pith citing papers

Observation 6e2d48bd-84b8-45e0-be97-51eef825d1f8 · inbound

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering cites this paper.

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering Open World Scene Graph Generation using Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T10:44:08.297552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:44:08.297552Z digest=sha256:531a2e4e2c26be28eb47cf62c3a3b25efccde00633f4d1bd0c3a7e5130e3f199

Observation 773895a4-abe5-4930-ac1f-05077e3c49f5 · inbound

SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models cites this paper.

SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models Open World Scene Graph Generation using Vision Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:58.209549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:39:56.448121Z digest=sha256:8252ea5406e8d61b96f5ff9ad190b7244138011a247399323c49a71e3bcda862

Observation 2f9aae3f-2b1b-492e-b041-e6c39371b4bf · inbound

GraphVid: Interactive Graph-Controllable Video Generation cites this paper.

GraphVid: Interactive Graph-Controllable Video Generation Open World Scene Graph Generation using Vision Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:04:59.338909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:04:59.338909Z digest=sha256:f66ae4fc546bbf36a845cb4068beafc8c0063e0fc3248c818bf965de12aa6195