Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:22:12.194352Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 3 inbound Pith citation observations for arXiv:2506.08189.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:22:12.194352Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:44:08.297552Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-14T20:42:58.205735Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c0fc7967-51b6-447b-b055-61b378c7e6ae · outbound
Open World Scene Graph Generation using Vision Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f25714bf-0374-43a8-a38c-dad42bafb371 · outbound
Open World Scene Graph Generation using Vision Language Models GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58085406-ab58-40f7-89c5-380f19210270 · outbound
Open World Scene Graph Generation using Vision Language Models Expanding scene graph boundaries: fully open-vocabulary scene graph generation via visual-concept alignment and retention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 12ae2883-1520-484f-a502-e06b47a57183 · outbound
Open World Scene Graph Generation using Vision Language Models Reltr: Relation transformer for scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 45(9): 11169–11183, 2023
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 347bb022-99ea-4d76-b760-7e683b7f98cd · outbound
Open World Scene Graph Generation using Vision Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1b86b8-b2a4-407e-8684-09326248bc2c · outbound
Open World Scene Graph Generation using Vision Language Models Prism-0: A predicate-rich scene graph genera- tion framework for zero-shot open-vocabulary tasks.arXiv preprint arXiv:2504.00844, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e752d72-4886-4e04-9aa2-c3206c38b5b6 · outbound
Open World Scene Graph Generation using Vision Language Models SimCSE: Simple Contrastive Learning of Sentence Embeddings
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4064992-17d0-47ec-9142-2b387f3f5749 · outbound
Open World Scene Graph Generation using Vision Language Models Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2effcd-34b2-470b-8d48-da7c88326306 · outbound
Open World Scene Graph Generation using Vision Language Models To- wards open-vocabulary scene graph generation with prompt- 7 based finetuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ebe5fd25-ee12-4f92-81b4-9fd2e32e8673 · outbound
Open World Scene Graph Generation using Vision Language Models Scene Graph Reasoning for Visual Question Answering
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39fd54ed-26af-4a45-a6aa-a5146ce61db2 · outbound
Open World Scene Graph Generation using Vision Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4b5ce4d-ae68-41f6-aab5-5e5c72e934fe · outbound
Open World Scene Graph Generation using Vision Language Models Enhancing scene graph generation with hierarchical relationships and commonsense knowledge
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 552fd6b8-9ebf-4403-a986-1561a10867d6 · outbound
Open World Scene Graph Generation using Vision Language Models Image retrieval using scene graphs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 696ec004-e2bf-40ce-9e50-a1f600cdf44b · outbound
Open World Scene Graph Generation using Vision Language Models Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e21e7d9e-d8eb-4955-8b94-c9698e1a00c8 · outbound
Open World Scene Graph Generation using Vision Language Models Llm4sgg: large language models for weakly supervised scene graph generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c181adb0-57cc-41bd-bad9-4d7bad47c8ea · outbound
Open World Scene Graph Generation using Vision Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 96942250-333e-4d51-94be-96c8e55e2d9c · outbound
Open World Scene Graph Generation using Vision Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 234f1add-b521-4887-b86a-e8b758effe78 · outbound
Open World Scene Graph Generation using Vision Language Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7060702d-f3f2-47d5-a9cf-5442e3405a86 · outbound
Open World Scene Graph Generation using Vision Language Models Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cc27b857-6ec1-4e1f-82c9-3807eda49d7a · outbound
Open World Scene Graph Generation using Vision Language Models Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b0f3a94-2ee0-4e36-8300-118e09f4fc59 · outbound
Open World Scene Graph Generation using Vision Language Models Sgtr: End-to- end scene graph generation with transformer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a643e086-5fad-41cd-8288-23c100320f15 · outbound
Open World Scene Graph Generation using Vision Language Models From pixels to graphs: Open-vocabulary scene graph generation with vision-language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d53b965-e8b6-454c-b67f-ab098aab04df · outbound
Open World Scene Graph Generation using Vision Language Models Gps-net: Graph property sensing network for scene graph generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e5deb73b-4c2c-4de7-bd78-7351c852f8e3 · outbound
Open World Scene Graph Generation using Vision Language Models Llava-next: Improved reason- ing, ocr, and world knowledge, 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba7951ae-a5d1-4c26-98c6-f0aad3cb06fa · outbound
Open World Scene Graph Generation using Vision Language Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68dbdc3-7d4b-4cf5-95da-3c6769641e7d · outbound
Open World Scene Graph Generation using Vision Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f4b881e-e4f8-4a2b-9b7c-e647a1c58216 · outbound
Open World Scene Graph Generation using Vision Language Models Relation-aware hierarchical prompt for open-vocabulary scene graph generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6637806e-20e0-4cd8-9a6c-e7ef2fca36c3 · outbound
Open World Scene Graph Generation using Vision Language Models Visual relationship detection with language priors
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 988bbf7c-5eca-4d04-8db9-ae985e5cb788 · outbound
Open World Scene Graph Generation using Vision Language Models hello gpt-4
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0eb9fa80-0d12-430a-a583-b5d2f23f67e6 · outbound
Open World Scene Graph Generation using Vision Language Models Learning transferable visual models from natural language supervi- sion
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99885462-b79e-4011-8808-e3941a838617 · outbound
Open World Scene Graph Generation using Vision Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 812dc2d4-3b6a-4774-8498-ebbc3dfc3623 · outbound
Open World Scene Graph Generation using Vision Language Models Learning to compose dynamic tree structures for visual contexts
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 503d2f66-2ac8-4fb3-8522-b3c0e83d734c · outbound
Open World Scene Graph Generation using Vision Language Models Unbiased scene graph generation from bi- ased training
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 65b3abb5-eb71-4ee9-b3a4-86a26d7bf75b · outbound
Open World Scene Graph Generation using Vision Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd5f0b44-2676-4109-b62b-4b17ab9b1d5e · outbound
Open World Scene Graph Generation using Vision Language Models Graph-structured representations for visual question answer- ing
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f63bdcee-c9a4-4717-8cd2-4cd0d6011ecc · outbound
Open World Scene Graph Generation using Vision Language Models Structured sparse r-cnn for direct scene graph generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 95efe330-ac5c-42a4-88bf-5db75ab49aa4 · outbound
Open World Scene Graph Generation using Vision Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de25351e-2698-4140-9294-4e3b9e37c319 · outbound
Open World Scene Graph Generation using Vision Language Models Scene graph generation by iterative message passing
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 61c721cb-7ad3-4e2d-a19c-af7d4c3059f5 · outbound
Open World Scene Graph Generation using Vision Language Models Llava-spacesgg: Visual instruct tuning for open-vocabulary scene graph generation with enhanced spatial relations
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4c8bca80-1123-4596-8c8e-04df8d28fe3a · outbound
Open World Scene Graph Generation using Vision Language Models Panoptic scene graph gen- eration
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9d37b875-eaef-4109-87f9-bdf8b9876244 · outbound
Open World Scene Graph Generation using Vision Language Models Depth anything v2.Advances in Neural Information Processing Systems, 37: 21875–21911, 2024
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8fa9a057-99ad-4258-87da-e326a9954e9a · outbound
Open World Scene Graph Generation using Vision Language Models Cross-modal rela- tionship inference for grounding referring expressions
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ba671b8e-4c7b-42d5-91d1-bb93b9c4bb5c · outbound
Open World Scene Graph Generation using Vision Language Models Visually-prompted language model for fine- grained scene graph generation in an open world
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d6453c6-b1c5-4294-8718-c055ff3cd14a · outbound
Open World Scene Graph Generation using Vision Language Models Open-vocabulary object detection using captions
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation db98507f-16df-4147-ae74-ecbab09f577c · outbound
Open World Scene Graph Generation using Vision Language Models Neural Motifs: Scene Graph Parsing with Global Context
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d9e044-a009-4de3-a517-797cc1b16b3a · outbound
Open World Scene Graph Generation using Vision Language Models Graphical contrastive losses for scene graph parsing
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eb752ca1-0f56-4403-8f17-e3ca98a7e91e · outbound
Open World Scene Graph Generation using Vision Language Models Learning to generate language- supervised and open-vocabulary scene graph using pre-trained visual-semantic space
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d22aa359-343e-48d7-ad0c-99ab81d9d6c2 · outbound
Open World Scene Graph Generation using Vision Language Models There is aXin the image
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6e2d48bd-84b8-45e0-be97-51eef825d1f8 · inbound
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering Open World Scene Graph Generation using Vision Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 773895a4-abe5-4930-ac1f-05077e3c49f5 · inbound
SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models Open World Scene Graph Generation using Vision Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2f9aae3f-2b1b-492e-b041-e6c39371b4bf · inbound
GraphVid: Interactive Graph-Controllable Video Generation Open World Scene Graph Generation using Vision Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.