Pith. sign in

Paper Citation Record · LEDGER

Synthetic Visual Genome

As of 7 August 2026, this Paper Citation Record lists 100 of 136 outbound references and 0 inbound Pith citation observations for arXiv:2506.07643.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07643 v1

Coverage vector

measured 100 of 136 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:34:56.485538Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 136 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved89
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9ca91f4-db87-41f0-9308-5317b565070e · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Synthetic Visual Genome Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.127801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.127801Z digest=sha256:4fa90f590e4e45af9590dfb4d3e7c82505853d8387f7929e6b4cb98be83f44fe

Observation 73fd8e3d-13fd-4bc5-91dd-16ced5c8517d · outbound

This paper cites Tallyqa: Answering complex counting ques- tions.

Synthetic Visual Genome Tallyqa: Answering complex counting ques- tions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.132663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.132663Z digest=sha256:9ffb023c677cd06d21e8e056c8a8670554935caa194fcf2c92f6251ed79f1a04

Observation 843123ab-78cf-4147-84f2-f2d8413c7631 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Synthetic Visual Genome Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.136278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.136278Z digest=sha256:d21264faad14807e1534ccf42078766df48bbfbcc9f1b57573e9031addd084ea

Observation af9a5cfc-9c79-4443-80c9-74e5fe85537e · outbound

This paper cites Qwen2.5- vl technical report, 2025.

Synthetic Visual Genome Qwen2.5- vl technical report, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.140227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.140227Z digest=sha256:b22df982c8bd2fb27c85ef4e58bfbb24f8398ebc0e00e8c1455cc4e22caa109f

Observation 66014a5f-486a-4ef1-92e2-077b1527a6a1 · outbound

This paper cites Recognition-by-components: a theory of human image understanding.

Synthetic Visual Genome Recognition-by-components: a theory of human image understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.144144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.144144Z digest=sha256:f56a8bb98dd13c404b3dfefccf352a0f63eb9617cb4085a3a8a03c164b00a17b

Observation 6f78d3f8-944f-4c9a-9ef2-f66756f1db8b · outbound

This paper cites Scene perception: Detecting and judging objects undergoing relational violations.

Synthetic Visual Genome Scene perception: Detecting and judging objects undergoing relational violations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.148080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.148080Z digest=sha256:140b480d0a17757c1d61a228fc914f1050867bd466df4982a97f7433b82b054e

Observation 621a401c-846a-4fd6-afa5-c5a7bb983f12 · outbound

This paper cites From machine learning to machine reasoning.

Synthetic Visual Genome From machine learning to machine reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.151726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.151726Z digest=sha256:766523f077bf935b79307af5b5b584e42494cc7332545822bfd4ba9acb6a57ae

Observation cd08e4b6-ef5d-4ff9-8482-d8609d25b6ac · outbound

This paper cites InternLM2 Technical Report.

Synthetic Visual Genome InternLM2 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.155396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.155396Z digest=sha256:8b01cc08b12a155e40620bbbaad67ce6107678340aab2cc339f6dd669872df09

Observation 35021869-bf88-43f1-9460-cb3d2a0777e9 · outbound

This paper cites Spa- tialvlm: Endowing vision-language models with spatial reasoning capabilities.

Synthetic Visual Genome Spa- tialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.159529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.159529Z digest=sha256:80cb0ad3a8c85bd074b67d3200881ab3e82f6db1e83655d5ec51c7157887f080

Observation 4f834ca1-588e-46b7-89c7-957a58a15f21 · outbound

This paper cites Scene graph generation with role-playing large language models.

Synthetic Visual Genome Scene graph generation with role-playing large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.163039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.163039Z digest=sha256:1e7704fad3131d506b8dfb55672f91dbeb50857e0794acdc057d3a76ca8ab6eb

Observation 8358cc76-c7f8-4885-b7fb-332f9af45017 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Synthetic Visual Genome MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.166607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.166607Z digest=sha256:7a16b4b45ebd4160c9fe49afdc0d3165952139cab8e7067a1c1ce622fd2fe27a

Observation 1fa95942-7b78-406d-8936-b133ca531c27 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Synthetic Visual Genome Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.170431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.170431Z digest=sha256:82630bfb7e38f745e8adf2cfd3d3e60b1263c4e532284fe1aa678d5d62a88da4

Observation d244ef11-44d7-4702-b2a3-f7b510f15722 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Synthetic Visual Genome ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.174628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.174628Z digest=sha256:6388389c756dbd12bf3a247e94894225dfd688e45907e80a6fbcafbeae0099ab

Observation 8c5458fc-0d9c-488d-89e8-0b4b86bb8343 · outbound

This paper cites De- tect what you can: Detecting and representing ob- jects using holistic models and body parts.

Synthetic Visual Genome De- tect what you can: Detecting and representing ob- jects using holistic models and body parts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.178591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.178591Z digest=sha256:641a1686d11b5c71656c09647077e3fb5d373de3d558ddee28fc0efe298ba991

Observation f691a715-d436-4257-8bcd-78060c30bf35 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Synthetic Visual Genome How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.182129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.182129Z digest=sha256:f81fda386e66dc602147e1c89e48d4a5efba8f05591635cf72fa73e38e7b3828

Observation eef70ba1-f99b-4921-844d-63719bb87971 · outbound

This paper cites Some contro- versial questions in phonological theory.

Synthetic Visual Genome Some contro- versial questions in phonological theory

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.185764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.185764Z digest=sha256:79abbfea3d82f61445bfe2ff779ae32a67ddb3cc2504b621bbb0d19770c47875

Observation 81756ea0-0dde-4d05-999f-575bd59ad68b · outbound

This paper cites In- structblip: Towards general-purpose vision-language models with instruction tuning.

Synthetic Visual Genome In- structblip: Towards general-purpose vision-language models with instruction tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.189572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.189572Z digest=sha256:e4f1f2332639f6f059865740c13f082a1b0c2165c22f64bc7f55738931502a4f

Observation 58f543eb-c494-422b-adc3-d91eb5814d1d · outbound

This paper cites Data Filtering Networks.

Synthetic Visual Genome Data Filtering Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.193696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.193696Z digest=sha256:e569d5a1ef2b9afabe38ab4bb812ae15d6d8cc5599364b8f522d7b16a1896074

Observation 395a5a37-0c82-48f3-a1a5-3d79e44675ee · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Synthetic Visual Genome BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.197159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.197159Z digest=sha256:e380e339d62a712f3a377b517a71bace01712394a5b19742bb1fbeb8e8ba327f

Observation b73cc2f4-d977-494e-aa5a-28d386d440b8 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Synthetic Visual Genome Datacomp: In search of the next generation of multimodal datasets

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.200921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.200921Z digest=sha256:66c6c0edc5c9fc28b3b9cd5dd5266d509f1188913ecc0c90058de78be1524a02

Observation d2d29edd-9cbb-4257-b5d7-93b45b8f30df · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understand- ing in visual question answering.

Synthetic Visual Genome Making the v in vqa matter: Elevating the role of image understand- ing in visual question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.204191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.204191Z digest=sha256:6f7bfedfbec0c0d1e2819ba7046941e7f3dbc419fb673794904f4ddd3ee61237

Observation 7fd1fe8b-843f-4466-af2e-9ff1de480b38 · outbound

This paper cites Agqa: A benchmark for compo- sitional spatio-temporal reasoning.

Synthetic Visual Genome Agqa: A benchmark for compo- sitional spatio-temporal reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.207859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.207859Z digest=sha256:a81b5f18a439616268f1b9f9b015a5327bd4944ad3688451e81da281036f7244

Observation a66fd0b3-7f33-4b05-9ecc-14e3e9fb574f · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmenta- tion.

Synthetic Visual Genome Lvis: A dataset for large vocabulary instance segmenta- tion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.211197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.211197Z digest=sha256:0e804a9f3ad6015acbef32bb620491b46fe69e3744ec5229ba299d79c587a614

Observation 154f2970-d1c7-477c-93c8-ddc2a9d89bf3 · outbound

This paper cites Dsgg: Dense rela- tion transformer for an end-to-end scene graph gener- 10 ation.

Synthetic Visual Genome Dsgg: Dense rela- tion transformer for an end-to-end scene graph gener- 10 ation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.214266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.214266Z digest=sha256:80a2d234ede92c08abe6ccd71131c7f7bcd23dc058dd5fb8c92f45005f046cf9

Observation 190fa380-8023-4ef3-bfb2-d265879ed3ef · outbound

This paper cites Partim- agenet: A large, high-quality dataset of parts, 2022.

Synthetic Visual Genome Partim- agenet: A large, high-quality dataset of parts, 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.217733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.217733Z digest=sha256:ea73f2852a56127eb496b05f62ff2673b72c1813cdaa63d52a6ff0f4bccb80dd

Observation d507d362-1e83-4122-8850-e25e00a3b9dc · outbound

This paper cites The curious case of neural text degenera- tion, 2020.

Synthetic Visual Genome The curious case of neural text degenera- tion, 2020

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.221442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.221442Z digest=sha256:23ac54b3b14adb5fe097e697f6e14f70d034df0dedc99954fd5827a9bd633b69

Observation a6b91b7f-3fb6-4f52-a138-ac68540b9a43 · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language composi- tionality.

Synthetic Visual Genome Sugarcrepe: Fixing hackable benchmarks for vision-language composi- tionality

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.224704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.224704Z digest=sha256:1830f8f6016626eb9c0709f7f4f4876ba2ff854b752df080079aacbd473ebe6e

Observation 8653c0d0-3872-4fc3-b731-81a48967ace8 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Synthetic Visual Genome Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.228027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.228027Z digest=sha256:f310d0a6ecaf82ebfbfbc916d43050fc5f6d94c5c9d1deb09b32ec8c638e9e91

Observation b55462a7-4705-4134-b514-5ab167e9c7af · outbound

This paper cites Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intelligence Research, 67:757–795, 2020.

Synthetic Visual Genome Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intelligence Research, 67:757–795, 2020

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.231437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.231437Z digest=sha256:5e7de652611457559ca91fb3192111a4dc6afec18fc34fef432052920263d327

Observation 38c3e5d5-f37f-422b-ab2c-fa3c6f24cdf9 · outbound

This paper cites Openclip, 2021.

Synthetic Visual Genome Openclip, 2021

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.234735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.234735Z digest=sha256:f0aa2fe3aa381365e730777f42f00a75e00851ed54f6c283ae531d9aef8c11d9

Observation 5eb09afa-beb7-4308-8720-fc42fe779bdf · outbound

This paper cites Composi- tionality.

Synthetic Visual Genome Composi- tionality

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.239075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.239075Z digest=sha256:429bcfcbaf6caa718217745b35dc64188cccf4b89fbd6fdf73bd5f277bea207e

Observation e390da5c-37bf-4155-a1ff-cf3b648340c5 · outbound

This paper cites Action genome: Actions as compositions of spatio-temporal scene graphs.

Synthetic Visual Genome Action genome: Actions as compositions of spatio-temporal scene graphs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.242711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.242711Z digest=sha256:5ba2f6d2793f0d11f426be57ce28f88b31e939eedf149a72a5d6114183682a64

Observation e680ed8a-74b7-43bb-a857-0c7330884c8f · outbound

This paper cites Dvqa: Understanding data visualiza- tions via question answering.

Synthetic Visual Genome Dvqa: Understanding data visualiza- tions via question answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.246134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.246134Z digest=sha256:0e509b3cbb211f5bdd2eb3c4aab1edf88551886b89e9b9d9142c471e0bb99c00

Observation 3637a4dc-3f5b-473d-ba09-dd7a31d99e13 · outbound

This paper cites What’s “up” with vision-language models? inves- tigating their struggle with spatial reasoning.

Synthetic Visual Genome What’s “up” with vision-language models? inves- tigating their struggle with spatial reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.249666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.249666Z digest=sha256:013ac7c903dfbee51a8810aec04efd3145601384edc9a08e4e67dca4397652ef

Observation 95fe4f56-daef-4ca4-a7de-5b3962301881 · outbound

This paper cites What's "up" with vision-language models? Investigating their struggle with spatial reasoning.

Synthetic Visual Genome What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.253961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.253961Z digest=sha256:8f657381121a91b5f6ba9ed5b0f0bd29126822adcc1c4c850e0249b74f212027

Observation 759036f5-8542-4cb1-bcc5-5065eaf3e1a7 · outbound

This paper cites A diagram is worth a dozen images.

Synthetic Visual Genome A diagram is worth a dozen images

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.258047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.258047Z digest=sha256:2b02c737583c5b870e0471b1634b3b05e73bb9553bd5107166a7b4658e64948e

Observation 08a2d4fb-fef0-4c76-9bfc-27edab966a79 · outbound

This paper cites Are you smarter than a sixth grader? textbook ques- tion answering for multimodal machine comprehen- sion.

Synthetic Visual Genome Are you smarter than a sixth grader? textbook ques- tion answering for multimodal machine comprehen- sion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.261392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.261392Z digest=sha256:e38918d367dfa8281660c5c2c338de9b3711138092a021ec7bb7fe8e57e79dae

Observation a6c30a0e-ed8d-4f06-b883-65e6aa4ccace · outbound

This paper cites Llm4sgg: Large language models for weakly supervised scene graph generation.

Synthetic Visual Genome Llm4sgg: Large language models for weakly supervised scene graph generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.264742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.264742Z digest=sha256:9ffec515f24daf18a1fb575155f64c69529e98a5a797b91e9296e4c62df17f41

Observation cb438b5b-3fed-4aa8-b934-e8b48cb7dc27 · outbound

This paper cites Segment anything.

Synthetic Visual Genome Segment anything

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.268064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.268064Z digest=sha256:bd39d0b51b46ca190f283dc0a001593bfbeaead9ff7ca42e377c981ed61e848a

Observation a750bf21-1158-4bf8-a17b-7c2422b5bf1c · outbound

This paper cites Vi- sual genome: Connecting language and vision using crowdsourced dense image annotations.

Synthetic Visual Genome Vi- sual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.271522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.271522Z digest=sha256:4d04b956b29121f1bf7e26fa06fc4deedf2440b3f1dc47cb65542548585c3fa7

Observation be22b953-e045-4b7e-bea5-142f24c67f0e · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Synthetic Visual Genome The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.275008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.275008Z digest=sha256:68ebf6e17a2a62c59eb3725c422163695bd0401497e80ff7f90f9b6a442656a1

Observation 7dcedcd2-9f0c-4a1e-8edc-af5cc08444e2 · outbound

This paper cites Seed-bench: Benchmark- ing multimodal llms with generative comprehension,.

Synthetic Visual Genome Seed-bench: Benchmark- ing multimodal llms with generative comprehension,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.278599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.278599Z digest=sha256:bdc3bb75f8d89544efd6ac16ff12b5150dccfd3d74ef539890a14b06f8bc03b3

Observation 0a0653aa-0fc1-4e56-941e-561e8d0613b9 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

Synthetic Visual Genome MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.282191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.282191Z digest=sha256:081eb8815b3f8e1a5103e3c5ee1f3e2f9d521324f962692cc3baa75b77c912aa

Observation c563c1f1-84d0-439a-8963-fdd760c74134 · outbound

This paper cites Semantic-SAM: Segment and Recognize Anything at Any Granularity.

Synthetic Visual Genome Semantic-SAM: Segment and Recognize Anything at Any Granularity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.285746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.285746Z digest=sha256:1334e9a139600aa4ab28e175b20c15ef22c5a99963ad1e5642c79e0f98dba37b

Observation 766c363b-2435-485a-9584-d459cccf1692 · outbound

This paper cites Sgtr: End-to-end scene graph generation with transformer,.

Synthetic Visual Genome Sgtr: End-to-end scene graph generation with transformer,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.289546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.289546Z digest=sha256:17e39d695fccd42188543bd83d6fea1616fb7d43676cbf459edd205494e4ec4c

Observation 6a265df5-0964-4b0b-b059-1afef75098d4 · outbound

This paper cites From pixels to graphs: Open-vocabulary scene graph generation with vision- language models.

Synthetic Visual Genome From pixels to graphs: Open-vocabulary scene graph generation with vision- language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.293393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.293393Z digest=sha256:d1f19326c300cdbb5eb09aaa80c76a80440ea3b1f5553fa72f7174f4628154c8

Observation 635def2f-8e23-415f-9fc8-08656c6d96d6 · outbound

This paper cites Vila: On pre- training for visual language models.

Synthetic Visual Genome Vila: On pre- training for visual language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.296832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.296832Z digest=sha256:bd8cd0e80a8cd59ee461f7ca0b13b61077e6e6df9b61463c99476ea2911edeed

Observation c89f1ff3-2383-41d9-a704-350217b7092b · outbound

This paper cites Microsoft coco: Common ob- jects in context.

Synthetic Visual Genome Microsoft coco: Common ob- jects in context

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.300186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.300186Z digest=sha256:bd6dd18a139df18cb0aaab5106133d97e62cd424a8af129e8212a04d0348871a

Observation 1bbfaad7-dd63-4ff1-bb67-0ea68d42726e · outbound

This paper cites Gps-net: Graph property sensing net- work for scene graph generation.

Synthetic Visual Genome Gps-net: Graph property sensing net- work for scene graph generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.303622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.303622Z digest=sha256:052bc9282e8556126c5835df5fe4dd779c76b03b4adf1b136a09f8e04681eee5

Observation f65d36ee-664b-42cb-9df9-cc0cc0047354 · outbound

This paper cites Visual spatial reasoning.

Synthetic Visual Genome Visual spatial reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.307221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.307221Z digest=sha256:63f1815ce9b35dcbf859ebe03b486ffebac0c59c29513ee8841b4e678f1292f9

Observation 5e7f6793-a6dc-4679-aff5-253e0caae909 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Synthetic Visual Genome Improved Baselines with Visual Instruction Tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.311317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.311317Z digest=sha256:271d25ad30a7c0218f7e1f58b5805e9ca4e5c9c484f716b1b4252c4f5f7ad924

Observation 15a46ba1-9ef2-446b-9248-f2b7ee21e784 · outbound

This paper cites Visual instruction tuning.

Synthetic Visual Genome Visual instruction tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.314772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.314772Z digest=sha256:2efecaaa7192885f4b2f4d8b892b902d46dc6c719b4d67bf8f23c1ec31105aab

Observation 70118e8b-c904-40d3-9813-714408d96a84 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Synthetic Visual Genome MMBench: Is Your Multi-modal Model an All-around Player?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.318290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.318290Z digest=sha256:0733c4cc50e6b13d734510d146f3666584e1f4ed69d89828a4c143dfdccb03fb

Observation 1869da5b-7dd9-49ec-a545-6c6675d21df2 · outbound

This paper cites Visual relationship detection with lan- guage priors.

Synthetic Visual Genome Visual relationship detection with lan- guage priors

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.322396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.322396Z digest=sha256:bf428d6575bc32d3f60c853039449aea44793a03df1e7ad6b6f20ea7e9589ec3

Observation 26960e45-f922-484b-b6c3-8607b1f17359 · outbound

This paper cites Groma: Localized visual tokeniza- tion for grounding multimodal large language models.

Synthetic Visual Genome Groma: Localized visual tokeniza- tion for grounding multimodal large language models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.326676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.326676Z digest=sha256:2a94f7f996a7472a63cb3a52c2888a56c468b118f0f8a6cc9c15f8cd1b0a5f5e

Observation e733bdb7-1ab0-4e90-8320-9be7a0267600 · outbound

This paper cites Crepe: Can vision-language foundation models reason compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10910–10921, 2023.

Synthetic Visual Genome Crepe: Can vision-language foundation models reason compositionally? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10910–10921, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.330051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.330051Z digest=sha256:b667452bcc2907c809e45125fb96c289acfd86e86164c61c8b94deb10374c2f2

Observation 31eff2e7-03e2-44ae-b24d-53656fde7692 · outbound

This paper cites Generation and comprehension of unambigu- ous object descriptions.

Synthetic Visual Genome Generation and comprehension of unambigu- ous object descriptions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.333507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.333507Z digest=sha256:92e86ce058cc0c64415c1491d3edc3faee59ec076d365b8e3f459ed0571c9706

Observation 5d3683ad-d4d0-4391-aef0-f33890610358 · outbound

This paper cites Ok-vqa: A visual ques- tion answering benchmark requiring external knowl- edge.

Synthetic Visual Genome Ok-vqa: A visual ques- tion answering benchmark requiring external knowl- edge

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.337199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.337199Z digest=sha256:64d42094ff17db151e619df44cba8bbc83f51d98f1d336da25a2676fafe7c1b2

Observation 9ce88af0-17bb-4942-a2a3-475cb895b420 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Synthetic Visual Genome ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.340682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.340682Z digest=sha256:05c93d026f2fa43fa0b27f9b25bafa94c6167219616a17136793fd2648dccd2b

Observation fb224d5c-7609-41b4-aaa3-bcfa16c183f5 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Synthetic Visual Genome Docvqa: A dataset for vqa on document images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.344278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.344278Z digest=sha256:c67e02dba7aaa4c2ee9137269d5872e49c48735dd9dace77d0cb191501c28202

Observation b38f56ce-fee3-45bf-a545-1ee82014b56b · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Synthetic Visual Genome MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.348041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.348041Z digest=sha256:23310e9e38d7866f49abd0ffcc6880e5dc198622d512dc08739b6bd0d540d1f6

Observation b6949f4a-3c2e-4c82-a8a2-dcd9a9d65efd · outbound

This paper cites an unresolved cited work.

Synthetic Visual Genome Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.351500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.351500Z digest=sha256:876b90ac2382600263d395ae9d9b78235eb6066902abf19cfa56588a29103d60

Observation 97576187-3383-4a1d-926a-33255125e53c · outbound

This paper cites Localized Symbolic Knowledge Distillation for Visual Commonsense Models.

Synthetic Visual Genome Localized Symbolic Knowledge Distillation for Visual Commonsense Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:56.947759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.355778Z digest=sha256:a0daf00928a48964480c5fc1f77de6d2cfe43cf9731423b0e6937f55d2fe0aeb

Observation fb21d42a-4bde-4359-b3ee-1eeef744921a · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Synthetic Visual Genome Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.359405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.359405Z digest=sha256:3bf94cc595969736d04b46d087a6e72e18f4f96e5522f5cbb5b35add75b1869d

Observation 80ba8dee-bad6-40fe-9d86-df044fb927b7 · outbound

This paper cites Plummer, Liwei Wang, Chris M.

Synthetic Visual Genome Plummer, Liwei Wang, Chris M

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.362597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.362597Z digest=sha256:2a0615592100c3fceea289942a8efa4270716413843ccefe60ea8bc8ef1ff77f

Observation f6f9ad47-2c09-4d16-af02-02e7b913a977 · outbound

This paper cites Filtering, distillation, and hard negatives for vision-language pre-training.

Synthetic Visual Genome Filtering, distillation, and hard negatives for vision-language pre-training

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.366288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.366288Z digest=sha256:331df6fdcfc049836ae80e6138451cbdd1c92d5cc7f840f7918c1dca47058d88

Observation 1bdbdeca-8d3e-48e5-9c34-bf342391da66 · outbound

This paper cites Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sas- try, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever.

Synthetic Visual Genome Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sas- try, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.369936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.369936Z digest=sha256:05688ecba67e2e35a2a335787c5a8e84ea4bbcf46bb7a46660a98d5448844cd2

Observation a5049823-3312-44d8-b173-26fb24d6fac1 · outbound

This paper cites Paco: Parts and attributes of common ob- jects.

Synthetic Visual Genome Paco: Parts and attributes of common ob- jects

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.373066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.373066Z digest=sha256:b37b20c51ef741d7c3808fe538fa0a38ae8e46ccfce76c98e7fbe990b4704831

Observation 22d2aca9-be99-406e-815d-a7e7aa6b1e77 · outbound

This paper cites GLaMM: Pixel Grounding Large Multimodal Model.

Synthetic Visual Genome GLaMM: Pixel Grounding Large Multimodal Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.376390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.376390Z digest=sha256:6169ed30a3b7e364f4947bbd1587db4ad64e4f3f48dbf5be88ea27afa7deb5ff

Observation 8c95e9ec-8b78-4775-bd19-2cbad3b59826 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Synthetic Visual Genome Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.381288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.381288Z digest=sha256:ee195e095bbfc95aa1575caa2df181d3336d90702a4c307a8a51c5228a395ad8

Observation 60fd5677-cd24-47eb-ae44-fc60fea3642e · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Synthetic Visual Genome Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.384658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.384658Z digest=sha256:94c5891e3804183fe92ff5160b2930f11422c99fa93ba9abe1cafc588e169b3a

Observation 99543457-7f41-4236-9855-6ac962b887b3 · outbound

This paper cites Scienceqa: A novel resource for question answering on scholarly articles.

Synthetic Visual Genome Scienceqa: A novel resource for question answering on scholarly articles

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.388383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.388383Z digest=sha256:b28237dc8c1db4713d58847e86e2b3cb8501708981cc56ef22659a243d5d60ba

Observation 2057fdbb-aa2b-4505-b9f1-b5910432c66d · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Synthetic Visual Genome LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.391967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.391967Z digest=sha256:b022d67d3049362c557c01e1c8195ddffc775fb483efdd3db65fbb22598693ec

Observation 1431c41e-45dc-44f0-b397-12802c2e6d23 · outbound

This paper cites A- okvqa: A benchmark for visual question answering using world knowledge.

Synthetic Visual Genome A- okvqa: A benchmark for visual question answering using world knowledge

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.395633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.395633Z digest=sha256:e8612b309b51e08265ba87e0c4ebe819871877d676e398184a392014b671ba5b

Observation 62c8d675-5ca3-418f-9b30-17d5781275cc · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

Synthetic Visual Genome Textcaps: a dataset for image captioning with reading comprehension

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.398911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.398911Z digest=sha256:12ada5118a10b464700fe2e3ca6f49fc5d3c5d7ccb1e6b5d4ea24ecb483d8637

Observation 932b6292-a847-48c3-bbbd-467b58400c35 · outbound

This paper cites Learning to compose dynamic tree structures for visual contexts.

Synthetic Visual Genome Learning to compose dynamic tree structures for visual contexts

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.402470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.402470Z digest=sha256:650dc3d8469ef03bae4daac232a41535054b4bfbc377e84d010e519fd4b219f2

Observation d0fea6cf-5322-4cb4-9a35-899011ecac55 · outbound

This paper cites Unbiased scene graph generation from biased training.

Synthetic Visual Genome Unbiased scene graph generation from biased training

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.405973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.405973Z digest=sha256:aad16d5821061dedda9bf6b338712e798b32aa76d7244d86a45b1001d3e34170

Observation 36dcad90-d463-478d-9e41-baf1443bfe0c · outbound

This paper cites Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models.

Synthetic Visual Genome Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.409537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.409537Z digest=sha256:f4e72a365ec5cee26f22ca2099988c94f6b0b90690dabbdb9df8cafd574afeed

Observation 5b104c79-3288-4e23-94d8-f6495d894c86 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Synthetic Visual Genome CogVLM: Visual Expert for Pretrained Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.413258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.413258Z digest=sha256:1977ce25d767c499a1dc79fb990f5fe928f8315ffa51107cc9c73067f189c450

Observation 9c70c81f-6cc8-4285-b0ae-3ba391976aeb · outbound

This paper cites Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters.

Synthetic Visual Genome Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.416946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.416946Z digest=sha256:c995b24fe9b4b4adb385aa21ad046efb3edfcc84ba3c69130efa32e14ef8e060

Observation e7760941-faf7-4006-b315-c3ce95d6a50b · outbound

This paper cites The all-seeing project v2: Towards gen- eral relation comprehension of the open world, 2024.

Synthetic Visual Genome The all-seeing project v2: Towards gen- eral relation comprehension of the open world, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:57.563117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.420682Z digest=sha256:86d3978cfcbcd2e0d49d71d2a3b7b0f0b234535735ddd0135f9dcac727d5cc37

Observation eb482150-f3de-4131-8e6a-9f0a12eecc9b · outbound

This paper cites Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks.

Synthetic Visual Genome Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.424085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.424085Z digest=sha256:09dd179778125372b7422e5c8004b73e0a9598e32d4d192054951604f51ca34e

Observation eff2ee90-61f0-4ad2-88af-dc5d121c97ec · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Synthetic Visual Genome VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.427639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.427639Z digest=sha256:6b95c82969f6ad72a35e2798c18813532b5f271a0b6b2dd57282bd879d01d2dc

Observation 87b5ebf0-12e0-4221-84d3-d315d3dda227 · outbound

This paper cites Scene graph generation by iterative message passing.

Synthetic Visual Genome Scene graph generation by iterative message passing

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:57.551650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.431221Z digest=sha256:df36e77713d4c76e6997742d354c16e0bd881cfd7899730e21b44ec4a580e26b

Observation e17624a3-2449-4eed-9fcd-06ae8670e146 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

Synthetic Visual Genome xgen-mm (blip-3): A family of open large multimodal models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.434261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.434261Z digest=sha256:d67bdb125a19d171bd481dda3a63adb0b9d121826e649375b64ec8bf411ff9c5

Observation fb319264-96bb-4058-8a55-93a1f50e3353 · outbound

This paper cites Qwen2 Technical Report.

Synthetic Visual Genome Qwen2 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.437648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.437648Z digest=sha256:35a8e3a55dc176bde6c8f0d5ef7a608920a316e9992b7ab912d2cf6f7032d49c

Observation 09a79a29-10c2-4eae-b2cc-344f781e72f3 · outbound

This paper cites Panoptic scene graph generation.

Synthetic Visual Genome Panoptic scene graph generation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:57.541562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.441070Z digest=sha256:bee37fcd7240389679fad99f545aff04ca1f130f49d6a6d0b4d043a913a02be5

Observation 490020ae-2146-45f2-9d49-c7f6978ec42a · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Synthetic Visual Genome Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.444424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.444424Z digest=sha256:e6a34bc0712867eb5f4fd8588cc47e497eaa4ea152d62d56cbcebeda48819dc4

Observation e564cc03-3c19-4c0a-bca7-aed2b9d4722e · outbound

This paper cites Depth any- thing: Unleashing the power of large-scale unlabeled data.

Synthetic Visual Genome Depth any- thing: Unleashing the power of large-scale unlabeled data

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:57.531233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.447428Z digest=sha256:cdffd67d23f5c3314146607a5bbc70056c72f3edb4c996cdd15cf7cb801c7d40

Observation 0b651cdb-e637-436e-93fe-b8a1eb327107 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Synthetic Visual Genome MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.450803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.450803Z digest=sha256:d6ec20079da38a87223710ba061f82ce96211636235105696c9e52980215c2f2

Observation 3ea739e7-06be-4d83-b0a8-81b1af7d4232 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

Synthetic Visual Genome mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.454152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.454152Z digest=sha256:42fd337ffffaf5bee085953ae4e44f4a4cc56cf811594d60b6a79781f6c7c542

Observation 48f06bb5-545d-471d-9d18-ede5f479399b · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Synthetic Visual Genome Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.457915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.457915Z digest=sha256:b3e4415368488a64018bd6e34cd50b85643b052a3a1937d645054fd825ff4a78

Observation fe124c37-9cd8-4e7d-93a7-4452f29765aa · outbound

This paper cites Modeling context in referring expressions.

Synthetic Visual Genome Modeling context in referring expressions

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:57.521049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.461550Z digest=sha256:18c034609b5de675e3422348a44a0bdfd7cd9d4c57b4ae1b22d52d20d9c7c88c

Observation c17ffcaa-4999-47a3-b668-0dc29a7a7116 · outbound

This paper cites Modeling context in re- ferring expressions.

Synthetic Visual Genome Modeling context in re- ferring expressions

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:57.511279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.464682Z digest=sha256:af63447cdde9f9a4d1262b29f4acae5d631a3234a8c357b840796c633c616f55

Observation 4f8809ee-4581-45f2-810f-dae7712da24b · outbound

This paper cites Rlip: Relational language-image pre-training for human-object interaction detection, 2022.

Synthetic Visual Genome Rlip: Relational language-image pre-training for human-object interaction detection, 2022

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:57.501079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.468011Z digest=sha256:0d4fa181e055292b7a2d8b139e96d0493b73ecc5c5ec0fa2df9bf60e1a9ad99f

Observation 3af32829-0125-4437-b825-8da297d5b11f · outbound

This paper cites Os- prey: Pixel understanding with visual instruction tun- ing, 2024.

Synthetic Visual Genome Os- prey: Pixel understanding with visual instruction tun- ing, 2024

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:57.490871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.471582Z digest=sha256:290e8d2a26a39992b49c9959ffb9d470a44e7c397407bd5a39b153d12224e8b0

Observation 8ad533fb-5a44-41fa-90ef-c9d01c30bbff · outbound

This paper cites Neural motifs: Scene graph pars- ing with global context.

Synthetic Visual Genome Neural motifs: Scene graph pars- ing with global context

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:57.480480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.474880Z digest=sha256:9f03ceea4e852fead8879db09dcf86f96e1ebd823c5b2766b6900f9c1bff67d8

Observation 20a541f8-b032-4737-8af2-a408fd3198b7 · outbound

This paper cites From recognition to cognition: Vi- sual commonsense reasoning.

Synthetic Visual Genome From recognition to cognition: Vi- sual commonsense reasoning

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:57.470192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:34:56.478428Z digest=sha256:be0dfd1fa7d6a44c122e52113bc552dfb1025472768eae6173def314582eef5c

Observation c210b42b-8c9b-405b-81ac-53a1ea037dab · outbound

This paper cites LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models.

Synthetic Visual Genome LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.481852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.481852Z digest=sha256:e02b9fbd34ab942c18910c55fc2994cd89bf8dd822ea972edb91a4235a21e49e

Observation 2b467c74-ff80-4c1a-951d-140c60b59b96 · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

Synthetic Visual Genome MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.485538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.485538Z digest=sha256:b60f812a90f38ae8179dff97bde51320ebe0e57bbb4bdfa74fa049d6b44fbe16

Pith citing papers

No inbound Pith citation observations are available.