Pith. sign in

Paper Citation Record · LEDGER

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models

As of 12 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2607.05716.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05716 v3

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T16:12:58.882161Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3dc51d32-dfc8-41b1-a004-1994cc86de1d · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:e2e5982b3c1a87347618dd0c6fad2407593fcb0da0710238d1408d19ec713bc8

Observation 6fcf583a-9dde-47f5-9976-532e3d5e6d1d · outbound

This paper cites Qwen3-VL Technical Report.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:0e4891a275f7728223401fa926d815f3e7cdd0b4737ec855a681055d24da85d6

Observation 93bb6b37-4803-4d71-afa4-fc20d61ac681 · outbound

This paper cites Unifying vision-and- language tasks via text generation.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Unifying vision-and- language tasks via text generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:75ff543e364e0b02bee42b27367239acbd3c22c6d7091d7c1ef2750af495e34a

Observation 621cc074-5131-41f5-b815-d1d515d337f9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:f51490eddda4a5a1a2ece400049f21027c6ead22b218a001f049b8041270fb5a

Observation 8c79a5ae-e61f-4544-b33e-d96d7766b491 · outbound

This paper cites GRIT: Teaching MLLMs to Think with Images.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models GRIT: Teaching MLLMs to Think with Images

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:03f4c09fc1dc679b6177c7dda464a405468c6a889e4fc224b0e2afad09db607a

Observation e29c0637-9707-4e2e-a854-96311e3071cc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:aacea5a802bc34ba2e94397e6416815e877c91a46604e7b8a1f612bf0af5bd86

Observation ccd58374-5857-4d06-bdb7-164b50a54d64 · outbound

This paper cites DeepEyesV2: Toward Agentic Multimodal Model.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models DeepEyesV2: Toward Agentic Multimodal Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:10a49f475b491cf8d786c5cafddb93b345af1872e0cfbf5aeb20c306a509e048

Observation e0a68e16-a5c9-4fa5-adf9-e85d34a0e569 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:765b4d692a7ef851e4897890f93c83eaff1b935771fc57ab7d9085bc77f7fa39

Observation 0c597887-75d0-4d21-ab67-7c58184babdb · outbound

This paper cites GPT-4o System Card.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models GPT-4o System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:860dbdae885365602180baa501683beba1b6c94dbbd28374d6ceff6125b2296a

Observation 144b49cf-a355-4511-8fe8-b615818d4be7 · outbound

This paper cites OpenAI o1 System Card.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:a9e636867eee00f002985b8428483749c775f502941f9bebaa386ca37325358a

Observation 73fea798-c620-4e7f-a31a-c404cc8c4a02 · outbound

This paper cites Referitgame: Referring to objects in photographs of natu- ral scenes.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Referitgame: Referring to objects in photographs of natu- ral scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:0864d0e71fd54ada71ab2df44a968b5f5f9ececa3aa37783ec5c1c36d2d468e0

Observation 8bfce708-47b5-4aef-af18-f2198e5e2961 · outbound

This paper cites B., Liu, O., Guo, P., Neiswanger, W., Huang, F., et al.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models B., Liu, O., Guo, P., Neiswanger, W., Huang, F., et al

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:a113aa4992e4f85179309efa9705c25b209bd1b134b71cd5c27b7d4e989770a2

Observation eaac7e21-34d8-4a19-9570-e6c8f58b5eb7 · outbound

This paper cites Latent Visual Reasoning.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Latent Visual Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:dafb80fa0c3446b4fafa2d12076e4cec3cf7c5c95d6a56410cb7d1e23a8021e4

Observation fe682f3a-5dde-4368-aa13-7a02f928b244 · outbound

This paper cites an unresolved cited work.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:62d9ca792875a7116b5b721059fa20a488b232133b798c44fe3c8791e0783e3c

Observation bfde5013-4da9-45c6-b6e3-ef7bc87b7c7b · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:40e55a6b1e3545a33b5ac9e9798bee87c41cc1e56bf9d67619ff50dd5556f311

Observation 28cde4be-7dcf-4d3d-893c-76029b9f4b60 · outbound

This paper cites OmniCaptioner: One Captioner to Rule Them All.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models OmniCaptioner: One Captioner to Rule Them All

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:526ea8c586ef010de3b1b8c1a36371137e7a583dcc54cdc5a7188d5f53393242

Observation 8d95e82a-12cb-40a1-92e0-b63c09812553 · outbound

This paper cites L., Tan, J.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models L., Tan, J

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:101ece8c9de414e3cb1f006cafd88cc45ca7d1ab6e5a9737f11652f1fa20b2b3

Observation 8d6d2def-5f79-4e08-a815-29d8a719e7e2 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:ef0bbb7de15dfc05e0b92b577720dd38d06e12c6b62afa35d9e5d0d3550ddb98

Observation 77e876e6-f453-4939-93eb-07ec8a36a1e0 · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:28d0ba3bf04c294e6d7cdc9f80e792c56bd81afbc728e2bd55785eb648ba103b

Observation abf6fdf4-2cc8-41de-9c11-c074a12d27eb · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:238b1d4e3c306cd9494b4f72e60c55dc23bf127b13df1bef50bdf8643449d7d6

Observation 410c8ca7-d0af-41d3-9183-ab82f413ac7b · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:2f7774fc25e08647ba4dfe7b8d543ba1622aca136086b7974faf1f6d362bb519

Observation 6fd10c4c-b8fe-45d9-a9dd-090016831bec · outbound

This paper cites Qwen3 Technical Report.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Qwen3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:0ffc2fd4e2d91b582b26e09ac65c1af271e089a8fd76c89a165b5678f9c340dc

Observation e2d3edb7-b771-4f98-82f2-678bc8a38447 · outbound

This paper cites Look-Back: Implicit Visual Re-focusing in MLLM Reasoning.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Look-Back: Implicit Visual Re-focusing in MLLM Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:e94b6304d8c462dd87902311126fcd1ce57cc092f752d233c9498be80a3984da

Observation dafdcad7-2020-4db2-8637-1f0acada9b85 · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:d1b074ef242e79d8a511af99dac754db5703dc2a659f70ce7b20932e37974342

Observation e371c26a-8b38-4a8b-a463-8ee95d19106b · outbound

This paper cites Thyme: Think Beyond Images.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Thyme: Think Beyond Images

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:ccf5fd61e85a22726f0b0db017134eaec2fa460fa00195ad2784dc47e83af64b

Observation 7c359854-d584-4012-99cd-f014de4eb92e · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:f1d4e446b3b4cb6f30877f0900c664d8121bdb0f99f49ae816a8cb990941d7e6

Observation 438da7f8-7e9f-4ca0-879f-6c20aa6c4425 · outbound

This paper cites an unresolved cited work.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:4bcaf3c0d2a86cc6fedae4b7be857b86ad70b5b70a24ab83cee092bdc9393e32

Observation 28840800-a621-431a-92d3-6c803894852b · outbound

This paper cites More Results and Discussions B.1.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models More Results and Discussions B.1

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:7423bd6be47fe5205a2a8e13253c3149e72a9aeb28994b3a6978a8e8767d5140

Observation 6727996e-04e7-4e98-945c-ab1ce426f8a6 · outbound

This paper cites bike on left\.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models bike on left\

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:982343da6258cdbf3614e95f8ebff475515c3f50e962565c36e69f581e167386

Observation c7436bae-73af-4a61-b7c1-c34cea9bf71c · outbound

This paper cites Together, they effectively promote graph-aligned visual reasoning.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models Together, they effectively promote graph-aligned visual reasoning

Reference 30

Resolution
malformed identifier
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:dd6736ecc1c73d705bcd2d23050924b88cc2f825a0df93f371b3c13dc2483cf6

Pith citing papers

No inbound Pith citation observations are available.