Pith. sign in

Paper Citation Record · LEDGER

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning

As of 18 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2505.09118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.09118 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:46:32.749086Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:11:02.464556Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:14:20.876165Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved65
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe2c70e0-cff6-416a-8dd9-27f12a166e6f · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.484993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.484993Z digest=sha256:c6ad56c6e8c619fa09f978d9fc1ff511dd219090f9cd665ecdcf478882c430f4

Observation ca777a10-efe4-49e8-971e-c61244dbff29 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.524290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.489879Z digest=sha256:aa237db0a0374d82ddb7466b18666936d43b57bbb8daaa1b3d3ba97661f9bcf9

Observation 64b013b5-42c1-4edc-86c0-dad766c0cf2f · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.511654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.494023Z digest=sha256:8e06527c9836a0524dced7baeab70185ca833645bae05ea1fd46a3cff1ec4736

Observation 7650f597-ab3b-4af8-a7d1-64b447d35949 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.498304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.498304Z digest=sha256:17b6e48558f7145fcff9ec3a309b09296388a9aae68996e0704c92a300366dfe

Observation 6c20f0be-8851-46a4-b78f-fdc4ce3987be · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.502700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.502700Z digest=sha256:35f396b95126e1af2fb7da9d6f4a00aa222597b6f37eb423116b086735b286e0

Observation 108af4b2-e4e5-4c5f-b8d4-4459901a9392 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.510672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.510672Z digest=sha256:d56aba0060d33b677c24da38a07b0c5bc9f9fbc88a029a366b3a0c5169d4e56a

Observation 693a4800-8943-44f8-9734-a9b836fa72bc · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.485429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.514846Z digest=sha256:0a871c31ccb1168229efd1e04b0b1810c423611df3ad23f31addfef0dc0e333b

Observation e9055d57-2b38-465e-89f1-8c0d20c63852 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.518348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.518348Z digest=sha256:5371fe1b765447b5f3c708b1724975b9cded421c3b8f34a3660b528a43fe9cfe

Observation 135417eb-5952-45b9-890c-037c4038ca25 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.522122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.522122Z digest=sha256:97ddefcd68f0e97a435fb2ae779cf9afdb45a0f78f0f1bbb18d3d0e6d7923488

Observation 5ceaa70e-6c74-46be-b9b4-f122b067ff33 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.526083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.526083Z digest=sha256:d58f3653c311375adbce722bed914a81a51b99325958e7f3f91aefb45bf5ab4c

Observation 86018529-ba52-4dd5-8e0b-69def92f17f7 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.454478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.533273Z digest=sha256:2d09919cfeb07463d1b5daf714b63ec4a31147a9458b21b9ce4b731f9b73e321

Observation a468037d-4d17-4671-b329-215a4293dbf6 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.536528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.536528Z digest=sha256:fe5af888792efd8fcda78256c69d4b06c563a6abca748b0cd8f9ed15013c9038

Observation 7a35c10b-5dea-44b3-ba4f-6b4e78c5b470 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.437179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.539807Z digest=sha256:ccc8f2becf1296029ad5b09d22fa4f500714b76bf47f514c562afd782edd7fb3

Observation 600bb012-2fae-457a-a976-47fd38fce154 · outbound

This paper cites Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.543218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.543218Z digest=sha256:2fbf145a593d11073beaacfd472529bfe86a8feeda39dec7e031901af964e8fd

Observation eba7cefa-5d9e-417d-bfe0-90e5b5b37306 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.426243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.546593Z digest=sha256:754d3f3d17ae69c1c12d811423ca6a10de4d7048a247e38c619473492951308c

Observation 5b2c63fe-e610-4d3b-af5d-652de6437c21 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.550215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.550215Z digest=sha256:980dc01218d7250ebfb03f9897f5ed2c8bec308aea74f9fa01739be4f79bd7c5

Observation 270663b6-84d1-4a17-b27f-740ec23feaa1 · outbound

This paper cites A Survey on Human Preference Learning for Large Language Models.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning A Survey on Human Preference Learning for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.554927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.554927Z digest=sha256:0b7df998bf8161d82d32c6d337c718692972c1a5c0ac18c7d934e0bdbc475047

Observation 9a6eed4f-86f2-4201-8207-b271bd33c015 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.559106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.559106Z digest=sha256:28e82d663962b1ef721f1a9e7045c79607e36db542f7e0d1049bbd063f1a3d2c

Observation 640c7e61-d748-4491-8ed2-6e70325b594c · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.563270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.563270Z digest=sha256:7f3acfc7f29e65b013b800d5ea8694eeceffe000e3e8f936d5ad68a2fb84371a

Observation 1ad0aa57-f6da-4ba9-8d63-5a999cfd9d35 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.571125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.571125Z digest=sha256:9f0501158f53baf529af7736fde339d7ea75c5580ff647e759777f0720fa9527

Observation 498ef88b-311d-48cc-86fc-f8fe5a54f288 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.377229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.575652Z digest=sha256:0bfbca5b7ef4f47e213de13f0733c9daf1280cacb0255763f8e1c9e4298d4f9d

Observation 9f394feb-2315-4413-93e7-aaaa9922f7b5 · outbound

This paper cites International journal of computer vision 123 (2017), 32–73.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning International journal of computer vision 123 (2017), 32–73

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.567261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.567261Z digest=sha256:450c35d160928ed8360f7e0967a211527ebcc5f173d0f84a6e60acc68286c8b4

Observation cea41c8d-645e-4e0a-8a99-a00ccf697ead · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.583403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.583403Z digest=sha256:bdf09d70aa9aa78e9bbaf5e3fdcaf5799fc45a570d55e20f06547ceff48d012a

Observation b0aca05b-7507-47bb-87fd-67883281d908 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.351466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.587506Z digest=sha256:a290b3a7ff9d9c1d27ba7a9e57bb27e009f1273ea4e9a6a0995ef86ecb8fafb9

Observation a91a7f99-530b-440b-b59d-91a0ecaf9af7 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.579501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.579501Z digest=sha256:f37ffd4811fafccba4a130a5fb458d8515ab967bfad4244689b94af57dc41fec

Observation 7724a97f-cfcc-4c98-8035-5d0dce267496 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.339568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.595681Z digest=sha256:f01e0dad229ee931a0222aa2d01bf9fd178c603767c79479e6622774236b0a8a

Observation f06a70cd-9d61-425f-bb38-756d8b196616 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.327445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.601176Z digest=sha256:bdcdd587710a03443277413641d253b52ef6c0e49cd9b4ed1f21c466bc964a98

Observation a29dd3bf-55c5-46f8-95f3-ea4ea4e223a0 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.591591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.591591Z digest=sha256:b347076f8eee116ba306a81bbfdac31f5e708b08652e0260d98965d14e695465

Observation ae71b79b-ec48-4838-b68c-6b51fd0aaf75 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.609115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.609115Z digest=sha256:fbc650b1d325593ee526e4a4fb79412fd3489a35b618230e8846dae37381040a

Observation c83e1f6d-cf8a-476a-a0b6-543752d6413a · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.616868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.616868Z digest=sha256:712fdea5354322e921504e9f98084cdc464e781e97a556cc8ee890e2879f485a

Observation 2499ede0-6221-44b3-ba0a-b9ef10247007 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.604995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.604995Z digest=sha256:9665e08bc9e0e4a0a5cdabe89d44d9e248b2e54906e6cfcbe523f7ad53fb87bf

Observation 6874933b-cdb5-4814-9102-2f4515047035 · outbound

This paper cites What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.624657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.624657Z digest=sha256:7568b436d6184c139ab2042b908df0cabb09b96ad2b9520672f6ff4b98983de3

Observation d2bb92f4-d370-4335-9e8a-06c5c3336ffa · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.612929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.612929Z digest=sha256:e3b54caa7ec05d3113e8218ad9460986e779daed1b78765107c562e03cf4741b

Observation 7c372cf1-b92d-4c13-9d85-821b14d4fbf6 · outbound

This paper cites #InsTag: Instruction Tagging for Analyzing Supervised Fine-tuning of Large Language Models.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning #InsTag: Instruction Tagging for Analyzing Supervised Fine-tuning of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.632306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.632306Z digest=sha256:154425aa4f53d1ca1c5b99e5dae588ab6f68fd2e5708469388acd898a1d7cb85

Observation 8e8dfa81-9264-4af9-8c5b-16ae82e793f9 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.294462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.620670Z digest=sha256:213114ef75594c5f3a8e641741aa7a39c5db9b60c7f6e7922528bebb99008693

Observation 2c10dfc9-4412-47da-9d5b-c01449dccc89 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.258723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.640109Z digest=sha256:00354efd482220149fcd845f5b44ec4c37a593850d6f7107ee55a1e9abd00c96

Observation 2aa3dd4e-4383-4ba8-81d3-fda8e3bf824d · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.628793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.628793Z digest=sha256:b262f26d67efa223a2a32328403d3bfa7d135ecea380f9758afc3e9a4aefd33d

Observation 48590838-8881-458e-921a-85c7cc21d896 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.648094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.648094Z digest=sha256:cbe8c5ecc58d7665b493b8dc6f4c1797ccc2fa226ad9ff34a4989100afd0e40e

Observation 3452de8b-0a78-4708-9f97-08585edcd770 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.272664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.636276Z digest=sha256:3dac8c63a1ac7bf728453c948c6a14787c3ada18087dcb350cad0d5a83ebfbba

Observation 6bbd4126-4876-4242-ba64-55788d2f6436 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.655993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.655993Z digest=sha256:76140f0f8665febd26d157670d6e25bde80460548845b291c8d8b22396cb1649

Observation f42ef5ea-dfcd-43ab-a68d-cc3e9331ff72 · outbound

This paper cites Instruction Tuning with GPT-4.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Instruction Tuning with GPT-4

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.643808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.643808Z digest=sha256:a40fbad4196a875abf78bb3218c2b2f8bea6a2589c51b1b0dc0256e2f953341d

Observation 8cd91289-a74b-44e6-9b00-ba0b3cb40256 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning LaMDA: Language Models for Dialog Applications

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.663414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.663414Z digest=sha256:8c5f9c07a82f5de0e850880edd739654f8aed5960575ce248781acd0343f8ba8

Observation 1ade2d10-616d-4f40-bea0-8e4fa24c10a3 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.244936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.652222Z digest=sha256:af70b75aba49e24c39ff9d0efc735e4600b0f665926db590b9f3ccb45866d2f1

Observation e9fbeeaa-147d-48ab-8af4-d1dc59b836ec · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.670495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.670495Z digest=sha256:41b1eba6eebb1b4fa167ec4f53dcca9a40abdf1095733af831868be86be2458e

Observation 14457149-d8c8-458f-8c05-403f4c7648bd · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Emu: Generative Pretraining in Multimodality

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.659594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.659594Z digest=sha256:4bb4b22fbb26e1a2e2a04e7a40d8ce970662667e593a4d2a284fd850abec735c

Observation 27ee7363-6615-408f-837d-95bb62da53f9 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.212870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.677492Z digest=sha256:20d383d5ca078a5a0cf5f10408818baf15d6d0ec3223a1719be68b5c36f8fc13

Observation 253cfeb2-28fd-4360-8773-a9d1dac4cc83 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.224849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.666925Z digest=sha256:9e58e205072edfa2b6b8ec2861532cb6b82abcfa128b309edf346f361ac92db5

Observation 4f558daa-f18f-4a5a-8bef-377f8df8026d · outbound

This paper cites Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.685395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.685395Z digest=sha256:50daf05d249c8633b143bcd76cd599e708c4e03ad2b80aa8600d3ff6123f9d1f

Observation fb679c5a-8827-4061-aee7-c07d5c49e649 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.673777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.673777Z digest=sha256:6c2a7455ac01270bcca2c47341c7be107dfcade45b9e66c339cbb185a1b7f754

Observation 8f1fc05d-263d-418c-a434-feaeaf035045 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.178332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.694451Z digest=sha256:58bc947b04080f473664355b05d453b4fc240e0e948daa090978a8d7d20668f5

Observation f44f3d03-16d9-4e16-aa5e-0d0a5e929529 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.201682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.681740Z digest=sha256:b5e4cf0332900e0fc5a4eaf1a96dbf0b4256dd2b5e90ba99bcee46ebb9bc3b94

Observation 6f35519b-5595-47c7-89a4-6ae29a98db82 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.702038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.702038Z digest=sha256:bbc5eaa60530fb4783ef8af8f1c19e2f790cdec980d6e05477ecfa4501a2693a

Observation 5700b233-dd1c-4d10-9b45-cc8772485ea1 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.189245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.690018Z digest=sha256:3e34d43387fbbd93994bf460e0f6fd4bce158638f43d5fd33775e48f40b53ac5

Observation 2588ec95-6ce6-4a29-9409-1f98ca454302 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.141108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.714720Z digest=sha256:0910e9e3892a287310738a8095ae5cb2d6f22f1b6f513216230b1f38ba265af6

Observation babc3aac-4f90-454b-a4d2-a38b77deee3a · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.698380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.698380Z digest=sha256:24149faedb9e1bb57850d7a2abd2ab1ec4069f791ec74d82a6c06fd06ec42ca5

Observation b0f6c79e-dae6-49e9-8ee9-25d18e881094 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.722238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.722238Z digest=sha256:0c4cd294f687ad2ae984999909b160818394b73fcf0ab6244df3e417d47fc041

Observation 0b0881ec-37aa-466c-807a-aa050aa91ef0 · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:46:33.153280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.706011Z digest=sha256:39003fb141435f7da06e76849b7c58eb0907abda84042146d7f97bcdc743c814

Observation 2220eef3-5d3f-46ad-a08e-eafd0a89da07 · outbound

This paper cites A-Bench: Are LMMs Masters at Evaluating AI-generated Images?.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning A-Bench: Are LMMs Masters at Evaluating AI-generated Images?

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.709744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.709744Z digest=sha256:b050ab09738159370c9d1b70c794cfdfe01cd4f3e97abe8904e4457c5826ab02

Observation 865fe1a7-9e18-4066-8ad2-cd2445e186a4 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 59

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:46:32.733861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.733861Z digest=sha256:bf03e302c943ec4a7e278b7d1ee0a60316c65765ce760cc39de7e8b1d20499f0

Observation 169f20ff-e6b2-420c-85ac-174e26a9ee8b · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.127762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.718445Z digest=sha256:820e4c187ef7b588b220df6728232f789fe0335946ba8a557281fd0149737899

Observation c41b6b79-53db-463e-b4a7-d793c781695b · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.725986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.725986Z digest=sha256:aa9067e5ae1b09e51549bd35669ddb20744407e26f83c83f1b65eb539a46e964

Observation 6555c831-6e64-4344-b131-1cadaf870e19 · outbound

This paper cites VLPrompt: Vision-Language Prompting for Panoptic Scene Graph Generation.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning VLPrompt: Vision-Language Prompting for Panoptic Scene Graph Generation

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:46:32.807084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.730168Z digest=sha256:cad282fc6da2689c019309330d6d50b2abd176a84ea74c3ac67fe7d55a64af5c

Observation 8e4da9c8-4827-4087-969b-f91145430de0 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.093460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.738040Z digest=sha256:680a9c14033885eb01c16227da761c3c4419f1c2e3d7dc371b86b96d20f0ec8f

Observation d3560ba8-ef8f-4027-83ad-cdeadba51043 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.079047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.741657Z digest=sha256:7bbd47a04cce6cd26980ff7602ccf1637d7c0095f320f56ec0bd98ccc63f7334

Observation af8f94e6-a77f-42c3-b071-1c317cf12495 · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.064723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.745351Z digest=sha256:2d9d66c0cae6ab4e296593354c80b16f21efa9c7c32a2c71c827c13da059ae90

Observation c3a686c7-0960-46b4-aff4-c96910834b0b · outbound

This paper cites an unresolved cited work.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:46:33.052510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T21:46:32.749086Z digest=sha256:d6376e59080734ff0290b8fc764fcc83f1f05673d8c94371c8a1fd36f00b4290

Observation 56fffcb5-3916-4186-951e-f17f3d735353 · outbound

This paper cites In Proceedings of the IEEE conference on computer vision and pattern recognition.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning In Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.529788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.529788Z digest=sha256:d83b370d6d37ef8e4d599f348dbf7f171c8a3d25a1a98c874c8d2f9e7a06db9c

Observation 11210261-7067-4124-8809-8b0f5a25ca79 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:32.506569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:32.506569Z digest=sha256:ecc4577eaeee361b702a79dd9d15944ccac718d8266b3083c2c1c7b7dc5d5c57

Pith citing papers

Observation d0d34ec4-6bf9-45c3-99e9-bda1de719a45 · inbound

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning cites this paper.

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:20.877492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T07:11:02.464556Z digest=sha256:c497198b285088eb25b1c46894849742288306db6d44b4efba76b06ed5d764b2