Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:07:36.844173Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2506.11571.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:07:36.844173Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T09:54:33.563266Z
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0abaabf4-c43e-48f3-8d8c-ed11ae2f1b6c · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Hallucination of Multimodal Large Language Models: A Survey
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2de84e-4f9d-4c43-bdca-87a18ce0ff6b · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Eagle 2.5: Boosting long-context post-training for frontier vision-language models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7637cb9b-dc85-4776-8ff2-9efccd7b9cd2 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adb9539c-bbe9-4d46-bc61-7e061c22ee24 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2f71a6-6d3e-42c6-bc3f-951120f45647 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eee8200-8caf-4adc-98e2-211589b90ad5 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbef4c86-23a9-4ba7-b7a3-cc1c5e026a43 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90ecb6a8-ce9f-4510-9908-270fc34ff20f · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a649fa7-a297-4640-82a1-1d40ae8fa644 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ef88ad-7914-4c91-a6e4-0f7584eb693b · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa453ad4-0ad9-4b46-8019-78d046e0a629 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Seed1.5-VL Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a20dea-c999-4157-9586-0d4a6ec543f5 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d92c24-9872-40ad-a9a1-082c9a2b7a66 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ff9ea26-f393-4506-9922-25216861501e · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57369070-914b-467a-9d5e-4de52cf965ce · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? GPT-4o System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a94db2-df0c-498d-a628-6311cf99e159 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? OpenAI o1 System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35009f1b-316b-4e07-a07c-a50ae55b903a · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? LLaVA-OneVision: Easy Visual Task Transfer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9876737f-06d9-4d91-8603-06bd20891a76 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Evaluating Object Hallucination in Large Vision-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a2ee76d-092a-4c18-a14f-371f0b13bfe2 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b759e0a0-aa7d-4e6b-bd61-b01086fb387b · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70c288f9-a8eb-4fdc-b88e-b5b704dabdd5 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e949cd-5a64-4d0f-8c61-d0672c991c69 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f578a40-6dee-4156-92a5-2592e763fa5b · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6091be3-7e48-42a5-8e95-8f6c51e6a581 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Gemini: A Family of Highly Capable Multimodal Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9024d70c-e031-4c04-90c2-80499c3da584 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c4133bd-d41a-42d7-8080-f873a69e81af · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38dc04e-cd60-4b4d-9a3f-f55c27fbb8e2 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Measuring multimodal mathematical reasoning with math-vision dataset
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2edb0dae-c5a2-4436-b1b9-2868fed6b484 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0793f6ff-3665-447c-b54b-3f85c43aa745 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Chain-of-thought prompting elicits reasoning in large language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2642020f-54ce-45c5-b412-27e43624b0ba · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1165e788-0faf-4ce7-b2d3-b886ed25850a · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b10a50c3-9317-44c1-8011-5dd93ad91905 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64848bed-8331-43f6-a7f8-50025735980f · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f848019f-2edc-47ec-a455-f1c34aded2ed · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab18994c-87b4-42b1-a5fe-574d3fcc4437 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 907799d7-17df-4531-9005-ab4ee4715382 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cbee560-f79c-4c3e-850a-a752a3295f19 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20999872-3733-4c93-b6a4-1859a808f77e · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f681d69-358f-4b4a-bfa6-f689294ea924 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98dcbe5b-8481-4df7-a5c6-8d29ffa9a212 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f842e7a6-5627-4b8f-b46b-b6a0001c49c9 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31826df4-9b47-4a54-976c-26c18e42758d · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aeb8bcfc-5433-4087-b4dc-e468a5675797 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The screen visible in the image is small and not a touch screen
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3176c547-2f17-4fcb-bf8b-c8dfe69166b4 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is speculative and cannot be confirmed from the image alone
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85f92ca6-d4c5-419e-9a24-4d2b0a23fd56 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is likely correct
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 223a4946-abc7-4129-9639-9c05df960629 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The image does not show any indication of this technology
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ef69d52-90a2-4fa2-aef9-a1de3c5d9348 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? <vcues_2>The screen visible in the image is small and not a touch screen</vcues_2>
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca1fb7ff-5e97-4513-9931-84e7dfb958d2 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? </vcues_3>
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03bc4ae7-ed92-4abf-b66d-ce84abb40fbf · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is likely correct
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93890703-e101-4190-9a39-d083b9789e5e · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? <vcues_5>The image does not show any indication of this technology</vcues_5>
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86ef2ce4-a475-454c-a7bc-9a658381ebf6 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? <vcues_2>The screen visible in the image is very small and not a touch screen</vcues_2>
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5ad5350-509d-4b95-a5f8-da23aa4ca87d · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Therefore, this option is incorrect
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47154189-fc59-433a-811a-67dce248275b · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is correct
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c073b7ea-f187-44a0-936e-fbcc93b2c171 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? new_option
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9f65da7-5d69-429c-b231-0e8f3e28b1f1 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7d41924-d39f-4128-ac98-7de16039e3d0 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The presence of dirt could be from the environment it has traveled through, such as dust, debris, or even road salt in some areas
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4f9552f-97e2-435e-adae-3957608376dd · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72575578-c4d6-4e54-93ac-36c4e9d7c5c3 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The train was caught in a rainstorm: While rain can cause dirt to accumulate, the image does not show signs of recent rain, such as wet surfaces or water streaks
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 655077ef-3a7b-479a-b7ff-2410614e6255 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f339e64e-a48b-4402-ab33-5d9271709600 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9a69d1d-bb67-4fcd-9848-5b7df5cd8ab2 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 885b32cd-a831-483a-838a-0b01f99498e6 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 958fcb89-84a5-4a7d-8d98-5019d2c34456 · outbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Given these observations, the most reasonable inference is that the setting is a small residential room, likely a bedroom in an apartment or a small house
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f88081b-4638-42e2-99c6-e46901bbed90 · inbound
Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.