Pith. sign in

Paper Citation Record · LEDGER

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2506.11571.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11571 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:07:36.844173Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T09:54:33.563266Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0abaabf4-c43e-48f3-8d8c-ed11ae2f1b6c · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Hallucination of Multimodal Large Language Models: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.587700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.587700Z digest=sha256:9b05704c4f623d5127d86f2de2adaadc595924be1be5bdf6b9f652dfab480a5a

Observation de2de84e-4f9d-4c43-bdca-87a18ce0ff6b · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Eagle 2.5: Boosting long-context post-training for frontier vision-language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.660752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.660752Z digest=sha256:c06967e25fed4fd4003ae8aca8126346b2f6e034fb2dc6e1bc20fddaff6915ca

Observation 7637cb9b-dc85-4776-8ff2-9efccd7b9cd2 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.703713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.703713Z digest=sha256:e2e471942c082f1321a3f2b86d47a7f0e069e9c59b127e97753718284d0b6b81

Observation adb9539c-bbe9-4d46-bc61-7e061c22ee24 · outbound

This paper cites MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.778408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.778408Z digest=sha256:22643c52333f2d1f1b1b6128a256b5358d6754765c555e78f0b2f685973d36fb

Observation de2f71a6-6d3e-42c6-bc3f-951120f45647 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.820419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.820419Z digest=sha256:0ae2944abfd3a22cec62bfa78fa32c7e2ccfb16419292d9af8d12448db16e3fb

Observation 6eee8200-8caf-4adc-98e2-211589b90ad5 · outbound

This paper cites See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.881851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.881851Z digest=sha256:53a70af8ad0421043bf7d08dc125783301cce9726664eb3c09122ad55c8f8228

Observation bbef4c86-23a9-4ba7-b7a3-cc1c5e026a43 · outbound

This paper cites Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.956173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.956173Z digest=sha256:f72223a8470ef853606d72e9998ca01fa05f56447061e503b0be5d9e3a656d1c

Observation 90ecb6a8-ce9f-4510-9908-270fc34ff20f · outbound

This paper cites Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.061606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.061606Z digest=sha256:a4d09d603c558e2b5a78273ca4e8fef63c4cfacdeb5d35c9cddfb7fced725890

Observation 3a649fa7-a297-4640-82a1-1d40ae8fa644 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.177092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.177092Z digest=sha256:b495d48a06d2f0faec75a6459f719a5d3156fc0fb7dffeebd1ceca2932230d3b

Observation a7ef88ad-7914-4c91-a6e4-0f7584eb693b · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.289278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.289278Z digest=sha256:7d624a07d981c5c0bf0346d5ff282f05f0274d435fcc41f1ece92dee9e997b07

Observation aa453ad4-0ad9-4b46-8019-78d046e0a629 · outbound

This paper cites Seed1.5-VL Technical Report.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Seed1.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.364515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.364515Z digest=sha256:d1c58d280a5fca330cc77095fedb18f7133ad11fbf6179ad47ae67e4a732a8ea

Observation d0a20dea-c999-4157-9586-0d4a6ec543f5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.480638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.480638Z digest=sha256:78f0a9aba0d2d84d5287514a5af5231ee552c7b69ac93c518441e039a0a5ddab

Observation 73d92c24-9872-40ad-a9a1-082c9a2b7a66 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.555893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.555893Z digest=sha256:8a14250bcb5c02d49498fc33986819a9b0c035369536aa09e74a368d67d72790

Observation 0ff9ea26-f393-4506-9922-25216861501e · outbound

This paper cites CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.651474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.651474Z digest=sha256:de5fcebb198fba9d9837ea86e7a1b0fb51aaa0388f2a0c49223404c65ac41deb

Observation 57369070-914b-467a-9d5e-4de52cf965ce · outbound

This paper cites GPT-4o System Card.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.738113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.738113Z digest=sha256:fe296116099e0bb281c286ea1603fd1199832615be3f8c0f3cdd997d942050bb

Observation 80a94db2-df0c-498d-a628-6311cf99e159 · outbound

This paper cites OpenAI o1 System Card.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.813269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.813269Z digest=sha256:5422217f9dee70d9c498765f2ca9c240605a1a2762c7d0f688698e79a8c90cbb

Observation 35009f1b-316b-4e07-a07c-a50ae55b903a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.918115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.918115Z digest=sha256:413415f344f624af18ba418a84ea0919a826b02a206e037b64bd3878d99b84bd

Observation 9876737f-06d9-4d91-8603-06bd20891a76 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Evaluating Object Hallucination in Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.024804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.024804Z digest=sha256:8f4b51c6ee92f525b5fadeb8635378305b69be76a4e074f8a76daedfc71ffd3a

Observation 9a2ee76d-092a-4c18-a14f-371f0b13bfe2 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.115300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.115300Z digest=sha256:a8be1b8e5612d6f657dc4dfa61b7988393ab74b3c60e31d717b7dd8b26ee84c8

Observation b759e0a0-aa7d-4e6b-bd61-b01086fb387b · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:42.342580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:33.208153Z digest=sha256:9693272c391f474e3c2b0494759b24e6c4f44f9cbecbf30904b069d9bef12f70

Observation 70c288f9-a8eb-4fdc-b88e-b5b704dabdd5 · outbound

This paper cites Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.318843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.318843Z digest=sha256:749b5d585a30f14b9414ffccdc6a367c6293004be71fcb9014ae22194d2dc158

Observation a3e949cd-5a64-4d0f-8c61-d0672c991c69 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.434821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.434821Z digest=sha256:a64f34092c39d93b955658fbc2c76ee9df11263f6087cf71c9b4890341d82a96

Observation 5f578a40-6dee-4156-92a5-2592e763fa5b · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.560796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.560796Z digest=sha256:8154abf271c576b4227104c632f2a711441264583377ada54abd805781f4283e

Observation c6091be3-7e48-42a5-8e95-8f6c51e6a581 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Gemini: A Family of Highly Capable Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.622232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.622232Z digest=sha256:56c06d57ffe80a0297b07542720756ce3ca7dfecf89981696f044e582d7d78e7

Observation 9024d70c-e031-4c04-90c2-80499c3da584 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.722511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.722511Z digest=sha256:52dcbf39c8ad90769bc2546b276ee9fa0b70d8f4e30bbd060794eb7552242375

Observation 0c4133bd-d41a-42d7-8080-f873a69e81af · outbound

This paper cites Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.839250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.839250Z digest=sha256:b7679736535deaf420fb21d5724b362fdcadbe4debb163ea6d09d107b5516382

Observation d38dc04e-cd60-4b4d-9a3f-f55c27fbb8e2 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Measuring multimodal mathematical reasoning with math-vision dataset

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:42.037028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:33.917505Z digest=sha256:9379e579608b5a6db2f44833d156ac95739be1ef1aa64a86e14582a4a3429825

Observation 2edb0dae-c5a2-4436-b1b9-2868fed6b484 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.024686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.024686Z digest=sha256:57d1c80d89ac8b2ea0a7f282fece19e1577777945a8ad78bfcc758a95d3d7de6

Observation 0793f6ff-3665-447c-b54b-3f85c43aa745 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Chain-of-thought prompting elicits reasoning in large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.166022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.166022Z digest=sha256:07be40a0371e84452aac3dc6be64c219dd0f711b5d809de2c81c29ed0b4a86e3

Observation 2642020f-54ce-45c5-b412-27e43624b0ba · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.276313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.276313Z digest=sha256:a32e70bb38180b52a92d29a1013992abe44f53fe98e291bbd16e096988b6d920

Observation 1165e788-0faf-4ce7-b2d3-b886ed25850a · outbound

This paper cites Valley2: Exploring Multimodal Models with Scalable Vision-Language Design.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.362773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.362773Z digest=sha256:c3ed82655529f721ec4d7265dfe95858f1ed54660adc81227c2087cfabe98c1e

Observation b10a50c3-9317-44c1-8011-5dd93ad91905 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.464471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.464471Z digest=sha256:e5fe3413f6a635584527e17d7cf65c71ac9a219f112874f25add1819e0fad307

Observation 64848bed-8331-43f6-a7f8-50025735980f · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.579200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.579200Z digest=sha256:7a26c7409cdc89d94b259329c5c781f37989dfc7bf36eb6c796e9a2bd51b83ae

Observation f848019f-2edc-47ec-a455-f1c34aded2ed · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.668498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.668498Z digest=sha256:728188391009f83f190e10bdf253e0a48690aeb185d8356e8044c275a8844a53

Observation ab18994c-87b4-42b1-a5fe-574d3fcc4437 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.753888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.753888Z digest=sha256:cce95af1d4fb64ba59f54600171e91a0323e9ca133b5fa4d34b5457276292fcc

Observation 907799d7-17df-4531-9005-ab4ee4715382 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.797292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.797292Z digest=sha256:c5c2f002e46ddd84ec07250f846295ccb8f1ece98fd574b5dd8d1ae23037d9be

Observation 3cbee560-f79c-4c3e-850a-a752a3295f19 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:34.864725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:34.864725Z digest=sha256:9627efae67ccfd0f301db3eaa2154e63c4acdad09dad4307fc90d74d05cf5f54

Observation 20999872-3733-4c93-b6a4-1859a808f77e · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:41.794677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:34.953774Z digest=sha256:546c08d3c255784b5b65c23c5cc7db62eee7cc6f6e6daf21063a248fdfc487bc

Observation 5f681d69-358f-4b4a-bfa6-f689294ea924 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:41.652440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.041699Z digest=sha256:6937932c7cf1596ab9011f1b91480da82c08a6a46deeecad37f50a777e65edc6

Observation 98dcbe5b-8481-4df7-a5c6-8d29ffa9a212 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:41.516516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.156688Z digest=sha256:3fc49eabdc517333ec4c0c560054c8a101dcd083f91d7aab230394a44b0296a6

Observation f842e7a6-5627-4b8f-b46b-b6a0001c49c9 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:41.348976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.228622Z digest=sha256:34b8e3efcdf07d473227a0b8183ef424f8cfc2dd046300a7bcc3906fc49449b4

Observation 31826df4-9b47-4a54-976c-26c18e42758d · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:41.133657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.271391Z digest=sha256:86a8c63d0ea789c99b302ff73f9742b980be9c4fe0bb156c6d773a6385a5c17a

Observation aeb8bcfc-5433-4087-b4dc-e468a5675797 · outbound

This paper cites The screen visible in the image is small and not a touch screen.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The screen visible in the image is small and not a touch screen

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:40.932629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.356982Z digest=sha256:92ed63da8bcdd37cf7417da128ebe3336e5f76ebb4131b731d88f580333608b9

Observation 3176c547-2f17-4fcb-bf8b-c8dfe69166b4 · outbound

This paper cites This option is speculative and cannot be confirmed from the image alone.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is speculative and cannot be confirmed from the image alone

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:40.667712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.443459Z digest=sha256:96de50ce9d109aa6388b967cebbd72b0c4868e0dc1b0f44cbdb2f1164ae6ae7e

Observation 85f92ca6-d4c5-419e-9a24-4d2b0a23fd56 · outbound

This paper cites This option is likely correct.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is likely correct

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:40.401958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.521699Z digest=sha256:c16d99b9a2923ac80b8be0a31b099d999d9cdeb4d8f333870876f97b706e1b82

Observation 223a4946-abc7-4129-9639-9c05df960629 · outbound

This paper cites The image does not show any indication of this technology.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The image does not show any indication of this technology

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:40.194944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.582676Z digest=sha256:c09660537d743006974b5a93f8e0dbac339100a9ce055f44316e891941dadb80

Observation 9ef69d52-90a2-4fa2-aef9-a1de3c5d9348 · outbound

This paper cites <vcues_2>The screen visible in the image is small and not a touch screen</vcues_2>.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? <vcues_2>The screen visible in the image is small and not a touch screen</vcues_2>

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.964892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.669343Z digest=sha256:bbeb3d0e3ba64640d7e9b1b4e99e12ecf487ce6d7f6f51745baeb66ab4cf2835

Observation ca1fb7ff-5e97-4513-9931-84e7dfb958d2 · outbound

This paper cites </vcues_3>.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? </vcues_3>

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.782624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.723483Z digest=sha256:50b46c7e176de20e71bd4506ae634b1c0b7ce11e7b9f04380b4745298f7c6623

Observation 03bc4ae7-ed92-4abf-b66d-ce84abb40fbf · outbound

This paper cites This option is likely correct.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is likely correct

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.570953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.756741Z digest=sha256:090f5352bc2fe890f9993231874aaf1c38815f9ad279841779ff17c1b29f2cf9

Observation 93890703-e101-4190-9a39-d083b9789e5e · outbound

This paper cites <vcues_5>The image does not show any indication of this technology</vcues_5>.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? <vcues_5>The image does not show any indication of this technology</vcues_5>

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.399955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.811204Z digest=sha256:943c9debd6f647bbc91545139e4b35ee4fc0dc54e0c9b01f0e0bb3b973946f5c

Observation 86ef2ce4-a475-454c-a7bc-9a658381ebf6 · outbound

This paper cites <vcues_2>The screen visible in the image is very small and not a touch screen</vcues_2>.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? <vcues_2>The screen visible in the image is very small and not a touch screen</vcues_2>

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.293227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.883434Z digest=sha256:24a7c2b271c4c0f40e8c0183e62d3c1e05e1e51c3388e54107e4ef040081c048

Observation a5ad5350-509d-4b95-a5f8-da23aa4ca87d · outbound

This paper cites Therefore, this option is incorrect.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Therefore, this option is incorrect

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:39.113386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:35.976409Z digest=sha256:6976f1569eba11829dd42c691289de36e36025f3508df1baf21b788f600ff624

Observation 47154189-fc59-433a-811a-67dce248275b · outbound

This paper cites This option is correct.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? This option is correct

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:38.971452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.012004Z digest=sha256:bdcfb2a4c8c72a0726d060020f5d6231d75fef34889ed74e13abef5aecd6bb22

Observation c073b7ea-f187-44a0-936e-fbcc93b2c171 · outbound

This paper cites new_option.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? new_option

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:38.796243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.086645Z digest=sha256:4db45abb104d24d3f76034fc076ecb372dbd1d0cf6a5ddae87cbaca168db6334

Observation c9f65da7-5d69-429c-b231-0e8f3e28b1f1 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:38.666183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.188559Z digest=sha256:6ec084a0f2bdb9400f794b1a9f3db2d773e853878f7113c6d0de07855a54d97f

Observation b7d41924-d39f-4128-ac98-7de16039e3d0 · outbound

This paper cites The presence of dirt could be from the environment it has traveled through, such as dust, debris, or even road salt in some areas.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The presence of dirt could be from the environment it has traveled through, such as dust, debris, or even road salt in some areas

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:38.554796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.276358Z digest=sha256:fd631fd641a698de7171bfe9aa5e98d6e7f7881f9488de1c08d117b423d5eb03

Observation b4f9552f-97e2-435e-adae-3957608376dd · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:38.421824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.364831Z digest=sha256:31adccbcd6fd96778caf87042edf09e09ed671cd1157f9f6506057fec77814bf

Observation 72575578-c4d6-4e54-93ac-36c4e9d7c5c3 · outbound

This paper cites The train was caught in a rainstorm: While rain can cause dirt to accumulate, the image does not show signs of recent rain, such as wet surfaces or water streaks.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? The train was caught in a rainstorm: While rain can cause dirt to accumulate, the image does not show signs of recent rain, such as wet surfaces or water streaks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:38.195822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.449471Z digest=sha256:ead682b63e8c88559077675cca0da98e5e184b4f8e5dad6a25ce2a182cdd5ff1

Observation 655077ef-3a7b-479a-b7ff-2410614e6255 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:37.953934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.536880Z digest=sha256:dec6f766e703bc0bcc61ed4e74ed2212c25e0ecbfc0c227abc0db19d7dbca8fe

Observation f339e64e-a48b-4402-ab33-5d9271709600 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:37.681408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.603461Z digest=sha256:b475e531ad2831ad85034114d7166456d00ff681cf4c5057a057ae86d0096b4e

Observation d9a69d1d-bb67-4fcd-9848-5b7df5cd8ab2 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:37.501653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.704682Z digest=sha256:9c1aa7c90318198691b0e584ba377dfab0b5741268925b6cb196d1cc6ad3e42b

Observation 885b32cd-a831-483a-838a-0b01f99498e6 · outbound

This paper cites an unresolved cited work.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:07:37.393645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.792454Z digest=sha256:14d16ab4d7d20c6948971dc76dc3f1b487ab01e56732e76905a01a653efa641d

Observation 958fcb89-84a5-4a7d-8d98-5019d2c34456 · outbound

This paper cites Given these observations, the most reasonable inference is that the setting is a small residential room, likely a bedroom in an apartment or a small house.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Given these observations, the most reasonable inference is that the setting is a small residential room, likely a bedroom in an apartment or a small house

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:07:37.321432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:07:36.844173Z digest=sha256:45671121e19af4c0fabe33e60399ceb05073e77d604db79b3697d2d7464190ba

Pith citing papers

Observation 9f88081b-4638-42e2-99c6-e46901bbed90 · inbound

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models cites this paper.

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T09:54:33.563266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:54:33.563266Z digest=sha256:00fbe081485f8eb0b9f431ae3749473cde50c97e8aa91be839b5e5ac8f28dafd