Pith. sign in

Paper Citation Record · LEDGER

Training-Free Reasoning and Reflection in MLLMs

As of 9 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2505.16151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16151 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:13:58.155456Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b4a03c8d-c0d0-4c32-b66d-099fb747db79 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Training-Free Reasoning and Reflection in MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:51.828346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:51.828346Z digest=sha256:12466b7c6c9eeaa262f0c9e89d095e3b5c834382c59feafbed9874052f833761

Observation 57845998-5e7e-4750-bcdc-a2068d408ced · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Training-Free Reasoning and Reflection in MLLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:51.879241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:51.879241Z digest=sha256:fa027fb7762050f45ad7e96f03fafdb48ab445a3b93242b436459ed6c29bb572

Observation 93e86eb5-d40a-4b50-8c8c-b65249a2b758 · outbound

This paper cites S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning.

Training-Free Reasoning and Reflection in MLLMs S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:51.984499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:51.984499Z digest=sha256:6e915a137973d83fcaf4ecfa42ada297989a0f92a63a1160d59d1bc812d120bf

Observation 0c42477e-565d-45e7-b736-146e083932f4 · outbound

This paper cites Let’s verify step by step.

Training-Free Reasoning and Reflection in MLLMs Let’s verify step by step

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:04.355711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:52.146034Z digest=sha256:767ea77744e1bf7fcfe14be33633045b997ee52b7382894793cfb252ce85d886

Observation beb3ea43-93c7-4bfa-92dd-cf5433b58684 · outbound

This paper cites OpenAI o1.

Training-Free Reasoning and Reflection in MLLMs OpenAI o1

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:04.149806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:52.295443Z digest=sha256:c104ac9ec7b5792cb9596c557d0c46fca33fefe8f52ba4e64829b0c9b23d4429

Observation 41caa1b0-7bc0-4854-b185-0b4f219ffe8b · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Training-Free Reasoning and Reflection in MLLMs LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:52.385935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:52.385935Z digest=sha256:35e01a7b311ef3868c68d16b6eb6a1d63d8af8dee011ea98c0794958a46d425b

Observation 14cf2302-bd83-4130-928c-a89068c677ce · outbound

This paper cites Insight-V: Exploring long-chain visual reasoning with multimodal large language models.

Training-Free Reasoning and Reflection in MLLMs Insight-V: Exploring long-chain visual reasoning with multimodal large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:03.936163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:52.502448Z digest=sha256:b1fd1ae4625ecbe81f32b3cb1d3abf1a301eb36ac805e3e8391ca11ead9ad4a6

Observation 9d634656-5c72-4ccb-9890-89369531f8f8 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

Training-Free Reasoning and Reflection in MLLMs Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:52.634235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:52.634235Z digest=sha256:e51e40a10e90ef12ccde4a0823c263bac8ce350d5f0d2d872f01ce8db29d3dd9

Observation c58e677e-b258-4034-85c2-3c79f7f93288 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Training-Free Reasoning and Reflection in MLLMs Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:52.747958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:52.747958Z digest=sha256:5d9ef1e88d3a95cc280e0c3fbb7a7796746d6665814e3cab631c7cacf43f524c

Observation c0044c76-f0d6-480e-bf76-6e1045f06c76 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Training-Free Reasoning and Reflection in MLLMs Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:52.912688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:52.912688Z digest=sha256:89117ec641e9a90639c2beb43e1234c9aa230e34d5c1030a3bca26193abd41cb

Observation c16d2579-a035-41ab-91e7-114625bcacae · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

Training-Free Reasoning and Reflection in MLLMs LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:53.084279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:53.084279Z digest=sha256:886c2c1e682ad40825a3c6ff26e2c7ceef914b59650294de83758894f5a575e1

Observation 0c3d472a-cddc-490d-aac9-d9e2ba56f329 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Training-Free Reasoning and Reflection in MLLMs R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:53.259417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:53.259417Z digest=sha256:597a1bae036ab582edcf4a1b4970ada003c4e8fe13f722e9c34cd1b09a1e0ef6

Observation b9901528-38be-4f40-ad0f-a6bc60af2194 · outbound

This paper cites Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt.

Training-Free Reasoning and Reflection in MLLMs Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:03.795695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:53.411747Z digest=sha256:8717f089bf6ffd712fdaeb712d641df9f40d6db4574c806cb092c4a2624940ae

Observation d55647b8-133f-4200-a45d-331eab01d5c3 · outbound

This paper cites Editing models with task arithmetic.

Training-Free Reasoning and Reflection in MLLMs Editing models with task arithmetic

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:03.631396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:53.532874Z digest=sha256:e4696f95b4eeec513fbdf0787eb689c9a9c2b8f487f25975bce6b6bb8e69bb33

Observation fd227eb9-e86f-44c5-9d0d-9148a56e3c05 · outbound

This paper cites Composing parameter-efficient modules with arithmetic operation.

Training-Free Reasoning and Reflection in MLLMs Composing parameter-efficient modules with arithmetic operation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:03.426559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:53.655470Z digest=sha256:480683a568296aa41e0c6617e62f1b8eca0348275c2eca9cd3ad2c91f24149a4

Observation 93053a11-4333-411e-9419-5bb4c56a9471 · outbound

This paper cites Gradual progression from sensory to task-related processing in cerebral cortex.

Training-Free Reasoning and Reflection in MLLMs Gradual progression from sensory to task-related processing in cerebral cortex

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:03.243842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:53.801760Z digest=sha256:47342a1836cb5c726e28a575b7cb3869746591b0a64f6da33b78573e5062f00c

Observation db001b72-5745-4abe-abc6-1b1c1e4d1378 · outbound

This paper cites Hierarchical processing of visual and language information in the brain.

Training-Free Reasoning and Reflection in MLLMs Hierarchical processing of visual and language information in the brain

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:03.048780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:53.915991Z digest=sha256:f57f1a0c18b90c90c7412e5bbe5ac40b033cec2543b6e8354a4ffa1ad02fc2b9

Observation 37983c83-55b1-4773-8a90-c17d4a67db5d · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks.

Training-Free Reasoning and Reflection in MLLMs Visionllm: Large language model is also an open-ended decoder for vision-centric tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:54.045942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:54.045942Z digest=sha256:22ec4af4563e877efff00015f9ec609ad1fdc08d4166af6b2402a65d86836b34

Observation ae20434e-d4d3-41ad-933e-ed70bf3a7f99 · outbound

This paper cites Qwen Technical Report.

Training-Free Reasoning and Reflection in MLLMs Qwen Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:54.232145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:54.232145Z digest=sha256:4b7d72a3b013273f48d99206e29bd09c2e64aad54ab709e3f40fdea2b5f01097

Observation 9bd3e0e5-35f8-4239-8ce2-90440b2bab58 · outbound

This paper cites Visual instruction tuning.

Training-Free Reasoning and Reflection in MLLMs Visual instruction tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:02.852006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:54.390407Z digest=sha256:28e8c74275a7e75ce454091f1b6b038b19c5c6d59cae9f0e6aeed586207f8769

Observation 6cde11fc-f8ee-44cb-91ca-8af5c0207f04 · outbound

This paper cites Emu: Generative pretraining in multimodality.

Training-Free Reasoning and Reflection in MLLMs Emu: Generative pretraining in multimodality

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:02.676886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:54.529176Z digest=sha256:69c27146b4c523bac55716a09c17067711b2bc38b47998f2552a118923f8ea14

Observation 7ee25ad8-f013-4c6f-b450-7203ce445ce8 · outbound

This paper cites Chi, Quoc V.

Training-Free Reasoning and Reflection in MLLMs Chi, Quoc V

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:02.505631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:54.654112Z digest=sha256:8bbdfe2c8605c170bc07e40f0ed703607da92b97535cd3ef68f4da5591f98ee4

Observation 7068f311-2234-4343-9571-58b2bc26b0ea · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

Training-Free Reasoning and Reflection in MLLMs Improve Vision Language Model Chain-of-thought Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:54.759630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:54.759630Z digest=sha256:e2cf30f5fb95b6ea02d6052354a66751772d151614fcc2bf8ae8ff95bd14af3b

Observation cd65305a-d2d3-4566-9141-65de8ebf159f · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

Training-Free Reasoning and Reflection in MLLMs Compositional chain-of-thought prompting for large multimodal models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:02.272805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:54.882809Z digest=sha256:106596e142e99389a92a4f78190578c961e74cedc1b5c33f24468047602f8ad3

Observation 14d2e6a5-43c5-4b83-9c53-e5550f74b49f · outbound

This paper cites TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding.

Training-Free Reasoning and Reflection in MLLMs TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:55.011386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:55.011386Z digest=sha256:4bb32c6a826e26386dc561abae37fb8aa162c77ab19f4b43db69e743c3746f5a

Observation 6727c625-bba4-4568-a169-45ca1b8fe1a3 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Training-Free Reasoning and Reflection in MLLMs MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:55.145082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:55.145082Z digest=sha256:3083d561f513f765ce998f3bc14bd42758d47f63656f8ac8d7e2c2e647863527

Observation 897af4c1-7e20-4dd4-878f-c34a3423183d · outbound

This paper cites Merging models with fisher-weighted averaging.

Training-Free Reasoning and Reflection in MLLMs Merging models with fisher-weighted averaging

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:02.178644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:55.288092Z digest=sha256:b204d03ea9f41c82ed510e32b06b81f8abd63cd209e2667ef965025949a6c1c1

Observation b96f9645-0ec3-4ef9-ac84-ff3177a291a9 · outbound

This paper cites Raffel, and Mohit Bansal.

Training-Free Reasoning and Reflection in MLLMs Raffel, and Mohit Bansal

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:02.038488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:55.377913Z digest=sha256:a8ddaa8ef27318a8700dea1dcb6ba1d877bda3db1de2f051e9b90730c17d5ccd

Observation 198173b2-c49d-4ca9-8ece-5f69c3f9a79a · outbound

This paper cites Metagpt: Merging large language models using model exclusive task arithmetic.

Training-Free Reasoning and Reflection in MLLMs Metagpt: Merging large language models using model exclusive task arithmetic

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:01.860402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:55.488112Z digest=sha256:1c04affe96599f5aa957e3f101c6b4673bfba4b6a627da4a9ce87ed460b21e00

Observation b0e58955-0b2b-44e3-99ae-a2e60fbe3d22 · outbound

This paper cites Dataless knowledge fusion by merging weights of language models.

Training-Free Reasoning and Reflection in MLLMs Dataless knowledge fusion by merging weights of language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:01.719617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:55.556366Z digest=sha256:6b5feb8298f0d18927d5a236a08f3e562f09c631c2436a04a3c7294a238976e7

Observation c95559bb-14b1-4004-9cb1-c43d4b8ecdc7 · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch.

Training-Free Reasoning and Reflection in MLLMs Language models are super mario: Absorbing abilities from homologous models as a free lunch

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:01.603696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:55.649661Z digest=sha256:e0327a17d71cbdeb5530b8f3c5081587189529af0b51013a91ad6a1759baa27f

Observation c9568ce4-e1e1-418e-97cd-37470082b171 · outbound

This paper cites Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging.

Training-Free Reasoning and Reflection in MLLMs Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:55.819405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:55.819405Z digest=sha256:185a7450451f0c86443368583848dc7a5c02674356455638ebf9f0813bd3685a

Observation babf7510-ac6f-4b55-b94a-f9be4d6483b0 · outbound

This paper cites Neu- ral Tangent Kernel: Convergence and generalization in10 neural networks.

Training-Free Reasoning and Reflection in MLLMs Neu- ral Tangent Kernel: Convergence and generalization in10 neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:01.423826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:55.972987Z digest=sha256:70ef898add51d43283b43e6b8bde9ae5f99fcc70d8b05eeafe44283bde57a078

Observation 5fdac022-bc8f-44fd-b5b0-c1eba68b3f94 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Training-Free Reasoning and Reflection in MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:56.060872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:56.060872Z digest=sha256:b3019a14dc769eb7bbb929e8709af64a3b72a556dd1e7aaf499d9fa5b899b61f

Observation 383efb5e-2548-4a87-a036-59f905451354 · outbound

This paper cites AGIEval: A human-centric benchmark for evaluating foundation models.

Training-Free Reasoning and Reflection in MLLMs AGIEval: A human-centric benchmark for evaluating foundation models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:01.259648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:56.214550Z digest=sha256:42994f4777263f083dd8c89dbe1379a5f56f470918fe2e20efbc4292362b2653

Observation a0006f44-3ad8-4b4c-bafb-60e3314e2184 · outbound

This paper cites Improved baselines with visual instruction tuning.

Training-Free Reasoning and Reflection in MLLMs Improved baselines with visual instruction tuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:01.109036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:56.305641Z digest=sha256:87cc1848dfc17ffbd096d9b170bf980812fa49f7e648e7c2a99fcc39f0dd732c

Observation 80fa561b-4be0-40a4-bc90-2158c8d5e01e · outbound

This paper cites Llava- next.

Training-Free Reasoning and Reflection in MLLMs Llava- next

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:00.942818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:56.410192Z digest=sha256:26da6b0cf0fb671889cea6455a7b4a87267afa785aaff1545f3f50e053640ec4

Observation 0a029e9f-43ef-4f43-a265-d67a00e0aef5 · outbound

This paper cites VILA: on pre-training for visual language models.

Training-Free Reasoning and Reflection in MLLMs VILA: on pre-training for visual language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:00.804085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:56.536836Z digest=sha256:1908b6536c1167e0c34f073c8cebafeb8a9ebc3606bb4e567b8c8945a195fb81

Observation 71c96e96-d19e-42a1-9d43-e83aa32072a7 · outbound

This paper cites Building and better understanding vision- language models: insights and future directions.

Training-Free Reasoning and Reflection in MLLMs Building and better understanding vision- language models: insights and future directions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:00.681087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:56.609209Z digest=sha256:f974d5288dd049adce7efb6442207168396b6b5300d04b3d4846e4c6897c01f0

Observation 79992ad2-8c98-474b-ba7a-df3a9f70a619 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Training-Free Reasoning and Reflection in MLLMs Sharegpt4v: Improving large multi-modal models with better captions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:00.508551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:56.679480Z digest=sha256:eddd70ac0fd483d11608250726e6161d1a01714d14c435333e9c9f1084631878

Observation a8dab4d0-57e2-4045-bbef-6ea7cf7dc34d · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

Training-Free Reasoning and Reflection in MLLMs NVILA: Efficient Frontier Visual Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:56.772463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:56.772463Z digest=sha256:a00b3be054ad95e8d6571f542a9f18c171031f1971d509c9c42bcf8e055bfe94

Observation 69526aa6-a77e-4f0d-8ee7-a0e8566f7696 · outbound

This paper cites LLaV A-OneVision: Easy visual task transfer.

Training-Free Reasoning and Reflection in MLLMs LLaV A-OneVision: Easy visual task transfer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:14:00.340628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:56.841558Z digest=sha256:11761229713f04274eca7aba0b4bc3a31cbc0f34d334458e10cbf643798303ce

Observation 52f83b67-50cc-4255-8a95-013e8a8d3ad3 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Training-Free Reasoning and Reflection in MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:56.924965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:56.924965Z digest=sha256:6c5372af820fd5f7dc29b7c8bd6154e624bc0a58e7f0527fc4b475c4138d4e85

Observation 2cdb1580-914e-48d7-af6b-f22e16ec8d97 · outbound

This paper cites an unresolved cited work.

Training-Free Reasoning and Reflection in MLLMs Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:14:00.145206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:57.024187Z digest=sha256:2a99ebbdd706c5e9108dd1d6d644e1c9db19cf51bc2efee4aeb28fd971ff2fae

Observation 294efcf3-6fe5-4ed8-bcf1-da716dec89ff · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Training-Free Reasoning and Reflection in MLLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:57.151656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:57.151656Z digest=sha256:1ec9f5741278111f9b2032d2fb4ae064b08f03b9a4e300bb6424dce64abcf68e

Observation ce5f00e2-6d81-4ddf-a05d-8b68d7b1f3bb · outbound

This paper cites MMMU: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert AGI.

Training-Free Reasoning and Reflection in MLLMs MMMU: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert AGI

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:13:59.979941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:57.319687Z digest=sha256:a14ba2ee8cb0f36b6a61d1fec7681a9ba903a9725460cac2d5af3bdad82e6705

Observation 7a260aa3-ee5a-4850-a16c-4fa8b2149eb9 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Training-Free Reasoning and Reflection in MLLMs MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:57.403852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:57.403852Z digest=sha256:e64a67f86a0ec6f40d4774d232fb7ce28114614cb40f742d9652335571fabc5f

Observation 598d3a4f-c541-4a1e-85f4-24484d68a512 · outbound

This paper cites MathVista: Evaluating mathematical reasoning of foundation models in visual contexts.

Training-Free Reasoning and Reflection in MLLMs MathVista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:13:59.801231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:57.534671Z digest=sha256:780f8cc38812647128fe1e4d7fbcb290d69febdc902b80b1a7b716572d64fd7c

Observation 1f77ec34-5c6e-427d-8fbf-dba8e89d4de6 · outbound

This paper cites Measuring multimodal mathematical reasoning with math- vision dataset.

Training-Free Reasoning and Reflection in MLLMs Measuring multimodal mathematical reasoning with math- vision dataset

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:13:59.680133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:57.670724Z digest=sha256:a96f72113f011c32b26094a0deff534048979621851f298caed6d8794704a439

Observation 237520f4-695a-42dd-bf94-e5bc8b55970d · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Training-Free Reasoning and Reflection in MLLMs We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:57.748180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:57.748180Z digest=sha256:1bc98cbdf69101b2cbda1b9bf4a3731232d4f36eb1ec784d83594e9ae4a84932

Observation c4dd5115-8240-4cc4-8aee-bf37bff505c7 · outbound

This paper cites an unresolved cited work.

Training-Free Reasoning and Reflection in MLLMs Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:13:59.569232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:57.824122Z digest=sha256:89796b7f3bd138496cd3a2aa5d02694ddb15060a9647e0569580ca3522e5d832

Observation f9cbc9b4-afda-41ea-be0c-aa0ac8f62cd1 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

Training-Free Reasoning and Reflection in MLLMs Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:13:59.427612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:57.957815Z digest=sha256:6233cd00c8d74ec3522234c85f7988c6580878abb704fd78a511e7903e302d36

Observation a7941d7c-dc4a-4bf6-96fa-b588987efdcd · outbound

This paper cites Non-Reasoning MLLM.

Training-Free Reasoning and Reflection in MLLMs Non-Reasoning MLLM

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:13:59.285079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:58.054094Z digest=sha256:5fe4972a914b1fd06bd800b173690d8e43481e654e60c7506785ea3b8c093305

Observation 35c9be25-702f-4eb6-99ec-11fa196a66ac · outbound

This paper cites an unresolved cited work.

Training-Free Reasoning and Reflection in MLLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:13:59.157793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:13:58.155456Z digest=sha256:7e78841f892ab1eed9bbb76ee9ef1cffea4b05e5de623f77559af337fd202e82

Pith citing papers

No inbound Pith citation observations are available.