Pith. sign in

Paper Citation Record · LEDGER

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

As of 21 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 10 inbound Pith citation observations for arXiv:2411.12591.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12591 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:35:45.274312Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:10:02.962883Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T17:18:53.747034Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved28
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e6706d4-3415-4aac-a73f-798fa6a258f8 · outbound

This paper cites GPT-4 Technical Report.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.000081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.000081Z digest=sha256:2d7d44c64887bc0ba41c7e8d415556c8332adbf2028a3d900d5e53246c4b33ef

Observation fccdf2c1-e8fd-4853-8af9-18c9613a300f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.006364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.006364Z digest=sha256:8aa19c455e5888a76c89d1d90b05e7bdd7fbca1118e594d715c0846d5497edbe

Observation 9df722e4-4254-40c7-8c40-9e6bd7845bdd · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Hallucination of Multimodal Large Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.012930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.012930Z digest=sha256:8adce105f8b71ffe52a926b841e52454d7df6e3ed42ea5cf56a1073ccadca377

Observation 418e0488-da97-4b86-837a-c4819cae40e6 · outbound

This paper cites Visual question answering on image sets.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Visual question answering on image sets

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:46.204962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.019073Z digest=sha256:252ab421d5cfa5dbe8f3f9e8379a4e3f22d846c293d55d77be35b6cc4670f20b

Observation d3a42c52-c14f-4e79-8351-059360ab9996 · outbound

This paper cites Rubi: Reducing unimodal biases for visual question answering.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Rubi: Reducing unimodal biases for visual question answering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:46.188257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.025714Z digest=sha256:a5181625ae0ad531d11371e623c2148b7c8e2e1fd3e600a53ad720f4d0d6204d

Observation e2bc77d8-2610-458f-b0f9-44570dff61e6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:46.171028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.031370Z digest=sha256:093798c42306471c6685d8281a94221a0d10a368d45b44acf73acac21518f8cf

Observation 003119e3-3d0c-488e-9647-a969312f1d80 · outbound

This paper cites Learning to prompt for open-vocabulary ob- ject detection with vision-language model.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Learning to prompt for open-vocabulary ob- ject detection with vision-language model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.042023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.042023Z digest=sha256:82e84b39cce321136f1bc42e58e4eca8bf730498d544e6e3a3db7cfa4bfb5878

Observation 0dddba74-901b-4687-bc9e-05c453ba19d6 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:46.109156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.049441Z digest=sha256:8635050beec85b37b9d52200a4495af8fe975e4821c2ee843eb3d5c2c61cc389

Observation 3a467f93-04ca-4118-8162-befb31904b1a · outbound

This paper cites Complexity-based prompting for multi-step reasoning.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Complexity-based prompting for multi-step reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:46.090078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.054933Z digest=sha256:800256eee0e8ff4b098885646680640724183840bf7737b65aa822b8f32bdf62

Observation 3ea81aaf-bce2-4f73-8bf6-94b8b0f89cb6 · outbound

This paper cites Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:46.071887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.060886Z digest=sha256:13136a0605e5c82b699191cbbc543e83151151f2c52d1f58f0230b1b02fd6970

Observation fe1b2a44-f5cc-431c-bf8c-c1c1ae72c523 · outbound

This paper cites Reasoning with language model is planning with world model, 2023.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Reasoning with language model is planning with world model, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:46.054274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.066142Z digest=sha256:12c53f904548ab2122f0909a37080e26ec65c9e2aa5ee69aed6741e8e0e3e57f

Observation b710a177-c817-44eb-acd7-d705215ae4f0 · outbound

This paper cites Towards reason- ing in large language models: A survey.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Towards reason- ing in large language models: A survey

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:46.035823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.071456Z digest=sha256:c66c6060cd8436db2645cc08988d8929513e4430efc5a917daa47af4b22e7857

Observation 7605b9e2-e2cf-4856-8add-bf8b6aba6039 · outbound

This paper cites FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.077022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.077022Z digest=sha256:581f41dd133f95e9413fe4cd6d401f223492da4bf378fbca0fa21a072beca4d6

Observation 48e1db89-2dfe-469c-8d84-36cb4c85c192 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Seed-bench: Bench- marking multimodal large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.082890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.082890Z digest=sha256:5c37d2468a21baf580cc6e56a24975812a38eae3bbb089916101e23c035e492e

Observation 71a2c00d-b5c0-405a-b7ca-4b598b2ab2be · outbound

This paper cites Making Large Language Models Better Reasoners with Step-Aware Verifier.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Making Large Language Models Better Reasoners with Step-Aware Verifier

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.088222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.088222Z digest=sha256:cbd520e2c2c3b2cc51552188e2e3e9b1a5507db38e4007ef22ac143e8b5a43a5

Observation e8ca54f3-666d-45ec-a171-3ede520f9401 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Evaluating Object Hallucination in Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.093144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.093144Z digest=sha256:b1b0f708fe372d8fac32406cee1a57a6cd438266041cbf7a1171d5e0c50c2270

Observation 8ba21dcb-707c-4579-b659-6f34f7ed9736 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.098207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.098207Z digest=sha256:f825767ac865640544b2e87ac89efb5eac6e48467c1c210ee1399eb0c5c47cb8

Observation 47b632c8-6cac-4f46-b534-b63e918fa968 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.103211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.103211Z digest=sha256:145fab0deafcc3fad7a9b54471b9e5c67b6e5756117a4f1a4c196eab1161c27e

Observation 1c6e49f5-f8e6-4e92-ac46-b989b08edcc4 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.109012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.109012Z digest=sha256:4e3ca07266cf89310965a0ff5cfa63cdbca53af7df617abd36ea4aeaf69d4f0a

Observation d4d8b8ae-df91-4644-9f1c-b54758c07750 · outbound

This paper cites Improved baselines with visual instruction tuning.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Improved baselines with visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.114224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.114224Z digest=sha256:340157c1b844831cd83e1baf8e589b1e5f08c90bcf255556ac8a6ac69c5983ad

Observation 90cc9f59-b33a-49a9-8b6d-bfe8d149ab19 · outbound

This paper cites Visual instruction tuning.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Visual instruction tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.119893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.119893Z digest=sha256:d01d2829182a7cbc38d6c3b3fb942dae9b5cafb90d1394b0e582999be03f0bd1

Observation bcc4d45e-8e47-4304-92e6-8782866ebf17 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination TempCompass: Do Video LLMs Really Understand Videos?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.126098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.126098Z digest=sha256:db687b7294aaddddce309d184296e689b87524eb56451b99a1620bf7a2fee075

Observation ed4a26d4-44b4-4ea9-afd2-a0c074532fe5 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.131502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.131502Z digest=sha256:cbcd8564cb3be1c3f4f858d63848e7642cd5046d5c23821f28256ccc95be52bf

Observation 4b294586-5ab2-4bbf-8dc7-ac70814d972a · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Compositional chain-of-thought prompting for large multimodal models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.972918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.137305Z digest=sha256:371e0d1674873b22056b632d4749ed0f2ca488fad6fc803ca22064df33eeca6c

Observation c2ab1977-10c7-4792-b7f2-9b8883e4bb6b · outbound

This paper cites Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.954577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.143347Z digest=sha256:7e7f40fb62e378c4a3ddd499a8ecd7a1d218e4056c9eb30e0599b79d82952958

Observation 5f0b499f-b78a-4530-8bb4-1132459c84e8 · outbound

This paper cites Gpt-4o system card.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Gpt-4o system card

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.937701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.148321Z digest=sha256:4df1dbd2fde761ba7b9641a00f73ec6eac450f0586f05524804d40211195961b

Observation d7611116-14db-4f4e-afc9-72a8cfe821ff · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Learning transferable visual models from natural language supervi- sion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.153643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.153643Z digest=sha256:ae1b67d34fe2a2d24132dddf03bda19afb1920482f3a7ed23ee62101e6860bf6

Observation b8c425de-b99d-4d42-b227-4a89295de079 · outbound

This paper cites Learning To Retrieve Prompts for In-Context Learning.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Learning To Retrieve Prompts for In-Context Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.160409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.160409Z digest=sha256:9df1409aa4a0784446e8a9c158480dfd93636ee4f88999cb2a351d9a04bc06f8

Observation bfbb3fde-d773-4201-84ec-34ab9fc8bfd5 · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.165361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.165361Z digest=sha256:e35fdbe06890bfb51f5c5535f2dc99ae1ddba08853c2363c09dfaa2c5fb2da66

Observation 651e6ade-92f5-40a6-a53f-b316294644c9 · outbound

This paper cites Cognitive psychology.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Cognitive psychology

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.903238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.170551Z digest=sha256:c942913f99584c1c930c243f5c4d48aa43dffea088d044255c64c796af89ba3e

Observation ad6ad5e7-c6b1-4451-8e09-448e4aee02c1 · outbound

This paper cites Expectation (and attention) in visual cognition.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Expectation (and attention) in visual cognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.882580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.175766Z digest=sha256:d2525f0c4c7cb626b077e305bb09eb8dbb22659d2638f0d2356ab91f3e482015

Observation e6511a7b-8b8f-4f33-848f-39170cec7c3f · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.180971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.180971Z digest=sha256:0fda223aa7639c5eea77c920bbe431b34f71415d7e6da55ca109303d625bcc77

Observation d6fa9169-ed14-42e0-b2a2-0d24ed339044 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.863860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.186255Z digest=sha256:f7e52d9e9e82b9b75a02b5d32de57a96f8ca75656508b57de7989a3e9e082068

Observation 7fc8af1f-7def-4fa6-95e6-40274a45617e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination LLaMA: Open and Efficient Foundation Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.191013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.191013Z digest=sha256:a759ac137c540f1cae0acfaa2d5786ce46f7b6ada3c6147a59f7dbb8d07e68da

Observation 47b4dfc1-bf0a-4a5a-b01b-aaf1c93cba43 · outbound

This paper cites Cross modality bias in visual question answering: A causal view with possible worlds vqa.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Cross modality bias in visual question answering: A causal view with possible worlds vqa

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.846308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.195527Z digest=sha256:d16cd8d363e6b8ca121ad30f72705c006d626fa3f2715cd4db20e1feb0489efa

Observation 21645259-89ca-4f6f-97d2-526a508de09c · outbound

This paper cites Vigc: Visual instruction generation and correction.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Vigc: Visual instruction generation and correction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.829114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.199952Z digest=sha256:1e93baf468c63d1cd8d1c2a0ac71629da215676da141f3330d7559837cc17908

Observation 063d4dfd-0678-4759-ab06-fcb52b1123ae · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.204331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.204331Z digest=sha256:ff726f5a19e2d937094ebbd1937d9bbb0a05d83d939b69b6325f419944721b34

Observation 27619d99-79b7-4f59-b1fa-1f32b372487b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.209840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.209840Z digest=sha256:c4cb7a8c031dd80bbac7c7ba84e62db649d79cdc4fef709abef5d88e613868d1

Observation 96814179-5438-4fe0-9f7a-c16c0cc22857 · outbound

This paper cites Self-evaluation guided beam search for reasoning.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Self-evaluation guided beam search for reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.797178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.214962Z digest=sha256:220fc4596d04cb6d2abd1f16e77ea26a70ebfdbd74e8c17faeb7915710d4ff8f

Observation 964677b3-b86d-47f7-97e1-7731716164d4 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Tree of thoughts: Deliberate problem solving with large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.221100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.221100Z digest=sha256:d95c6fc405851a8bc8c7264b7c37c6eda3ab7d4b3dd91b65fa7c9210df9073af

Observation 7a2c6989-198b-43ee-987b-7b279199e7ff · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.226470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.226470Z digest=sha256:1938502fb4cc4ae33b2c02dd6b0a5cb5d483377b9b197089d77eb37a70bcc55b

Observation 593b5a7f-0b12-4f86-a585-7fab7a251aef · outbound

This paper cites Contextual object detection with multi- modal large language models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Contextual object detection with multi- modal large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.767096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.232123Z digest=sha256:18526a401fe15d7b70eb766e17f7e0dad8eb118b1ed268ab0085d592c63ada12

Observation a76c9d70-7b37-41ec-b452-3fbc1b93f645 · outbound

This paper cites CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.237845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.237845Z digest=sha256:41366ac9c6bca7648f3d476801142429667cc1fa150db0fcd9e0ec762fd42cbf

Observation 0d98664b-cc5f-408c-aca6-3128f8ff547f · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.251046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.251046Z digest=sha256:41a3980d0c4c1650ac096e9a6e4c5bc8f357fa732461749e6c404d625e11a087

Observation 9f52fecb-9ad2-48ec-b77b-d136b0a4f160 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Automatic Chain of Thought Prompting in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.257211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.257211Z digest=sha256:0fbe6f78746e4a5f80744437828a17b41a5ba62f6562ed380fb6915214c40e85

Observation 2f2804e1-aa08-4b14-80e7-368bba383291 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Multimodal Chain-of-Thought Reasoning in Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:35:45.263587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.263587Z digest=sha256:593b4a4b43eec89e113e98cb27c7963ae8784d3b1b5dcc514d45dc5a8ba97b41

Observation a3f11512-b227-4cc5-875e-26cbd892782d · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neu- ral Information Processing Systems, 36:5168–5191, 2023.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neu- ral Information Processing Systems, 36:5168–5191, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:35:45.745985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.269303Z digest=sha256:124171873953ca52c0bd2022ec5bc94366911f7ab4e155d3975364c0f086169e

Observation 4de40864-9826-49af-bbc7-6f277bab5103 · outbound

This paper cites Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

Reference 48

Resolution
malformed identifier
no resolver link, observed 2026-08-12T19:35:45.274312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:35:45.274312Z digest=sha256:3502567ca68338352efd581871a91cacf96fc94f6d896ba64150ed1a30ad8c0a

Observation db9bf6ac-8065-4ad1-8cda-dc6bdc911ad9 · outbound

This paper cites an unresolved cited work.

Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T19:35:46.152884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:35:45.036827Z digest=sha256:1f715d219aebb727458da05177164fbe53ca4e120bebf1cce564829249843b45

Pith citing papers

Observation 50ea1800-5255-440e-951f-995cfe6c9911 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 220

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.443213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:8cb638314c30f75cbec2b112daf2cc58a0fe5e0c78db5862b49b2ff1007cb0f9

Observation 3b478f3d-7e97-4776-97be-1a2b1dc81a7f · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.749212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:5f0329a9e33796a1ea7e7d30e014cac9dcde3846dd007c796a62ddda8bf327be

Observation 7f35f66e-4cbe-463e-a511-e8f97f085ad1 · inbound

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM cites this paper.

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T12:33:13.663322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:33:13.663322Z digest=sha256:599689b82b5a019d7c6feb9efd2683457b2d2d9712b43f575914e7edbf7dae27

Observation 1235f923-a65d-4a8e-b541-feeaed84209a · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:00.846334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:00.846334Z digest=sha256:debdc402516a28de3bab000e9350da3cfccc19fbfc56e5199316e2a7ad5c429e

Observation 0e2aae97-76fb-4127-a7c1-4530bfc3e7ae · inbound

MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment cites this paper.

MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:10:02.962883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:10:02.962883Z digest=sha256:03d20f495c6a97f196fb24bb56633db19091a56ac3896968e1218ce9f1bbb2de

Observation 140f773a-c2e3-468a-b8c4-9dc3f4a975b4 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:01.812797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:01.812797Z digest=sha256:288cb3b6429dab1e61f025b172988863c0c464e11ce70a6cecff36e71e3fa155

Observation 74692ec0-193e-4668-9bc6-eeab11ca3654 · inbound

On Semiotic-Grounded Interpretive Evaluation of Generative Art cites this paper.

On Semiotic-Grounded Interpretive Evaluation of Generative Art Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:52.681366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:35:53.179645Z digest=sha256:b0a467652a9e92fe28907df75528ae3d49ae536fdc914d06ed68985b32bebdd9

Observation 86f32005-f045-4120-a6e7-e2ddc4732861 · inbound

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training cites this paper.

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:22:23.500754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T06:17:57.264809Z digest=sha256:66061a6d49a04f23c8ef3531b06b226c7053e835c2dfdcdadf0372fcbea54ee2

Observation a612626b-f8e3-49ce-ba22-58413842c0b3 · inbound

CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection cites this paper.

CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-12T06:01:58.407747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:01:58.407747Z digest=sha256:25dbb31d2a2ccc2d41cf98b1f4061e63a75840007b62d856d1887842d5e5e959

Observation f4cdebd0-27c0-4833-8107-4148bfc07050 · inbound

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models cites this paper.

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:48.425006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:35:48.425006Z digest=sha256:5acdd57a513671927d9c56555ae310aa40244dad44b0e95e3f1cc5418c6fa07a