Pith. sign in

Paper Citation Record · LEDGER

Visual Hallucinations of Multi-modal Large Language Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2402.14683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.14683 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:53:57.505100Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:56:29.276137Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bf07f7af-9c6a-4294-a552-7cd19cb80e1f · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey Visual Hallucinations of Multi-modal Large Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.817257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:8fd65d016f34fcbb0199e75db82b0cb2d081498ee6889191c69c1e722a70eaa4

Observation 0cca5bdc-380e-4862-a3fb-e35c7739b617 · inbound

Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering cites this paper.

Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering Visual Hallucinations of Multi-modal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:10:15.470574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:10:15.470574Z digest=sha256:94730bae7b74b73909ca853fd6761b2a5317b9d5f02c7c1c7524c8a30e3486b5

Observation 86984983-2499-4fd6-ae75-bc43128e2e14 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs Visual Hallucinations of Multi-modal Large Language Models

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.149903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.149903Z digest=sha256:c6611021f00cc5ee042bb6a0a0f87987554c3d4e02417542cd9073dd85218a0c

Observation 92d78a79-3cd0-48a6-a544-68334969cf80 · inbound

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection cites this paper.

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection Visual Hallucinations of Multi-modal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:50:08.127179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:50:08.127179Z digest=sha256:d858ad89cb114894ae3ed43ba70da9f40de3a6608c0f3b5f6551fbf8941b707e

Observation 920249d8-c86b-4e6f-b7c4-f0a867405d35 · inbound

ChartLens: Fine-grained Visual Attribution in Charts cites this paper.

ChartLens: Fine-grained Visual Attribution in Charts Visual Hallucinations of Multi-modal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:19.398932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:19:19.398932Z digest=sha256:c7b7c50765c15c1fce654a6725b9bdcf9d57b20261c26ac98cd575d332d0eaa6

Observation 862b44a8-62ee-49d1-b434-f352787d9f27 · inbound

Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model cites this paper.

Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model Visual Hallucinations of Multi-modal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:57.776021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:57.776021Z digest=sha256:dda47094e13ebc2402d3d9c7fb773e59bbbfa776a334f129ec8208e57f3e2ef7

Observation 467d2227-2375-4c1c-98e8-3e3dbcb4cde6 · inbound

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models cites this paper.

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models Visual Hallucinations of Multi-modal Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:15.573134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:15.573134Z digest=sha256:1583d518377fe2ce0ecdf60fbf457f955a82239e006d0a2a675293e3aa18611c

Observation d0bd4249-50b2-4801-a136-0565748610a2 · inbound

ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM cites this paper.

ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM Visual Hallucinations of Multi-modal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:53:57.505100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:53:57.505100Z digest=sha256:02ee1e5000847fcc3d5e0f0a158e74001ab4236e87ff3dd8b77cda9c573fef5d

Observation bb510fe0-5957-4869-8512-70d7d52aad82 · inbound

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs cites this paper.

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs Visual Hallucinations of Multi-modal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:44.314634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:44.314634Z digest=sha256:b403502d16b97fa59890109225915def53c912645e251a9d8dea947bcca9e860

Observation 441644d5-d89c-431a-83b2-81a162f614d1 · inbound

ReFrame: Rectification Framework for Image Explaining Architectures cites this paper.

ReFrame: Rectification Framework for Image Explaining Architectures Visual Hallucinations of Multi-modal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:59.754622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:59.754622Z digest=sha256:3c029d86058f0c4f2f5e0984b74145ac88f03c201c91e79dfa36b4b07316521b

Observation 490c5c7a-309b-4952-9dd1-f72c7038452c · inbound

Unified Multimodal Understanding via Byte-Pair Visual Encoding cites this paper.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Visual Hallucinations of Multi-modal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:54.374403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:54.374403Z digest=sha256:09be3a2b2ae2dc30be5b501cd6a258ee0c60ebb9fc687dbd6f136a099dbfb25e

Observation 7c577c89-7a9d-4e85-a3cc-d7e6067ac6e4 · inbound

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs cites this paper.

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs Visual Hallucinations of Multi-modal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:12:23.562660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:12:23.562660Z digest=sha256:462105b59248f3230fe29c1d2e96762c6d2ac5847b81fc079d672944cb8737af

Observation 22d363ff-ec68-4982-b136-9c4ba7dffbba · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Visual Hallucinations of Multi-modal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:44.432144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:44.432144Z digest=sha256:9caca5fd9daa6a9b9e51bf60cece519921a937ba591430fa64bf2b30f0cfe1e1

Observation ec4e3dde-8ab0-4c98-b4cb-7f4132f7ba17 · inbound

Mitigating Multimodal Hallucination via Phase-wise Self-reward cites this paper.

Mitigating Multimodal Hallucination via Phase-wise Self-reward Visual Hallucinations of Multi-modal Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:43.337592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:10:45.144421Z digest=sha256:8a2dc131d598466961cb8a24f60a6c3b7ed919622bbb97db5664e844599822ca

Observation d8d4980c-0baf-499f-9a1f-34b238e14fa3 · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Visual Hallucinations of Multi-modal Large Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:05.974790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:2fc6f5b34389592f1bf8c21f5a299b2f84e812f40ac544836165c3807ad248b8

Observation 4671d905-cb61-4458-ae74-f0ea239706a4 · inbound

When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs cites this paper.

When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs Visual Hallucinations of Multi-modal Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:05.264219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T01:55:07.172228Z digest=sha256:a3cee4dfe4dcd482c4d4090dabc03f296c1024740d5a6745173090d5a87deb04

Observation d80d52b0-69fa-4bf6-b027-362ec319cc4b · inbound

What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness cites this paper.

What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness Visual Hallucinations of Multi-modal Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:46.607805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T23:21:50.254449Z digest=sha256:5d2b1187cab95645e8b1a20cabc85b3bd8697e501c63364edfd24b36d549bf4b

Observation a6fdc414-622f-4fc3-82fa-f89423275927 · inbound

Investigating Adversarial Robustness of Multi-modal Large Language Models cites this paper.

Investigating Adversarial Robustness of Multi-modal Large Language Models Visual Hallucinations of Multi-modal Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.581523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:11:34.152223Z digest=sha256:ee0e34af2812f62dc9310faa10bb7b05910f1418ca14262e6ebb6550de44de7e

Observation 89f0e3dd-7a66-4585-8667-6330cf9a416a · inbound

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation cites this paper.

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation Visual Hallucinations of Multi-modal Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:56:29.278185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T10:28:41.952330Z digest=sha256:6f1b8d98c96860afa8fbae69e74310b86e177c72d9f614bddaee73d005de3aa0