Pith. sign in

Paper Citation Record · LEDGER

The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.12787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.12787 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:30:22.181322Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.441655Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0fc46da7-1f83-493f-b96a-7791f73734d0 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.231593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:e7fa42e12923e9ddc18ed895129de1e1222f16a486ce205b616bb1c793de1037

Observation a9af95a0-7df2-4c73-8114-69135c20ee99 · inbound

MLLMs are Deeply Affected by Modality Bias cites this paper.

MLLMs are Deeply Affected by Modality Bias The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:22.181322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:22.181322Z digest=sha256:a811015bf53b97dc824ad8f88b526681db25795071c230df772d07dc87546d65

Observation 430d8b40-28c3-485f-a683-4845b1a48d5a · inbound

AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models cites this paper.

AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.761053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.761053Z digest=sha256:0794061401df4fd52ad253a1d5940429d950b3774b98993761d11d322caf668c

Observation 833c9254-754e-4661-92f0-21e24bb4c782 · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:45.280415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:b0b1ce08f807ff5c8f059ecbce37fb8b7a0977b207c0f55d4a7b43939faf0e23

Observation 70a1b00b-d496-44ea-9da8-771c397e9079 · inbound

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models cites this paper.

MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:06.449230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:06.449230Z digest=sha256:9dc624d9a1996ae1e821f607bbdc0d2737c15724cd9d5cf6120cd2c184e3c7d4

Observation 3b48526d-9f14-44c8-a57c-2dc57b7492f2 · inbound

OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination cites this paper.

OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:22:19.020115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:22:19.020115Z digest=sha256:c8b1723ac8c856564d5e4fa01baf06618c2ceafa593470698acf7d8c563617da

Observation cb2f6a73-a98a-4d9e-8f62-d4c276e99bd9 · inbound

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models cites this paper.

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:29.026513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:57:47.356373Z digest=sha256:86232463a27306035d8719fd79b9470af45a0b41bb23ecccb84b9564b4036ece

Observation 3dba48bc-6a31-4745-8e45-20087f00ad48 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.174926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:84a8d698b3fc7c3793ae017d924f175a1e54cd1235818bca86352fa39af59da5

Observation af26b459-ad70-4ae0-926e-665a5b91c5a1 · inbound

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation cites this paper.

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:27:44.948394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T12:04:50.483329Z digest=sha256:edad522b170143e1dd8230bf4409147e020dd35350c4589b091bf7aafc0fc9c8

Observation 4dff1acc-d9ae-4ea4-a3f6-1993317d2aba · inbound

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models cites this paper.

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.236037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T07:28:54.731659Z digest=sha256:07fc46fe150b2632f87475eded02d56163f4acd68bdde1dd04cc9232a4a2c96b

Observation 3a63a04c-2f71-4ad5-81c3-040a4a07d6ed · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:06.444991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:ace66361aab388eec6efbc024abcc7370c82327e021e0a0690912f695981e08b

Observation 4195d20a-1f75-4cf4-8761-4f07df3a19d3 · inbound

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding cites this paper.

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:06:02.491723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T01:18:13.657975Z digest=sha256:9c979f82b2ad8a4787a267bf3f5b45511fd54b8930cdad7c5ebc0e62ef0048c2