Pith. sign in

Paper Citation Record · LEDGER

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings

As of 22 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 3 inbound Pith citation observations for arXiv:2511.05017.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.05017 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:35:18.935031Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T16:34:34.452430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T10:53:13.112211Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d35fa9a-aab1-4ad8-932a-d20e71b3dccc · outbound

This paper cites Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.891961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.891961Z digest=sha256:4e6716dca8c0f0376a749294a1666eb5673135aa302a500c8aba765a9a56173d

Observation f088ccac-837a-4e8c-bd5f-478b8ee0fe35 · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Efficient Multimodal Learning from Data-centric Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.894715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.894715Z digest=sha256:b04985758b1cbaedded8e7645e7c6f82d94e16c044359750dde0e725f25ab0fc

Observation 494c1152-5089-467c-ab79-b9d9791b3a7f · outbound

This paper cites Faith: Faithful and informative textual hallucination detection in image captioning.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Faith: Faithful and informative textual hallucination detection in image captioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.900175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.900175Z digest=sha256:f7e087550f90784899741f14b5212f1b50029453df71d14ad750ccf9dfa58ae9

Observation ab1d5f5b-58db-4c43-96dc-a236dae7fbf0 · outbound

This paper cites URL https://openaccess.thecvf.com/content/CVPR2023/html/Jing_FAITH_Faithful_ and_Informative_Textual_Hallucination_Detection_in_Image_Captioning_CVPR_2023_ paper.html.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings URL https://openaccess.thecvf.com/content/CVPR2023/html/Jing_FAITH_Faithful_ and_Informative_Textual_Hallucination_Detection_in_Image_Captioning_CVPR_2023_ paper.html

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.902642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.902642Z digest=sha256:0e58d1bba2da2a3d24dfaa97f9e228859032e1a4cee65e587b8cba9f5b50ea63

Observation a6840d9a-8e78-4a18-b8ae-34c1947da9df · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Evaluating Object Hallucination in Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.905058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.905058Z digest=sha256:23c0153b300d9e8660945225f95ede80c1ac2f3c8ec51235c2326cc2b8d0efdc

Observation 066a815b-5d7f-4b47-b465-4e705632a3ae · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.907951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.907951Z digest=sha256:73f432ed28410c8ecf67ec6ae28d8640a6d766dec42442a8989ed05b435583b6

Observation 6513c6f1-6fa2-46c8-8b5f-2dc6ed8e91f0 · outbound

This paper cites URL https: //openaccess.thecvf.com/content/ICCV2023/html/Lovenia_NOPE_Evaluating_and_ Explaining_Negative_Object_Presence_in_Image_Captioning_ICCV_2023_paper.html.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings URL https: //openaccess.thecvf.com/content/ICCV2023/html/Lovenia_NOPE_Evaluating_and_ Explaining_Negative_Object_Presence_in_Image_Captioning_ICCV_2023_paper.html

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.910544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.910544Z digest=sha256:7b072fccf8f08b7d75100a991d1c8683c608485455f9ede7564f70f951a16700

Observation d1720cf9-477e-437e-8ee6-b2508fcc5b00 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.913235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.913235Z digest=sha256:8b4c02918138b5a29c19aeb311626dd4fe57c0d9278479777ff38fa8961b8d97

Observation a871f417-af15-47d5-9f2c-c4a47cde47be · outbound

This paper cites Aloha: Assessing language-only hallucinations in image captioning.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Aloha: Assessing language-only hallucinations in image captioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.915633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.915633Z digest=sha256:f8da457e5130f2edc4c3ef3ac423156d9e029f4395f62e3329842296756cb608

Observation 96f42c52-834a-4756-ad4a-1f47a3c891dd · outbound

This paper cites Moments of the Poisson distribution of order $k$.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Moments of the Poisson distribution of order $k$

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.918106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.918106Z digest=sha256:d6445985afc3557b7cb5e8702fec63a2275856e53c5bc7ce7590209fa63c9833

Observation c1f68f37-682c-48da-bb34-d2e721074214 · outbound

This paper cites Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.920899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.920899Z digest=sha256:17df7314635f4cef0d37ca0d1a68c774c016c0c9dc8ad3b3c23f9648106a64dc

Observation e52a1569-1195-4c37-a629-46c10e323306 · outbound

This paper cites EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.923981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.923981Z digest=sha256:325b3710d882a94d63dd042645ca021474e5d1c95befe3a090e0fc5b79f999e3

Observation 8b18b037-8b02-4f4e-86ca-244c91950496 · outbound

This paper cites Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.926817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.926817Z digest=sha256:b087e0896d9b579616472b5bbe7d881dcd8a5a02c34188ca09721bfcda5d28c1

Observation 3d575566-45a9-419f-a1f3-548f984800a2 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.929446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.929446Z digest=sha256:cd4509197f6412ba97817490e38217e430dc271f4fb85ae77872c4e7c735dd75

Observation c59ee11f-0d63-4839-b71d-d11eb97f3df8 · outbound

This paper cites We evaluate the individual and combined effects of Visual Contrastive Decoding (VCD), and VisAlign.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings We evaluate the individual and combined effects of Visual Contrastive Decoding (VCD), and VisAlign

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.932326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.932326Z digest=sha256:38990542f081af97fa0f27d5a6bd86057170bbaa7c44866ec8516573d145a29d

Observation e1256842-56d7-42c9-8bc9-d7a31a482739 · outbound

This paper cites an unresolved cited work.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.935031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.935031Z digest=sha256:85fcd7eae96c62107a87cbdef0d58e6b3cc8e75aa04830e20ea0390eb58cf156

Observation 0ef1a681-eb6e-43c6-b5d3-5aae4a26879d · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.885878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.885878Z digest=sha256:acb6b7e884e5c708bc18dd6730b4d9d1a449ad855d605896336334ac65d3928b

Observation ebb7a317-260c-4608-843f-a1d651a4ee3c · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.888887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.888887Z digest=sha256:cb0e76f843b8f8629ddbba4474db58036e58d67f61581db1ca10a2470b360a41

Observation f5abee2a-0f4c-46ad-88f0-8ad0a3f1288e · outbound

This paper cites Mitigating object hallucinations in large vision-language models via attention calibration.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Mitigating object hallucinations in large vision-language models via attention calibration

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.897419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.897419Z digest=sha256:20983dd896c678290678adf38fa97acb2a9cdf879d3ab16a1ce60f762228eaa1

Observation 55ad7066-1187-4b0e-98f2-ea60b2651ddd · outbound

This paper cites PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.882276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.882276Z digest=sha256:22cc928ce5c1800889409d9d0056f47fa9002982cd2c684c1256ed5208b139f3

Pith citing papers

Observation 6424e6d0-f282-4c6f-8ebe-d7823008e478 · inbound

Improving Vision-language Models with Perception-centric Process Reward Models cites this paper.

Improving Vision-language Models with Perception-centric Process Reward Models Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-09T02:06:07.552179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T04:33:36.634359Z digest=sha256:15985e9fcdb58e80d9cd94ba13b97add9c1295c98d5c66082d1f753ea41ce874

Observation b2bad249-c48d-4360-862b-2468e6206f8f · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:07.552179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T10:52:32.238241Z digest=sha256:e76149f3cf3bf88f786ba48e5d6f276357461e61ca8aace68baf9b7fb920b640

Observation f04da37e-6bc5-4ef1-8ca2-18823aaf51c4 · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T16:34:34.452430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:34:34.452430Z digest=sha256:248b97ea9a72ec4bbc89cf5d4c654fda2cae92ceeaad01c5e8bafeee0a62f406