Pith. sign in

Paper Citation Record · LEDGER

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation

As of 15 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2510.22067.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.22067 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:14:37.817377Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ae79580-080b-4527-abfb-919b202f836d · outbound

This paper cites Qwen2.5-VL Technical Report.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:35.768024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:35.768024Z digest=sha256:75050dcd2135fac0de25d17641fd74d88cb415f7ac85fbc8dd24a378229ed850

Observation d8977dd1-3c51-4d54-a3fc-f62988cf15b8 · outbound

This paper cites HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:35.849953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:35.849953Z digest=sha256:3c05253b1337e602b3be91f54ba70957cc5b177670d9877efcb3d92224512a61

Observation 44510d44-ff8f-4659-942a-4ce9cc822da4 · outbound

This paper cites Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.018675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.018675Z digest=sha256:0b6f29341ce1c700a31e1e910b56428e7cdfb09c64d4db2f1656dcd1e25249da

Observation f9d43695-07d0-4e16-ba21-541b64c32ff8 · outbound

This paper cites GPT-4o System Card.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation GPT-4o System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.169734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.169734Z digest=sha256:c75ae960e22d0f9c4f79dc3986564750cb96c750c1e73c11539c37be5f6db53d

Observation 863d9f65-ccbb-4e1e-9ae2-fd0b085c9787 · outbound

This paper cites See What You Are Told: Visual Attention Sink in Large Multimodal Models.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.218482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.218482Z digest=sha256:71f68e8c4eed4177ed444ea51bd466b8c86478694f4f27768508faf0830a2f99

Observation 6cf07e3d-9b3d-4692-b185-ff93c934dce8 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.In- ternational journal of computer vision, 128(7):1956–1981,.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.In- ternational journal of computer vision, 128(7):1956–1981,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.259983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.259983Z digest=sha256:bf8c2bd80f3a558091d6aa4bd45f9ae0d8af8085e32a4c0906534e4e8ee12dd1

Observation 2238118a-95e6-4ae2-a752-ed872caccf5f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.303147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.303147Z digest=sha256:5576738f27c5de81b560dba4a67281900dd3287b137818e8d40366c2a0be2ce4

Observation f7c9ead3-b030-400f-8aee-7eb155274e66 · outbound

This paper cites LLaRA: Supercharging Robot Learning Data for Vision-Language Policy.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation LLaRA: Supercharging Robot Learning Data for Vision-Language Policy

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.362519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.362519Z digest=sha256:46fa18053ebab8feebddd59a04091de1777b8ee85b5ee0f6f563b492c05d34d0

Observation d2be8d0d-1efc-48f7-90dd-cdf54dd17cbd · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Evaluating Object Hallucination in Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.395839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.395839Z digest=sha256:8f9023c199e1c0e412c0cb40db3b6100a5a7c5c182f5bd0cbff9052fb2004373

Observation 8117fa5c-6801-4783-9d2a-43967664a79a · outbound

This paper cites Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.541664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.541664Z digest=sha256:6b1a237af37235546c89134b395d62779b0ea159edccda199083c4a49057a54a

Observation 5ef232f3-ed38-462a-8175-99bae75ff279 · outbound

This paper cites Object Hallucination in Image Captioning.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Object Hallucination in Image Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.635360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.635360Z digest=sha256:bca99599c1a9ea24455b53af87f11e5f5c5a2a4fac3e04aa9e09a0ae353dbc07

Observation 6455772b-8bc2-4e39-a866-c54cd2bd2bd2 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.745595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.745595Z digest=sha256:edf44036cd0d41a4e1a88223160852d5a6d091bfe4d3589bef98e9f374fca521

Observation fd86c048-156a-4a5a-8069-44b7abf78f3a · outbound

This paper cites MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.819533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.819533Z digest=sha256:0271fa19931834b2e7fa7cddbd0a454dc0ea49826c74b13c816013781323d86e

Observation 6f147be7-e61c-4e8b-8103-9c192ac89fc8 · outbound

This paper cites MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.931127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.931127Z digest=sha256:d35a7b18ac8719310d93a4a3d5879e0277c1e1980e5a48ccebcbeb3dc933718a

Observation 1657e0f0-7d95-470a-bb6a-b6c83753092a · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Efficient Streaming Language Models with Attention Sinks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.022240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.022240Z digest=sha256:1fc2d2f8482eb5ae9720ebf4566d76cbb479f6f90ca60b8864f01018a98b01ec

Observation 0925bab1-209a-4a2e-baae-b823c54a58ab · outbound

This paper cites TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.107192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.107192Z digest=sha256:6d5fe5d2c6a6820f608034e5dc9f8c1c7011fd820f616966e129436440b8d3d3

Observation 71750f2c-0f0a-4adb-a554-14ff6c96b1d3 · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.271830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.271830Z digest=sha256:d8e421dccfada2c96412e1144b1f4e5763af9db5f2fc73182304a52276f05116

Observation 2c38204a-6773-4a42-8f8f-36161a6ecb54 · outbound

This paper cites Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.418270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.418270Z digest=sha256:7665be21d30c352bc355d229a2f2f4fdb250ba58efc2a753e298d341a86e101a

Observation 62835c2e-a1ee-4832-8abc-d6737303a147 · outbound

This paper cites Debiasing Multimodal Large Language Models via Penalization of Language Priors.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Debiasing Multimodal Large Language Models via Penalization of Language Priors

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.477584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.477584Z digest=sha256:5adf4e34526b142d52ce375d74e607be6606a75f30cc20c322f9eb481a61c5fe

Observation 25dd4e4e-ab37-4922-b807-c21f269f41e1 · outbound

This paper cites Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.532330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.532330Z digest=sha256:b8c48951be1d1ac4cd47b4c54cdba10e71461f8525a87538a15ffa65983ec835

Observation cc44c118-708b-4dd8-bf3a-ebcd0d0a57b6 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.580891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.580891Z digest=sha256:4e63788b37c325b93022f35f7d76ffd10eadfdbc759394c5dea875403d011a63

Observation 3d8c921f-02f7-49c7-b9f9-fbe8cb314f8b · outbound

This paper cites Is there a frisbee in the image?.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Is there a frisbee in the image?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.624957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.624957Z digest=sha256:9549270c4293dbb66a94a16ddb9b2e036a10d5ca14bdf4125ec36817a11c5b80

Observation c95a9c96-069b-4f01-9bf5-c8eef3349d5f · outbound

This paper cites an unresolved cited work.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.677470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.677470Z digest=sha256:3bffd69a942567e196eda1c4e58eb16cb8d242e8f4fcaa38217accfaf6f7938f

Observation a1ff3dfc-3f15-446e-98d6-3eb5779162d9 · outbound

This paper cites Please just answer yes or no.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Please just answer yes or no

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.733137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.733137Z digest=sha256:7e32061b9a3dd8de919a982cab2d760179acb6bfa9b81686297ec25d70c4b1ce

Observation 1010fef4-3055-45a3-9bd6-836145d28101 · outbound

This paper cites Layers for Cross-Modal Fusion Enhancement.Since the benchmarks we consider lack dedi- cated validation sets for hyperparameter tuning, we follow Kang et al.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Layers for Cross-Modal Fusion Enhancement.Since the benchmarks we consider lack dedi- cated validation sets for hyperparameter tuning, we follow Kang et al

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:37.817377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:37.817377Z digest=sha256:c9ac757bb5a998dea9f7c20e9bcaaeb1feeddf56a7f08578e2d3e147010c0a45

Observation c248c1f3-1199-4590-9a48-c1ca94ec102e · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.452212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.452212Z digest=sha256:30f724323495b5e837bd1f0264fe09eb841c1e51a6107d1ecbb254fc7dd2014c

Observation fec52c2c-783d-44ee-98a5-99ef68d68da4 · outbound

This paper cites Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.113235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.113235Z digest=sha256:4c2eb7266191f9e347e458f77b128c1362c7a1923666adb64832bc770f3285ec

Observation eb5943d7-59fd-4b9d-b3c0-40c2b3307e46 · outbound

This paper cites Drew A Hudson and Christopher D Manning.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Drew A Hudson and Christopher D Manning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.069949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.069949Z digest=sha256:552b345c2b78f310e9788b4b68a11582ffe64113f8a44ffe0cff8974d85fdd51

Observation 19dfcfd5-bdc8-44fd-a210-9d3f803fef97 · outbound

This paper cites Information Flow Routes: Automatically Interpreting Language Models at Scale.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation Information Flow Routes: Automatically Interpreting Language Models at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:35.922218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:35.922218Z digest=sha256:c734019ff7ef64d52edb6403dcb3000e9d5e99e88e74830a9f26f3d79f94eb7b

Observation dc905e3d-6cc1-4f57-9794-b725bce5e1e8 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:35.967483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:35.967483Z digest=sha256:c8f63259ca6cd5e76dbc30e0a6bdb086f08227df8fa21b37ecad823b7e3968ce

Observation 3183994d-7ed0-4758-b10b-f3d17259ae19 · outbound

This paper cites PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:35.798130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:35.798130Z digest=sha256:297bac3a3935947b31427823af6ad305c12419e6e8d97e55c38d219878f0fbed

Pith citing papers

No inbound Pith citation observations are available.