Pith. sign in

Paper Citation Record · LEDGER

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations

As of 9 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2505.17812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17812 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:45:32.061891Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 882aa6cd-9854-4dfc-8100-a9006e0ca4f8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:26.861270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:26.861270Z digest=sha256:2173ef044aba91ae51ae99f9c1ba23f0acad86084f2770444aa093d60aa1f724

Observation 737a10df-2d22-40cb-bfce-2668d47766e6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:26.929012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:26.929012Z digest=sha256:d03716a805498b139b001efbd6dac6cd46db27e907ab1d3facb7aa40566b38cb

Observation 81c64c90-6c7e-4622-bfbb-60dfd6c91d37 · outbound

This paper cites Visual instruction tuning.Adv.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Visual instruction tuning.Adv

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:39.249729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:27.001968Z digest=sha256:18fac89ab9857062d99a1d92d866c11999c65331eb72fba01d244f59824fd46d

Observation 4455e4e8-22fe-4da9-847f-c24901bf3ee4 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Improved Baselines with Visual Instruction Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:27.094656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:27.094656Z digest=sha256:d48ea2b77fad01baf0f6165ceffc9bab54357b8e91c90a920df027e477e1fc46

Observation c8a21e09-834d-476c-9bdd-064d5f45d427 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:27.170843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:27.170843Z digest=sha256:bbdb0040c43db803c39349812405a105e14f87c817cb576912ad7eca3bac7555

Observation 8f887f16-e771-46b9-8839-be6850db0e26 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:27.251904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:27.251904Z digest=sha256:b92db9cabbe55f2f7b3427fa6d7df3f91fb1a802a460771b1716c38f33e3cdf7

Observation 9e209bd6-89d7-4769-830f-56caefcb9c6f · outbound

This paper cites Qwen Technical Report.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:27.325359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:27.325359Z digest=sha256:a33a4427c677d653a45d64be9224f0569ba3a3a7104272a734b89e26810f7642

Observation cfe75358-ff73-4c7d-8cbe-b2cfef5b7b6e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:27.396913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:27.396913Z digest=sha256:006a7766a5d1a7235481c34f34277db5ecbdb89f3ab72fb0de0843614632dcd8

Observation f1d4b41e-4c28-4897-82a9-ea75f1ccdfe5 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Hallucination of Multimodal Large Language Models: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:27.497184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:27.497184Z digest=sha256:da9fb614ae63e91b44a63985bdb00df2c2f1f527815e7e24559f70e3cd4cc945

Observation 66e2ef9a-7aaf-41da-9155-14405f3ba299 · outbound

This paper cites Nullu: Mitigating object hallucinations in large vision-language models via halluspace projection.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Nullu: Mitigating object hallucinations in large vision-language models via halluspace projection

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:39.061415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:27.582588Z digest=sha256:b512567a6c784074162e98e19099696d8cdba05b2483fa2ad9e60d557ec37bf0

Observation f12fe377-086e-4152-a5f1-8a0c152760c5 · outbound

This paper cites Truthprint: Mitigating lvlm object hallucination via latent truthful-guided pre-intervention.arXiv preprint arXiv:2503.10602, 2025.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Truthprint: Mitigating lvlm object hallucination via latent truthful-guided pre-intervention.arXiv preprint arXiv:2503.10602, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:27.675017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:27.675017Z digest=sha256:47e673d4b7319fb1ccb3952f1c2a1b7f6d4e8ec439dac85cddfa9afadbfa011d

Observation 0473397b-fec5-45da-b3c0-13bccfee3515 · outbound

This paper cites Analyzing and mitigating object hallucination in large vision-language models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Analyzing and mitigating object hallucination in large vision-language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:38.889223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:27.767172Z digest=sha256:8d4b6f4eeeaeb3989f77dd1364b745277016a13ff96e7da07a6a65ebde2877bf

Observation 7bd6d8df-457b-42e7-abf7-ea2e01419e83 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:27.838292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:27.838292Z digest=sha256:7b5ff24ef81c33a45a2b6776c976bc14e59ef780e13ab6325e4f6131e9ba12d4

Observation c1a8a98d-b379-472b-9222-2525fcf3a5c7 · outbound

This paper cites Hallucination augmented contrastive learning for multimodal large language model.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Hallucination augmented contrastive learning for multimodal large language model

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:38.723563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:27.945183Z digest=sha256:620f998ec9a67d7da9cbe7f771e6af661f6234752b771ec80e0fa09be3fa947d

Observation 4c94a0d3-e0d6-4f2f-b22b-47cf3c77c1a1 · outbound

This paper cites Exposing and mitigating spurious correlations for cross-modal retrieval.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Exposing and mitigating spurious correlations for cross-modal retrieval

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:38.528588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:28.024239Z digest=sha256:9a74d4d91875582703ea4c03e102700ef5328bc178ce9b98b69c628d19310a33

Observation 5671c12d-a4aa-4613-a229-587492166ecc · outbound

This paper cites Mitigating object hallucinations in large vision-language models through visual contrastive decoding.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Mitigating object hallucinations in large vision-language models through visual contrastive decoding

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:38.339332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:28.107255Z digest=sha256:f9a2f20a2800118e28e149504806822dfd61d9f0ef869b9924c4aa132b4023bd

Observation 196a4003-abf4-40d3-8f5e-25063d5a3327 · outbound

This paper cites Debiasing large visual language models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Debiasing large visual language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:38.147222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:28.181188Z digest=sha256:9b29e2e7a3951ccb0e8a1713bd349a099ff2e845c522a44a0570fbca247a0e4e

Observation 65bf3d5c-5958-41c4-ae01-28e195126a20 · outbound

This paper cites Halc: Object hallucination reduction via adaptive focal-contrast decoding.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Halc: Object hallucination reduction via adaptive focal-contrast decoding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:37.936158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:28.261744Z digest=sha256:74d411b98a1a41fd4a5174132b0e704dbb1f748e2c4e1c84f7bfac2d65fd4e53

Observation ad860591-dd20-4f99-9944-9c9f1db318f6 · outbound

This paper cites ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:28.344201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:28.344201Z digest=sha256:bdda0337c527e6c2a0913ce7cc553bd2f8cc331e17f18e52bfffcd40403e80ec

Observation 25ab535f-ea73-4a22-8646-e069e69f639a · outbound

This paper cites Reducing hallucinations in large vision-language models via latent space steering.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Reducing hallucinations in large vision-language models via latent space steering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:37.709424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:28.404847Z digest=sha256:630fe7f9230e860ba1a9681361503eda8b5568ea9062e7b5a51cd992adcbf59e

Observation 0cfe0c03-bf00-4cff-95a9-3fae186b43fb · outbound

This paper cites Where do Large Vision-Language Models Look at when Answering Questions?.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Where do Large Vision-Language Models Look at when Answering Questions?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:28.465442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:28.465442Z digest=sha256:d944bfbffb0cfddecfc71d80a74fe176fa9d50975cbfad1cea475aa987924cb9

Observation 49f3903b-7a7a-4ab0-9e11-feaee20efb4a · outbound

This paper cites Lvlm-intrepret: An interpretability tool for large vision-language models, 2024.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Lvlm-intrepret: An interpretability tool for large vision-language models, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:37.516328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:28.517311Z digest=sha256:3df32030cc725adb023081af94faf213bbc72780dd487e13a009480334f6b497

Observation 26f19368-69be-4297-ad00-dc9f67a84f3a · outbound

This paper cites See What You Are Told: Visual Attention Sink in Large Multimodal Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:28.622401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:28.622401Z digest=sha256:88e376253ffb85f528ef697425e211f98b1a18b8b96ba98c9e07c226fb611102

Observation 3f9dcb33-210b-4d5c-bbad-5ea4f33e9034 · outbound

This paper cites Vision Transformers Need Registers.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Vision Transformers Need Registers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:28.704062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:28.704062Z digest=sha256:f1d002a3617ba4ae4590caf4884d7104689d9c7fa954628471c946a42252900f

Observation bb0f94bf-7c2a-4275-9518-11fbd28bcacf · outbound

This paper cites Massive Activations in Large Language Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Massive Activations in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:28.759191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:28.759191Z digest=sha256:bb6ffcf29838866d9089fea0166fcccdd794737dcdee8a0b53fc66d300a65816

Observation ac162876-881f-4b81-8f13-701f6b5eeb37 · outbound

This paper cites Object Hallucination in Image Captioning.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Object Hallucination in Image Captioning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:28.823154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:28.823154Z digest=sha256:1d202b54759dbcda417f789b4b742db45082e9e917779e3884dbb74a706d1e50

Observation d6fcddd8-1821-40b3-b54b-4f0175019e62 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:37.289552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:28.894823Z digest=sha256:17e31fd2c1802e3fac743ab601365271d88dda5fb366324c77ebf424d9565f5c

Observation de3e5d8d-fea9-4b4c-892f-d99344a2f540 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:28.975925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:28.975925Z digest=sha256:bd535f5b3e1a982528ac59c504f1b86932d7b2ebd0685b2b2949474e8d3ac01f

Observation 9ab16929-c20f-4955-b6bc-ff8ad009c7d1 · outbound

This paper cites Llava-phi: Efficient multi-modal assistant with small language model.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Llava-phi: Efficient multi-modal assistant with small language model

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:37.107790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:29.029305Z digest=sha256:942234d678bd36aea703d1aa2e02cd742ba2726439d0c474bdef6bf523be0510

Observation fc8a6a40-c744-42de-9ae0-c7cf38e76412 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:29.111032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:29.111032Z digest=sha256:2fa518cc94c33595645f55a97d18ba147bd6e608ada8cf3c95cedf96c0e5033a

Observation cdd349ec-b545-4bf8-9229-43c1bfdc52b9 · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Detecting and preventing hallucinations in large vision language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:36.895414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:29.202872Z digest=sha256:666b8bc7729ba0f666efb08957748cbde80351db0c0fb2449014e5a68b725bad

Observation 125f43ae-65de-4b25-be32-532fb592a6ec · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:29.287602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:29.287602Z digest=sha256:01c7dd7f4543e5f8262bbf8d2a0887e9dcaa6078a6914d4d1d3b0356a41b1fd3

Observation 85fa9ad1-866f-4ebb-b530-46f4fd57c7f6 · outbound

This paper cites Dress: Instructing large vision-language models to align and interact with humans via natural language feedback.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Dress: Instructing large vision-language models to align and interact with humans via natural language feedback

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:36.712636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:29.363088Z digest=sha256:8eed50df5fd8de58c9dd6bf2952af034b4ef255c7a91a6739f4ccd38bab928b3

Observation 7f9b0304-50f3-4f2e-b0ec-d8cabe019404 · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:29.439784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:29.439784Z digest=sha256:bb5654132c5b3097d2b2c2964e29fb12d42edca655331dea09473bd255e26a16

Observation a0f958ef-718f-482d-b0f9-39b5dd6e23b3 · outbound

This paper cites Mitigating object hallucination in large vision-language models via image-grounded guidance.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Mitigating object hallucination in large vision-language models via image-grounded guidance

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:36.500425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:29.524121Z digest=sha256:133104ac898e2537124929f04ca847e63a21758f78c77fe25bfd85aa4a9cdf46

Observation ff3e8d54-95f9-4d80-b404-ff620b82690d · outbound

This paper cites Paying more attention to image: A training-free method for alleviating hallucination in lvlms.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Paying more attention to image: A training-free method for alleviating hallucination in lvlms

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:36.337308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:29.593024Z digest=sha256:293c48263090132a8c06cbe7c2d4de43a657b3d30b5089c9536044b1fea73f19

Observation b6698032-96ff-40d5-a3aa-b68b7600a43a · outbound

This paper cites IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:29.672408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:29.672408Z digest=sha256:d47c8d718787b0c8ceac6be33736b84bc23458a5d2f33b6f05fff214c5f84d61

Observation a24c3d4f-df58-4cf0-bcd3-cb50f4e2da2d · outbound

This paper cites Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:36.102530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:29.758823Z digest=sha256:df0f8995caddd05d7dec04e53c639459c0b912b779f8fe10c16e342984a86057

Observation 77a176db-9161-4150-9360-879d127009a8 · outbound

This paper cites Multi-modal hallucination control by visual information grounding.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Multi-modal hallucination control by visual information grounding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:35.874270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:29.857326Z digest=sha256:a2dedf59ce22b7e4078f2a53e6b08383ae26645c507419fde9f2b2cb7cf8b9a1

Observation a06fcb51-2132-4696-a792-026aa2d63c38 · outbound

This paper cites AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:29.943737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:29.943737Z digest=sha256:3c9e2b2ce622758d9c0ece6b72ef2fe0bd8128f708783add7760c168423d8347

Observation b2eb7736-2804-407e-9555-9576ea099f9f · outbound

This paper cites Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:30.042461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:30.042461Z digest=sha256:14f2b288402cdef4151cc0c0341c93630944b376c469e45853b36d81bc2836c7

Observation b4fefc28-d1a6-4b60-ab67-38a9e25413b6 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:35.675866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:30.110923Z digest=sha256:a8f696273b2b23f9506c0348c8e2991303b94fd6bc6785c65d73054994f4e3dc

Observation 847e83b9-e414-4ccd-9632-e2aecb5eb7ab · outbound

This paper cites Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:35.484548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:30.219446Z digest=sha256:63418bf7a2d9bc71eda08d2a701fb9c893a64e9530d4985c6848944da11e7c0a

Observation 3284945c-49ca-4b6d-9391-ef63178ae38a · outbound

This paper cites Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:35.230180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:30.320579Z digest=sha256:2fc8fc21edfb3c3dc7c0cda0f623ee61b555a7ef7fc2296146dc1cc20d3c9965

Observation 8272d023-731e-4d73-b974-2645efdb56d7 · outbound

This paper cites Transformer interpretability beyond attention visualization.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Transformer interpretability beyond attention visualization

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:35.077456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:30.399480Z digest=sha256:6a274df6ffcb22a95f3ae816bcbbc5589fd90acb6378d66d344abb200cf0add0

Observation 82447b30-d45b-4a50-a20a-83d9ff7a2806 · outbound

This paper cites Vl- interpret: An interactive visualization tool for interpreting vision-language transformers.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Vl- interpret: An interactive visualization tool for interpreting vision-language transformers

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:34.838376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:30.463305Z digest=sha256:c38ab54891e700fd4ab754b2d675c579c5f9af0270521f13924956099555afa5

Observation 2c82cd78-bd42-4df1-8c17-3405ca1d6741 · outbound

This paper cites FastRM: An efficient and automatic explainability framework for multimodal generative models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations FastRM: An efficient and automatic explainability framework for multimodal generative models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:45:32.849467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:30.527268Z digest=sha256:117d46485355a1dfac73c5d78b54cf4fc0057f0bc0ee61e7243bc6ed7e446052

Observation 3519c456-6106-4f20-b0df-74c2d34808f7 · outbound

This paper cites Explaining Multi-modal Large Language Models by Analyzing their Vision Perception.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Explaining Multi-modal Large Language Models by Analyzing their Vision Perception

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:45:32.614357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:30.624312Z digest=sha256:413ff91429b4291a4d5dd96642689bb552099de1eae5be899f6cb344a243a1b3

Observation 7089ab6c-d0a4-4247-8caf-b136bf755174 · outbound

This paper cites From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:30.728861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:30.728861Z digest=sha256:767b7116ee24d3da225f7605fc9429fb7a2a5bbf03ef37d7d6cd84dc2f649f38

Observation bc37ab4e-748e-45aa-aa7f-6b302367a133 · outbound

This paper cites Finding and Editing Multi-Modal Neurons in Pre-Trained Transformers.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Finding and Editing Multi-Modal Neurons in Pre-Trained Transformers

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:45:32.303375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:30.773041Z digest=sha256:94842d4cbb0a7a4a22d7e1b63824faf6bdcecf82555bd004944b8ea8ba909cdc

Observation 6ccb7e24-eaa5-4926-accd-cfd30a528366 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:30.862604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:30.862604Z digest=sha256:72cea6b934842fa9c7a921f3827a20d21082c6db4d4410faaa5a97b6a6c755c1

Observation 1889192c-40e8-4d30-96c2-56425dc377f8 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Evaluating object hallucination in large vision-language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:34.654011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:30.949942Z digest=sha256:81e02980a2e4352ac572d554cb4f1a9ad6c018264f76379119475c99ddc45338

Observation 3e5446db-90a3-421c-a606-095f94019fde · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Aligning large multimodal models with factually augmented rlhf

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:34.471205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:31.024274Z digest=sha256:ffe90bc705b0a7e7e6f8a03de69b9da6150a12bfcaa99a96ffda1e0163024ae2

Observation 707ed996-fdbb-4892-94d2-a0dcd63b2e9c · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Adv.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Adv

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:34.299454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:31.070353Z digest=sha256:f16b1388cb2bd6377c541119da75536d193012525e86fdaa8133d89c68ff908d

Observation 6f82f695-ebfd-409f-8796-bc9795f604ae · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:31.150819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:31.150819Z digest=sha256:30ece62cfeef139fb421c692be37fa9ae55e645b01d113575da9421cf538407e

Observation 4fe76543-d92b-4d9d-af14-a3bec28a87d2 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:34.114986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:31.203608Z digest=sha256:67bc58d3b30dc9f164e56a2e5ea96dd0bc3139d40a8c25492cb4ad064f6f4aeb

Observation ebdf2516-b466-4eaa-ace5-f91f494677cb · outbound

This paper cites Beam Search Strategies for Neural Machine Translation.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Beam Search Strategies for Neural Machine Translation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:31.289807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:31.289807Z digest=sha256:6e658f26ef525d6a77f86654f57a25aca27667e4c9a5cab2a00404b8195cea48

Observation 4de1a212-ef38-42f1-8a30-cf3c92771f3c · outbound

This paper cites DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:31.400584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:31.400584Z digest=sha256:52723794af47f2a2203fb49f6592875838ca68009f0b88dfc45008f6691a0cfb

Observation f8bab591-d094-47bc-931f-2d936769cd25 · outbound

This paper cites Blended diffusion for text-driven editing of natural images.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Blended diffusion for text-driven editing of natural images

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:33.960642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:31.506177Z digest=sha256:0931413512d6963bcd33578913e5c8be5b15fa7f5f0859294dd3ad72198ceb9c

Observation b87fca78-67d1-42ec-9d67-d2cc6903bcab · outbound

This paper cites Unveiling typographic deceptions: Insights of the typographic vulnerability in large vision-language models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Unveiling typographic deceptions: Insights of the typographic vulnerability in large vision-language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:33.708879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:31.619307Z digest=sha256:a8aa9703649ec0f177c4ab5db8f329a6bc712e6967db852a41224b50448cda5a

Observation 2e495d30-7f69-4a63-92fe-e137a6360e93 · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:33.567854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:31.774717Z digest=sha256:e76267445c1cbc0fbcd751422e18f37e51ab67cc83dfec2eb4fe3a1356dd7b58

Observation f7264abe-3342-48d1-8953-f240882fb7ca · outbound

This paper cites Towards interpreting visual information processing in vision-language models.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Towards interpreting visual information processing in vision-language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:45:33.440316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:45:31.888895Z digest=sha256:e34f10511dde54c89576c59d04db2e1cdeebc177d7be369a2ffaa591abeadd3b

Observation bf173a22-c628-4ba6-8068-16487a728e86 · outbound

This paper cites GPT-4 Technical Report.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations GPT-4 Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:31.991900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:31.991900Z digest=sha256:f89f0e24aa89a659b323008cdc3bd2307a223533437d79b5a02f8cc85ae5a4ad

Observation e4e60157-5b8f-4138-bf26-7e5b1f2d07bd · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 65

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:45:32.061891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:32.061891Z digest=sha256:2d0a847b81ef7eb6cd9197c33e9843877c43b1893032f356373cbed7be904e07

Pith citing papers

No inbound Pith citation observations are available.