Pith. sign in

Paper Citation Record · LEDGER

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2505.19498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19498 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:41.573398Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc270e6b-1411-400f-89ab-9b53f6f0d004 · outbound

This paper cites GPT-4 Technical Report.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.245378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.245378Z digest=sha256:4516288355d956ec339dd8ad676ba38f3ecdf6ba0cdd826ef62bd90c21c87b1e

Observation e16b32d5-3e7b-4711-95fc-a16660d8f7ff · outbound

This paper cites Qwen Technical Report.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.382282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.382282Z digest=sha256:a03d32556200b6bf60770957fd3061cdf1f7f06e999c9735fe28bac3270a6718

Observation 079fbe18-00ed-4260-9f94-42331fa2bf94 · outbound

This paper cites Qwen2.5-VL Technical Report.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.488464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.488464Z digest=sha256:a3b7e1b65537161dc762fd69667a410a5800c24c2e3fe2bfa437f1c9d4e5fe60

Observation a0f77696-68c1-46ad-8bf6-664d8bf419c4 · outbound

This paper cites Grounding everything: Emerging localiza- tion properties in vision-language transformers.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Grounding everything: Emerging localiza- tion properties in vision-language transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.882005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:34.559262Z digest=sha256:2194d7994a5f3927c82ffb9237769e2b04aa479a60b5d787ac8b73db42f15965

Observation 0a1960cd-3116-4762-ba08-3a0cd9bfd393 · outbound

This paper cites Evaluating a large language model on searching for gui layouts.Proceedings of the ACM on Human-Computer Interaction, 7(EICS):1–37, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Evaluating a large language model on searching for gui layouts.Proceedings of the ACM on Human-Computer Interaction, 7(EICS):1–37, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.616784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:34.663382Z digest=sha256:a7e12eb10ce26845af08c8a1cde01f83e4b8d78268581e67c3098c8739ccbf89

Observation 46070d0c-e37e-4745-88da-737228c6686a · outbound

This paper cites Lan- guage models are few-shot learners.Advances in neural in- formation processing systems, 33:1877–1901, 2020.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Lan- guage models are few-shot learners.Advances in neural in- formation processing systems, 33:1877–1901, 2020

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.757066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.757066Z digest=sha256:670c687279c61be158b50da3dc6eb3e3018ae5c163345804ac1e015ce265a423

Observation 5b69f746-4314-4683-be6f-53aba037e581 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.892057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.892057Z digest=sha256:06a6fe840a50d34b1aa38db815c59f19ad047e39d65b62ccb4a10bf8e344c734

Observation 961ad80d-293a-44e4-8a39-0d432ccb3e2c · outbound

This paper cites In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.034387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.034387Z digest=sha256:2367b6254c9b33b71dc7e9a6bcb5ebc55ab1a88611da53edc49e2ee78de9d029

Observation e49fb8e4-fbaa-47e5-8d17-208c0bae6518 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1–113, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1–113, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.325338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.106132Z digest=sha256:95c95d71909690a0435a015c9c351f8afbfbc9d4d65562fcb771fe74324b18cb

Observation 78c35447-c8c8-4279-84f3-d10040e6ade7 · outbound

This paper cites DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.174567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.174567Z digest=sha256:8a6cf05510dc86e8e8d5ada9fcb5d86295d393627a6529e7922f8c02e64b0404

Observation 13a97e78-3cda-4907-89c5-726b66c2a0a2 · outbound

This paper cites A survey on multimodal large lan- guage models for autonomous driving.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A survey on multimodal large lan- guage models for autonomous driving

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.069525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.275408Z digest=sha256:0cfa367c367efd091f85257cec7e8c64637ab01e399e16d6d2bae52c6700a9cb

Observation 32cefc74-1e4c-497f-be3f-85dde206a9b1 · outbound

This paper cites SADL: An Effective In-Context Learning Method for Compositional Visual QA.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models SADL: An Effective In-Context Learning Method for Compositional Visual QA

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:18:42.326741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.400317Z digest=sha256:e88bc05a2ae1c553e3e58cf3cc752063601adc4ac3a0ec0debcd8d108165a1db

Observation 9926fc8f-9672-4a0a-83c8-5015e96e5818 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.528263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.528263Z digest=sha256:852514f2f528be96d014ce0fc9ec4312004158824eceeefaaf957f98a3effb5b

Observation f114d59b-79f9-468e-b88c-805afc2873fe · outbound

This paper cites Chatgpt outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Chatgpt outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:45.809696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.635936Z digest=sha256:07c5feff4f2d709a14e450a521872e1a8ba614fbca19eccc10cb321b2b45ccb0

Observation b21a7905-088d-4403-bd70-a8c730e38936 · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Detecting and preventing hallucinations in large vision language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:45.537969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.765639Z digest=sha256:388a4d58fda0c57358a50a157b64bfa4542f9c3c6c2268cb88257cd5fe7979fb

Observation ff67bef2-f80c-4967-85dc-e5019a4c77a5 · outbound

This paper cites A comprehensive survey of deep learn- ing for image captioning.ACM Computing Surveys (CsUR), 51(6):1–36, 2019.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A comprehensive survey of deep learn- ing for image captioning.ACM Computing Surveys (CsUR), 51(6):1–36, 2019

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:45.264134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.872329Z digest=sha256:2ce46ec01cabbe4c04a918319021b2ddc78fadf73ae398900fad44f6c8f78d38

Observation 9881e557-33df-4395-a2b0-ffb5bfd6efa9 · outbound

This paper cites Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.982612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.982612Z digest=sha256:07e4870badecfafeba0ee82a96f9ebe65bd8319b8320fd49dc82612483d5b8b1

Observation c627c9d7-cc1b-4f45-8b2d-cd7c1ed60648 · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.945636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:36.149903Z digest=sha256:50f5f013c155fa1b6a97390cbdd140948e60e82a82e85779a31ca6a4ccabbff5

Observation 0f48cfdb-27ee-4522-997d-c56fd9557b5c · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.743928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:36.382436Z digest=sha256:9d956679b28f76f1843b4eeca7d9d64fd8388c010457eed4f58b823a0344049b

Observation de4fc015-8fd4-401e-be1b-5ab70560631c · outbound

This paper cites Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:36.503540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:36.503540Z digest=sha256:3214a10dd26dda69d3bd27bf43b512fbb8aaf048eead79fd23a95bb689a8d84c

Observation 7b122b0f-7388-4b0f-8273-825b5c7e9fbe · outbound

This paper cites Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:36.671574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:36.671574Z digest=sha256:2629e172578b234a919a236e87608d819a1fa4dd6478269cb9a59862d9cfa66f

Observation ac0ec77b-7f88-42fe-ba60-129f0d97725d · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.477176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:36.847427Z digest=sha256:d18bf520a3dad8183e40b1e65b6fa23e6a1124830d25cd58c01f798ed3d50fd1

Observation 8fb4c05c-aad5-4269-bf03-3a6bb44e59eb · outbound

This paper cites Contrastive Decoding: Open-ended Text Generation as Optimization.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Contrastive Decoding: Open-ended Text Generation as Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.002385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.002385Z digest=sha256:c935ffceb3ea08ee8d89aaece463337cec7477e3e4836ee07b45db61704238f8

Observation 6056d74f-aeb4-4c4b-9880-abbccfb11d05 · outbound

This paper cites Clip surgery for better explainability with enhancement in open- vocabulary tasks.arXiv e-prints, pages arXiv–2304, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Clip surgery for better explainability with enhancement in open- vocabulary tasks.arXiv e-prints, pages arXiv–2304, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.249543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:37.112543Z digest=sha256:78c59b1d6c05bbb9bfbd44259d24f12b6f20a6cebe8547e152ee110e8f573865

Observation 698de6d0-4453-4026-8f79-5ae428ea9567 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.259357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.259357Z digest=sha256:dd7976b2955f2a099de9fcfe539663a05663ebef823b5906f9bc6cb209f3f0b4

Observation 4902dd3d-c455-40a0-9263-3443816fe3a8 · outbound

This paper cites A survey of multimodel large language models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A survey of multimodel large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.028181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:37.405063Z digest=sha256:2bb8dedca2cec759f20f72515bb0a288cb0cf795dd0867eda642ac7d4f54c7f3

Observation 91c838c3-8ed9-4ecd-a6ac-01b8b5f84ab9 · outbound

This paper cites Microsoft coco: Common objects in context.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Microsoft coco: Common objects in context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.549073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.549073Z digest=sha256:735dd79b292a1327699afee7c9919b0d7e6b39f57f131520be69333965de8460

Observation d7781a78-38b7-4e9b-a9f5-f481e94ab301 · outbound

This paper cites Aligning large multi-modal model with robust instruction tuning.CoRR, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Aligning large multi-modal model with robust instruction tuning.CoRR, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.800320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:37.677362Z digest=sha256:524722140c2f9ba8c7d606d9c5b875a246f6fa252868e79dd988eee64a10f788

Observation 110de30e-dc64-4d62-9592-269f3d2449ba · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.836220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.836220Z digest=sha256:a9a91eb86cf20b3a2aabc38f328e1a19d4c9ca4b9e04b684a8dea78b0e07fd80

Observation 4e188057-66eb-47e9-bfa5-30fd4b3b69f2 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A Survey on Hallucination in Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.954172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.954172Z digest=sha256:8c78086d845cd38e30a45a7e8df1f632585acca689c4b354c6b6ec3cc44255d9

Observation a711fd13-8a82-4043-b4ba-07d4253993ce · outbound

This paper cites Improved baselines with visual instruction tuning.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Improved baselines with visual instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.643823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:38.129342Z digest=sha256:7279ad375bce8ad0f12cfc59a718c3c3e3184e186354a83ab5646f804754c893

Observation 80bc08a3-7f37-4130-8c97-cd9203b4f55b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.252873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.252873Z digest=sha256:8e874b089b88c6a7d2d90a21ac426672c216c18f802dc59b9cc8ce7fea40bc90

Observation f94bb474-2d7d-4270-a942-aaf26ec30034 · outbound

This paper cites Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.394009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.394009Z digest=sha256:26d5d073aae58530c864b057522abd6a3713ef57dd90f02c62b3299fed937678

Observation 56dd1117-b617-4ec6-8a8e-e51e528039df · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Learning transferable visual models from natural language supervi- sion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.505601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.505601Z digest=sha256:2c81a26da3c0d5cc003fed61ce9008f9bbcfbf0f3015f9179c90815d64441f9d

Observation 3cc9d383-ae89-415a-9aba-f8c93672ea7f · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.667164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.667164Z digest=sha256:7f8a0551432280f4cd5d63111f70db8af570daeafce484bc8786c758140c7e2b

Observation d5f9ce9b-ebf9-45c7-a4af-4c68bbd6e8a2 · outbound

This paper cites Object Hallucination in Image Captioning.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Object Hallucination in Image Captioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.796832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.796832Z digest=sha256:3ee4ab6fda7a26d3f40c700673739dabe399d23c3d3e19ba4f2bb271c69e95ec

Observation 999ebe78-13d1-4a2c-a2d0-c68b26805c3d · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.478030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:38.915835Z digest=sha256:b9f4dc1a8ff758c5cd0decff861818f08234ae7e8304ed6fb3cf9f44b5a729aa

Observation d01c58a5-2696-445d-ba40-1613dfdfca06 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.047457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.047457Z digest=sha256:f345d2eb5dc1f66fe9a1e4fcc6dddb92521c233cb9259084eea3dca4b7d18e95

Observation ea9e165d-095d-46dc-a28b-3a0875f96296 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Stanford alpaca: An instruction-following llama model, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.196522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:39.198903Z digest=sha256:cfb6c1b21a8f7be91756784844ca6e9d787c09287ae11f774a7c6bb095f0795a

Observation f82bb665-485a-4e88-94fa-ec945cfab64c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.327647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.327647Z digest=sha256:b641dd8da698227dd8dddc4d34d5eb6bdfd92f6e1725ed442b249cf830a04815

Observation d284adb0-9e97-4de5-903e-d4142296ef5c · outbound

This paper cites MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.471606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.471606Z digest=sha256:21e6baa388c5b1f7bf253444cc4d585af283d920245f1a3be3b7aeee877094f3

Observation 23b19e25-c978-481b-a472-cd78fdf6645e · outbound

This paper cites Sclip: Rethink- ing self-attention for dense vision-language inference.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Sclip: Rethink- ing self-attention for dense vision-language inference

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:42.910544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:39.629227Z digest=sha256:e28838004abd4f3f96245c01ff2ac0b7af534c00bb145ab5854d9e4e006ec9f1

Observation 8d8cab3d-7066-47c8-a081-242256095dca · outbound

This paper cites Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.809780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.809780Z digest=sha256:24721470a3b25bf9bbee85d7d2f0710381f0c1abba2d96e28186028fa86b0f6c

Observation c5d2aac4-a924-4f4d-be0f-d31272f4e1a7 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.958342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.958342Z digest=sha256:f2efae6b831caaef712e92c155ed1f56486fc8baf16f6f1bf25206726763bda4

Observation 235c3cd4-3505-467a-94fe-1cedd90abd9c · outbound

This paper cites ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.068458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.068458Z digest=sha256:d3b29ebd8f5a0ab39046b20a29f4e8fe3afb57d71b9b5c720a2d00c6f4c9eaa0

Observation 59e984b1-7277-4775-aa63-e510b2a2e889 · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.218185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.218185Z digest=sha256:abac6580f62372b37b04050653f90b04bc4cf18e2cd5136fc298ed84fdd72845

Observation 13453da0-ec46-4982-b78a-013dd73a6d8a · outbound

This paper cites Drivemlm: Aligning multi-modal large language models with behavioral planning states for au- tonomous driving.arXiv preprint arXiv:2312.09245, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Drivemlm: Aligning multi-modal large language models with behavioral planning states for au- tonomous driving.arXiv preprint arXiv:2312.09245, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.379779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.379779Z digest=sha256:e588e2fee2dfb27e9bcb8abadfa630232a057cf55495fc893bda42b1393f5476

Observation 9a54c286-d0e4-4615-9f80-2810d89316d9 · outbound

This paper cites Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.494427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.494427Z digest=sha256:943b889b6443705571854d3106aee11d9249321636d458bd58c8a1ec1b0452ad

Observation ad36b43e-9f5e-4e51-8c2c-75c1979abe93 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.650777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.650777Z digest=sha256:d7f44aa9382981f14d504265e257d23466dd9984de9687d282bf45ece59cd0e6

Observation 17801c05-cec1-42a6-aacc-867ccdf304dd · outbound

This paper cites Woodpecker: Hallucination correction for multimodal large language models.Science China Information Sciences, 67(12):220105, 2024.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Woodpecker: Hallucination correction for multimodal large language models.Science China Information Sciences, 67(12):220105, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:42.622977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:40.775096Z digest=sha256:4936b3deb964db7ef5d88b91bfae370f5cbffabc18a96c1a779b41b49f6b01f5

Observation 92015d10-50dc-43c7-a817-b1ebbaa634cb · outbound

This paper cites Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.942389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.942389Z digest=sha256:4b645aba4ce162ab63c7e64235e79ac6593abb0f9edc44af34ae172ee4b2e893

Observation 0e1a60c6-4179-4a62-8d22-b860337d70a4 · outbound

This paper cites HuatuoGPT, towards Taming Language Model to Be a Doctor.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models HuatuoGPT, towards Taming Language Model to Be a Doctor

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.054437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.054437Z digest=sha256:2cd6dfb7a28e5df0525a86d7de94db65a9f6ecc883c21f4c833252b0e780c84e

Observation d29c0b0d-3e93-47a7-9713-1f3b465572b9 · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.235492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.235492Z digest=sha256:aba4ae291f55925cf11a4049ad0a8663590835b72325d06eb8d61ee37bd60f30

Observation 60c69dda-7610-4e4f-99bf-92abdeafcb71 · outbound

This paper cites Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.382368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.382368Z digest=sha256:1ba7a25d8d380e7a14a2570b75e83daa8a8a5223de739a664d76b7d07de0fe58

Observation 6c42e6af-2ab2-49b8-ae91-db9e666073d1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.573398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.573398Z digest=sha256:ef956b4924e4c49f762711d5c863cd70c4065440576fa01711afa9e98df5767a

Pith citing papers

No inbound Pith citation observations are available.