Pith. sign in

Paper Citation Record · LEDGER

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models

As of 16 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2505.19498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19498 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:41.573398Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc270e6b-1411-400f-89ab-9b53f6f0d004 · outbound

This paper cites GPT-4 Technical Report.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.245378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.245378Z digest=sha256:0be84215158447268f86921e41748c6e051de2ba678d896c0ad0db6084722418

Observation e16b32d5-3e7b-4711-95fc-a16660d8f7ff · outbound

This paper cites Qwen Technical Report.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.382282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.382282Z digest=sha256:c577706091034f0050996385ff8deb833066b8db83451f225ccc0e8762cc79a4

Observation 079fbe18-00ed-4260-9f94-42331fa2bf94 · outbound

This paper cites Qwen2.5-VL Technical Report.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.488464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.488464Z digest=sha256:d6a00afd2a53511481e13f1dc2c593d62138473b819396ad676c32acbc6bb1c5

Observation a0f77696-68c1-46ad-8bf6-664d8bf419c4 · outbound

This paper cites Grounding everything: Emerging localiza- tion properties in vision-language transformers.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Grounding everything: Emerging localiza- tion properties in vision-language transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.882005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:34.559262Z digest=sha256:5407c99e5dec0f9485d7fa5ca1ac702be7216a4e226a6836c5731b4071227332

Observation 0a1960cd-3116-4762-ba08-3a0cd9bfd393 · outbound

This paper cites Evaluating a large language model on searching for gui layouts.Proceedings of the ACM on Human-Computer Interaction, 7(EICS):1–37, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Evaluating a large language model on searching for gui layouts.Proceedings of the ACM on Human-Computer Interaction, 7(EICS):1–37, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.616784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:34.663382Z digest=sha256:4a8c00ced4a08513ca8092becc6e553136cf47dcd0bc12900b4f8d44109c0139

Observation 46070d0c-e37e-4745-88da-737228c6686a · outbound

This paper cites Lan- guage models are few-shot learners.Advances in neural in- formation processing systems, 33:1877–1901, 2020.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Lan- guage models are few-shot learners.Advances in neural in- formation processing systems, 33:1877–1901, 2020

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.757066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.757066Z digest=sha256:86f83c2acc7f061afd00a4577d69228418f9ee79087951e0948d3feccafd3a94

Observation 5b69f746-4314-4683-be6f-53aba037e581 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.892057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.892057Z digest=sha256:dbedf8b9930f4ee48d8c0fa04dc1a8f098b3ad3c1534a04cd9dd9f20b06c0f3b

Observation 961ad80d-293a-44e4-8a39-0d432ccb3e2c · outbound

This paper cites In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.034387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.034387Z digest=sha256:407f5704a3fb899f436850aeabd2747af7effcb0deb287972d02dbd9501496f6

Observation e49fb8e4-fbaa-47e5-8d17-208c0bae6518 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1–113, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1–113, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.325338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:35.106132Z digest=sha256:6277cbfef62b73a18429390648e9d2626566036b12748d7be1765bfb5b286cd2

Observation 78c35447-c8c8-4279-84f3-d10040e6ade7 · outbound

This paper cites DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.174567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.174567Z digest=sha256:84321b2e6a0709a2177e8e50cdc68939f5dcebbb5d83e7e21dbbbf2bc0e28f2e

Observation 13a97e78-3cda-4907-89c5-726b66c2a0a2 · outbound

This paper cites A survey on multimodal large lan- guage models for autonomous driving.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A survey on multimodal large lan- guage models for autonomous driving

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.069525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:35.275408Z digest=sha256:be98ebc865483b94a648c1ff81b29f2cd06ee3943140137ad60afbba959a7f8e

Observation 32cefc74-1e4c-497f-be3f-85dde206a9b1 · outbound

This paper cites SADL: An Effective In-Context Learning Method for Compositional Visual QA.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models SADL: An Effective In-Context Learning Method for Compositional Visual QA

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:18:42.326741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:35.400317Z digest=sha256:f7eca8593b771b3ca9d467bfba330b49e8964ab69a7ea300aa3b77bb5a5b28a7

Observation 9926fc8f-9672-4a0a-83c8-5015e96e5818 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.528263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.528263Z digest=sha256:746b2e9c309d82cadcaa13ee8eea7ddb4bf6a45fe1ce69275f05d7451f96e6d4

Observation f114d59b-79f9-468e-b88c-805afc2873fe · outbound

This paper cites Chatgpt outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Chatgpt outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:45.809696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:35.635936Z digest=sha256:6624de5026a82b5045e8aeb0793c5558b28bef708ef4a5e65eeee63694669558

Observation b21a7905-088d-4403-bd70-a8c730e38936 · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Detecting and preventing hallucinations in large vision language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:45.537969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:35.765639Z digest=sha256:595f1e4a1393e61d9fcde01796ee01ca92227a201d98bd6b6dd49ee9e58774a6

Observation ff67bef2-f80c-4967-85dc-e5019a4c77a5 · outbound

This paper cites A comprehensive survey of deep learn- ing for image captioning.ACM Computing Surveys (CsUR), 51(6):1–36, 2019.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A comprehensive survey of deep learn- ing for image captioning.ACM Computing Surveys (CsUR), 51(6):1–36, 2019

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:45.264134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:35.872329Z digest=sha256:c540037083de17a7711bd0952b0ea9607defe3b044ad9fcee9d32d84bb9b574c

Observation 9881e557-33df-4395-a2b0-ffb5bfd6efa9 · outbound

This paper cites Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.982612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.982612Z digest=sha256:0c75b5ab72be201d1e98b2cba33ac6732220f6b4ed741f41647647d8a0189930

Observation c627c9d7-cc1b-4f45-8b2d-cd7c1ed60648 · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.945636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:36.149903Z digest=sha256:ce4d30a7e0576bba77d159f8e780ec37ddda4ff53f5692d11cc90f22cd6c143e

Observation 0f48cfdb-27ee-4522-997d-c56fd9557b5c · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.743928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:36.382436Z digest=sha256:97c382f304a3066ff3536c188dc9de1024b83857c9d2a16a795cb43ab5456ae3

Observation de4fc015-8fd4-401e-be1b-5ab70560631c · outbound

This paper cites Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:36.503540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:36.503540Z digest=sha256:1cc4b9ae2fa5d6a82be3a16ea8ad4cccab99a435f7abe702c524e6e9e3817827

Observation 7b122b0f-7388-4b0f-8273-825b5c7e9fbe · outbound

This paper cites Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:36.671574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:36.671574Z digest=sha256:4ce92b908acd7b960149d3449ae81a9f87a6282d4b6c8cd8975dca560a5ee2ca

Observation ac0ec77b-7f88-42fe-ba60-129f0d97725d · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.477176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:36.847427Z digest=sha256:8d89131868c4dcc6025f1b9c6f1fa94e1db0aa2f3f2a63c308f3f583640dbac7

Observation 8fb4c05c-aad5-4269-bf03-3a6bb44e59eb · outbound

This paper cites Contrastive Decoding: Open-ended Text Generation as Optimization.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Contrastive Decoding: Open-ended Text Generation as Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.002385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.002385Z digest=sha256:cf400b632a00ad49fb73a17a97bc3526eb0647934f37e272d8e080743e16ad78

Observation 6056d74f-aeb4-4c4b-9880-abbccfb11d05 · outbound

This paper cites Clip surgery for better explainability with enhancement in open- vocabulary tasks.arXiv e-prints, pages arXiv–2304, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Clip surgery for better explainability with enhancement in open- vocabulary tasks.arXiv e-prints, pages arXiv–2304, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.249543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:37.112543Z digest=sha256:762122924a3c16484466a12f85b6585452fb4dedc01318f4febf14c91f2a3a79

Observation 698de6d0-4453-4026-8f79-5ae428ea9567 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.259357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.259357Z digest=sha256:cc55f291dedfb7d00134e83eb2240d60f25ee08e23fafb8f1c46b67e3e7da30f

Observation 4902dd3d-c455-40a0-9263-3443816fe3a8 · outbound

This paper cites A survey of multimodel large language models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A survey of multimodel large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.028181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:37.405063Z digest=sha256:7846b994d00262a721759a9392173fef8a891c91148efbcba2b65064b5935fb3

Observation 91c838c3-8ed9-4ecd-a6ac-01b8b5f84ab9 · outbound

This paper cites Microsoft coco: Common objects in context.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Microsoft coco: Common objects in context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.549073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.549073Z digest=sha256:acc9fab61a52553e10e3ec283744065dc689f419854f24234ee06179e96b298e

Observation d7781a78-38b7-4e9b-a9f5-f481e94ab301 · outbound

This paper cites Aligning large multi-modal model with robust instruction tuning.CoRR, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Aligning large multi-modal model with robust instruction tuning.CoRR, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.800320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:37.677362Z digest=sha256:fac356c7a65bfdcc11cff0b2ed81c6515eb6b5ec1ca2da0b271e6b78950bdf53

Observation 110de30e-dc64-4d62-9592-269f3d2449ba · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.836220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.836220Z digest=sha256:e07084be5f9e110f3b7ea15cd3a2e7a4a2c58a871610f2d61a05530b1968d92a

Observation 4e188057-66eb-47e9-bfa5-30fd4b3b69f2 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A Survey on Hallucination in Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.954172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.954172Z digest=sha256:7d7581c5a28cb39fe37d6e9ab6d15884607ffa02d87de445ba171116badaee19

Observation a711fd13-8a82-4043-b4ba-07d4253993ce · outbound

This paper cites Improved baselines with visual instruction tuning.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Improved baselines with visual instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.643823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:38.129342Z digest=sha256:dc349ba6e02fa20aa8cbd79b538425198f89fcf95f80c76c7d8cf6bbefd44424

Observation 80bc08a3-7f37-4130-8c97-cd9203b4f55b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.252873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.252873Z digest=sha256:60b825771bbc4fffeb34bc71b495837ad772bbdb19bcedbae47094dac712f944

Observation f94bb474-2d7d-4270-a942-aaf26ec30034 · outbound

This paper cites Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.394009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.394009Z digest=sha256:15d9348e052df10a8fe63eb9e7b684625aed8fa51085cc397f8bc6aea20fbd9d

Observation 56dd1117-b617-4ec6-8a8e-e51e528039df · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Learning transferable visual models from natural language supervi- sion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.505601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.505601Z digest=sha256:e3940bcffbfa1b6363d1fe573b548a314d8b84b9af6b31d81f8723fe7a7291da

Observation 3cc9d383-ae89-415a-9aba-f8c93672ea7f · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.667164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.667164Z digest=sha256:bd7dec376e953839d9aa70c17170c83a724b1ed41b622da9f09db36546d1216b

Observation d5f9ce9b-ebf9-45c7-a4af-4c68bbd6e8a2 · outbound

This paper cites Object Hallucination in Image Captioning.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Object Hallucination in Image Captioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.796832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.796832Z digest=sha256:fa8461cdee395324953bb30af5630575d9ef3ed4d2cd4c066a08e4bc3df66c6f

Observation 999ebe78-13d1-4a2c-a2d0-c68b26805c3d · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.478030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:38.915835Z digest=sha256:6d517ad4b5e4ffe6936906c622853bb6d4056a28966e38862cc09bc188c7555c

Observation d01c58a5-2696-445d-ba40-1613dfdfca06 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.047457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.047457Z digest=sha256:9e84be033ac75e3f8d56dde3c5d1c9b7d7e688578cf7428dfc11f3aa8f0366cf

Observation ea9e165d-095d-46dc-a28b-3a0875f96296 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Stanford alpaca: An instruction-following llama model, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.196522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:39.198903Z digest=sha256:7a98a23a42f4ec7428f2c73c268c4ed97754216c748e085d8271711908bbb22b

Observation f82bb665-485a-4e88-94fa-ec945cfab64c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.327647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.327647Z digest=sha256:7203860e6f14f5a560dba793ed7a8f174c9a4f61189a4cd9fd0401715bcdd3cb

Observation d284adb0-9e97-4de5-903e-d4142296ef5c · outbound

This paper cites MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.471606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.471606Z digest=sha256:70c7d8eba209a76ac58e35154c146fbdcf78d48fcebf51daf18d9b169e17b244

Observation 23b19e25-c978-481b-a472-cd78fdf6645e · outbound

This paper cites Sclip: Rethink- ing self-attention for dense vision-language inference.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Sclip: Rethink- ing self-attention for dense vision-language inference

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:42.910544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:39.629227Z digest=sha256:91bf4711c5e7365e159e63ca9f36b524eefe2eb61bd5fa9de5893e0841ae1027

Observation 8d8cab3d-7066-47c8-a081-242256095dca · outbound

This paper cites Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.809780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.809780Z digest=sha256:b4e3bc8fc89ec94d7dea7c8ff490a5621e9bd3410d1a2aac74f4c6e284528213

Observation c5d2aac4-a924-4f4d-be0f-d31272f4e1a7 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.958342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.958342Z digest=sha256:d263846a8e0ca9e26bcf7e93fae9f98804421ac61a56dfc234ea49551983398d

Observation 235c3cd4-3505-467a-94fe-1cedd90abd9c · outbound

This paper cites ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.068458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.068458Z digest=sha256:6cf8c6fa420a6a58c26fa919becb802d41b647b5df8c6c7b01ffae5be89c3cb7

Observation 59e984b1-7277-4775-aa63-e510b2a2e889 · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.218185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.218185Z digest=sha256:9f69492f06509973345cf9f73360d9aaec64fd5e4004a579fe578e9b3d04195d

Observation 13453da0-ec46-4982-b78a-013dd73a6d8a · outbound

This paper cites Drivemlm: Aligning multi-modal large language models with behavioral planning states for au- tonomous driving.arXiv preprint arXiv:2312.09245, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Drivemlm: Aligning multi-modal large language models with behavioral planning states for au- tonomous driving.arXiv preprint arXiv:2312.09245, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.379779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.379779Z digest=sha256:4ad0575e636bff29718625914583e96c7dbd1a34621796b6f06e6b1ec60a77ee

Observation 9a54c286-d0e4-4615-9f80-2810d89316d9 · outbound

This paper cites Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.494427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.494427Z digest=sha256:85cad03725c1aef0bb847631f0907b2b88b5121030e651329709b21ce310350c

Observation ad36b43e-9f5e-4e51-8c2c-75c1979abe93 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.650777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.650777Z digest=sha256:290ede4d527f5b08d7c10f6444af5dac92d9290c09c171c1f9e096d6d3927b82

Observation 17801c05-cec1-42a6-aacc-867ccdf304dd · outbound

This paper cites Woodpecker: Hallucination correction for multimodal large language models.Science China Information Sciences, 67(12):220105, 2024.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Woodpecker: Hallucination correction for multimodal large language models.Science China Information Sciences, 67(12):220105, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:42.622977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:18:40.775096Z digest=sha256:0513d9d955338ccb0eed63b2276b7c803f354e8510bf94388d3753a39b297dc7

Observation 92015d10-50dc-43c7-a817-b1ebbaa634cb · outbound

This paper cites Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.942389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.942389Z digest=sha256:aadcd425c7b9854b03a17745f83061d0b762cf29367c857d8403dac9e1fa1dcd

Observation 0e1a60c6-4179-4a62-8d22-b860337d70a4 · outbound

This paper cites HuatuoGPT, towards Taming Language Model to Be a Doctor.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models HuatuoGPT, towards Taming Language Model to Be a Doctor

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.054437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.054437Z digest=sha256:fad918a0cf05131dc0a8038a357b96ed93f0b9f9d48a0348e1639b7e4e8d8dc6

Observation d29c0b0d-3e93-47a7-9713-1f3b465572b9 · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.235492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.235492Z digest=sha256:ea8ec0ad78e66bf7ec52f93bc243c235c872130dcb9cdb733845f62b78afaeba

Observation 60c69dda-7610-4e4f-99bf-92abdeafcb71 · outbound

This paper cites Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.382368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.382368Z digest=sha256:4c70457965ae9c6bcaee8342e5136a34aa13ddea308f60dd632f282660581ff7

Observation 6c42e6af-2ab2-49b8-ae91-db9e666073d1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.573398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.573398Z digest=sha256:a580451602dd0aac3c7ff0d36d1cbe647f852437ad066c20e9845cc458728e28

Pith citing papers

No inbound Pith citation observations are available.