Pith. sign in

Paper Citation Record · LEDGER

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2505.19498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19498 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:41.573398Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc270e6b-1411-400f-89ab-9b53f6f0d004 · outbound

This paper cites GPT-4 Technical Report.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.245378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.245378Z digest=sha256:5a09c1fd89cd5204d7454cf5880a02e02b54279fbac73bec246d435775708480

Observation e16b32d5-3e7b-4711-95fc-a16660d8f7ff · outbound

This paper cites Qwen Technical Report.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.382282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.382282Z digest=sha256:980f1bbcff4e6b6658283233b6126bbd032474eb4febb7039d17fc90139d5242

Observation 079fbe18-00ed-4260-9f94-42331fa2bf94 · outbound

This paper cites Qwen2.5-VL Technical Report.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.488464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.488464Z digest=sha256:1a54e1a2e38ee80b93f4dc7506fd45d6db8445c118107048b82aa8a905711744

Observation a0f77696-68c1-46ad-8bf6-664d8bf419c4 · outbound

This paper cites Grounding everything: Emerging localiza- tion properties in vision-language transformers.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Grounding everything: Emerging localiza- tion properties in vision-language transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.882005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:34.559262Z digest=sha256:4396ce8a89af71ae142384114a4c2161f4d245c361bbd8380cc5bab04f53f17b

Observation 0a1960cd-3116-4762-ba08-3a0cd9bfd393 · outbound

This paper cites Evaluating a large language model on searching for gui layouts.Proceedings of the ACM on Human-Computer Interaction, 7(EICS):1–37, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Evaluating a large language model on searching for gui layouts.Proceedings of the ACM on Human-Computer Interaction, 7(EICS):1–37, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.616784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:34.663382Z digest=sha256:68c845f8e4739348835ec2943b11a9c47b64159a651519bb416e19fef9e6eda5

Observation 46070d0c-e37e-4745-88da-737228c6686a · outbound

This paper cites Lan- guage models are few-shot learners.Advances in neural in- formation processing systems, 33:1877–1901, 2020.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Lan- guage models are few-shot learners.Advances in neural in- formation processing systems, 33:1877–1901, 2020

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.757066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.757066Z digest=sha256:1b6081a1ef2c4c7d28493b05f72a99fecee273ff3272cfaa603c9907c8ea5186

Observation 5b69f746-4314-4683-be6f-53aba037e581 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:34.892057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:34.892057Z digest=sha256:6229f3f0029a7a4538945a4ca3b7357cacab56631bfab74de203ddbf68e7481b

Observation 961ad80d-293a-44e4-8a39-0d432ccb3e2c · outbound

This paper cites In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.034387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.034387Z digest=sha256:832b6cd49f8bb4f920b1f00e99f85ac2f988185ea4eaba251bebb5048db93ec1

Observation e49fb8e4-fbaa-47e5-8d17-208c0bae6518 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1–113, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1–113, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.325338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.106132Z digest=sha256:cd3b4d36ce7d148c32936f3006e12df17e881495c20df9d7c696384b8807877d

Observation 78c35447-c8c8-4279-84f3-d10040e6ade7 · outbound

This paper cites DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.174567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.174567Z digest=sha256:9e9e52804a9c148fe8528891bc6e21def37fe06e011eae08c4f8a976c2cfa96d

Observation 13a97e78-3cda-4907-89c5-726b66c2a0a2 · outbound

This paper cites A survey on multimodal large lan- guage models for autonomous driving.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A survey on multimodal large lan- guage models for autonomous driving

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:46.069525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.275408Z digest=sha256:29d855622c844f67be283df36ebc89212f5f437706f092a20e5e439539766320

Observation 32cefc74-1e4c-497f-be3f-85dde206a9b1 · outbound

This paper cites SADL: An Effective In-Context Learning Method for Compositional Visual QA.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models SADL: An Effective In-Context Learning Method for Compositional Visual QA

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:18:42.326741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.400317Z digest=sha256:9a71714a5126a78dd5cf3d3492e2bac8a383d26851440d4328b880762761910e

Observation 9926fc8f-9672-4a0a-83c8-5015e96e5818 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.528263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.528263Z digest=sha256:dbed65d6fc3a9cdaaf5cf32b16c30c2381ed331d7fdd4cad31f2ecc7b32b2fc1

Observation f114d59b-79f9-468e-b88c-805afc2873fe · outbound

This paper cites Chatgpt outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Chatgpt outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:45.809696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.635936Z digest=sha256:13d1b2106a313865bc3f8f14c2acdd8952a6464fa026865ed83e27f1983229f7

Observation b21a7905-088d-4403-bd70-a8c730e38936 · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Detecting and preventing hallucinations in large vision language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:45.537969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.765639Z digest=sha256:57c8d905f9adf83b2b0c14e3a1d89db28e7285b85914e94297eeb97c739d812d

Observation ff67bef2-f80c-4967-85dc-e5019a4c77a5 · outbound

This paper cites A comprehensive survey of deep learn- ing for image captioning.ACM Computing Surveys (CsUR), 51(6):1–36, 2019.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A comprehensive survey of deep learn- ing for image captioning.ACM Computing Surveys (CsUR), 51(6):1–36, 2019

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:45.264134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:35.872329Z digest=sha256:946f2ce0f0ee57b0dccbf340e5b50879f466e403c4724c41d1f0ae3fd3584264

Observation 9881e557-33df-4395-a2b0-ffb5bfd6efa9 · outbound

This paper cites Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:35.982612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:35.982612Z digest=sha256:0250b5f295f3085d967ae95434d2805ebba4856fb4e1233bbb247dceb0b89f19

Observation c627c9d7-cc1b-4f45-8b2d-cd7c1ed60648 · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.945636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:36.149903Z digest=sha256:4736de341c05a6033f06292e31d0098e5514cf312e702496708abb7da13e1adf

Observation 0f48cfdb-27ee-4522-997d-c56fd9557b5c · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.743928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:36.382436Z digest=sha256:ea2d68b313f3a87d5f448a51abf4577c644f4beee74c13b70ae9f549e9556be0

Observation de4fc015-8fd4-401e-be1b-5ab70560631c · outbound

This paper cites Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:36.503540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:36.503540Z digest=sha256:bb6557295a96e517ada754c526ffe1699642b9228f99d1e3f84e14123db5762c

Observation 7b122b0f-7388-4b0f-8273-825b5c7e9fbe · outbound

This paper cites Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:36.671574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:36.671574Z digest=sha256:b047431d6d4274bf6004bbc3a28f71d20e556be9b125fffc4143fc08f242585d

Observation ac0ec77b-7f88-42fe-ba60-129f0d97725d · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.477176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:36.847427Z digest=sha256:d6c4d2b7fb89262bd04f77514af91942ad52bfbc9f4f65c5ba55c6adcc9f5e45

Observation 8fb4c05c-aad5-4269-bf03-3a6bb44e59eb · outbound

This paper cites Contrastive Decoding: Open-ended Text Generation as Optimization.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Contrastive Decoding: Open-ended Text Generation as Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.002385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.002385Z digest=sha256:bbf5e524fbba09d881b7b27301c809635fb91957c59ec621860b5ca629e14902

Observation 6056d74f-aeb4-4c4b-9880-abbccfb11d05 · outbound

This paper cites Clip surgery for better explainability with enhancement in open- vocabulary tasks.arXiv e-prints, pages arXiv–2304, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Clip surgery for better explainability with enhancement in open- vocabulary tasks.arXiv e-prints, pages arXiv–2304, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.249543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:37.112543Z digest=sha256:2ec61886a3ae8e95c802b3331040eb4865e5f55bcf40dd0c2989758371e84f9b

Observation 698de6d0-4453-4026-8f79-5ae428ea9567 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.259357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.259357Z digest=sha256:e8a0baeedf4a7351a71b43f970dee0d6165a769dbc27270e57eda1d0c16fec42

Observation 4902dd3d-c455-40a0-9263-3443816fe3a8 · outbound

This paper cites A survey of multimodel large language models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A survey of multimodel large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:44.028181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:37.405063Z digest=sha256:7d7701ba2e19c2561a4c42f65d336266a793e9f3c3de93e5aee7fec60486c1e4

Observation 91c838c3-8ed9-4ecd-a6ac-01b8b5f84ab9 · outbound

This paper cites Microsoft coco: Common objects in context.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Microsoft coco: Common objects in context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.549073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.549073Z digest=sha256:d2318173b13cfc286b71134130446681a35abc11781f1ac19f2ef7c83c93d82a

Observation d7781a78-38b7-4e9b-a9f5-f481e94ab301 · outbound

This paper cites Aligning large multi-modal model with robust instruction tuning.CoRR, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Aligning large multi-modal model with robust instruction tuning.CoRR, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.800320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:37.677362Z digest=sha256:e60d6fff996a7162a6972627415d933c316b8b2204180d9a291d3b1e916e51ad

Observation 110de30e-dc64-4d62-9592-269f3d2449ba · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.836220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.836220Z digest=sha256:05c3d8c08ab1ecafe81ba44877c0de643457aa166b3607ea168023563b1dadcc

Observation 4e188057-66eb-47e9-bfa5-30fd4b3b69f2 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A Survey on Hallucination in Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:37.954172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:37.954172Z digest=sha256:358ce8ba08777255abee272286d66e9ec1e0c35ff875b4fd6a2fb8014139e20b

Observation a711fd13-8a82-4043-b4ba-07d4253993ce · outbound

This paper cites Improved baselines with visual instruction tuning.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Improved baselines with visual instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.643823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:38.129342Z digest=sha256:a4659dcfca7bcc9fbcb240417d18d670ac74636a6702751745bee6f5ba93d9c1

Observation 80bc08a3-7f37-4130-8c97-cd9203b4f55b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.252873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.252873Z digest=sha256:a0610b74506cd4f85b1921d314aeee79f6e069aae42837471301f1ef179ab26a

Observation f94bb474-2d7d-4270-a942-aaf26ec30034 · outbound

This paper cites Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.394009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.394009Z digest=sha256:bae1b026106c33023c97a16e7569bdecfaff1923d7346f49b8db658ce2af8328

Observation 56dd1117-b617-4ec6-8a8e-e51e528039df · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Learning transferable visual models from natural language supervi- sion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.505601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.505601Z digest=sha256:c444307735a245663edb34c12ae4ad47ca755552a9f4a05c628a7cbaaec53aab

Observation 3cc9d383-ae89-415a-9aba-f8c93672ea7f · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.667164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.667164Z digest=sha256:11eeae04691bd0835601a91abd39d4f995079d1379caa98a1fe01dbec69f5ab9

Observation d5f9ce9b-ebf9-45c7-a4af-4c68bbd6e8a2 · outbound

This paper cites Object Hallucination in Image Captioning.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Object Hallucination in Image Captioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:38.796832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:38.796832Z digest=sha256:e94dee0eac1c67f117bcea66edc7f2ca7c827d6abc2b0f707601561d7b0742b8

Observation 999ebe78-13d1-4a2c-a2d0-c68b26805c3d · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.478030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:38.915835Z digest=sha256:18ba361011c82d8e4c4abb2745bfb76e91dfa5705e8669deedb8746d3ef74e6a

Observation d01c58a5-2696-445d-ba40-1613dfdfca06 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.047457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.047457Z digest=sha256:069c3875a089139dbd90a3b36e69a6a74d9fb37167c0882df5903f3a6887a18d

Observation ea9e165d-095d-46dc-a28b-3a0875f96296 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Stanford alpaca: An instruction-following llama model, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:43.196522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:39.198903Z digest=sha256:24ccc9576041a1514775e0b161cd847f1ae6e9f019f848cdebfc6ea9f4151578

Observation f82bb665-485a-4e88-94fa-ec945cfab64c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.327647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.327647Z digest=sha256:8412b0974f670e1908dbf14f024a5a818548b6cf9b4f5e887978f7254ae09fdd

Observation d284adb0-9e97-4de5-903e-d4142296ef5c · outbound

This paper cites MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.471606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.471606Z digest=sha256:8338913239a46d8343ca01bed432ce8ad5866256c73d0078be6382b851043ebc

Observation 23b19e25-c978-481b-a472-cd78fdf6645e · outbound

This paper cites Sclip: Rethink- ing self-attention for dense vision-language inference.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Sclip: Rethink- ing self-attention for dense vision-language inference

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:42.910544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:39.629227Z digest=sha256:c199051cac0e5d4f7193b76bb3414734941fbca12fdaca2a6089138b76594876

Observation 8d8cab3d-7066-47c8-a081-242256095dca · outbound

This paper cites Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.809780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.809780Z digest=sha256:80b05d56591e89dd25a6dd27e4bd9f660d0933ef309caf563659e5aa9e2abe46

Observation c5d2aac4-a924-4f4d-be0f-d31272f4e1a7 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:39.958342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:39.958342Z digest=sha256:a3e14dc79ce047d54cc051de1b80ad7e5e635e75f44c74740b02c08b4db61e95

Observation 235c3cd4-3505-467a-94fe-1cedd90abd9c · outbound

This paper cites ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.068458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.068458Z digest=sha256:8da348daa14af47a22d3dc1cd82a95afb3b6625af618b9c363fc79fcefef9101

Observation 59e984b1-7277-4775-aa63-e510b2a2e889 · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.218185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.218185Z digest=sha256:24e8f4c7993026d1839f7cdeedf2b77661881d70884310cccd5570ed33f4f5df

Observation 13453da0-ec46-4982-b78a-013dd73a6d8a · outbound

This paper cites Drivemlm: Aligning multi-modal large language models with behavioral planning states for au- tonomous driving.arXiv preprint arXiv:2312.09245, 2023.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Drivemlm: Aligning multi-modal large language models with behavioral planning states for au- tonomous driving.arXiv preprint arXiv:2312.09245, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.379779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.379779Z digest=sha256:522f1537900c79c2faf429b6eafc6714a510c53de3bb5b7946790bc516b0b8f9

Observation 9a54c286-d0e4-4615-9f80-2810d89316d9 · outbound

This paper cites Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.494427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.494427Z digest=sha256:70497c8afecaaf7e380bc1d4942dd932051c318827468fc7423d9a4be46b1877

Observation ad36b43e-9f5e-4e51-8c2c-75c1979abe93 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.650777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.650777Z digest=sha256:dc305a87ae6d6dcac8a036f6b9604866e4cd3870dcaacd5b7450eca83b084400

Observation 17801c05-cec1-42a6-aacc-867ccdf304dd · outbound

This paper cites Woodpecker: Hallucination correction for multimodal large language models.Science China Information Sciences, 67(12):220105, 2024.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Woodpecker: Hallucination correction for multimodal large language models.Science China Information Sciences, 67(12):220105, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:42.622977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:40.775096Z digest=sha256:aa760d226b55286dec7ecb833cd6ed11f9fb68116fd2ed4a26e6097ae59b908c

Observation 92015d10-50dc-43c7-a817-b1ebbaa634cb · outbound

This paper cites Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:40.942389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:40.942389Z digest=sha256:374e8bd54d5f8bca99dbc6977064b861a1c71591cb8e89aafaf1b452b3ebdff4

Observation 0e1a60c6-4179-4a62-8d22-b860337d70a4 · outbound

This paper cites HuatuoGPT, towards Taming Language Model to Be a Doctor.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models HuatuoGPT, towards Taming Language Model to Be a Doctor

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.054437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.054437Z digest=sha256:ec17876fb978ba0321fc4c5055c67d4d9d97f58d39a59ab8ecc3ac8116997573

Observation d29c0b0d-3e93-47a7-9713-1f3b465572b9 · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.235492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.235492Z digest=sha256:d844058136c3cb360e6f999447c0d49d91fd773de87aa7ed2dcd9def68bd8004

Observation 60c69dda-7610-4e4f-99bf-92abdeafcb71 · outbound

This paper cites Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.382368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.382368Z digest=sha256:8e803b6082eb0a63db9752ed67c33d91bf621d56efac9ba89eb13326074b9cc9

Observation 6c42e6af-2ab2-49b8-ae91-db9e666073d1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:41.573398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:41.573398Z digest=sha256:8d34c41dea2b324969a29e6d4020591ae12165ae4b8fe0559feeef516e9883c0

Pith citing papers

No inbound Pith citation observations are available.