Pith. sign in

Paper Citation Record · LEDGER

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 6 inbound Pith citation observations for arXiv:2501.02699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02699 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:12:46.203510Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:39:46.817981Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:33:33.048939Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c7b9e58d-426c-4226-af10-28feab2a975b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:44.251736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:44.251736Z digest=sha256:bdc76255ecc4c1c4b1faf55d3771773945b1951d97cbe01c9c5486df9ef0cdd4

Observation 0a343b1a-2c71-4d75-9b27-31a539697c31 · outbound

This paper cites From colouring-in to pointillism: revisiting semantic segmentation supervision,.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models From colouring-in to pointillism: revisiting semantic segmentation supervision,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:52.954875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.296963Z digest=sha256:995d270eea851f570d4d06e49a6092fafeb4deef011e4f133500ed566d324c97

Observation a7042ddb-c0cd-432d-84ac-f696ee083c86 · outbound

This paper cites Lan- guage models are few-shot learners.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Lan- guage models are few-shot learners

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:52.765074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.364762Z digest=sha256:cad3fe702f87c44928af8c9929fdf938c476400a1bcb624a62b46bd1c61ec93d

Observation d3008bc1-1a52-4ac6-bb86-461bc590c0b5 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models A simple framework for contrastive learning of visual representations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:52.584897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.424825Z digest=sha256:a69fa7d6848c10a717caa907f474e6751958baa17d8547e4c23adb48d47a87d9

Observation 05bad6a0-b0f8-44e1-92d0-4b8b5337f409 · outbound

This paper cites Pali: Scaling language-image learning in 100+ languages.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Pali: Scaling language-image learning in 100+ languages

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:52.392965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.453927Z digest=sha256:33909543c1826c826f73697538bec092dd4f6051c59d3b066b6c2e99113ea48a

Observation 295dfee2-7742-4d67-bf2c-1a7c9b24c125 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Reproducible scal- ing laws for contrastive language-image learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:52.254841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.463204Z digest=sha256:c3c7fac967fcb6e85a63251a9dea97755a4a5f6d20aea176c767ef7d750ddf28

Observation bd63c6de-4d5c-4948-b8bd-8fc567f0c520 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:52.083232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.484447Z digest=sha256:501f4f75e8428a19ff5229a01a0595c86dc60ddb948bbd38f56abb02905f8af6

Observation 1b7ca14f-55bc-4981-9cdc-d76196d4bf43 · outbound

This paper cites Scaling instruction- finetuned language models.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Scaling instruction- finetuned language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:51.920778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.544990Z digest=sha256:a73bcc822141d2deace8a51ddce0d9ceab94abe4c33c0db91f8e20e96750a82c

Observation 3ea23f45-5b55-4806-a84b-674e585aac84 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:51.833134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.584748Z digest=sha256:db34602a1a35d8b1dd798f58db132f0998c0b90d598157017e24b740f08d02f6

Observation 0b528abc-5b6d-470e-bfd7-fdf62a34c9b4 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:44.634750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:44.634750Z digest=sha256:597025188887a7544a71c35f4af798e61974db516eccf854f4a1ae62d6a41ffb

Observation e3a372a6-26b2-42c9-be24-1ffdb16a0a21 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:44.674747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:44.674747Z digest=sha256:54911e1a7541edac20570a05e0616c3678dbbb1d2e1dcd474a63d637299a4b23

Observation 42ec67c8-c742-411e-9e4b-e5bce5ad1869 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:44.724730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:44.724730Z digest=sha256:9266c9c1997af9be5ebd424e574639f9f06833881178688e4b6339c59437fca1

Observation 480426e8-a53b-436f-8222-3f5ac6131c2e · outbound

This paper cites A unified continual learn- ing framework with general parameter-efficient tuning.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models A unified continual learn- ing framework with general parameter-efficient tuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:51.685406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.775478Z digest=sha256:52bee9ce0f5504ae682d1fb8e56dd230e3aaf4e376307e67c6b81b2200e5364e

Observation 17fd81d1-ee78-43f7-a47d-00aef5545268 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Bootstrap your own latent-a new approach to self-supervised learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:51.544757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.824750Z digest=sha256:25d01fd0f2cde1cee1562589a9fa1b4fc9134347ab43844bca6d17c0c114e496

Observation 6fe60b39-e4c5-445c-9b7d-12051a3fbf12 · outbound

This paper cites Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:51.382895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.864868Z digest=sha256:08930c73fde35af055b14e307ffb091c8faff6233373c5b51fa1c54b4da7d1c9

Observation bb3d1da8-92c7-442b-aa18-140b45367798 · outbound

This paper cites Dimension- ality reduction by learning an invariant mapping.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Dimension- ality reduction by learning an invariant mapping

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:51.255103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.920819Z digest=sha256:d9346174de446a0c1a5b2c2009643b1cfe3b04c8224d8113071aa7c54084a21f

Observation 300379b3-abc7-4e0a-bc4e-018c85f0604a · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Momentum contrast for unsupervised visual rep- resentation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:51.148550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:44.974859Z digest=sha256:c41aa6d68acaad86b6993cf054e412c5991bc62efb5e1665ce073bb68f28296b

Observation fe3f92e1-3b6e-4386-b31a-43b7c5d11564 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Lora: Low-rank adaptation of large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:51.051988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.035048Z digest=sha256:b1b6499082749de231ed16ff99c0b645f30891aa72cfe86501bfb09792bae5db

Observation 06141fea-6bb7-4ecc-a2b3-a91be3ae543f · outbound

This paper cites BRA VE: Broadening the visual encoding of vision-language models.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models BRA VE: Broadening the visual encoding of vision-language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:50.986785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.086231Z digest=sha256:d2efb4b0008cbb58f115a1e8639aaecf9025dfd76266724ca95fa9e2e6dd9d42

Observation 5a82ff62-5666-49b7-a6b4-b69f49c4fde7 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:50.828960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.134839Z digest=sha256:bbd7f5782f47843345b1eb18b304598758f60a3327d0a8553ca233d85e72179a

Observation 5475cd02-fe9e-41d1-a807-c216b85c3f2d · outbound

This paper cites Evaluating object hallucination in large vision-language models.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Evaluating object hallucination in large vision-language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:50.684823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.162801Z digest=sha256:5905e7918d09b2e0ce58ea8842e648e7d3750bf249939230d93d8fd6fc2afc69

Observation 9242b92d-a487-4efe-9c13-01176c3c2c61 · outbound

This paper cites Microsoft coco: Common objects in context.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Microsoft coco: Common objects in context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:45.205998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:45.205998Z digest=sha256:96e69f5ef590efe3952bb22f47e896dcb7120d8f85be396c0869a9dd738f7efb

Observation 0ecc45de-0f82-4f83-b926-50dd3a274eae · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Improved baselines with visual instruction tuning, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:50.184915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.236334Z digest=sha256:b98fa1f9570f78fbec30de9a86d19849e1bbf7a7cc589eb22dcacc9abcbd2932

Observation fdfe1b4f-0fff-4b23-93c4-10d09b7fdad3 · outbound

This paper cites Visual instruction tuning.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Visual instruction tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:50.013967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.247018Z digest=sha256:ebbd47ba5497e6ce32a0d57e50f92f2efc5146577dee4c39a2a8e447c4ed7d5f

Observation 2df76d0f-031b-485e-93e0-ae15b865f0ee · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:45.264752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:45.264752Z digest=sha256:60d1a70c3752ddafd8f1ac4045d2184c8d8c074e7db6085e5edddca77415ea83

Observation 63674953-6d78-4576-887f-d3fe1d4463bf · outbound

This paper cites Gpt-4 technical report, 2023.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Gpt-4 technical report, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:49.938858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.304752Z digest=sha256:478eedc564219703140f6c7da3199cd46813d5b1664d787a185504b170bdfbba

Observation a788af3f-93ba-4716-844f-0c764e310eea · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models DINOv2: Learning Robust Visual Features without Supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:45.347445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:45.347445Z digest=sha256:eb058ef2326def5ed28c4714ac68f1ced46cd1dae62f9bde957522a14043cdef

Observation 09d6a444-330f-4814-a7b1-7066a29cd35c · outbound

This paper cites Grounding multimodal large language models to the world.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Grounding multimodal large language models to the world

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:49.784833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.382543Z digest=sha256:d85574501dacfeb51cbc774decda6a2a06c9f1448dd71f74f680c16df3ead185

Observation 32d80f44-e30b-4224-9dce-ec26c648ec6e · outbound

This paper cites Learning transferable visual models from natural language supervision.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Learning transferable visual models from natural language supervision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:49.624754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.444757Z digest=sha256:fa9f61b120e267df0c493c6d79742260ad054ec310bceb8e826c1b9cb968974a

Observation 84eeb2b9-2e1b-4566-a8b2-b566d629f2c2 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:49.464845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.494753Z digest=sha256:833df032431cca35844f0a21ca8a9ad3f539b005c26c0ea67c972abe0c2801cd

Observation f04221c6-25e9-404c-a1fa-84982be34d33 · outbound

This paper cites xgen-mm-phi3-mini-instruct model card, 2024.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models xgen-mm-phi3-mini-instruct model card, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:49.311266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.546305Z digest=sha256:d58846de4a45206de1a1b951632d7b7632646bd3e1a602343eefc3a744a1842f

Observation 5af1e5ca-f966-4c73-a321-995674b0c369 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge, 2022.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models A-okvqa: A benchmark for visual question answering using world knowl- edge, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:45.588051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:45.588051Z digest=sha256:30d3dc52fddbaefd7127dd6536420c794296b1dbfbbb4a294597196e4bf2dcf2

Observation 635290f4-9207-4838-9dbf-00fc11950316 · outbound

This paper cites Eva-clip: Improved training techniques for clip at scale,.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Eva-clip: Improved training techniques for clip at scale,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:45.634749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:45.634749Z digest=sha256:d9aa9546a48f6c5c707b4c7cac8f0f57873909f8802d5cfadb2c8816fba48e42

Observation 5ef7d0e3-9af0-47fa-aeee-ede1cf3e9863 · outbound

This paper cites Eyes wide shut? exploring the vi- sual shortcomings of multimodal llms.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Eyes wide shut? exploring the vi- sual shortcomings of multimodal llms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:48.905413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.677014Z digest=sha256:a450ab515816715ea789f653eeec9cc470ebbad4c258a93e1babacb93ba8f3e2

Observation b95bbb3a-0605-47f3-9421-11dd07e27506 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:45.716145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:45.716145Z digest=sha256:72fa4eecf0390e95371f92f5f4a32effef60db96a13b995bba76e31b5b82b582

Observation 96c09293-1ff0-4048-9783-5fa2c93009f0 · outbound

This paper cites Pivot: Prompting for video con- tinual learning.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Pivot: Prompting for video con- tinual learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:48.744839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.764756Z digest=sha256:2b10d9aa66d2c137417258f50960d3a7ffffc25795180812677eff02e406b267

Observation b1828c33-1c9c-4cc2-b7f8-29247a18e67e · outbound

This paper cites Behind the magic, merlim: Multi- modal evaluation benchmark for large image-language mod- els, 2024.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Behind the magic, merlim: Multi- modal evaluation benchmark for large image-language mod- els, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:48.564753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.815414Z digest=sha256:93f015278e4badaf1b5e734b09321e536f7bf311c6fc638baa118f9d0c93a041

Observation b63a8b17-9a5d-4957-b9e6-21f37885b11d · outbound

This paper cites CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:45.874755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:45.874755Z digest=sha256:9474a98d3f26272bd9cfc65564dbfc965d71abbde03eed94246406bf13eb1984

Observation 701f9a50-6733-426d-8c1a-a1422a1dcc10 · outbound

This paper cites Sigmoid loss for language image pre-training.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Sigmoid loss for language image pre-training

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:48.405323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.944753Z digest=sha256:1c9496849e9f2c4324951e8a85320454ee9cb5208b9c707c0908c55b7e16e9b6

Observation f5d1b0f0-9725-4153-b827-08613a188c29 · outbound

This paper cites Galore: Memory- efficient llm training by gradient low-rank projection, 2024.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Galore: Memory- efficient llm training by gradient low-rank projection, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:48.238813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.994195Z digest=sha256:1b8e4714e9fb96c7ecc915c91247fcc1b310a5defea95745d6a29e59797c036a

Observation 8a0cf7ee-cb42-4269-8d08-5d482c3ed60e · outbound

This paper cites Analyzing and mitigating object hallucination in large vision-language models.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Analyzing and mitigating object hallucination in large vision-language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:48.065807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:46.042073Z digest=sha256:4df01c3703658defe53970487baad1ba33be8e7e708afb2647233d49fd580fc6

Observation f597d41b-bfd7-46b8-a88d-c284693e4134 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:46.074760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:12:46.074760Z digest=sha256:7633bdf0413a6c150552445f15d2a6afb90977978a65360434a4675fcecf71d8

Observation ad788c7b-c3b7-46c7-9ac5-2784493283e4 · outbound

This paper cites LLA” (LLaV A-1.5), “LLA*.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models LLA” (LLaV A-1.5), “LLA*

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:47.914917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:46.104763Z digest=sha256:c35606720472b9c86cd5f40bda77c69a22d1ffaa1e16c2184720552543961351

Observation 9afe98ca-a354-49c4-a965-ac45f47ab924 · outbound

This paper cites We compare the zero-shot and linear probing performance of EAGLE-tuned VLMs against the original models.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models We compare the zero-shot and linear probing performance of EAGLE-tuned VLMs against the original models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:47.714755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:46.140752Z digest=sha256:58e9dba508ddfe811b63e20701a30fb86c8620796fa3e9ad9d8b39cfc785ee8c

Observation d41bdf56-ba47-46dc-9a91-7014699d54b2 · outbound

This paper cites Figure 3 presents 3 additional scenarios when EAGLE effectively reduces the hallucinations of the IT-VLMs.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Figure 3 presents 3 additional scenarios when EAGLE effectively reduces the hallucinations of the IT-VLMs

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:47.555717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:46.185025Z digest=sha256:5c7ecf6c70c93345c73ebdc40678a1e608642530692df0164b5ad2d54c74231e

Observation 2d98a4e1-12c9-4405-a4ea-9d474b4d1b83 · outbound

This paper cites The same hyperparameters are applied to both VLMs, EV A01-CLIP- g-14 and OpenAI CLIP-L-14-336.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models The same hyperparameters are applied to both VLMs, EV A01-CLIP- g-14 and OpenAI CLIP-L-14-336

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:12:47.404903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:46.203510Z digest=sha256:7ffa4b6393fdb65c13c4d0155bf4d3a197c2b19f437ba10e01a1ff0f62e2a998

Observation b7eceb0f-73eb-4269-b15b-d19ff7b11c95 · outbound

This paper cites an unresolved cited work.

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:12:50.460160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:12:45.183653Z digest=sha256:33245fa6be6601270e42285a3da99a7722ed78d1af413eabbaf88f55f2f9857e

Pith citing papers

Observation 3f6fa8ed-8e71-4294-bb90-12668f424ada · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.052044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:0ec1de5720177b03c01390e6cb6eb440367d8c2044bdd60b059c4dd8c653ee99

Observation 83d1695a-9fc1-4e70-86bc-45c46c44e8d2 · inbound

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs cites this paper.

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:46.817981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:46.817981Z digest=sha256:543905daed4c95d71ce954eaf52b953da518621298cbe652c05826cb49ea53b3

Observation 9e0cb9c0-6a7d-45ad-bf9a-db4528cb194b · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

Reference 209

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:03.976981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:03.976981Z digest=sha256:5f9db73f4c9975ebed9c8b11056818d2bba06a534cb89e09efc53b60258e38de

Observation e52a1569-1195-4c37-a629-46c10e323306 · inbound

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings cites this paper.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.923981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.923981Z digest=sha256:ceb4294c49164bfe7e5dfac62aba31d3fc13fa1ad709c6760ffa5cf458470360

Observation 65cc5cbc-6c77-4dca-8e7e-da991fdcbdbd · inbound

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation cites this paper.

CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T08:57:18.134481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:57:18.134481Z digest=sha256:40a0e7109b7d9702ec82c3158abc84f39922b3c3932adac26c89813bff917508

Observation 7e809645-c5ef-4a63-bc2b-21e91d2d3166 · inbound

DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models cites this paper.

DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:06.683252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:32:06.683252Z digest=sha256:36332e6954c733f731bbe49934c704cc98b7c0eafa1a2dd3e2b65895a087ddf8