Pith. sign in

Paper Citation Record · LEDGER

Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2312.02219.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.02219 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:10:23.063436Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:33:33.035396Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5719a4c0-2b08-419c-aa78-cc1eab742cb3 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Reference 158

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.046661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:8fdaf7412e4036117f2222d0c4fd58eecb7ac52e597580398e635de01eeaed16

Observation 751b6ba8-c280-4357-b9ed-8cee9cb374ff · inbound

Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects cites this paper.

Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:10:23.063436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:10:23.063436Z digest=sha256:fbcc4ff9a8df66dbec1b580026434c7efa71dcadb86e168d18e4c50a5d274330

Observation 4d07bc38-9034-4e72-8b4a-265125a337a1 · inbound

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs cites this paper.

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:46.564120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:46.564120Z digest=sha256:84d5c08e48287dd0b3ed235073011bf46f9a21f3fba7056160deaaea8b205d7a

Observation 0933451b-f5f2-4c78-8abe-84314733f6a8 · inbound

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images cites this paper.

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:33.934149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:43:33.934149Z digest=sha256:e3e803d72675f28d998845e0f6551943b4edc0950dd4325a65fa41900774070c

Observation 0c4133bd-d41a-42d7-8080-f873a69e81af · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:33.839250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:33.839250Z digest=sha256:a6a2acad046b03c4565e989a55ea39adc8b46e52511cb75defd89582d467013e

Observation c1f68f37-682c-48da-bb34-d2e721074214 · inbound

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings cites this paper.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.920899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.920899Z digest=sha256:17df7314635f4cef0d37ca0d1a68c774c016c0c9dc8ad3b3c23f9648106a64dc