Pith. sign in

Paper Citation Record · LEDGER

Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2308.04152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.04152 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:19.405502Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T13:48:48.737973Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9b8c2625-e0dd-4836-9b36-f21edba8c619 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.612655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:1cdbdf8c0d9f0e9467ae257ca1a7bfc6619ebe78e37e3f0e06046c7ca8504f6d

Observation b2a9dea2-6d23-4c30-ae80-9c3a6e17766e · inbound

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition cites this paper.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.740728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:a27141ae8c25f59cf0f77dd81088e6c96592c3a1b9ae8dfd54a5c21e1a8bbd48

Observation af34bcc1-16cc-41c3-ac41-7e8054bfa2de · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.618766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:2fe4e30a9749eb241a8837e9c7a6fab8cf7f8684211049f48c79e46cd4dd28f2

Observation e7821c05-6750-4e3b-91b5-d3fde00a313e · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.405502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.405502Z digest=sha256:f8c6cf711133476cbc6aba71f04ac334a5d063b75c25a0c7c8b38bbbe3ba54f2

Observation a5d08ef2-8fb8-46da-aa02-8e3f1342f3db · inbound

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL cites this paper.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.526429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.526429Z digest=sha256:60b6389de6beb90a1c2722d4ce591249a7707b48b09e8c72241d0d8f71c79270