Pith. sign in

Paper Citation Record · LEDGER

What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2307.02469.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.02469 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T13:48:48.661566Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T13:48:48.809031Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 97fe064e-72d2-436f-a61e-1e6fef227178 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.330760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:e98d0ff7ea1000777ec2e0716963ab59b240b21c3da7389cf868c8e62757dcfa

Observation cf6f323a-1b96-494b-a7ea-2a486b3c7a8f · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.572284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:1768d583df863b52bb67a0799cee7997b3d07a38443781c3b74a796e68f3170d

Observation 25237c60-43e8-4f83-8214-bd179d311a9d · inbound

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models cites this paper.

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:57:09.061703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T23:57:08.818348Z digest=sha256:6f4566ac335ee836a6e4472f932486293296f14c62969d89f174e4a91c67e1fe

Observation 5e3faddd-0351-409c-b533-c105526e204d · inbound

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition cites this paper.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.813327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:eb3aeda9a635214a415bee58e88ba8f1232df611f547cc0fe26b961d5724b0ac

Observation 8b0dbc90-91ca-44d6-bf23-fb31328afd03 · inbound

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models cites this paper.

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:22:04.195502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T01:22:04.035994Z digest=sha256:f0b573033d257f8590aa17a1507184915e21324d6d444a66f3d7209898e26a54

Observation af676ff4-e473-4d76-a556-4a79dd83cf88 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.224847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:7c1f6ed2cb2e792d84538d9f1d470fe5bd65f7570bc6ff976ad939b7b3d77740

Observation 00a15118-ef80-4fd6-9517-4d71f41dae92 · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:22:35.390815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:baf64e209357feb88c48ffcc5210559021ee2114a7ad0f37c8ee10c4604bdd68

Observation 1563d9f1-f958-4f54-bf58-411fca5787de · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.349839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:57ccd2ba7db752887ee29504238dd3f58f37403d547c97ed58757d687baa37bb