Pith. sign in

Paper Citation Record · LEDGER

What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2307.02469.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.02469 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T13:48:48.661566Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T13:48:48.809031Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 97fe064e-72d2-436f-a61e-1e6fef227178 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.330760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:abea1b8b1a23138ace523d8f8a30dd343e4905c938e303dea063d448a99e7ffd

Observation cf6f323a-1b96-494b-a7ea-2a486b3c7a8f · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.572284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:4f4d69d11d3f536241756b9cb6301c73cf7d4578ca3b3e60dfb5e6fdef0d0c9e

Observation 25237c60-43e8-4f83-8214-bd179d311a9d · inbound

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models cites this paper.

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:57:09.061703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T23:57:08.818348Z digest=sha256:d496ee6a4b8f066b274b7ce322ffd80a61d60a26faa0865ebe7c16f6713ed691

Observation 5e3faddd-0351-409c-b533-c105526e204d · inbound

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition cites this paper.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.813327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:72158572152d900b47063192ff0590df7fb93bbda699b3f05d96acd3fca565c2

Observation 8b0dbc90-91ca-44d6-bf23-fb31328afd03 · inbound

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models cites this paper.

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:22:04.195502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T01:22:04.035994Z digest=sha256:14e9414e0caa354e099cd6eec78acf008dc8735459c7a793db4230b780f9938c

Observation af676ff4-e473-4d76-a556-4a79dd83cf88 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.224847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:162958cb9100530648a555be1505968cf6a01de3a8d8ffc8bb86c601b47c8a1a

Observation 00a15118-ef80-4fd6-9517-4d71f41dae92 · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:22:35.390815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:4952cf5db714245de6e4efd9027d2a6d4fc4a029433f68e3c1ee001e1f73152c

Observation 1563d9f1-f958-4f54-bf58-411fca5787de · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.349839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:2b0efa5e0244b0d8fc4ab1df19d8fc6f8807198aa835650f9f16a254b814d8e7