Pith. sign in

Paper Citation Record · LEDGER

MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2305.04790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.04790 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:29:27.951706Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

65
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 36daf7e4-f826-41e6-b75a-17c9a5908b9f · inbound

Evaluating Object Hallucination in Large Vision-Language Models cites this paper.

Evaluating Object Hallucination in Large Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:44:09.992326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T13:44:09.626361Z digest=sha256:66e020249fb390fa1bc144ab4b501cfd55ba0cecfe6edd4308d67b12867ee562

Observation 1fc21cfc-c9fd-4b51-bde8-4f916e6eab78 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.533899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:63b58d45b2250aff812e04e378ff9c3c614d53bc6cd38500a8ec5ccdd7838241

Observation 86dbe26a-9234-4a96-bcfa-090610f47140 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.661294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:bddfe6cef636762e3149a75e86b2a631ea495001bec384d67ab690f894a29c6a

Observation d4a385d0-b631-4b8f-a8a9-6f49cd240e8a · inbound

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning cites this paper.

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:34:56.925616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T17:34:56.836034Z digest=sha256:bb6b46afc4a10dc011a2ffba2992924092f11118a3f5c6d9b3bb455ee731601c

Observation 6d0cf446-ed1d-46b4-bd04-78d0f90a5ff8 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.218129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:ea934b652c289700f392ddd6f1b7050b92cfa01c1dae03597b9fc3c63956c6b3

Observation 8d63810e-811b-41fe-b39f-90fe5ddeca6c · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 290

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:54.086446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:e249c5122d07cb824f261ef12f67395b0e52f5f1dc4cddaa0b0006329d8a50f4

Observation 5d60023b-8956-4dff-90c9-98a8381dbbe8 · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.871504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:29048b513ec988b26632d21499aeaacabc1bba3e23646c46c06df714952f28b7

Observation cf8190b8-135e-4957-980f-dc37a0eb712d · inbound

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning cites this paper.

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:13:08.967416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:13:08.867745Z digest=sha256:712a0efd9eb8f825631bcd82d66b4f80098087a0a7a07da60e1170f954fad321

Observation 8f444185-6efc-40b1-b169-19812663166e · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.645922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:ecc98e1410114a0972fb8300cba1f7821e51e785fe253cec9022080f2ab5805e

Observation d08b84ba-c7ed-4b61-934a-b1d3d5ad9029 · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:08:01.377521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:6b40ae5a1ab935b7ba020a07410abb6cf21b1f1f843ceca67c1f8e7c5adabf8d

Observation 3f44ad43-9dd2-4b8c-8bd0-78657d63d7b0 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.127729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:f0e3803d8029b58d21b4cd64f6797b6e34c2d25e3f5ac021240e32d3f0ec8f87

Observation 57dce32c-9cbd-4f79-92b6-bde2d0787588 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.287822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:f8990d4c579485a8d9afca7d8a2860d8e315f39004bc016688d2c6a8e20ed589

Observation 5c10f305-4288-4827-a656-462842bf09e0 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.115905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:24e234eb261edab671eb85ae9b8217c5b26587ddf1b9696fa2dc19347374016d

Observation ad5d363d-ffeb-4beb-ae4f-f7c953c45272 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.719546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:23f0f9f97899e2333ab2c4c39335b78c4aff1bfb8c2362b54c563996f430d3e1

Observation 4f0c569c-2947-40d3-b5a5-39c76f571b76 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.388847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:867c2197a2272a09c2c1834531f2e648ad5308652f73b2855933bc25d639cee4

Observation c90589bd-17d0-4111-86b3-031d14f5aefe · inbound

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions cites this paper.

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 255

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:55:50.308479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T21:54:26.670284Z digest=sha256:15f335a959ac0fa5f7fc31463bb2621e17fde20a61204c47be4c41bd3db79f90

Observation 4d673733-08c1-4cc5-85e1-c3dc09dff0a5 · inbound

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving cites this paper.

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:35:42.170256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T16:35:24.063578Z digest=sha256:17e8e26985dd7e10025bad1b8d54b7412c0f44f684f414036e824bf0e95029d6

Observation a6e18ee9-723c-42ce-ae64-2d9d3a993120 · inbound

Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink cites this paper.

Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:29:27.951706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:29:27.951706Z digest=sha256:094fd286083506a102bb6d03bb9e6dd6102fb4bd5f3e547796ab3873cc6744d9

Observation b78d3394-86f5-4115-8bee-1f12f53155fc · inbound

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective cites this paper.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 142

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.290744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.290744Z digest=sha256:cee937e67d10a6aeab8e32e0cd65fa902179cce502515e9abc2e6d724b4e8a8e

Observation 088ced41-4caf-4971-9b92-5c631bac88c9 · inbound

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models cites this paper.

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:40:33.039816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T08:38:02.944230Z digest=sha256:74e194cf9ce1e9172a90da3ecb8a740941c0acae65a3070c59dcb69385ef8339

Observation b6714ece-d05d-4107-9b94-f2872aac866a · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.592607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.592607Z digest=sha256:54c4c0d15ebf7495ff0d223f0cdebd544139be7e43b04c77274227e689158e66

Observation 2292cf2d-e482-49cd-afe2-1bcd858ada29 · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.397197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:dac90e668bbf8c5d80ce7e2593b6411c8896d35d498ad91dedb340628ddc3865

Observation 45aeee9e-f93b-48e0-8059-9d25601cfaa0 · inbound

ZINA: Multimodal Fine-grained Hallucination Detection and Editing cites this paper.

ZINA: Multimodal Fine-grained Hallucination Detection and Editing MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:12:14.614209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T10:07:18.162776Z digest=sha256:f9948cecf27d8c1da0eb1f270bed15c8b846a00e40d0914bd37b809311f33fbb

Observation 400765dc-0cac-40bc-a2aa-1021cce2aab4 · inbound

MDSAM:Memory-Driven Sparse Attention Matrix for LVLMs Hallucination Mitigation cites this paper.

MDSAM:Memory-Driven Sparse Attention Matrix for LVLMs Hallucination Mitigation MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:55.367724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:55.367724Z digest=sha256:45863bba701f39c1e9a1383472a05db5e5677bfd75eaea26e9a08050b661bbac

Observation dc773c05-af7f-483e-8f26-a00cc3dd7477 · inbound

EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices cites this paper.

EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:56.603817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:56.603817Z digest=sha256:6bbb826752ae551286fa4c8618ae75ac4e44d054dc9ff9ee0f1385b8f4acdfc3

Observation 5a8ca3ef-4922-450f-8dbc-eab82778dc6d · inbound

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models cites this paper.

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:13.222612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:13.222612Z digest=sha256:761d405ccd0cd55c91cf73b89dfd43fb93a3c42865332908663430f6b2f1456b

Observation a17860a5-add1-408d-b48d-f976276139fc · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:41.699694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:41.699694Z digest=sha256:03ce7bf404080f4affec73ae2d1720d94601d98150091571881f4ac8f08a326b

Observation c873d276-6d9c-4465-aca0-e444a94ae04b · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:34.629056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:34.629056Z digest=sha256:0a71e13ab0681503801c1cae6860512095d038792eb3e74fba881e7987c80591

Observation 108f2783-adb1-4ad8-bd18-b35c1718fbf9 · inbound

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment cites this paper.

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:27:28.673021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:27:28.673021Z digest=sha256:fafb493d565fd21e17b9f348bbc76c813d7f7275becfdf3162f76c1c2e714079

Observation 8a31ba8f-7f14-441b-b87f-39e522c77927 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 292

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.434747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:696b9655fcc5bdb91f8cd5e359f506c52c44779f004a70d53edcf23c883fb924

Observation 149b6ae8-df44-46f2-b627-27211285b6a6 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.569658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.569658Z digest=sha256:40e61a615c37adc3f93a5897e7faa85f3c10e8aab519dc6f0f1ad174602e25fb