Pith. sign in

Paper Citation Record · LEDGER

MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2404.03413.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.03413 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:59.428434Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.386392Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ccdaa848-0a74-47df-a5d9-781b71649ec2 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.511432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:414e491339926341f06fdc389dc4bd05805acb705616e7ff3bf2274650af703b

Observation 1c5b397b-649d-401b-a0f3-626d2b80f660 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.352805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:c328bcb35de3c37ceaa1bf269a6a1689006b39f93192b22bf3521d8833e49ae7

Observation d75d13d1-16a3-4873-9478-01a542fae4cd · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.610283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:697ab2e451913a29e2c7eeceddfd8295e822702216d855ba76ddaf53d9dba6cc

Observation 7bccb503-6199-46e9-9652-c5ee814d49e5 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.306020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:8baa7e6a6b3eec2f3aaa04e047f33e6606f20ace80a6737875ad6840529f0aac

Observation d5a8ee7e-ac4b-4091-9af7-9c75d21190cd · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.070072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:4fbb91d8c2f6d7eb04316d08022ce3794571e42746c6039858f02d87c079e8c8

Observation d2f6c73f-2372-44a3-b160-0537fb465047 · inbound

EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language cites this paper.

EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:59.428434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:59.428434Z digest=sha256:6a43fd9641e7276e4c7a89136fabacc3c6abd084986b520ac2d89f1fcd9e0d48

Observation 1b4b6cbc-051e-458e-86a3-e7dab822e61a · inbound

Multi-modal brain encoding models for multi-modal stimuli cites this paper.

Multi-modal brain encoding models for multi-modal stimuli MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:30.173213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:30.173213Z digest=sha256:3dff8f8b0d003b03151662075b1d948a75e5f42b94338fcba24e342406ad0dd8

Observation b324e589-68a1-4714-8d08-3deb4559a04d · inbound

VUDG: A Dataset for Video Understanding Domain Generalization cites this paper.

VUDG: A Dataset for Video Understanding Domain Generalization MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:13.398642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:13.398642Z digest=sha256:c16f170f93d61774ad841123e40b26dd653e52737d1c3ad8d1e64aa9cca07a75

Observation b24fe65a-5f5a-4cb0-ba0e-6cc9c46d922e · inbound

From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control cites this paper.

From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:00.125179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:00.125179Z digest=sha256:91fd54f2e212c29280dc114a580abad91d74accebf4e0dfd833192a609a4a164

Observation 90c069d5-78fa-4889-9b20-ecc8e2e66d6c · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:51.919841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:51.919841Z digest=sha256:30dab0911bbaa8920dc602032cdde4f5a8c2f8a0eae767dcbe842e9956bfc1da

Observation 8be7d67e-4201-4abc-acb6-e92aa4708ba6 · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.106252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.106252Z digest=sha256:23ce5a5477db5a65c19a00fdcd92834d8b98c62a9d45ee646d6c73c73770b967

Observation 95c4507e-4f56-42ef-bd11-4cb240c38ea9 · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:25.241174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:25.241174Z digest=sha256:a6401a204cfeb9847971566cf79939d4799ee88f886c91e2697b130db8c3436b

Observation 33c76992-50e8-44ba-9246-b3c7cc996789 · inbound

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding cites this paper.

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:51.936727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:51.936727Z digest=sha256:adaae80b971f9f8f301fa5eb9c0fa609a6c763bef9b0f6a0c5809b3201385a0a

Observation 53afb53d-4905-43c0-bc65-1518f0ee4249 · inbound

Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations cites this paper.

Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:52:11.947587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:52:11.947587Z digest=sha256:b73d32858fe239d75e58262f4a5f3b43cf0de3bf4fa53711e2b0335a3c2f239c

Observation 5c40a1ee-0707-4abd-8809-7b85169751e1 · inbound

LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition cites this paper.

LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:42:48.068160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:42:48.068160Z digest=sha256:cbb796080fe0d1bb9583f24c2f8612e0e25c7ae3f3239a436ba441058bb17d8c

Observation f9885207-34f9-4eef-b420-a4429e96f580 · inbound

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? cites this paper.

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:39:06.191591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T05:36:09.208754Z digest=sha256:d4d5dee243118d1d002310369994c088c1bf5a7c1e71b16deec540b069b42b71

Observation 8aba6b87-1732-4856-9e30-2136e5deb11d · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.978281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:ed80a100473faea8b326269874d1ba4708ed6cc12383d8aa7b23ec6122dcc48c

Observation 255a21a5-7f8b-4120-acc6-f40f65ea49f1 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.941287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:29419eea05be0aeaeb340ac8582afe84e520d07ffe4d5b41f8a74e0059d03262

Observation 7a72c2ee-6597-44f7-9e06-bbaebd2e3e40 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 158

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:2fb2f619b76d38ee0368f162b9ca57db013ee67bd3bdca15a23ebf2175033ec2

Observation 660701bf-6df2-4b57-801e-7e7628952385 · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.574338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:fde2039bc1a7c90a7caeffc3b81609fc4141090e9fd7ac89956bc456cf9119e7

Observation e44ff72d-fe3d-44eb-90f9-764ea821fa25 · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.876383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:e592ce2093041f1ba0d45dcc6bbe321e94bd6549e50cd975135c2895b458f056

Observation 850ea5ce-203f-4cd4-ac50-d4befa26a7f0 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.192029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:92919b29d014e03095af4ae8534786d547b022fc4b10df0c822502648c300d8d

Observation 670b309a-0678-4391-96da-8b21ed8b59ad · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 147

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.387790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:f2b5b425b1e7226d6e1ee0bf3425f0c689b347e84bed81fb2a1091dfa814034c

Observation b7bb3f1d-09d1-4fc2-bb2b-9f0d9867f14d · inbound

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context cites this paper.

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:04:21.289121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:01:49.904752Z digest=sha256:9200043504ff9fffc663cc4190b1d9433cb340a524f9cb970ee67a330e1d7e34

Observation 20a41ddb-143e-4c2f-a16e-6dcf548dd89c · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:b273d404f432086f52eeac4dce0adca23d8787010b918b63d1487d2e2e091998

Observation b5bbbc41-1e21-43ea-bae8-2169698fb5a4 · inbound

Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding cites this paper.

Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:45:40.479683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:45:40.479683Z digest=sha256:5bce6cd0a6fc0080c7311b3c77abb1fd3453408e2ba0447ebf624ac60f333590

Observation 556337f2-1942-4cbc-93e3-483748e20b39 · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T11:36:26.082445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:36:26.082445Z digest=sha256:d4a9d9eace37c6be1e1f67060fd42807ca3d353801d96df0eec24eba0c9f1b95