Pith. sign in

Paper Citation Record · LEDGER

Apollo: An Exploration of Video Understanding in Large Multimodal Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2412.10360.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10360 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:08.568366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:02.918791Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 51649a7d-f4f2-4ec0-8e87-21456cbedbd9 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.833571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:b2d11dc77063cc7f5235c96dcd845496cd7b61ced390da90045075500995661f

Observation 3ca642b1-742d-46e3-9920-106d019fb4c1 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.181502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:21abc1aec09cd1a0a5f9398ae0f1bea49c34723ebcbac5ac4e5b7c3ce3cb9386

Observation 73db5c87-2b89-4b6e-abd9-9da8866b837f · inbound

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? cites this paper.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.568366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.568366Z digest=sha256:368fd70f7c62b377a8a21ec1c8f0d7eba4e4e5511d31b61a29fd60389c0aebf4

Observation 208b340b-771f-4eda-aa3a-b4fde362b2f0 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.930004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.930004Z digest=sha256:2a979d175dfe05f5605753acce5de65988f8a1a8d92d8e5c4f6e3b49d7b580e3

Observation 3e01e98a-c288-4951-bc8a-ad6f1c0fdcf2 · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 120

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:14:47.811975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:47.811975Z digest=sha256:431c273d68a45f2b7fa481a6a08f481d097393b910ee5b7913989bf243b05665

Observation 447bbede-b0be-4fbc-9feb-22827ba5a7bd · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 84

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:37:24.588694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.588694Z digest=sha256:d131c5efda5319d56c3f7eb49642d1c5b55aa6eb757612d3a69106226f8bf0db

Observation 0ff19277-4d81-42ea-a6ca-e0b1fdb91644 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 53

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:59:11.354129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:11.354129Z digest=sha256:c2d75d3dd284dafc32f0a2597f70442a931ab31be5065cfb66ca2bb0ad61f29e

Observation 1ab9b6da-ae28-476b-9d98-ef9bd5fe884c · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:16.401147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:16.401147Z digest=sha256:fdf6a29c59f3555d9adaf9f12cbadbd9b761ab6411756d71f34dc4c606a6462a

Observation 1645717c-29a1-4eb3-8e9f-e7e2978c2325 · inbound

ARGUS: Hallucination and Omission Evaluation in Video-LLMs cites this paper.

ARGUS: Hallucination and Omission Evaluation in Video-LLMs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:39.827734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:39.827734Z digest=sha256:2f61dbee2dfafcbb53aa2012f2537178b2415644e4ea47fde95c4d565c3af694

Observation aff8f26f-bed7-41c0-a3c2-a8a286242dcd · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.976093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.976093Z digest=sha256:3ef9d4162493ce824409f494f046d850af8a1996791343a642966a9b1d7053c5

Observation 0a20b61a-a85a-4611-93dc-11fbe6e68ff3 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:57.031289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:57.031289Z digest=sha256:1bfdcbe44938bef0d88575b10712615ee0507536594f13688683785fb05cf37f

Observation 5f9be675-8665-4051-ba81-0ff362f1473c · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.498173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.498173Z digest=sha256:6d10c4831a72f082a00c8bc11890d7f02ad00721dcd01d7faf07ba0384789891

Observation 44079027-0ad8-4f25-8acd-0306ad0d4fb3 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:43.575145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:43.575145Z digest=sha256:f4674e28ef52c671b6c234a0c0041982b548a36d41043e9b54b553cfe0d972aa

Observation 6cf8de07-4242-40f0-a445-23d5848dd7d5 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.512195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.512195Z digest=sha256:f287f71e75035575b99423f8cd735492f010584f4e06562b478307564dc05246

Observation 4ebc1d82-7cb1-49f5-9264-9aaf78b8426d · inbound

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? cites this paper.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.145067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.145067Z digest=sha256:ad423e8659a4216f767b77232fa7a1644ddfe218816797e27176b9c7d19194e4

Observation bf372c0d-9bca-492b-a2ab-78d4e97b3cb7 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.860880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:95e032176c6b0dbaa93852192aabff97f10824bf34cc1eb420b764ca87bd0098

Observation 318777ff-820e-4174-8411-3e46e8365194 · inbound

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models cites this paper.

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 79

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T06:46:37.520929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:41:59.641410Z digest=sha256:c27fa5435ffb9fd7f2d2b28aff6fd621c793c19d5d49952549df995ccc40e352

Observation e639179c-8052-4189-8f40-7fd587b4e7c3 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:20.639077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:b423502b1deb640891cdfdba675796ff48fd56c1f178d793b8147f86aa537bae

Observation 0a37124e-a73b-4485-b966-0c8d1371122d · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.596352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:bc93341e55f7020baf2745ed7df195c0239fefa0484989066890afa88415b35a

Observation 3113e5cd-019c-4b5e-9caf-e93928e905c7 · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T15:34:37.002001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:34:37.002001Z digest=sha256:e6a68d838bcebebc688d26051e504c1245b517fc9cda666bc12bf7242e33a350

Observation 842d6468-15ae-4852-8695-f90881a5b040 · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T18:39:24.915547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:39:24.915547Z digest=sha256:06b47a8441cb98b255cc153644b2fdcdc86210d4253301c3c8b96b2569168b28

Observation e62a9e4e-ff01-4793-b538-68d25de14127 · inbound

PEEK: Picking Essential frames via Efficient Knowledge distillation cites this paper.

PEEK: Picking Essential frames via Efficient Knowledge distillation Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.314765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:46:53.704892Z digest=sha256:91036526993feb468d1f1899617c1f89594130c3227b2a0c9f9ba1678cac58dc

Observation 1b33d8c6-4bb1-49b1-bbfa-ae9e82327cbc · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 263

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.920378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:8c55ccc2db5a89ea96e3460535ea299eed2c0b171f57d162ab1eab931e796271

Observation c9dbe76f-b26c-4668-9b61-2ec3dea2696c · inbound

Agent-Computer Observation Interfaces Enable Dynamic Computer Use cites this paper.

Agent-Computer Observation Interfaces Enable Dynamic Computer Use Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:04:21.211164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:59:20.295818Z digest=sha256:fec0f509270205f2cf044556a8f518a2d735d55bc3a90924096f79c8e4907b5b