Pith. sign in

Paper Citation Record · LEDGER

Apollo: An Exploration of Video Understanding in Large Multimodal Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2412.10360.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10360 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:08.568366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:48:02.918791Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 51649a7d-f4f2-4ec0-8e87-21456cbedbd9 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.833571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:def76746d639fc1a22d94585947857f420a7a278ae647a5934d4a4cacec26c8e

Observation 3ca642b1-742d-46e3-9920-106d019fb4c1 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.181502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:d1ff3f9271e910d8bfbf5d0600e93ae7aad2e662c01f25acfc6d6440ffeaae5d

Observation 73db5c87-2b89-4b6e-abd9-9da8866b837f · inbound

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? cites this paper.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.568366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.568366Z digest=sha256:0468d37f774724681a0bb2505acfa4db918dced44b707d513047a1a2c2fd61d5

Observation 208b340b-771f-4eda-aa3a-b4fde362b2f0 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.930004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.930004Z digest=sha256:d8867a32786b8801a65f8c7f04bf74d847c5293fe3399e337e36d3212c0c4aac

Observation 3e01e98a-c288-4951-bc8a-ad6f1c0fdcf2 · inbound

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval cites this paper.

Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 120

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:14:47.811975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:47.811975Z digest=sha256:2bc1be6817137605f56fd87c3a2b81f91a7e8f166218978d144c9c8868379850

Observation 447bbede-b0be-4fbc-9feb-22827ba5a7bd · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 84

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:37:24.588694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:24.588694Z digest=sha256:9226a2aec6f3cf3ef4afa9581954b4ec43d654505ce5df7f04abfc5f44467eda

Observation 0ff19277-4d81-42ea-a6ca-e0b1fdb91644 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 53

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:59:11.354129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:11.354129Z digest=sha256:ffaf46d96c552297efd5ba478ca4fb56f95424a24c600ec937af4a2532454f32

Observation 1ab9b6da-ae28-476b-9d98-ef9bd5fe884c · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:16.401147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:16.401147Z digest=sha256:8080d59ed9801e2ae90cdd39b18dc4343f36bd2df84d488f6b4334e4d61a40bf

Observation 1645717c-29a1-4eb3-8e9f-e7e2978c2325 · inbound

ARGUS: Hallucination and Omission Evaluation in Video-LLMs cites this paper.

ARGUS: Hallucination and Omission Evaluation in Video-LLMs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:39.827734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:39.827734Z digest=sha256:dbc1c56aedacafffe88bebccc31881c6df9f51319494ca455f6c55a6a8cd5d9f

Observation aff8f26f-bed7-41c0-a3c2-a8a286242dcd · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.976093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.976093Z digest=sha256:1d66db5b1dcfe653bd55f6843e3d964da20993ce529cbdcfdd9a1a940d3318be

Observation 0a20b61a-a85a-4611-93dc-11fbe6e68ff3 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:57.031289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:57.031289Z digest=sha256:1bfdcbe44938bef0d88575b10712615ee0507536594f13688683785fb05cf37f

Observation 5f9be675-8665-4051-ba81-0ff362f1473c · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.498173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.498173Z digest=sha256:aaad387ac8831e61e446d3e1d248f3f7392396226391c5e41abfd00ca781a509

Observation 44079027-0ad8-4f25-8acd-0306ad0d4fb3 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:43.575145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:43.575145Z digest=sha256:abeba33f4bebdbe831dc548b2dbc3df6ef613afc20a2014c1d58d102cafde3a8

Observation 6cf8de07-4242-40f0-a445-23d5848dd7d5 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:38.512195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:38.512195Z digest=sha256:36a03a1f1406e3dca5a1deda3f0c7880a03026426b8c195462ace8a28e8eda26

Observation 4ebc1d82-7cb1-49f5-9264-9aaf78b8426d · inbound

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? cites this paper.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.145067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.145067Z digest=sha256:a60e3282346421d964af56b777e52a26f18ddf1bd89a33234b29ee8476e0f22a

Observation bf372c0d-9bca-492b-a2ab-78d4e97b3cb7 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.860880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:9ccf0f4e33ed97cf27983320c9b56a43f9f2d54cf3f8cbffe45e42bd08aa637c

Observation 318777ff-820e-4174-8411-3e46e8365194 · inbound

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models cites this paper.

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 79

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T06:46:37.520929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:41:59.641410Z digest=sha256:1515f020ac17adca08f23a449bb87087deac6b3347ed3af8e33c1cd3d5a67ccf

Observation e639179c-8052-4189-8f40-7fd587b4e7c3 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:20.639077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:f1ae4ecf12f912675f045a2de3edef37c3bad69f854a12775ed3bc0f1e761f96

Observation 0a37124e-a73b-4485-b966-0c8d1371122d · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.596352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:df54eb298eece7634bb142b3ac13af1dc25caaa1722bfed9a99e995e3bc38538

Observation 3113e5cd-019c-4b5e-9caf-e93928e905c7 · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T15:34:37.002001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:34:37.002001Z digest=sha256:e6a68d838bcebebc688d26051e504c1245b517fc9cda666bc12bf7242e33a350

Observation 842d6468-15ae-4852-8695-f90881a5b040 · inbound

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs cites this paper.

Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T18:39:24.915547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:39:24.915547Z digest=sha256:06b47a8441cb98b255cc153644b2fdcdc86210d4253301c3c8b96b2569168b28

Observation e62a9e4e-ff01-4793-b538-68d25de14127 · inbound

PEEK: Picking Essential frames via Efficient Knowledge distillation cites this paper.

PEEK: Picking Essential frames via Efficient Knowledge distillation Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.314765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:46:53.704892Z digest=sha256:c1c480bc1d8f179633ed0c045143a06fbfa5ea32fe49851649aad94e09883c4a

Observation 1b33d8c6-4bb1-49b1-bbfa-ae9e82327cbc · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 263

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.920378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f0394f594d86fdda83c0473a5832be125c3d52c2e084b50415210c58ce903a5b

Observation c9dbe76f-b26c-4668-9b61-2ec3dea2696c · inbound

Agent-Computer Observation Interfaces Enable Dynamic Computer Use cites this paper.

Agent-Computer Observation Interfaces Enable Dynamic Computer Use Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:04:21.211164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:59:20.295818Z digest=sha256:265bffe4c0229b9912747aaa7c41b41aa1a8e878b4790170b7d8b3d3dff9fb3e