Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:48:07.259130Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2412.14965.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:48:07.259130Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd9a2d0f-4f81-454b-86d7-ca7e8b50d817 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Flamingo: a Visual Language Model for Few-Shot Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6629e209-1133-4450-b9eb-37a1247dfd21 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 665c0f6b-3883-4577-8704-1dba32b7c9c1 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ddfea5d-024b-46a4-b01a-fb5db9d6ef34 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b93bcddd-58ff-49db-b644-c6af165c04fe · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7e4111d-1e26-4fea-a0c8-528596a46b7a · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f96e29d8-7ca7-4892-82c4-dc466d4740d6 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef7a6140-d6ea-41cd-84e2-b95b63334213 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de6d5932-b665-4bb6-92a6-5f3642cdbeba · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text PaLM-E: An Embodied Multimodal Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed0f38a-da12-4fa4-9290-714370f0ffd6 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e19bf70a-4876-4679-aec0-124eb466a99b · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Mistral 7B
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f4e0708-b829-4151-9a3b-7db13f38056a · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text LoRA: Low-Rank Adaptation of Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c77b2f8-6403-4a6f-8bcc-1d2915e1130e · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text pyannote.audio: neural building blocks for speaker diarization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9533073b-bed7-41ba-8f87-3d69b11d8777 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Qwen Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 918f1167-b37f-4b81-88a2-be8fba7523a2 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text GPT-4o System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f9a7a9-556d-458a-bfbb-362be4585dec · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1587efce-0860-4235-a2f9-c4753eca11b1 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 918c887c-7417-4ee4-8c50-3483b9cf65c3 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d037024-1051-4a3c-a188-cd172d282873 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b37fe7cd-2232-484e-bdc8-a2fada3b1a5d · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text ImageBind: One Embedding Space To Bind Them All
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f95db44-06c5-461b-b93d-2fb0f6537abb · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text The "something something" video database for learning and evaluating visual common sense
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89da694d-c378-4495-bbaf-48d4a02e93df · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Language Is Not All You Need: Aligning Perception with Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed91dd6-bca6-42cd-a65d-adfc38ce79bd · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text The Kinetics Human Action Video Dataset
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2507dc6f-705b-48cd-8927-98653dea6df3 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57a20458-e932-4b76-93c1-4b894ad55745 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text NowYouSee Me: Context-Aware Automatic Audio Description
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d52e13b-9464-4109-8aff-0babfde8d622 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddbc04f6-4c79-4f35-b983-679adfe59a15 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Kankanhalli
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa077860-7905-48e7-87c5-3891165760c8 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text VideoChat: Chat-Centric Video Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5eb1154-0f5e-4f3b-97a6-4c520c7c4bb5 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 234b8113-a46c-43d2-a947-c42230ea32d9 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8279db1b-2f6a-461b-b949-46621b2e2bb9 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7417c79-a758-4d7a-af3d-68c16309da6b · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Visual Instruction Tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cef6858c-35bc-46f4-8f41-19f74e14c841 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text MMBench: Is Your Multi-modal Model an All-around Player?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96e9d748-0816-492a-909a-bb559c1a83cb · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ec3c927-8318-4f6e-9ceb-87ba9fa6ae6d · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caed6f36-bebe-43b8-9854-1642dccd3482 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c9274df6-1dd0-4f2f-ab4c-c2e9eca1e918 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b6ac0b-271d-4707-9b02-efa9b456f11e · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Perception Test: A Diagnostic Benchmark for Multimodal Video Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea0ada39-3cd6-43c0-a629-3b7cd5c75c09 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Robust Speech Recognition via Large-Scale Weak Supervision
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ecd06f-889b-4095-9958-d7e4bc40c222 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 556e9dc9-89ad-4f62-bd1a-07b39e84537f · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text LLaMA: Open and Efficient Foundation Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d677305-df6c-43d0-a5ed-c76897f128eb · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9cfbe0f1-63d6-407a-8981-1ef078e0e4a7 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ce6ac81-ac32-4b5c-ac18-878f9ed0616d · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a182f44f-cc8e-4c4e-b4c9-72f1d94c21d1 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cebcaa12-1d22-47e2-9e6b-74f8d28985bb · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text FunQA: Towards Surprising Video Comprehension
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22177569-1448-45b4-a04e-406c37531f4a · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 543570b2-49ae-4006-8152-d932ba4c8155 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e2f40d3-0671-4b39-95d9-7cfe102ba48b · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db9eb1f-fce8-40db-ab3f-3c67398d39fb · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee578662-6381-4d09-a55b-038b0767eb72 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text What is YOLOv8: An In-Depth Exploration of the Internal Features of the Next-Generation Object Detector
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3856e7d6-087c-4f6f-99e5-df0dc02fc371 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260020f6-a8cf-48d9-98b7-faa274e780c4 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5325b058-cfc7-444b-8ad9-06d1bc24f3c1 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f1a2a0-8a7a-4c6e-8f91-c50e1ee5aad4 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text Distilling Vision-Language Models on Millions of Videos
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9afde8d9-d360-4f23-b161-c02b8ba82897 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4705e8a-b71d-4bba-a84e-73e898a89bb1 · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text online" 'onlinestring :=
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce3b0de5-1f1a-4ee4-8842-fb92cebbe38d · outbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text write newline
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.