Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2407.15841.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:44:05.836761Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T03:19:29.898625Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 05397695-4940-4979-8b87-c9bb464a2e02 · inbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 213
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 38b01930-0c6e-4c15-8049-3efbe895413d · inbound
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d96dadec-8a78-48f9-950f-320a2629dd55 · inbound
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a008773-a7cf-4d40-8821-176c5e918590 · inbound
A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1574aa84-21a1-442e-9207-b8afb5df4826 · inbound
NVILA: Efficient Frontier Visual Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 47d7345f-2cd6-49c2-ae8a-22677989dc35 · inbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d0de7b-e66a-447e-bf0d-8fa8f8cb0c7b · inbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebe804c1-4aeb-4ef2-81e3-a8dc5116d457 · inbound
Apollo: An Exploration of Video Understanding in Large Multimodal Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2645119b-22bc-472e-8c14-06d5518652da · inbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5252730d-c4ac-44a8-9b52-b2d50ffadf6b · inbound
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cdbed8d-6423-4f4c-9c10-2ec71ced3420 · inbound
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d84ca329-0445-4314-8988-8bf20ebd8055 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 147
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8e6a65e6-3602-4ea7-b3d4-6eab3cc07ab7 · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7a46c902-f8b5-4033-88b7-c49360ce7c9a · inbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0121dee9-6d54-4503-89a9-2066e849809e · inbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91fb4939-f0b4-4ccc-a4c0-9fc639f510d9 · inbound
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0a170583-a703-4bcf-9e43-4138824d1160 · inbound
Clapper: Compact Learning and Video Representation in VLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cba55f33-db7f-408d-99fa-78317f5baaca · inbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ca677bd-85df-4f23-bf31-c6e4b77df4bb · inbound
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 489e65a4-c689-416e-a146-ba4a2a8813dd · inbound
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec28918-801c-48e3-a1b7-284e41643535 · inbound
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52342c69-8d84-47d4-908e-36077b7c5952 · inbound
MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53ec807f-a90b-4c6d-a069-f22338da95a8 · inbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e49916f7-c529-4d14-8a85-dbf10b8ef374 · inbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d067ec-0901-4417-ae9a-7df704bc5cd6 · inbound
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b41929b7-75f7-4fcb-b25d-25142ef52b0c · inbound
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e54c71a7-b589-40e4-8698-aa6b8ab7bf57 · inbound
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 164
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e40e8b56-6ef9-4885-8737-9f6dc53236f1 · inbound
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4016aeb4-078b-482d-bb10-a143f7678108 · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 835a5c98-cc5b-483f-a37e-9daf7baccd97 · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5962e419-38d9-4350-b3ba-319120e1c9f6 · inbound
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 746f123b-cbf3-4c36-b8e7-f66d6fec7f4b · inbound
EgoSelf: From Memory to Personalized Egocentric Assistant SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ed0375a-ec45-4b75-ba15-d609e1622694 · inbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1e76134c-26d1-44d4-bacd-125336067b52 · inbound
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 265f85fd-529f-44cc-b304-b4fccd0d919c · inbound
Linear Scaling Video VLMs for Long Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e6c52ccc-1d76-4dfb-a9aa-00019a82e63c · inbound
V-LynX: Token Interface Alignment for Video+X LLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d43f3e47-6db7-4ff9-a3f3-42971a4175bf · inbound
MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e9a935e4-350f-4f6a-90ff-8778fee50540 · inbound
Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cf97c963-6366-4263-830e-b9b627f77316 · inbound
ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 16129752-a127-47ff-a819-6c64e0f6fe8d · inbound
Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ce06e3c-e938-4b4b-ab75-acf4c575595c · inbound
SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7103d0d-f290-430d-9da5-604b896cb835 · inbound
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.