Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T03:51:25.396887Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 87 inbound Pith citation observations for arXiv:2408.10188.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T03:51:25.396887Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T16:42:55.931881Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T10:48:02.935874Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cb1dc0fc-5209-47a9-9942-7ae5ee921114 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 56094f5c-d9cd-42fb-a593-1b42955fa160 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos RT-1: Robotics Transformer for Real-World Control at Scale
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 54bf625c-a741-4d60-81c8-4c46b7298ba1 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 90b7ab22-bff7-4ed9-8eb9-7a162f75ebdc · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Language models are few-shot learners
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 37de3bd8-d552-455b-8ae0-d85ac3b9b3dd · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 32defd73-02ed-41fe-8e63-0400eeb118d8 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 53a3aed6-ce6c-4011-99c0-46e83767ebd5 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ea921a89-baf1-462c-9244-d255d3a4832d · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos PaLM-E: An Embodied Multimodal Language Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d6ac9486-1e28-4040-97af-101056ef9f39 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Towards Event-oriented Long Video Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a9abefde-f27c-4d42-99c0-da002f77df0a · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f2405f20-9ccc-4d71-a894-2529394534bd · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos VILA$^2$: VILA Augmented VILA
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8ce1ff7b-9ec1-483c-9524-5056486179ab · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9cbf8cc3-7eff-4bb4-9520-b3411ddf9d44 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 170e0518-604d-4cd6-a103-28c2ad1cb0fb · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 46e5f56b-e178-4905-841f-e22a74343bf8 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b1216bd9-e3ee-4891-9ce7-10f6d56fc121 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 80893e8c-77c4-44cb-bc06-d1d12a13a3b0 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fa92da15-b858-46b0-976a-a1aa3a29e093 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2680417b-f22a-4d16-9b97-a82fac67b69f · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos LLaVA-OneVision: Easy Visual Task Transfer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7f547896-5c53-4ada-a1a9-90af91d4c2f1 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5154cb12-06db-403b-9760-a7a14a26b185 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 871aeb76-d8d1-44a5-bdfc-d6499bad7856 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d3b4c877-6572-4ee1-9d34-78eaa540acc5 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 765286f6-65f9-4b05-abc7-e4e392fbc8c1 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d0e5c6bd-88f0-4615-8af4-d2bfd8d928f9 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos LongVLM: Efficient Long Video Understanding via Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8f25afe6-c526-4608-ba71-a17de0a9feac · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 13225eb0-71a5-49ab-a703-09f0e60b2f0d · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a786d3c0-69ef-421c-8c9f-00faf4e84fca · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos X-VILA: Cross-Modality Alignment for Large Language Model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 16ce91fa-49dc-4fb9-85aa-a081ec43c810 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8bd723d1-b7bb-4e9a-b7fd-adcf22ab438a · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 346a20b5-cf8a-4d4b-9201-da914dea3e20 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos Specifically, the average scores rise from 2.00 to 3.26, highlighting the model’s enhanced capability in generating accurate and rich captions with more frames
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3d3d13dd-38ff-4899-8d5f-5dd411117d62 · outbound
LongVILA: Scaling Long-Context Visual Language Models for Long Videos We found that FSDP offers more efficient memory management, which led us to select it as our default configuration
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 94698ba4-ab28-4e1a-8f5d-e9bab638eed4 · inbound
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b516edaa-cfb4-4a59-ade6-277f40e74f61 · inbound
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b221bf69-f184-4f8b-b50b-59585d1ed87a · inbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24a736c-29db-4215-baa7-c81b097b72d5 · inbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c7d9bd-4841-488b-97db-f65a854fa955 · inbound
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 071fcb59-3f8e-4437-acd7-d53341c590dc · inbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea14d6be-01bf-48f9-b97d-92803853f4b6 · inbound
NVILA: Efficient Frontier Visual Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7d32a76d-932e-4a78-aa7a-52758bd63697 · inbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1565d455-bcb9-44bc-bbf5-f81d2e97bc0d · inbound
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 139
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18562d97-8e0a-4fd6-ba9e-ff9b5af20b43 · inbound
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9587f3f-bdc5-4272-8752-2b9c817eaaea · inbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242a68af-21f3-40db-b0b2-456ba3e78a2a · inbound
GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f1a4405-493e-49cb-9d05-c54036e8e2ef · inbound
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a100e30c-9c51-4c45-9148-50e2b78a265e · inbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b9e8cedf-e6b2-431a-b91a-31560e1cfde3 · inbound
Cosmos World Foundation Model Platform for Physical AI LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 227
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6a07cc5f-fb89-4dc3-ab11-3c6d811ce8cb · inbound
LongViTU: Instruction Tuning for Long-Form Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9383ea64-b995-442b-96be-1c8be2918c18 · inbound
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2170d37-63ab-44c0-b67a-6e242c29aca4 · inbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d942cf6c-18c6-4c23-948e-73da7d12d445 · inbound
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01335963-1e98-4664-9da6-dd7359c035fd · inbound
CoS: Chain-of-Shot Prompting for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab372280-d145-4bd1-91db-11abde4aaaf2 · inbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5387d609-cfcc-483f-b3da-0427aa9013b8 · inbound
MAGI-1: Autoregressive Video Generation at Scale LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f52197b2-3d0b-4a6e-8276-e907aa5a7935 · inbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 504d6f4d-54f0-4a77-8703-264b2636ec34 · inbound
Clapper: Compact Learning and Video Representation in VLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cccd9eeb-d8cc-44b5-a363-03d7c376bcdd · inbound
Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ba363a-5032-4daa-a114-44dd7d18c80d · inbound
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe79650-f316-4e3d-a8fc-6f3cc3a44e93 · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2509b013-fa07-4a8d-a82d-2f87c2901745 · inbound
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 27e4ed29-6118-42b5-b76a-559978453cd2 · inbound
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a27a710-b5f4-41e5-9d5d-e78a4c032497 · inbound
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fbfa54c-6909-484a-814c-ee2ecabe4087 · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 388f79e1-50ab-4210-b808-69698edb3de7 · inbound
EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b80eb3c8-54d4-45d1-975d-01eed4caa809 · inbound
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e0288bf-ef1a-4adb-b60c-c54ecd6bf42a · inbound
UNIC: Unified In-Context Video Editing LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90256300-b541-4ad1-a4c1-06973ecb44f1 · inbound
TextVidBench: A Benchmark for Long Video Scene Text Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea34d88d-cd06-4fee-838a-1704492c1e8c · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38da4126-715b-4717-a507-f8506b4447bf · inbound
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e82aa700-434c-4361-b9f2-dc66b162c905 · inbound
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef0ce03-ebad-4ebd-8418-b74ff83892fe · inbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b2f519-101d-44cc-af46-640da2b1c2df · inbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c29c6b-51ce-4a07-9a67-99f57cf17501 · inbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5a20d2-bc5b-4ea3-b24e-2cd0facb5bf6 · inbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee615172-d0e8-4d34-9f80-f1e44dcfb74a · inbound
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 339e7bb3-dc67-4a6f-8af4-aff75a9dfa3c · inbound
Infinite Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c56e3819-24d1-4035-b1c7-27b463c20ac3 · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c43c753-2702-410d-bea7-f066b6126247 · inbound
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09ded1e6-3c85-45f7-881a-d26649f03049 · inbound
MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1de4e52f-cc90-4a48-85e1-263603334207 · inbound
Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efb52306-5104-48a7-aef7-4735da755979 · inbound
DATE: Dynamic Absolute Time Enhancement for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb80cbe-af67-4928-af62-45ab11c8a3cf · inbound
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 53f3abd4-b80e-44cb-9b3c-03b818c2f57d · inbound
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 121
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0359ec22-6c24-4a19-9de9-8416ba0dd06c · inbound
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c6e27b-b2b5-4a5e-9fd6-39b89f3ecc79 · inbound
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 191946bd-210a-474b-a7e3-1cc0d365ad4c · inbound
Internalized Reasoning for Long-Context Visual Document Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 71bcfe3c-955f-4c7d-8ce3-230af1bdf7d3 · inbound
Internalized Reasoning for Long-Context Visual Document Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f4e2245-1fdb-41a4-a997-73f0a05b38f7 · inbound
Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ebdf246f-1770-4559-aedb-8b6979b1ba8a · inbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cd1c2ccc-c33b-4c39-87d2-0714eabfa0bc · inbound
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 74a47e25-cc35-40a5-87bf-6e7caa239538 · inbound
EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ccde0297-1dc5-463d-a8fe-859d7fd49494 · inbound
EgoSelf: From Memory to Personalized Egocentric Assistant LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bc06bbfb-44d3-42d0-82f5-508ddf1552ef · inbound
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 046bbe7a-7f19-48b8-99b0-d3d23e3ab290 · inbound
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation be87e1d2-e415-4020-937e-7552ae94f7f5 · inbound
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ab0367f3-5bfe-4ba2-9448-26f708c880c0 · inbound
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1e989343-92d8-493f-8b4a-b669ce1a35b8 · inbound
CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 029e2748-8498-4793-8a2a-405c3a0a1f05 · inbound
CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 31788ae7-27eb-43cf-8cd2-821f2e3b9366 · inbound
CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a1b24b-8d8a-49dd-a613-dd3c3806ec66 · inbound
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c5c71c59-4de5-49d0-aa0b-a0df6e477d6c · inbound
Swift Sampling: Selecting Temporal Surprises via Taylor Series LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 07f73bc6-f0ff-44e6-ac0a-f94993f37b21 · inbound
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0130ce51-e65c-48b7-abb2-c525fe0a0439 · inbound
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4f05b4b-83af-4181-bc86-8bef9461eabc · inbound
Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8e6a09a3-8e35-4240-b24a-47ed9e3baea7 · inbound
UNIVID: Unified Vision-Language Model for Video Moderation LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bcd0124d-f8db-4577-b579-82827431bebb · inbound
GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e9005378-3d41-4bbf-8e38-b64a4b307b1e · inbound
CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 74c4334e-a3ac-4704-9e27-81e3771cb963 · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 250
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 93d4162e-506b-402c-82c8-b3abd1f5afca · inbound
VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e5a4a78c-c2e6-4cd1-b2a9-ad8f673a58f8 · inbound
STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 486bd35f-d07e-44b0-a4db-4a328a3b95d8 · inbound
Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31af0a0-7e39-488b-9e68-b986d5e74fa7 · inbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8de799c9-32c9-4146-a16d-4d1b0ec32444 · inbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4e6034e-a351-4ec5-a00d-a4b7e8a57876 · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 132
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e611d802-01d6-4310-a637-78b7910859ed · inbound
TimePLE: Rethinking Temporal Representation for Video Temporal Grounding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a34c98e-320c-4300-afb0-5367261abcf7 · inbound
Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc18cbab-4a4b-40ed-a815-ea95be6e4c11 · inbound
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cec09a60-c794-45e3-8958-0ab05708261a · inbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a542cce-ab43-4480-b6da-baf3ac9d6403 · inbound
Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.