Pith. sign in

Paper Citation Record · LEDGER

LongVILA: Scaling Long-Context Visual Language Models for Long Videos

As of 14 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 87 inbound Pith citation observations for arXiv:2408.10188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.10188 v6

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T03:51:25.396887Z

measured 119 of 119 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 87 of 87 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:42:55.931881Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T10:48:02.935874Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact25
  • verified fuzzy2
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cb1dc0fc-5209-47a9-9942-7ae5ee921114 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.426441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:02e18842c73dbf56bc0b2840aff9534237ba94fe776c56aa5567cdc7f2b521bd

Observation 56094f5c-d9cd-42fb-a593-1b42955fa160 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.434788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:cfe763196e0fb994080f15fd4ae1bb146fd71858be97320bd645633a04fa58cd

Observation 54bf625c-a741-4d60-81c8-4c46b7298ba1 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.440421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:80dff92265add98d96965926ed35bfacc0d7569be52973c91bf5f0f2ed271215

Observation 90b7ab22-bff7-4ed9-8eb9-7a162f75ebdc · outbound

This paper cites Language models are few-shot learners.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Language models are few-shot learners

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:51:25.562568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:25afc6f3ec42cf212db0ac35b7e8ed3f17c37132291905696eda96af571683bb

Observation 37de3bd8-d552-455b-8ae0-d85ac3b9b3dd · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.468836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:ea0117c2e293981e328d7560f0b86c2b2d8f29e6b33ca6f8743d9078a833eee0

Observation 32defd73-02ed-41fe-8e63-0400eeb118d8 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T03:51:25.473222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:fc6fa580cac843132af11e5a433ff559dde0b8d9d5182fde1e77db28966eef19

Observation 53a3aed6-ce6c-4011-99c0-46e83767ebd5 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.477865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:2cacf5e737a6d81ed3c6ce5c10e2c53dc6e8d3ee3d7b27ba00e6e848d6015e41

Observation ea921a89-baf1-462c-9244-d255d3a4832d · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos PaLM-E: An Embodied Multimodal Language Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.482472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:893e3a9b29eeb2fd7edbd6676111d33d0d68d058c1684e3383a59f478c8b69d0

Observation d6ac9486-1e28-4040-97af-101056ef9f39 · outbound

This paper cites Towards Event-oriented Long Video Understanding.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Towards Event-oriented Long Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.487005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:e3ca4101fb13e1db52d354ac98c16b90e175cc0c0075095511e0eb7d2cb088f7

Observation a9abefde-f27c-4d42-99c0-da002f77df0a · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.491822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:d3cd0d200e8779ef5095c874d94c5b440fa0aba66c8f7358e5400f895aa2b6b6

Observation f2405f20-9ccc-4d71-a894-2529394534bd · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos VILA$^2$: VILA Augmented VILA

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.496419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:34217996710514f9c513dc6ce7e906c8fafe2669680657956894b418f12e7c08

Observation 8ce1ff7b-9ec1-483c-9524-5056486179ab · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.501093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:f01500e9f1ed50bc246c1c1bca9ad540c6fcb0932f5e3fc7dbc1d3f10c853a99

Observation 9cbf8cc3-7eff-4bb4-9520-b3411ddf9d44 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.506156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:e4afbdc3ae0f5dc8862871ab28a3c4a6684dece843ac7b2fa23d01e309363234

Observation 170e0518-604d-4cd6-a103-28c2ad1cb0fb · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:51:25.511626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:929ecfb29b4380ff98610ca058f1c3043990484d4cd526191d5c61beb9762353

Observation 46e5f56b-e178-4905-841f-e22a74343bf8 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.516837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:37ac0ac9203efb4e099b111f43e55b9f8eab49dea18381c953b7f17fce8b26f4

Observation b1216bd9-e3ee-4891-9ce7-10f6d56fc121 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.521450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:bcda505f7e60b0af9253be478a11b8cf48fb37ca1e67959270e422f925b5e1e3

Observation 80893e8c-77c4-44cb-bc06-d1d12a13a3b0 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:20:36.653799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:276ff950d150145303fee6494b6d74d1f52782f40c2ef6fc390b06e302327ead

Observation fa92da15-b858-46b0-976a-a1aa3a29e093 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.529874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:333ada9f3be52fb382518aebcd94b8f4f8fc159aee88040108005fe5f2fa6e9f

Observation 2680417b-f22a-4d16-9b97-a82fac67b69f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.534046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:c695494fd756da767fa9060015df5e25598f3fde8b1c09bb2dcf91bc6b53bcbe

Observation 7f547896-5c53-4ada-a1a9-90af91d4c2f1 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T03:51:25.538073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:f310beff5ec2b943620ae6e1441410b02d2d5daa5ec9f85ff9794144fff36c5c

Observation 5154cb12-06db-403b-9760-a7a14a26b185 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.542171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:8de494c1c5838166e6817f364a452572bb78c51f89fd65967a6265d24ba8c74d

Observation 871aeb76-d8d1-44a5-bdfc-d6499bad7856 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.546496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:c004a7e6a6b0f07b7555e65d49e1b68bf5aeac539b47e70dbb7c2e4d3ab17fdc

Observation d3b4c877-6572-4ee1-9d34-78eaa540acc5 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.550887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:3d435b8b386293e35c21f6b93d2bfedd30f0255702797f6433993bdafa68a796

Observation 765286f6-65f9-4b05-abc7-e4e392fbc8c1 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.554718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:f7306ac6aa5c1054da2ef06049bb550cc448087f73d0275c02513e306ef3b36b

Observation d0e5c6bd-88f0-4615-8af4-d2bfd8d928f9 · outbound

This paper cites LongVLM: Efficient Long Video Understanding via Large Language Models.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos LongVLM: Efficient Long Video Understanding via Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.559116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:8b4207e20f83dfe29ee9176736a19375e9d6bfc27749a457df1ea611007c1516

Observation 8f25afe6-c526-4608-ba71-a17de0a9feac · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:12.519784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:7824b29d95fc6e72d6da99145f38a88f61eb8fe01f153dd77c8f874ab67aede9

Observation 13225eb0-71a5-49ab-a703-09f0e60b2f0d · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:25.452760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:283bd9bb4a8b76c2d76d3a44000bad67b72af86b489edd511ecdba73c529df59

Observation a786d3c0-69ef-421c-8c9f-00faf4e84fca · outbound

This paper cites X-VILA: Cross-Modality Alignment for Large Language Model.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos X-VILA: Cross-Modality Alignment for Large Language Model

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.458475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:e14826f3ffa8d71e28de3522bcde96ccb797430bbe50da6cc32e90dd433f7349

Observation 16ce91fa-49dc-4fb9-85aa-a081ec43c810 · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.463933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:fe750b7e2b3cbe18d0e4f84a918d436fec3f41cb4bc3bf047bb7e3da6b9239e4

Observation 8bd723d1-b7bb-4e9a-b7fd-adcf22ab438a · outbound

This paper cites an unresolved cited work.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-17T03:51:25.568554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:940b16c691253b4f6c22c4f589c68e6bfa9688fec6c88b0b71af48e4bf3fb878

Observation 346a20b5-cf8a-4d4b-9201-da914dea3e20 · outbound

This paper cites Specifically, the average scores rise from 2.00 to 3.26, highlighting the model’s enhanced capability in generating accurate and rich captions with more frames.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos Specifically, the average scores rise from 2.00 to 3.26, highlighting the model’s enhanced capability in generating accurate and rich captions with more frames

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:51:25.571713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:c1af1f112f87f8bde4213158cd54e8df8e8b58d7189411315d72c4e156a727d9

Observation 3d3d13dd-38ff-4899-8d5f-5dd411117d62 · outbound

This paper cites We found that FSDP offers more efficient memory management, which led us to select it as our default configuration.

LongVILA: Scaling Long-Context Visual Language Models for Long Videos We found that FSDP offers more efficient memory management, which led us to select it as our default configuration

Reference 32

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T03:51:25.565650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:51:25.396887Z digest=sha256:320e4aa7e5a5f24a357ef7c00b3b7c87879e9e06c7e38d2d63c6bc4666ec4b5d

Pith citing papers

Observation 94698ba4-ab28-4e1a-8f5d-e9bab638eed4 · inbound

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation cites this paper.

VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T16:42:55.931881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:42:55.931881Z digest=sha256:ef4da62d7e522f128d249db2a54c7bc5150997c8b661669f12f8de5d311f2a1c

Observation b516edaa-cfb4-4a59-ade6-277f40e74f61 · inbound

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching cites this paper.

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:44.137558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:44.137558Z digest=sha256:23554d6786bfc0a276b283d2e6b03e27ddd37e41e1d557ee429e1f71cf277554

Observation b221bf69-f184-4f8b-b50b-59585d1ed87a · inbound

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis cites this paper.

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:04.502907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:04.502907Z digest=sha256:a99dadc5d3f9e724113c3c86a3024ce319ab040527d8cd32c26b6b67863edcef

Observation b24a736c-29db-4215-baa7-c81b097b72d5 · inbound

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding cites this paper.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.872055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.872055Z digest=sha256:f30f4d378fa9fbbdec8826a4eaf56172e43c3cec01e265dc154edcadb86e8eb5

Observation e6c7d9bd-4841-488b-97db-f65a854fa955 · inbound

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation cites this paper.

VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T04:56:01.474234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:56:01.474234Z digest=sha256:b8a5682fd307a5b3286df0c5719c53b082e656ddc460a8ab93414237f4aed3d5

Observation 071fcb59-3f8e-4437-acd7-d53341c590dc · inbound

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning cites this paper.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:34.850695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:34.850695Z digest=sha256:91309d270cf814739e9168b8e6a69e95db8ad3f82763fa4c1b4badcb2abedca9

Observation ea14d6be-01bf-48f9-b97d-92803853f4b6 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.247372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:625b877d91d66a84fe447305c47f555a77fedd4cb325383a31f082731387ea96

Observation 7d32a76d-932e-4a78-aa7a-52758bd63697 · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.556012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.556012Z digest=sha256:8e46930b32a6e315425b00565c19a81560709acf3a624c6b5d61f1ea404491ee

Observation 1565d455-bcb9-44bc-bbf5-f81d2e97bc0d · inbound

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding cites this paper.

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:03.571836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:58:03.571836Z digest=sha256:4da4e70c25d956f522f0a7cf91ff1d72cb15b1c1fe32f4869b8115ef89b8f5e6

Observation 18562d97-8e0a-4fd6-ba9e-ff9b5af20b43 · inbound

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs cites this paper.

IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T16:39:51.002585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:39:51.002585Z digest=sha256:f1e7768ca91ebdc60ed6f10c5568c8662e1ea8b2b4c4605a427cdfda2b6352d1

Observation b9587f3f-bdc5-4272-8752-2b9c817eaaea · inbound

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding cites this paper.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.041669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.041669Z digest=sha256:39115d91bd620a558359ed7482aa9ed02d08abdd0fd4522357c31d1f834a703d

Observation 242a68af-21f3-40db-b0b2-456ba3e78a2a · inbound

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models cites this paper.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.193369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.193369Z digest=sha256:2f70dc2530cb46fe9b07625e1ca4db4c362f8cc793bf3c3fb62934c54fdebde5

Observation 5f1a4405-493e-49cb-9d05-c54036e8e2ef · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:27:43.998546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:fe68a5a31187e82f6f783d9ba365a47214b16a211a19d63b19c47751434183bb

Observation a100e30c-9c51-4c45-9148-50e2b78a265e · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:02:43.569786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:ed322ca681a24dba4ed03d93cfddb366239432b2c856e42a41bf6b6fed03765d

Observation b9e8cedf-e6b2-431a-b91a-31560e1cfde3 · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 227

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:cde08e6fd60398b969cce44078787a57781f8809e544d5c684b653c8b612dac6

Observation 6a07cc5f-fb89-4dc3-ab11-3c6d811ce8cb · inbound

LongViTU: Instruction Tuning for Long-Form Video Understanding cites this paper.

LongViTU: Instruction Tuning for Long-Form Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:58.026866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:58.026866Z digest=sha256:9db551feaf9bd8eae17a45a9ae18f25e81c7f22d732366bd99b083295bbc2c60

Observation 9383ea64-b995-442b-96be-1c8be2918c18 · inbound

Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models cites this paper.

Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:22.822399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:22.822399Z digest=sha256:c8d4bcf03cf0b0e930d4471fa4c24402d7e71adc6fc98fda3160dc63651d6658

Observation e2170d37-63ab-44c0-b67a-6e242c29aca4 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:d5e7fa31e5f80d00b4547e90df2907cfb3be9bc20fce05e781dbf26a36279c87

Observation d942cf6c-18c6-4c23-948e-73da7d12d445 · inbound

Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile cites this paper.

Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T16:36:00.816958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:36:00.816958Z digest=sha256:fd0f55886e2de887ee35571046fb9f6731bf6d5dfe59f3d3685d521aeb524fb5

Observation 01335963-1e98-4664-9da6-dd7359c035fd · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.983567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.983567Z digest=sha256:75353aa47b0d58fb19dd3922c237b8e406bd2b37682668fd380a814359ad7c8e

Observation ab372280-d145-4bd1-91db-11abde4aaaf2 · inbound

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid cites this paper.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.741468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.741468Z digest=sha256:c9169895d0dbb818874a13c359bd8dde3ea30446e663976916345b09898507d9

Observation 5387d609-cfcc-483f-b3da-0427aa9013b8 · inbound

MAGI-1: Autoregressive Video Generation at Scale cites this paper.

MAGI-1: Autoregressive Video Generation at Scale LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T20:31:15.700943Z digest=sha256:7f1f541c36bf1fb46c42da70a5035e3f9412382cd73860c286e13c8f5b089f03

Observation f52197b2-3d0b-4a6e-8276-e907aa5a7935 · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.848837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.848837Z digest=sha256:1b27ada110757e1024baa192c12a69932ee38335c1548bab039f9c308b75369f

Observation 504d6f4d-54f0-4a77-8703-264b2636ec34 · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.582468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.582468Z digest=sha256:ed090d97e982ea5c6a8b93e97e2c0e6ecc15cd76a7e5f74f06e9e9d5f6c54cb8

Observation cccd9eeb-d8cc-44b5-a363-03d7c376bcdd · inbound

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM cites this paper.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.159249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.159249Z digest=sha256:b9a2b542ad14f3f824ea2dca5442f3d382c69f9ba15e03ef85291bed9b4e10be

Observation a4ba363a-5032-4daa-a114-44dd7d18c80d · inbound

EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection cites this paper.

EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:02.442537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:02.442537Z digest=sha256:27b43f31867a38668931b95d7ba8d7ca5cdd1cf816917e2e826b96bbbdd7d5ea

Observation 5fe79650-f316-4e3d-a8fc-6f3cc3a44e93 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:87b85b28d625236b7afb46918fafbcc263e55a0c1f40e204986b0a2e3b00844d

Observation 2509b013-fa07-4a8d-a82d-2f87c2901745 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:00:51.343379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:36b28dbc5541ce9b0ed310c75dd7593d5bde92d36e5cf25660a48bce4b6864af

Observation 27e4ed29-6118-42b5-b76a-559978453cd2 · inbound

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders cites this paper.

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:16.941581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:16.941581Z digest=sha256:b727254e47ebb903b1757f8cd8ac1a6307f987887187aa09a5fec88be3ee1686

Observation 1a27a710-b5f4-41e5-9d5d-e78a4c032497 · inbound

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames cites this paper.

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:19.402144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:19.402144Z digest=sha256:5483bd6cc3fd4626d910989d16e061ec31429d6b6806493486a1e1612d00c0bf

Observation 6fbfa54c-6909-484a-814c-ee2ecabe4087 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.898445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.898445Z digest=sha256:b844f5a1bba356bc51ae62f32c45347eebbfbb702984bb2ba9fd70a876e7458f

Observation 388f79e1-50ab-4210-b808-69698edb3de7 · inbound

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models cites this paper.

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.565742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.565742Z digest=sha256:2582e1c5a79dd4a169fec15ff7cf1151cb1081e514ef2451ca1b3d41654176af

Observation b80eb3c8-54d4-45d1-975d-01eed4caa809 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.389890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.389890Z digest=sha256:55129ced2d3d3a82b525d053c1bdc980581c8a8e42705063185344792c9f16d4

Observation 6e0288bf-ef1a-4adb-b60c-c54ecd6bf42a · inbound

UNIC: Unified In-Context Video Editing cites this paper.

UNIC: Unified In-Context Video Editing LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:42.980673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:42.980673Z digest=sha256:f694215f9f18798491fbf14defaac97fc1d97fa8ff203ca4d8cbeb5e5d324c02

Observation 90256300-b541-4ad1-a4c1-06973ecb44f1 · inbound

TextVidBench: A Benchmark for Long Video Scene Text Understanding cites this paper.

TextVidBench: A Benchmark for Long Video Scene Text Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.818281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.818281Z digest=sha256:e034444dda355bab198820fc877e421729fd84f836a8d7c1826efdede9c1f57a

Observation ea34d88d-cd06-4fee-838a-1704492c1e8c · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.411481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.411481Z digest=sha256:6c402d231b2d9d3247263039ab818f5a5b6f29cb3ec6c14564f66bdff852c6d6

Observation 38da4126-715b-4717-a507-f8506b4447bf · inbound

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs cites this paper.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.765313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.765313Z digest=sha256:2b809f2e3887821cf201a73dd93f23fce6878b1e4dde197a7f834584a247b261

Observation e82aa700-434c-4361-b9f2-dc66b162c905 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:28.324818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:28.324818Z digest=sha256:fda37953036daefc10130a71a236c866fdea7c21bcd99316bfecf0b436a4ad47

Observation 4ef0ce03-ebad-4ebd-8418-b74ff83892fe · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.317216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.317216Z digest=sha256:9430c8e32195b8d83df583cf3abafc42618ba320081fd2886e56577888285c7f

Observation a4b2f519-101d-44cc-af46-640da2b1c2df · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:04.382683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:04.382683Z digest=sha256:7a6bbe536adc01e3a69c936f74e86b7ab0e7ba8da830aa04a286a111ebf0baf6

Observation 93c29c6b-51ce-4a07-9a67-99f57cf17501 · inbound

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams cites this paper.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.873755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.873755Z digest=sha256:bd27ec238497b53a211cc29195f3b7ce3d8788eaeb9d1366f464aee449339bea

Observation 7d5a20d2-bc5b-4ea3-b24e-2cd0facb5bf6 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.468866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.468866Z digest=sha256:eb941eca8cfb16b7549edfe8b32f0128ac7c8eb71f153c0cbbb0b3d0628d500d

Observation ee615172-d0e8-4d34-9f80-f1e44dcfb74a · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:da4272635dc1202aecc006dd57d43b641cb03c5388c1bbaf3952a7cd50a5a486

Observation 339e7bb3-dc67-4a6f-8af4-aff75a9dfa3c · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.313230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.313230Z digest=sha256:216787689d7533d3a1a724374ee80cca2c8c144ef21fba092dbbbc0717581c53

Observation c56e3819-24d1-4035-b1c7-27b463c20ac3 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:10.656199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:10.656199Z digest=sha256:ac0e1a6e17caa3a93fa6354f191b33b2ef99eccdc184d40fdabbc340523595fa

Observation 4c43c753-2702-410d-bea7-f066b6126247 · inbound

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding cites this paper.

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:16.712705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:10:16.712705Z digest=sha256:57cffdf8ea67bd97ea35fcfac1c8cba479246535449ecff3ad20efc0289b647e

Observation 09ded1e6-3c85-45f7-881a-d26649f03049 · inbound

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs cites this paper.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.525180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.525180Z digest=sha256:98f4828b09b250b29104dd678e9019353b73c104cc511547da9b39898fd95edf

Observation 1de4e52f-cc90-4a48-85e1-263603334207 · inbound

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment cites this paper.

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T13:54:52.015747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:54:52.015747Z digest=sha256:376875f5310722f5456fd27e8c2205666a43865e9d27454da6da5ee5fc692d93

Observation efb52306-5104-48a7-aef7-4735da755979 · inbound

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding cites this paper.

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T19:28:34.352539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:28:34.352539Z digest=sha256:12b109c07f6bd8f91ab8c7cf92ccb756f4447e4d089f44ff23212ba85bc7c9e5

Observation bfb80cbe-af67-4928-af62-45ab11c8a3cf · inbound

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning cites this paper.

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T12:17:42.135851Z digest=sha256:14ae805a80b14847f2a2ca31a8d6f0eba2fe4601052e0452e495dedebad63582

Observation 53f3abd4-b80e-44cb-9b3c-03b818c2f57d · inbound

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources cites this paper.

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-03T06:50:28.622236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:50:28.622236Z digest=sha256:e3609d45251020fda969317f534e933c4384c5e07d4dd589b0c95af91edb60ae

Observation 0359ec22-6c24-4a19-9de9-8416ba0dd06c · inbound

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding cites this paper.

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T17:54:04.513726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:54:04.513726Z digest=sha256:42b0bbcbbc8248831277231f88b2b80ff65f0653d67d513b3281aeeb920bee65

Observation 10c6e27b-b2b5-4a5e-9fd6-39b89f3ecc79 · inbound

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning cites this paper.

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T22:04:18.591594Z digest=sha256:74f3dfcfb73a66085d47568e779bc415e18c51d0d0297049a5a03d340d8b089e

Observation 191946bd-210a-474b-a7e3-1cc0d365ad4c · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T23:53:19.148407Z digest=sha256:d6cbe5e676a02027d4510f7c76ec8e6a2678ed0977eab4994bf4c9803d9231ae

Observation 71bcfe3c-955f-4c7d-8ce3-230af1bdf7d3 · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T15:50:49.083652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:50:49.083652Z digest=sha256:8fae8c4cd623d2ae2e8b9ba8ac6b37ccd764085e8350aeaaeff6b6c422baabda

Observation 9f4e2245-1fdb-41a4-a997-73f0a05b38f7 · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:66de9821517fcf431e69d8ed838d74a3e561d4b75e59c0d2398ecb64eaa606b4

Observation ebdf246f-1770-4559-aedb-8b6979b1ba8a · inbound

Small Vision-Language Models are Smart Compressors for Long Video Understanding cites this paper.

Small Vision-Language Models are Smart Compressors for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:deb30b8682d2a45f4cba88f02155ab6af7a4c367ae2dad485cf4dff603fbdc94

Observation cd1c2ccc-c33b-4c39-87d2-0714eabfa0bc · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:5ab3362d0e1346cbce6746d6890a72002be518152482e94dcaf03503e35b3376

Observation 74a47e25-cc35-40a5-87bf-6e7caa239538 · inbound

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling cites this paper.

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T07:00:36.870817Z digest=sha256:f50cd2274824bec1841ff65cea22108cd52acc46f299653c89b092c6e1abe985

Observation ccde0297-1dc5-463d-a8fe-859d7fd49494 · inbound

EgoSelf: From Memory to Personalized Egocentric Assistant cites this paper.

EgoSelf: From Memory to Personalized Egocentric Assistant LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T02:23:21.119521Z digest=sha256:ea9dda6cbc7f14d7a699c1e1ab6553f256ed45e5f9633f12f3658585e856f210

Observation bc06bbfb-44d3-42d0-82f5-508ddf1552ef · inbound

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing cites this paper.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T01:30:19.531699Z digest=sha256:db2e240e5355b7440f82e523ed5e4c5ae4b4e0bbad692a6f9be3001130d8db49

Observation 046bbe7a-7f19-48b8-99b0-d3d23e3ab290 · inbound

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing cites this paper.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:92ec05f0a805428a12ffc6717f01972d20f3bdd6bb4ccefa1ebef2cb303f56fa

Observation be87e1d2-e415-4020-937e-7552ae94f7f5 · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T14:48:39.444933Z digest=sha256:4aa251383079550875733e2bd5d0330f16128e4bd5b3afd8821a97d9c40e215b

Observation ab0367f3-5bfe-4ba2-9448-26f708c880c0 · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:57:42.822121Z digest=sha256:6355f205ed86b2de74c1a466aa109dcc3ebc394fc3a6b224b68933dd651b2633

Observation 1e989343-92d8-493f-8b4a-b669ce1a35b8 · inbound

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding cites this paper.

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T03:06:09.753634Z digest=sha256:4a8cff19b3b1458dfd5e86b463f09d93ed04aa44b4824200f70a42de6afe9b06

Observation 029e2748-8498-4793-8a2a-405c3a0a1f05 · inbound

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding cites this paper.

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T13:45:46.088920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:47:19.742035Z digest=sha256:259e855e5692b56b95a94f3a3cb06ad55dc25cb610a1f9daee2f65219ae13bc9

Observation 31788ae7-27eb-43cf-8cd2-821f2e3b9366 · inbound

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding cites this paper.

CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:18:46.366043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:18:46.366043Z digest=sha256:162ab02c161c1caeaf2e9d304ef7f18cfb1d85f6831703e22ac840d4c41eab48

Observation 47a1b24b-8d8a-49dd-a613-dd3c3806ec66 · inbound

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context cites this paper.

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T19:16:07.851098Z digest=sha256:dd13199186542ca8788b431ff8eab342daadd46757eedd62998fa7e70099da61

Observation c5c71c59-4de5-49d0-aa0b-a0df6e477d6c · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:56:07.961469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:a1bdc900fb2fb33aa19cb48ad3b00e7df689d3ffeae6deb0cfa173a83880b7ac

Observation 07f73bc6-f0ff-44e6-ac0a-f94993f37b21 · inbound

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning cites this paper.

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.945623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T18:15:28.261185Z digest=sha256:f76409db061e3e64607517cb5eb212a21bc2b2ac073c37e7f4d3a26f6eeb2cdd

Observation 0130ce51-e65c-48b7-abb2-c525fe0a0439 · inbound

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning cites this paper.

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T15:54:12.909974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T15:54:12.909974Z digest=sha256:1400f1e23533296ecf6dec80b58abd183b8e759cb16b3b31e356cc88ea27bf69

Observation b4f05b4b-83af-4181-bc86-8bef9461eabc · inbound

Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining cites this paper.

Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T00:02:50.363523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T23:16:42.001361Z digest=sha256:fc138486856942efe7634ccada231f13c773b2b6675dc91d63839013f71e98fc

Observation 8e6a09a3-8e35-4240-b24a-47ed9e3baea7 · inbound

UNIVID: Unified Vision-Language Model for Video Moderation cites this paper.

UNIVID: Unified Vision-Language Model for Video Moderation LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T16:07:09.203772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T22:56:26.674841Z digest=sha256:249007598f75af796c6504eec568e213079c96cc214a1c8edc5b145fd6c4cbaa

Observation bcd0124d-f8db-4577-b579-82827431bebb · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.355750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:ca256f15827196cffad071c1a587c979e96d2ff683ce132a39d6acd2800976cc

Observation e9005378-3d41-4bbf-8e38-b64a4b307b1e · inbound

CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence cites this paper.

CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:57:38.941419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T14:17:08.713863Z digest=sha256:b6b5ca57daa9877c7276ecf1d4b6e749f92116aca087f41ae90fa73937f88cf4

Observation 74c4334e-a3ac-4704-9e27-81e3771cb963 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 250

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:48:02.937090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:979625eb7e2d488270a7f97004268c3a6dee7d5e7f6c5c9fa15a70dda82517fe

Observation 93d4162e-506b-402c-82c8-b3abd1f5afca · inbound

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context cites this paper.

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:04:21.232831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T06:01:49.904752Z digest=sha256:a9e95378353a3f627632e75f6dbb472bacf63d985ae8c24112127ac694ec8b0b

Observation e5a4a78c-c2e6-4cd1-b2a9-ad8f673a58f8 · inbound

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation cites this paper.

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T06:06:47.233814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:06:47.233814Z digest=sha256:77de2b757214b1accd45c1406ecc447113552ce9bc651e1dedc633811683f133

Observation 486bd35f-d07e-44b0-a4db-4a328a3b95d8 · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.056803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.056803Z digest=sha256:f38abbd4a50168d57c8e218d2e2c05db99e62c3101c39b3e36351e7787c7058e

Observation e31af0a0-7e39-488b-9e68-b986d5e74fa7 · inbound

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA cites this paper.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:42.330014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:42.330014Z digest=sha256:d82133ecd3c2d434c2c5ca9155e7c71c529d5267d66426925ece5d602259734c

Observation 8de799c9-32c9-4146-a16d-4d1b0ec32444 · inbound

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? cites this paper.

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T09:09:24.840911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:09:24.840911Z digest=sha256:983811a2329175b44d0f49d79fcc848329ffda1200d7618048d7d653b82887ee

Observation c4e6034e-a351-4ec5-a00d-a4b7e8a57876 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.649073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.649073Z digest=sha256:85bf2c645611447cc8428e611baeb8ff42f68bf731327731d9781671167379c1

Observation e611d802-01d6-4310-a637-78b7910859ed · inbound

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding cites this paper.

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:49.920089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:32:49.920089Z digest=sha256:2830e64cc5eb28c3ec7642e826ea3fcbd9a8debd693eb12189a5de33e431bd57

Observation 5a34c98e-320c-4300-afb0-5367261abcf7 · inbound

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding cites this paper.

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:00.531111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:00.531111Z digest=sha256:4f20f6bd8fab8edb2231843c923769ca828c95a8a3500c453bdabd44682b0dd7

Observation dc18cbab-4a4b-40ed-a815-ea95be6e4c11 · inbound

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding cites this paper.

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T00:45:17.837663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:45:17.837663Z digest=sha256:0522dbae5257f2424f7bbe01337ecb467cf0ca63254e1797db8e8d92c5332df7

Observation cec09a60-c794-45e3-8958-0ab05708261a · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.611444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.611444Z digest=sha256:946f8a4e65ddcfa3744f5e60fc807333b0c500e6ce078e6ce48242c91b3f29fe

Observation 5a542cce-ab43-4480-b6da-baf3ac9d6403 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.091644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.091644Z digest=sha256:d4e3eabc84baadcd2408acfabaf65083d87903675b7084fa8d89e01a6086e17f