Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:38:49.654073Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 3 inbound Pith citation observations for arXiv:2604.08120.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:38:49.654073Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T12:34:49.503345Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T16:39:58.308940Z
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5016fc86-0c22-4adc-9087-09269cfb3ac9 · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15654cec-49c7-4dcc-9a56-4e5efd52163e · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23dc9d6e-a508-41c4-8002-33e500dbf78f · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Token Merging: Your ViT But Faster
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ebdf246f-1770-4559-aedb-8b6979b1ba8a · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bbb71538-5646-4cfd-a29b-a210acff541f · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bac8378f-1dcf-4ea1-80f5-17e308c2061a · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b5c2f86-b199-40b1-8833-b6aca8febd81 · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2f84104-be0e-42d7-a0af-20076b606867 · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d1c7aabf-cc5c-4e9e-adc7-4a15863bba87 · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e82f47e-e342-4b7d-bd14-ba9df894202e · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b54d1c26-5e3f-447c-9cc0-dd3b4fc58135 · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 272b9161-9183-4ac1-86ec-d7f8351168c3 · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Kimi-VL Technical Report
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a83564dd-7166-4d12-abe7-a59f17ccce1b · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e9615c1b-4e71-41ac-8a87-31575d9d8efd · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f3d8f39-a4e4-4d9a-b896-eb99a658f92e · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1afa61d7-cf5c-414e-8e85-80ff3c064867 · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Video-llama: An instruction-tuned audio-visual language model for video understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d36e27f-b624-4232-8a0a-6456a6de7a6a · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Long Context Transfer from Language to Vision
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 827b03d8-05d5-45ca-aff9-982f7ca22373 · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 37b4dd1a-7f7b-415c-a9b8-f17d61feed4a · outbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Because the sampled frame count for LVBench is fixed atfmax, the resulting upper bounds are4096/1024 = 4and12288 /2048 = 6tokens per frame, respectively
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 11f6c408-8efc-44f9-9362-086b0115b8c7 · inbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Small Vision-Language Models are Smart Compressors for Long Video Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f30a5792-c856-4639-8d9c-b4eeb21cf9ec · inbound
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Small Vision-Language Models are Smart Compressors for Long Video Understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e7564e6-cb44-4e9f-bbf1-0467e2f29157 · inbound
BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics Small Vision-Language Models are Smart Compressors for Long Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.