Pith. sign in

Paper Citation Record · LEDGER

Small Vision-Language Models are Smart Compressors for Long Video Understanding

As of 5 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 3 inbound Pith citation observations for arXiv:2604.08120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.08120 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:38:49.654073Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:34:49.503345Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:39:58.308940Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact13
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5016fc86-0c22-4adc-9087-09269cfb3ac9 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

Small Vision-Language Models are Smart Compressors for Long Video Understanding LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.698662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:bbabe97f3402280cd34c95575b5d290db0fb893fdd3972bf76ed6f9df05e9ec9

Observation 15654cec-49c7-4dcc-9a56-4e5efd52163e · outbound

This paper cites Qwen3-VL Technical Report.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:15:51.411232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:e1de2ca54450f9fd5880f5b4423966eb731726e5567a1a3d09658cf81c97e754

Observation 23dc9d6e-a508-41c4-8002-33e500dbf78f · outbound

This paper cites Token Merging: Your ViT But Faster.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Token Merging: Your ViT But Faster

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:52:10.541797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:67a83fca35ce0cc5c6ce0d0e7de06ca819a366ccfd63e5f4a9a8f9c223090492

Observation ebdf246f-1770-4559-aedb-8b6979b1ba8a · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Small Vision-Language Models are Smart Compressors for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:51:25.572615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:e2b861ec5417e76d9a012fbd77c832d668878caa0a4adb61ae5f127cbc55df3d

Observation bbb71538-5646-4cfd-a29b-a210acff541f · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Small Vision-Language Models are Smart Compressors for Long Video Understanding VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:08:19.955297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:322238b5df2a5ba48d0787cff2922965e9b5edc6661863ed6b8441e40e9b3e2f

Observation bac8378f-1dcf-4ea1-80f5-17e308c2061a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Small Vision-Language Models are Smart Compressors for Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:15:51.503186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:d2d4990d889fc61268a1836cf2bf0e89cdb77dc8de9307539481c04410a1a848

Observation 5b5c2f86-b199-40b1-8833-b6aca8febd81 · outbound

This paper cites Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:36:21.368649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:af3fc1dc6b35278c7115a96ef01c63ca0d47c28e4eadb0e463f8827134a4b30a

Observation e2f84104-be0e-42d7-a0af-20076b606867 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Small Vision-Language Models are Smart Compressors for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:986ce32f822e229f6e455fff190450c0f0aa3a440d90326a6d51f86149893ad0

Observation d1c7aabf-cc5c-4e9e-adc7-4a15863bba87 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:15:51.479344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:1dfd1235d667bf49c80e62c50fb987a74793b869c311b6db564da9f7e354fb7f

Observation 9e82f47e-e342-4b7d-bd14-ba9df894202e · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Small Vision-Language Models are Smart Compressors for Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.725971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:167e1d79b35fee7f01366795ebe0577013033fc607c9e66de5bea3d3800077d2

Observation b54d1c26-5e3f-447c-9cc0-dd3b4fc58135 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:15:51.348826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:7651db32b58fcae3c168866953bbc8ce9d3aee53b9410c1fa9c7c408fc446730

Observation 272b9161-9183-4ac1-86ec-d7f8351168c3 · outbound

This paper cites Kimi-VL Technical Report.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Kimi-VL Technical Report

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:6f2dc4a4f691f0f4a91657b5edbfecd3db95326031c6cfb55bd39c961992c6a4

Observation a83564dd-7166-4d12-abe7-a59f17ccce1b · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Small Vision-Language Models are Smart Compressors for Long Video Understanding InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T00:15:51.305763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:8dbe6d7ff88b461026dc5c0d020207197ab5268aeeb259523d3097e6d4ed162b

Observation e9615c1b-4e71-41ac-8a87-31575d9d8efd · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Small Vision-Language Models are Smart Compressors for Long Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:58.165262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:3d1f8301c3679cca6b9dd599dc735af865c2ce76e84028b405a534c67c12a61d

Observation 9f3d8f39-a4e4-4d9a-b896-eb99a658f92e · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Small Vision-Language Models are Smart Compressors for Long Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.737157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:e727bdae82d36b23703ac04eb52e3d93efcca1f6d42848580b051bc24df2f86e

Observation 1afa61d7-cf5c-414e-8e85-80ff3c064867 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:58:08.257334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:ddcef30700e801f5f7d7dc9f515ca62fb9a26a7f539c0e7e6ad9c265a9720f94

Observation 0d36e27f-b624-4232-8a0a-6456a6de7a6a · outbound

This paper cites Long Context Transfer from Language to Vision.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Long Context Transfer from Language to Vision

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:08:36.726593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:f7354e716e4d6b92dadcf63ea32e2cad00c142824b3b477e89a489a518439d3b

Observation 827b03d8-05d5-45ca-aff9-982f7ca22373 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Small Vision-Language Models are Smart Compressors for Long Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:15:51.364401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:7202d7194117cc78aba807a4bfc331dd838056764ac9f1b212a3dea40c70f76b

Observation 37b4dd1a-7f7b-415c-a9b8-f17d61feed4a · outbound

This paper cites Because the sampled frame count for LVBench is fixed atfmax, the resulting upper bounds are4096/1024 = 4and12288 /2048 = 6tokens per frame, respectively.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Because the sampled frame count for LVBench is fixed atfmax, the resulting upper bounds are4096/1024 = 4and12288 /2048 = 6tokens per frame, respectively

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T16:58:08.260550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:d9964c21a5a384e084898ca8b738a316867b7cc5879d03d50f27a43b41e9840e

Pith citing papers

Observation 11f6c408-8efc-44f9-9362-086b0115b8c7 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Small Vision-Language Models are Smart Compressors for Long Video Understanding

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.629109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:2f7f97d4176cf982f33c63ce877b5cd6ee9abe5249775c2ca1ae9b38d5e53e89

Observation f30a5792-c856-4639-8d9c-b4eeb21cf9ec · inbound

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding cites this paper.

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Small Vision-Language Models are Smart Compressors for Long Video Understanding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:39:58.310762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T00:19:26.153682Z digest=sha256:5badc142253f9c188c2645028e36df5911dddae963e20efa262d64425062a2df

Observation 9e7564e6-cb44-4e9f-bbf1-0467e2f29157 · inbound

BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics cites this paper.

BLUE: Semantics-Preserving Video Compression for Efficient Vision-Language Surveillance Analytics Small Vision-Language Models are Smart Compressors for Long Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T12:34:49.503345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:34:49.503345Z digest=sha256:dd3ee3a62f0925fa798b6f471e7bbdf4b79cad69ab899060b354255cb5716faa