Pith. sign in

Paper Citation Record · LEDGER

Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2504.02438.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.02438 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:34.994738Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:29.022097Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d536abaf-3295-4fbb-a432-666bc4a20e34 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.098310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.098310Z digest=sha256:cd47f2aef53834e52ea57b1e083c5a15ca30b16638eed4e1fa07f26dcbc2cf00

Observation b4e04e31-54a1-4856-ae77-3dd741cf5db7 · inbound

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding cites this paper.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.855545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.855545Z digest=sha256:4cad0d7adfc04c2d0893a8f94e5127a8fe9872d478864111637f67531f90ffd5

Observation 809b2dcf-8bb2-479b-964a-3488ccdb6144 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.568750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.568750Z digest=sha256:800a23da062b210c61809d33a670cd897dafceb069e184d507c148bb0a7c31b2

Observation 46aa2c03-e352-4525-9cdc-8d2919322937 · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.166776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.166776Z digest=sha256:d670f8ffad1a4c7daa5122af7a55a74de3528ca67f65d0443a9fac8d9346eef4

Observation 0ca175e9-d415-434b-a2af-cc830cb9d6ec · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.064377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:92f7466d537a5cf00c8aa559b8939bfd3da03ce13de82499490dfe56317764db

Observation 29279496-f3cf-4c73-8508-c90752bb7e55 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 128

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:4ed8aa242b5acc912d13087b2493433f0834cb6e39bc88a286a8becf8808a04e

Observation 9b09393b-1b9e-482f-8bcb-fdd8bcb23fa0 · inbound

MedHorizon: Towards Long-context Medical Video Understanding in the Wild cites this paper.

MedHorizon: Towards Long-context Medical Video Understanding in the Wild Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:16:07.499176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T12:28:35.008604Z digest=sha256:9c5af3048feadbdea6caa586f2452ff5e4779f44404461e5e1870f7e541018bc

Observation b3c77a53-5529-4596-bf03-d20fc9750306 · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:56:07.829198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:dccb7756353347f555df9d3f4ad4cdc52d0155e43ce8b44a9c63efb3953a8ab7

Observation 9bc7cd0e-d49f-424d-948b-e164c258782e · inbound

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset cites this paper.

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:57.175123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:05:47.810096Z digest=sha256:84f547841e958ac90eebb33757152195ddc5aa527dd3f609d29c24f4134cd234

Observation 78c047b2-28ce-4fbb-b0f0-dac2af1b1ee5 · inbound

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding cites this paper.

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:29.023381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T17:35:07.667016Z digest=sha256:b416801e73ed4910b5aa81db5391e7940508244786e9b9b423947cc94f04020a

Observation 9d159b4b-e366-4cc5-8ef7-38491711865f · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.291857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.291857Z digest=sha256:aa2f8b1376788f71c9db1ec878cc4b4e592eb753abddf2f92b2d029bb44d83c0

Observation 1468d980-9e42-4d3b-9c1b-9749f003a400 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.994738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.994738Z digest=sha256:d0bc015ebaec261e9caa7ad17eccaffa2d15a3c3b05ef489a684a45c8cc07e8e