Pith. sign in

Paper Citation Record · LEDGER

LongViTU: Instruction Tuning for Long-Form Video Understanding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2501.05037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05037 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:59.162625Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:16:45.171051Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1445af28-cabc-47dc-8c25-d4d3be3d296f · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.162625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.162625Z digest=sha256:b7a8d3398ef2cdf2ccb6ede13f5822466ddd88689e2109179d0258e203695d79

Observation aea3764e-83dd-40b8-9999-1eb2fca7070d · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.416661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.416661Z digest=sha256:cad9538f3809abb717b76e96c99a9870375803706f6aa9eeb3d64763ca9b97ca

Observation 88627f00-7b73-4104-9607-25201e2aaa50 · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.003545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.003545Z digest=sha256:6ceba393e52c614bf9e8180d940e2064fe296fb856afa277340f03b930d38582

Observation e39ccf5c-5ad5-4945-9d88-75828525e7ab · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:216312e1443bd00dce1cf1334f88282b1237caed00d8776adb39055f4f3fd26f

Observation 2c57ab1e-4cf6-4588-a651-db91dc8b6080 · inbound

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding cites this paper.

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:16:45.173173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T07:00:21.192082Z digest=sha256:1ecca8e1fad536b6ea33d0e5bef204f36537f640f964bb95e3bcfae83d180702