Pith. sign in

Paper Citation Record · LEDGER

LinVT: Empower Your Image-level Large Language Model to Understand Videos

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2412.05185.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05185 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:51.011685Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:56:47.335501Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e2bdee07-8166-41f5-a9a7-78f04a2aaf8d · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:51.011685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:51.011685Z digest=sha256:ca1913835ec149f9e431bde36a4a270508a8eb22be542cfbd58fc8bc2e01b1dc

Observation 7a28d731-d390-4e5c-9c8d-1465ae651bdd · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.759496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.759496Z digest=sha256:eeea32cc9b22810dccc987ec4802cd3f58325f263039c55ff1f5c63d153e7afd

Observation 47912e91-9c6a-4594-8839-40ba677a6da4 · inbound

${\mu}^2$Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation cites this paper.

${\mu}^2$Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:24:45.561143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:24:45.561143Z digest=sha256:49c23e334823cd63c7409696274a5bddfb2a379a0be72f4d1be3a927efe3f271

Observation e555be81-c022-47df-b305-a1faed9b8205 · inbound

AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multi-modal Information Between Reality and Videos cites this paper.

AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multi-modal Information Between Reality and Videos LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:25:02.720820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:25:02.720820Z digest=sha256:fd58abbfd0d639867aa47626bbd30226d943f493c468dfe01b291ca367776646

Observation 9b54c22b-409b-4082-828b-5cac0d08ff71 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.511427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:d2ab19c8acdac1813e8caf81eaeb46387e0f170293f4e4433fe973a81622344f

Observation 944fc3e0-f602-4b05-be63-7a75b5014222 · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.337270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:6e13698503f2a4dd8acb5f454e06818369a79d9e1045681c86ec436f21d5b07d

Observation c3ac952d-456c-4709-bf4b-b1a34cc5cead · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.394593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:84d052147fa3fa053e5b1ef94e19164fbb25e8c3c814a6076482216ad12bbb4c