Pith. sign in

Paper Citation Record · LEDGER

SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2503.18943.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.18943 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:55.856714Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:14:02.175777Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 765079fd-8673-4ee6-80de-35b19f652f4c · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.856714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.856714Z digest=sha256:6a925bf1af81e980e4a967fbf7e6ae0d4dfb369fb7923fa54ea64be759ac2835

Observation fdd0506c-935c-43ec-b472-bfb3ddaf2094 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.935958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.935958Z digest=sha256:1088e904c96c83682c7389344fb807627e01442a824115d14a71ecc6bd4287ad

Observation cda3bfca-1bd2-4a86-b1ff-576b3bf1202b · inbound

Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance cites this paper.

Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T05:16:57.525652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:16:57.525652Z digest=sha256:8b95ef7e671345aea03bc7c2312266e2be4b33c5a0116a212425e7e85b3d1a0e

Observation 8cab5815-04f5-4877-b350-280990ddda3d · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.162651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.162651Z digest=sha256:89be605f6d6abb7b8ac8fe5b155896f6e6651f2c681f12fba9aadfe55f7216c7

Observation c26e1f9c-aca7-40f5-96cd-ea1a1892c99b · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:862bc09286dead1b7140c6f3b01260b1f21568ab7435c247dc1c65067f6e79c4

Observation 9b7ca5d7-6d06-4dcf-b72c-144ae5a59245 · inbound

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute cites this paper.

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:40:52.137771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:36:50.406353Z digest=sha256:e3795499dea8e8ff63a08b7e278ae4bfbb7a9f61e7ac95b6a055809ccbcd09d4

Observation d36f150d-e770-4291-9c13-db06a1cf51b9 · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:56:07.903630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:879a521843db4c00868cf1009032dcddcef0662197b4e067f76bf553567728de

Observation 9c8bb1ec-0283-48c4-9bae-2a968b2deed9 · inbound

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models cites this paper.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.177596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:e10696e50ad3b8eef618375f1dd8bebc9823baa32d10c13b7413fc6c4e44aaaa

Observation 9589498d-9658-46b0-9588-7730b46fa04e · inbound

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams cites this paper.

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:51.049541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:24:57.881644Z digest=sha256:fd578d00564fe6c7c4d17fb0ee3ecaa71eb911478e981d2d4c2b5350299f6823

Observation b07bc466-d2f3-4511-bdac-08875e0c8160 · inbound

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer cites this paper.

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:25.906481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:01:42.880738Z digest=sha256:28dabc71893bee6b969f60112911587ec94763f904ae054893bf9b3a3e152cb0