Pith. sign in

Paper Citation Record · LEDGER

TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2504.09641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.09641 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:47.309529Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:31.402219Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6903ccf-ac39-4c50-b14e-b931bf6a40e2 · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.542292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:4f47b204476c205c016dce3ee723f1741d8c954b50f481af4f7b1f60a041190f

Observation 3203da8e-6f23-4b56-81fa-9d2d711bda85 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.563294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:04f5b11e2520c274ee42ae19384f9a1806edc903c12684a2c0959dd71ea127e0

Observation ff5ef0de-8880-467b-b83f-1b11574f8eaf · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:40:56.049230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:bb85095c7905cb85e9dadad2159156f03a0775d22100caaa717cc32c07837300

Observation 6b93e37e-a4f1-4806-a581-13cf263521b9 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.309529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.309529Z digest=sha256:93b6f29bdb269552ae26aa55f9447f22d9b7481db5f5c84cc2632da690413bf6

Observation 3aa44a55-2906-47d5-8c55-b4bd0248f676 · inbound

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought cites this paper.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.150407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.150407Z digest=sha256:d387caafa0f6640279bd28ebb1edceca27d9a5e04895577bbc2937b7a84f932b

Observation 505ef48c-82e0-40ed-9fc6-f2d6663dd08c · inbound

OctoNav: Towards Generalist Embodied Navigation cites this paper.

OctoNav: Towards Generalist Embodied Navigation TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:43.895513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:43.895513Z digest=sha256:d5666ab19cfde8f781f53a73cd216c165cfb5f1b06ef1d406767042e31ddeacc

Observation e28a9912-3e8c-4a5e-9e30-659ee8039dc9 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:08.007028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:ec0508e23226546c193f2399f8dabb3cf7563197877253f1fb01902c987a497e

Observation 7b24de6e-9992-4c1e-8ff4-c7fef8c2dfda · inbound

Compressed Models are NOT Trust-equivalent to Their Large Counterparts cites this paper.

Compressed Models are NOT Trust-equivalent to Their Large Counterparts TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T19:00:30.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:00:30.163389Z digest=sha256:5bbddefaaa9a711db3d4e177addcf6036fd5b2b00b17518c4b8dc1d65d642236

Observation 4ee90811-fdf6-4cea-836f-44f2efa11dba · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:47.267163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:47.267163Z digest=sha256:c53c470871712a5e73587de1da254a9c75f788998b1079d584d12497177626e3

Observation 4b2b738a-4799-4f6a-8943-d902c017ab8b · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:09.855828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:09.855828Z digest=sha256:a416b7d728c9d188b057aed871cd9b3cd31789fc99439d2707cdb2ded6012cd0

Observation f503adc7-1cf5-4832-9b23-ae3118e44777 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.954750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.954750Z digest=sha256:eb7406aec37e2bb890e57bf6b5bc5e7a0cfeed5a2eb80c5738ca5573da8c2aeb

Observation 17f3cfb5-ae42-482d-a31d-a47b4f0a720d · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.544726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:ecd1bdd4f4e2fef41df835bd461134fbfef236a9f4b65c24fc0aee1644024b96

Observation 7e940b54-6378-4f98-b9df-28dc33c57448 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.321509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:aebd9b7480c1ff4a8f135e41d87e68cb10687ab969fa8aa5a3332699bdcbf115

Observation d0ab1da0-ad2d-4dc3-bbac-69371a37e1cb · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.581348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:a6110fd746978b68d843f7e441f070f21496d84e105e5402c8b5e5feb3d00501

Observation 6705c224-50d7-4c86-bc5c-852e24f6ea36 · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.514057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:4414c0b2c6e31f47314a7fc96b78dae31cc5f99747fa40bace1d5fc502cd09db

Observation b3520bef-61bc-4d47-8a43-dc30f82f04d7 · inbound

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs cites this paper.

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:06.511080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:30:56.297653Z digest=sha256:2ba751de093efcfb61a7b0b82d89c58c83a3c0fedd8a9c80fad1357d3b568b5f

Observation 542ce35d-b672-4081-af8e-a288041fe7da · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:23.112134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:54:23.077914Z digest=sha256:332690cc74b516dfec078b7b0b5b4139399c56ca9a9fb0866e8f9a5c93f549ba

Observation b2b4267c-3ccb-4aab-aad3-7ef8381ec0d5 · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.555385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-04T00:41:02.284215Z digest=sha256:8c5bd16761f26fe6f2a872dbd23b31a3ac86b35084a2933bee57c2352fb9628f

Observation b2cf85bd-4380-42a8-9f96-e2c95434cb17 · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.741199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:9044e558742affc2782c40bb25ab38c5a8f88760484899a49e2a52f3268e1d10

Observation b5854cbb-01ac-42e2-9818-3d7e90e1c590 · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.404380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:150dc2e34696888b335ca03d78938d80e43e9a0dc90c57871ad3b5ef327cba2e