Pith. sign in

Paper Citation Record · LEDGER

TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2504.09641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.09641 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:12:18.725011Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:31.402219Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6903ccf-ac39-4c50-b14e-b931bf6a40e2 · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.542292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:dc9b07fc9ea8adccef9fcfefe1dc26835e44bbb469198f22179b8e9167054dd6

Observation bf64e574-9a75-44d6-8ac1-a2b85d8ffcc9 · inbound

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models cites this paper.

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 148

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:18.725011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:18.725011Z digest=sha256:6e3f6aca68e1136d7d88841c56ab46e1087b7a8ac775f764f3e5c05e7101c24c

Observation ece43dd0-ab59-49e3-80e4-0ba2d1f4d93f · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:18.656306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:18.656306Z digest=sha256:109ce46c19a0e6377973bdec0394ba01b2aad231ce8d16f57e20b2df847eb015

Observation 6c710c76-cef8-41fb-ad94-db1d04df5e4e · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:15.937040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:15.937040Z digest=sha256:aa839ce4b597b4668305e93188ccc4cd63973440ba790ecabd864c7f3e801b0d

Observation 3203da8e-6f23-4b56-81fa-9d2d711bda85 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.563294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:b33814da5ce9c322a91a45812cc9fc0d4a277a32135093bddd88d1fb1bc59203

Observation ff5ef0de-8880-467b-b83f-1b11574f8eaf · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:40:56.049230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:2ef2d5c3a9baeffe97d87f9b3d3ee73b0f0af67804ee986029cfdd580f4a70a8

Observation 4afaad58-44d7-4e64-8e01-b8477d86472e · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:05.300590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:05.300590Z digest=sha256:541cc18ba590cf3c317447d2ea9078522eba11f1c85a37a35814a51db37dcbe9

Observation 9024fe1b-a3c3-4883-a50b-a5445afee8c8 · inbound

SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization cites this paper.

SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:28.598898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:28.598898Z digest=sha256:43153a408246b3069c7891326946b75c5f4ad8b9ebaf6bdcf2687ba460363278

Observation 6b93e37e-a4f1-4806-a581-13cf263521b9 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.309529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.309529Z digest=sha256:5cec8873a7513d184fe0b20ade3d88ec319bf02f1120cc5965a6ca50eb7d031f

Observation 3aa44a55-2906-47d5-8c55-b4bd0248f676 · inbound

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought cites this paper.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.150407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.150407Z digest=sha256:79368e7090e057c0137db31c9221b15ec7c5705ddd919c6c065fb87687b62e67

Observation 505ef48c-82e0-40ed-9fc6-f2d6663dd08c · inbound

OctoNav: Towards Generalist Embodied Navigation cites this paper.

OctoNav: Towards Generalist Embodied Navigation TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:43.895513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:43.895513Z digest=sha256:36afb5fe0c3a4a365222cc2c68cbde64925d8d725082f5cbea4314c1b7794671

Observation e28a9912-3e8c-4a5e-9e30-659ee8039dc9 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:08.007028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:5853320c4553e2815c7d0e51eed27c2368db2c43f01ba7f06e69412995ee75ac

Observation b782d167-d4c3-44e2-8bf8-c9b727fbe744 · inbound

ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models cites this paper.

ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:10.352493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:30:10.352493Z digest=sha256:3bc80ba8aa38f3acbc8216f9ce38b202bc19a0ef56e9231c497e9c5c704997e9

Observation 7b24de6e-9992-4c1e-8ff4-c7fef8c2dfda · inbound

Compressed Models are NOT Trust-equivalent to Their Large Counterparts cites this paper.

Compressed Models are NOT Trust-equivalent to Their Large Counterparts TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T19:00:30.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:00:30.163389Z digest=sha256:e654261a55912b9f0b61cfa2acbcbe385616256ac5f8d7f2441e3ab057ae7736

Observation 4ee90811-fdf6-4cea-836f-44f2efa11dba · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:47.267163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:47.267163Z digest=sha256:7088c2d83a8678be54b34732091bece4fa2d66a2e16215f63caf4c0415a407dc

Observation 4b2b738a-4799-4f6a-8943-d902c017ab8b · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:09.855828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:09.855828Z digest=sha256:54b0c46a536b16c069f3662cb2a1de23644756448ed25c333fefdf7fd6819320

Observation f503adc7-1cf5-4832-9b23-ae3118e44777 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.954750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.954750Z digest=sha256:98e5055f7705dd168b685b881583ce28fbc453f3d3933830d0056f7de1d88cb4

Observation 17f3cfb5-ae42-482d-a31d-a47b4f0a720d · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.544726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:eab19a9a58c29078e1552f25bc4c9c41f002c4298e0d3f6cd58d4243e676accc

Observation 7e940b54-6378-4f98-b9df-28dc33c57448 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.321509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:d1b3b289a763f5c2457ee4ffd90454650f7fa7c50113671e3c8e6c6a25003f01

Observation d0ab1da0-ad2d-4dc3-bbac-69371a37e1cb · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.581348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:9f1d5ef8c3b547c92177934fb55f60aa68641f6f9ea14fc67b02160e4455b8e2

Observation 6705c224-50d7-4c86-bc5c-852e24f6ea36 · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.514057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:512c57ff57a5d90cb6271ac39d4fb82051e34f3d5b6ecba76ec14e565a257c34

Observation b3520bef-61bc-4d47-8a43-dc30f82f04d7 · inbound

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs cites this paper.

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:06.511080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T14:30:56.297653Z digest=sha256:7e709e0e0622a3e6c2e2f41ea46a4608a786cc2ee36d343f4316c78f48a0705f

Observation 542ce35d-b672-4081-af8e-a288041fe7da · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:23.112134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-25T04:54:23.077914Z digest=sha256:27edd6e0713f6be017fc1d27667d34a01ec84ecd0125e0c2506c3b41f710e717

Observation b2b4267c-3ccb-4aab-aad3-7ef8381ec0d5 · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.555385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-04T00:41:02.284215Z digest=sha256:4c6bbc9c87c9e0e0794bfc92f905a26b8f9367780609d0ac4ac18caefd2c289b

Observation b2cf85bd-4380-42a8-9f96-e2c95434cb17 · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.741199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:5a36ddfd88f93e1e930b676fdb228dd5975216f8e77462aec86b4d0bd764d481

Observation b5854cbb-01ac-42e2-9818-3d7e90e1c590 · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.404380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:239f35a02955f5d3ddbad0d90ed38a1fcad90e6fed062ccbd5dceb1e91c110b5

Observation e0e37da4-0640-458e-8d61-44194b7922c2 · inbound

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning cites this paper.

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T04:59:13.789253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:59:13.789253Z digest=sha256:6648bdfe79f44755193c882f987c6589f8b2600113aea4d8d791b6c225193f99