Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 14 inbound Pith citation observations for arXiv:2506.01908.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01908 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:35:48.103145Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:03:29.800596Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.219030Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e2ed415-a459-42f9-a09f-d03a53b8e48a · outbound

This paper cites Localizing moments in video with natural language.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Localizing moments in video with natural language

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:44.652453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:44.652453Z digest=sha256:90ec737d71abec64f8f48485468c559b820f052446803955e818eb161eaaacbf

Observation 5e6a877c-0211-4de8-ac75-94c96c2d6796 · outbound

This paper cites Qwen2.5-VL Technical Report.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:44.713622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:44.713622Z digest=sha256:82d7c26c7c28d3ec8c96468eb1bdf4f6a7c18a4486c37d3ce699e92360154be5

Observation eccda8be-c02a-419c-b524-25ba3ae25d04 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:44.824339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:44.824339Z digest=sha256:0fefab4c6b34a9620a2727be2466aae9c5d9b061fb7a8c335c6c2cc569d78f37

Observation 8fc8c328-42a3-459d-8d07-088e5c15919c · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:44.928073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:44.928073Z digest=sha256:a95986f34e1a3bf1bb191a74ed68b01bf5fb4cfba3bb072f0b3edf1f247d22ba

Observation 8569e3b4-44da-4cba-a19a-2d97d92b075c · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.014401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.014401Z digest=sha256:7108d285ee19b82442d14668581e9d4d660cf251856b250484f7addf5461b78c

Observation 45f9f087-ac3c-4db6-88b4-1b9f4814e1cb · outbound

This paper cites Tall: Temporal activity localization via language query.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Tall: Temporal activity localization via language query

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.104221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.104221Z digest=sha256:3ccdfd205231fe84d95857a865a8363a5171856d588ea5447008c887f6a85389

Observation df619bf6-970c-49cf-b621-4114a27adc8b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.203654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.203654Z digest=sha256:3168d1da3320012a44afb452f4af26c33057fa0b2d62a6aff8e5d2b1132ba11d

Observation d07b57b7-0ada-4364-a2cb-8ef7e1fd7278 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:49.620578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:45.285208Z digest=sha256:59d9276a95af2dfc3084b5e41aa6e09a8d3ee73c24ac95c5df3309bb8cde08b2

Observation 2771756b-e8fa-4632-881e-95c545e86d39 · outbound

This paper cites VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.372866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.372866Z digest=sha256:fbd2a9130e91c809de457e7bcc44df7d6edd662b353bd4220f111566cedbc9c9

Observation 9da80ad0-f63a-4bc9-a3db-7b3d1e21690c · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Vtimellm: Empower llm to grasp video moments

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.455429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.455429Z digest=sha256:e83b7442077edf801f867f13e85eb3155831610f19d2cd21fc2515d97e50c3d1

Observation 43b7121e-3cac-4838-88fd-b8361a7a9dc0 · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Lita: Language instructed temporal-localization assistant

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.552030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.552030Z digest=sha256:dba6eb4f672179abdee35dad5ab91fc665e0aa5fc7905fa618d629f7faad1509

Observation 45cfdbd5-8265-4231-a9f7-f3e8f5fd40f1 · outbound

This paper cites Unleashing the temporal-spatial reasoning capacity of gpt for training-free audio and language referenced video object segmentation.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Unleashing the temporal-spatial reasoning capacity of gpt for training-free audio and language referenced video object segmentation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:35:49.420703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:35:45.640676Z digest=sha256:9e509a08ef251a62824ef45d3cf21d62792de7a3b5cdde5bfd50edaa5ee9fba4

Observation 3d982fbe-eb3f-4b05-8f52-62c741a52b0c · outbound

This paper cites Dense- captioning events in videos.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Dense- captioning events in videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.715578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.715578Z digest=sha256:19e83bdbce1be12f8fbcf4c49e94065213f0f595aa763ade3d88134f798b67cc

Observation 4a6e2bcb-30ea-4ec2-9188-066139becf08 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency LLaVA-OneVision: Easy Visual Task Transfer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.796764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.796764Z digest=sha256:cc318b73cb8b76bb549f80fc4392183dcffad3e1ee844784659eaafac3fc8a96

Observation f79c5a1c-78e4-4915-8011-72b2ac8f5279 · outbound

This paper cites LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.865632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.865632Z digest=sha256:ada416a27cd5b516f1eab4d7068e51c5cb20522502991023c177d77238261d35

Observation f4f70289-08d2-46c8-ae70-8bcb2c2b402d · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency VideoChat: Chat-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:45.945951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:45.945951Z digest=sha256:42d4e8f4ac176aa070c88a804ec363defb185b5847e517e3fbc1f8b51fa712c3

Observation ea610f96-09fd-40b8-9665-c8b5486b81d1 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.014402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.014402Z digest=sha256:21f2f47e9f547c628f141b20cf64fb2246b88298b69cc003e7191451f4422882

Observation 4a53e36c-eb1c-4e6f-8fb3-3e6f8f8694de · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Llama-vid: An image is worth 2 tokens in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.091779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.091779Z digest=sha256:4c15f706e0a8c16b8212e4cfd81d7fc97eff8c156ea7d0f300bed8cc61aac065

Observation c7ba70d3-efd9-41c4-bb52-345f72de1a67 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.174675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.174675Z digest=sha256:b135e5e709fef7aba66033c15c0f55252106d1abc47e6cd372daccc041d846de

Observation 79e47c18-6a64-4be8-92c6-176f2ac99756 · outbound

This paper cites St-llm: Large language models are effective temporal learners.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency St-llm: Large language models are effective temporal learners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.261016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.261016Z digest=sha256:cd8ba6445fc6ddd8383183d7b41db7c82b3d855be150ab6c3fe4005b79332fde

Observation 56dba7d3-10b7-4430-87a0-afb694e0234a · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency TempCompass: Do Video LLMs Really Understand Videos?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.327864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.327864Z digest=sha256:db5c55432fd52168ab42801a9c9b38a2061031a9660a5c0fe1da244beb5a3375

Observation 123f7ac3-f57a-4fee-8903-251e01aa0edd · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.421128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.421128Z digest=sha256:b556bcdea7cf62e25387bb1bb842cd40227ec72b1dd4a8fe6630be88eceeadc4

Observation 573628b1-4839-4208-820e-ad9f759dbedb · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748–42761, 2023.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748–42761, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.515892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.515892Z digest=sha256:c26c8c710e30c2f81512b37e82e9e5930192bfd18f0399a00c0686cc58da54a6

Observation a10982dd-14f0-48e5-a377-08aea53acc44 · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.597000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.597000Z digest=sha256:7bca584979146b7d3f7ef759b848a55d099b3344322060ed7091a59063911484

Observation 2e07006d-1dbc-4672-bfbc-14f921321bc9 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.692867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.692867Z digest=sha256:6a3adbac23fcfa11504d1bdd80c8f6c413127e4c828cc4947abe409b948c5e9c

Observation 886ba38e-c3c1-4fe8-ae62-1a6c3912c6d6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.785592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.785592Z digest=sha256:d68a2fefc2f9e5af6b0da53be998ac6521a5d3157d0d8c37f376d645e676a0d0

Observation dc1fead3-176b-4331-8547-5c910e8d96d6 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.872216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.872216Z digest=sha256:1f5187955c62c31daab3e9e125ad59ecc00f9f38e8616e78eee1420a223bdaae

Observation ac4dc91a-c40b-4533-9d0e-c72615597896 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.943484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.943484Z digest=sha256:9d3c9ef2bcad8acd20a207db4224a2d67429d72b6316da44e3f8748034d394c3

Observation 921a662e-7183-44a5-b656-983f70e2ff53 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.027169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.027169Z digest=sha256:4c38d5938350391dc7ba7beba5fec17ff688b95c82413e7929b4fbedf53d68a7

Observation c8498e34-385e-46d5-b571-f9cbbe09859c · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.149960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.149960Z digest=sha256:b7c23b15934186ec666a6cc9dfe8a99300895ecdb1c8a4446ac1533515722083

Observation 85ce6aae-bad9-42a4-9987-1de562e8821c · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.240790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.240790Z digest=sha256:e96ec9593b4b8de8a6d30d5b711ae325c9ed90e22cb1fc571b71a0663d7ecfd4

Observation e97de0ce-7f61-4652-b68c-917365ddae96 · outbound

This paper cites Next-qa: Next phase of question- answering to explaining temporal actions.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Next-qa: Next phase of question- answering to explaining temporal actions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.308312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.308312Z digest=sha256:225fe8f806738be3e01f62583b6bf1fc778c75ec150e92147c5a2be0d2eff512

Observation 295dcf46-87a4-4baf-b49c-c5725e07db3d · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Can i trust your answer? visually grounded video question answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.389258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.389258Z digest=sha256:14dbb5634ee4002d7515fc180eb2daf69fb5652332e23b000414c56281acea60

Observation 0bb155bc-f509-48e6-9723-679ee5d23c62 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.480823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.480823Z digest=sha256:708fb761d7b6550c9aedd8726d8e919876ed29e9cdb4ff4fabe62c294dd48491

Observation 8493c7c0-257f-4882-ae9a-c2fa44dea18b · outbound

This paper cites Unhackable Temporal Rewarding for Scalable Video MLLMs.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Unhackable Temporal Rewarding for Scalable Video MLLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.573465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.573465Z digest=sha256:045fcef12c359ab116f2c2af12738bc846e3beebded028b5f1219d7efa401509

Observation 997cf40a-d098-4e74-9803-f8343e4597cb · outbound

This paper cites TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.666720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.666720Z digest=sha256:8dcdfbcf158f50a847c9a47bc99190ea671e87f6bc6d996d1ce2208012690646

Observation b6525d68-4675-45d3-ad5f-a6f602a39e3a · outbound

This paper cites Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.736748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.736748Z digest=sha256:5868f8b01ed90c7068aa622502c5c54c74f4d6504344d60687d5b74e37730f73

Observation 03c4b6b9-57e9-42a5-ae6d-9513ed906ae0 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.821566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.821566Z digest=sha256:99931753fbe0dcf795cce6cab975236dec5a22e82f527164d5cd84b1199b2065

Observation cc55fd6e-4219-447b-b325-f6fe6f62535c · outbound

This paper cites Long Context Transfer from Language to Vision.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Long Context Transfer from Language to Vision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.916580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.916580Z digest=sha256:04a7e050c4db2db9133ed70b7f4bc281507203272564d344a28b4c265077d6a6

Observation 81a57794-523a-4f6d-aa33-8a70e34563d8 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:48.023662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:48.023662Z digest=sha256:bc3367ac494601ee279f7d964e0fa81c9feabcd8c772b84a5df598f9ca8e9083

Observation 68083eb2-8d87-452c-9426-e39db36d209f · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:48.103145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:48.103145Z digest=sha256:160e05d03d0fa9e05433b8437ecc863fecce172e958b6bc0022d65b952ac3a3b

Pith citing papers

Observation c2f4e65e-d940-42b8-8d43-1f68641fb498 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:29.800596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:29.800596Z digest=sha256:ad6dd92c59f03c6eaf2ca812c9b64f6507a9c825bb31da11de406ffab3de6633

Observation 9ab09eb4-570a-404a-9b79-903c34a99536 · inbound

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding cites this paper.

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:01:25.936958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:01:25.936958Z digest=sha256:ab50ce7b53ef20ddec27d94e48a755b5d65153f85d7f1b9efdf9e0a70a3f506c

Observation 050345ff-e308-45d7-9f82-44827a2917c7 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:26.630505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:c3e9503ca868e377c8954f7564942711118f3790b57a9a12cfef83fdf718b590

Observation 5d6fde9d-8de0-49ff-b154-7fd71dcbe877 · inbound

AdaTooler-V: Adaptive Tool-Use for Images and Videos cites this paper.

AdaTooler-V: Adaptive Tool-Use for Images and Videos Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:34.346222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T21:23:33.598026Z digest=sha256:a5c52e9e0953b1f0922d398267d017ff71a601b91548568ea2e6f1f4a39c5371

Observation d62ebeaa-5f0d-4f4c-bb7e-d320650c0026 · inbound

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking cites this paper.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.320193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:8f7d8e098d70ddd051a59261e3bd11b7bf7a53e47c7e693d39dbc4dd89c517e2

Observation 5f329e02-4e84-42af-94a8-3b9914d312ba · inbound

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding cites this paper.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.135590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:0af36495f7c14a86db9409b0d12bec3d781eb42a3f0ca8da5cff10a32c837b3f

Observation 2673ee9c-ee9e-4103-a9ff-3e0560d47143 · inbound

Temporal-Aware Reasoning Optimization for Video Temporal Grounding cites this paper.

Temporal-Aware Reasoning Optimization for Video Temporal Grounding Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 122

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:17:29.087471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:21:30.638275Z digest=sha256:e086e5450c8c81a11d9f8dc1f350012ef39c53b2a139643e986cb52ed36ea735

Observation 78b1981f-1631-4e84-8e43-db819365475a · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.749200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:113c470f8d7ffff4d845d18ecd70a087b57f0e5d2e756829d240c6476a320967

Observation 5fe7afff-ff8f-4d04-bc81-d73212b62804 · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.391135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:0710a80e7d229fbac175674c8acd7f26414278161a1b27789ca1a493975e94c9

Observation 17baca14-be29-4760-8abd-79b8b54c4837 · inbound

DramaDirector: Geometry-Guided Short Drama Generation cites this paper.

DramaDirector: Geometry-Guided Short Drama Generation Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:59:55.979608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T01:57:15.550335Z digest=sha256:c7adf4a92321979ed4be16063bb6d4d8955be2951617f542c57c2625bc603e93

Observation c728d01b-29e6-4990-bdbd-fdb58bd8a711 · inbound

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards cites this paper.

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.220535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T00:16:58.697810Z digest=sha256:b0202a527bb8ca9b47ed586cf5484a43e83aac481f68991895be8d302971a4f0

Observation 2ba040e6-b12e-4f6b-a703-ede10224c1a3 · inbound

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning cites this paper.

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:13:52.951272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T04:53:25.259840Z digest=sha256:6d0bc6683715fc82cbde263021834ec94b2e3cb5f39fd8314aa3db66cf08b768

Observation 2c46cb5e-b4d1-46c7-a68f-ac8bc4892e0c · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:13.530716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:13.530716Z digest=sha256:15f09b81b04839ce4ae2bb76823195cb80a77fdb6d81cf617265dee9737c2fc8

Observation 1ed0b2c7-1ed2-42c3-be13-e58a000da904 · inbound

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning cites this paper.

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T17:24:14.453570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:24:14.453570Z digest=sha256:2524239871fcc7e9d43fd75a574eed8af7de0f6beee01abf36585385b3fca8ec