Pith. sign in

Paper Citation Record · LEDGER

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

As of 11 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 100 inbound Pith citation observations for arXiv:2504.06958.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.06958 v5

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T20:56:07.247122Z

measured 142 of 142 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 105 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:41.669333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:49:57.434938Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact32
  • verified fuzzy10
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa300db1-6c7e-4328-b28a-2e9b467ab154 · outbound

This paper cites Qwen2.5-VL Technical Report.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.773391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:8fef574bb145d79fe554fdbbf3cf195693f48c721014702d5abcc49864af93b8

Observation 3be6774c-fefb-404f-8bcc-7cad0151f320 · outbound

This paper cites FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.664542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:ad71ce014889f35a0cc762f3a1f258a523036afab84ecbf36517e69765a10e9b

Observation 2c866938-5ec6-4e64-b313-5723a5276481 · outbound

This paper cites Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.673376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:6439d61872986916be7667a8828c281f8a5634cec21bcea10de45396080fa531

Observation 8a4e1c4c-5e79-476a-99f5-59a5023b75de · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.519833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:15f76b05622c4ada47333b4425607e868dbcc80d98512eb650af73c5710f262b

Observation 565583f5-7cc3-4243-877d-b51e82a91195 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.684571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:c9016f3e2279795570a23ad2cfe563d3384a3cfd1673e6a58cb3b1598f195caa

Observation b1ed0ac3-259d-40bb-88a1-bd76ffd01727 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.695329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:46d58db206416963feaab9030505e2316d64050df1e8d29fedd9591091ca2834

Observation 04fc8b7a-aa16-4854-b090-bd69702cb522 · outbound

This paper cites Tall: Temporal activity localization via language query.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Tall: Temporal activity localization via language query

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:56:07.888845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:da838d6f86eac931f920118e90d053e864e7f925b5f371d23c729007c67901e3

Observation 49cb3267-5edb-4867-b5b5-dc8809246fac · outbound

This paper cites Saliency-guided detr for moment retrieval and highlight detection.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Saliency-guided detr for moment retrieval and highlight detection

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.779989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:25632baf15d8afb3c681a40fbffcffa49f0aab7a7648dadd1773b369b7d577fe

Observation 0949eb1c-3dcc-4606-8923-aed3099d416c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.786208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:64cb2538b5f7beb049f1ff58800b249c5cf2f19a81b5ff8a18b1a44c28251ec1

Observation 6ace0399-f7f9-47d5-8d00-07ca26340e47 · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild.IEEE transactions on pattern analysis and machine intelligence, 43(5): 1562–1577.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Got-10k: A large high-diversity benchmark for generic object tracking in the wild.IEEE transactions on pattern analysis and machine intelligence, 43(5): 1562–1577

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:56:07.901787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:f1ff7f6a0c7ac65990538517b607b399a3f3f4c1d8a89cd1d10c60499cdd1305

Observation 8c402674-59cd-4bb8-8c71-e557319a38ff · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Online Video Understanding: OVBench and VideoChat-Online

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.794133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:11e64b8f2cfecf38e7811ea09f0bb7dd632ec1aed795ba6165cc5af3ddd8bd41

Observation 4b2d7e38-b69b-412e-852b-4cdb9c56169f · outbound

This paper cites OpenAI o1 System Card.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning OpenAI o1 System Card

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.800834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:bb255061a2058048cde0272e34a5dc9cc942d580dc9ee001cb63a9480d0eeca3

Observation 7365848c-3171-4ddf-af3c-d18e296cf342 · outbound

This paper cites Dense-captioning events in videos.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Dense-captioning events in videos

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:56:07.917926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:759d0f9d278ee0c1f825bcd11ede472c92d17b5f0f8edb268ead6140f96bbe81

Observation 660889a1-6a46-4982-a916-8f1ef2e247ea · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoChat: Chat-Centric Video Understanding

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.809487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:3f25c834a8f1e3aaf3958b40bafd0b7cdc55f73d2483bf661fecf1bd651557fc

Observation 38592194-23db-41de-bebe-606b597e70fa · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:56:07.926756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:25185574d27ec5d8bffbebd5238a146bc9c5b4fd0417daf917f031f5d1564ead

Observation d74d232f-e4f9-40be-b7c5-f9d0342000ba · outbound

This paper cites VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.815156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:3fa6d05df3636047a90355350de57dad974ae8c9db9d1b5b004cf5ac896d35eb

Observation 304f2a4f-3440-41b9-a78a-5c5e63b912ff · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:49d2c1445fd3f307c9a40b9d96f1c6466176caf26dda01e358801479481a3a65

Observation b1f451b9-5f8d-4ee7-9797-66ffb66a0549 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:31:43.641168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:cdb1871c55bd149622ff6d0c235c6ea2126dcdd086b2ee1914c817c2c47bc3d1

Observation ec234bcd-674c-4608-888f-c610fe3f33f1 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.836042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:90ce8b16af9382a27070428029aa40bf9c4076f08ba6b292419c8d9243fa926d

Observation f5ee1669-04c0-4714-a61c-7722cfde4a38 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.844768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:7651229d055af11a91da8760758dd8a61f4345de18317181d268c6c9bad259f7

Observation 552140d7-f2db-42cc-b8c9-085333b56e8d · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36: 42748–42761.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36: 42748–42761

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:56:07.922359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:38ea012405d9ceeb87d5c0915e672b6e3e89ac33dbe334af786bb14cec9ba7ad

Observation 60953261-4d9c-40fd-a447-9f79c74b7298 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:15:46.589538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:57873ea9f42bda6c1a56064805b5dbbf3da9bd41c8fc1e961847913435f09fd9

Observation 829f1d38-6a9c-4983-ac44-fafe008a183c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.856503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:d0497e9c154f087478c65756adff31417b89cbec96b0e7fe9ae519432a9e8a1a

Observation c409a1c8-e932-48b3-aa22-b95800169138 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.862745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:9c1add64e14e06795a29296ceed98f9c1c6c1f81ec4ed326971337d14cf09fcc

Observation 1ba35c61-2595-4c3e-81cc-798898ead23d · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.869285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:3dd6c10131069849ebb1029503cf9e6ae3f9f52bb531521a6df481958c35fc0f

Observation 2887b8fb-5f10-4006-aa6c-aac2821363c2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.874721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:eababa9ddafb8ec62f3dda95a4378d11c6d51b44d3316983f0310a6f55998da4

Observation 8470cac6-2907-4aee-90c7-e6e76052aa6a · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:accc8dbfadfb5c9dd2ef6ef50607bce4fd1cab7e2bb6a3c0941dbe75e01322c1

Observation 08e619e6-26c9-4e84-aa76-f1dfd6212508 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Internvideo2: Scaling foundation models for multimodal video understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:56:07.893482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:4d9fab15978bb4221b2bf4022091a348cfb35370b7614cb09ad01455f2ddac27

Observation a0ccc2c1-0ac2-406c-8c64-f23499a54bad · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.835573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:7120769de62b77e256e3fde5f39d40671f3b5d84efd5200308a274575d075d44

Observation bd8ec7a5-ec15-40f6-b610-69e30b6d4c39 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37: 28828–28857.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37: 28828–28857

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:56:07.907161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:bc9e0e1e5cd5a3b8cb024ce92ba2ac7b817bcad6f209771eb6ab7fc951357ae5

Observation 04a206cf-250a-48e2-ae51-bf2eb5a759b9 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Can i trust your answer? visually grounded video question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:56:07.912757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:a5c2a1a2b1311ccb3ebc21413cf393c2275ffa7340f3d68a189393080a8249b0

Observation 262fe47c-c479-434d-ad18-cf9fbb39a787 · outbound

This paper cites CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.717041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:52a839d28c34a898aba6279b10346e61c0b8758599fcf924689906f4b0c8b903

Observation 38640935-a25c-466a-b623-b70ebbdb310d · outbound

This paper cites Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.724355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:73f56b9d6e248477ab4d34b74df16a87469dfaa4c0ecd4f9092c231d61d0e761

Observation e3e3247f-29a4-4ae2-a1fb-2f6d60552615 · outbound

This paper cites Qwen2.5 Technical Report.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Qwen2.5 Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.729810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:53394b726aea0401f91cda047aae63b3fcb71afc7f5231ce226215521c548fdb

Observation ab973397-268d-4593-8393-2e207b42a010 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:19:20.780377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:fd5010edd8612527aee96b89a538c5687ab9124ebda55efb9aff13f0ca4b7e36

Observation 4eb0ca53-8467-4220-afe8-7fa34cd6c1a5 · outbound

This paper cites Merlin: Empowering multimodal llms with foresight minds.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Merlin: Empowering multimodal llms with foresight minds

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:56:07.930594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:c4e4031c770b1816f7a2d2137757f1dae2358afeacfe8d681d66059ed8fafbeb

Observation 159f87b7-00c8-4b9a-ac7a-9274e28fa6f9 · outbound

This paper cites TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.743281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:16883e8e63ef569f9110c44f0a1bf27a964a3ade29a2ca45a37cf785837d9998

Observation dd411d1f-0a5f-44d6-b5df-ac076a45f325 · outbound

This paper cites Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.750683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:8474d2f58cab56b842e6a5a10516ce3337a3b5ec677e6d7ee5b5b1852c13f5aa

Observation b1cca0fc-c2a3-455a-86fd-09981ad80142 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.925544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:8b330ad5e1a7b4d43e28a4669a233e6aab8ef745ebbf0687903d3a11a26d23cc

Observation 3caa9af6-8bcb-4af9-a24a-5ae85ccc00ba · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.764632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:19ebebd2e2c96ad2665d51dd67287122aceba284c0616c95f314724e87f4c3f4

Observation 01cb1b88-b7a3-439c-8d34-65dd1858338a · outbound

This paper cites R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning.arXiv e-prints, pages arXiv–2503.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning.arXiv e-prints, pages arXiv–2503

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:56:07.897597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:87dbab5a3066af1277b4473a7ace3701ecab1043cf1a44a0892cb072a4396aaa

Observation bbcc9512-14cb-4621-bf6e-37c7f8efb5be · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:13:47.672124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:6d550f0f22d1288fe94337c1c3eb61e167f0f3bd0f162e3bd93fa0e529ae2cf1

Pith citing papers

Observation 34216a65-fbda-4284-9c4c-1b6854591650 · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.669333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.669333Z digest=sha256:7b716bc427b5a5c74cb3ac11cced6468c7a9b74a9d6f69de4e3df8d0fe1dce2b

Observation 5c0dda25-23c1-44e5-9915-fb6ca1751178 · inbound

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO cites this paper.

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:52.695914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:52.695914Z digest=sha256:c512b4028163ad860404430f091a4d2b8e8f6b6ab3e130293f5fcf716224eb61

Observation cda6905f-c5b1-4793-87b6-7c120fbe169d · inbound

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning cites this paper.

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:04.492737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:04.492737Z digest=sha256:9b2317f32dd3fa3af84da23548e1c8c019183f4f63a60197baf70b236b5a5a49

Observation f889743e-cbc9-441d-a88a-fbc561be813c · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:18.582137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:18.582137Z digest=sha256:5d2d2e7bac086cae69031ee3a5ac5dc89e570ef8b58690bbf3965bb1bc762991

Observation 553478bb-aace-4b35-baf5-b767d1820f7c · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:09.154751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:09.154751Z digest=sha256:9ce4aae58644f9a9d7b8016bff3ac9cdbdcb054ffee7d6dc4fb5ff4c62db8e6d

Observation 5f7eead2-aa70-42dd-8e9d-189cd6d4bd3a · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.164163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.164163Z digest=sha256:010b218e499c505d63a4ec5363f6f64481d5c850209033c6a1986499be17a77e

Observation 2411361e-59bf-47fc-a2d9-edc8ed537e64 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.558866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:8b5a3c0be253bbe127a7d65de9df29c50c6e5a8c3a882a082b61e1af9abf1448

Observation be058646-4150-404a-a6f0-1aa8b1fe742d · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:40:56.067041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:16815f1a6f59b91c29a5e5ad189dd73f54145647e2ab505212e7585f2d030b01

Observation 0a9a00ae-3fd0-46c2-be4e-a04520f0240e · inbound

Reinforced Reasoning for Embodied Planning cites this paper.

Reinforced Reasoning for Embodied Planning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:23.131565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:23.131565Z digest=sha256:4680ff04f2c21373569441eac77debea7a327460c51c85a313d9b15ec6ce17fb

Observation 0ed9d90f-79aa-4cec-9f55-8a61297044ca · inbound

VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning cites this paper.

VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:06.206753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:06.206753Z digest=sha256:6802085fe3865a6de06216b6ba9672dbafe9b6c4b791c530d229400d2e7d0599

Observation bd845bc2-083f-42e8-b64c-5e2e92bbb9f5 · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:05:52.011365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:b6f74b79965ce5812c62525a31d300186180b8087f192291561408dcd8e8fea8

Observation 059e7857-0160-4348-a977-ab5a86d5b89e · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.976247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.976247Z digest=sha256:e2a57194ed2258f76a7242273e5524e2d20d548bf34a99a6d087e8a5738741eb

Observation 7c4361f0-b713-4b46-a04a-8ee3ac1a028f · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:11.468450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:11.468450Z digest=sha256:e0a0ce8a17dcb08060dbb4b590ebde9346fa45cd7b37d64905b8d4267095d43c

Observation 2c600c9c-ab3a-40f5-b0e5-da9bb6a2016c · inbound

SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization cites this paper.

SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:28.670761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:28.670761Z digest=sha256:4f029c93c74f2f81e0c305c4bfc7be66f76afe805d61a885eb9dbaabdf2a66d5

Observation 00491f7a-f816-44ec-89d2-873dd68e7f75 · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.120821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.120821Z digest=sha256:51185ae248ca2a221f45d6c9d56468e6e0e07209c09992f1dbb150200594179d

Observation 79e302ca-b92a-4cb5-9391-af5c5397bfa4 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.594716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.594716Z digest=sha256:5a2a4130673f47f97097f25ff598d3cb51bc471dc2a22ce34463de28a268513d

Observation 7e2c5b33-baa2-47ee-8ec1-9c08b265e0d4 · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.275756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.275756Z digest=sha256:3d5d0d71818f66da97eabe89022dec03b98c8772902822a0e6c90187b177b2de

Observation 9da0316d-eead-492c-9689-d5a407928dfb · inbound

Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward cites this paper.

Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:50.358423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:50.358423Z digest=sha256:4f1029d939caf53f8e091594cf1824ffd332da5d24bd319b91e0ca1e10260627

Observation 56be8d9d-0980-44d9-8a08-da5276953aa3 · inbound

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation cites this paper.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.648654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.648654Z digest=sha256:ebdc681964aa9d12c4045538c3966ba0dd127d8433e6baca129133a2a87a3eb5

Observation af25d190-2708-432b-b466-11303fba6d59 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.157304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.157304Z digest=sha256:21d5fcae2905f44404afaf6ba1a07bbf69bcecc8b2a9c7e37415050736caf32a

Observation 37320b77-74e9-4433-921e-1f4164d9d5ac · inbound

Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli cites this paper.

Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:04:27.053881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T00:02:41.893373Z digest=sha256:3a2d46fc61aa6fa2a613af8e879e0b76d2c3db9156f9add926bda6aeca837a24

Observation 85e6105e-b847-47d6-90f3-150fec9a4228 · inbound

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought cites this paper.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.381964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.381964Z digest=sha256:508c1b662b19c9fe43063fe6fc71b224851eebb077d2b68fd4def12c26ef377c

Observation 96a7f6a6-d097-497d-887f-4378e5386efb · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.898375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.898375Z digest=sha256:90ef5d573113ac94af7a232a21f91d306a926ce8c1191a5d5c4146a9f81c2675

Observation c4bc7e62-d089-4692-9099-0801bb3d77e5 · inbound

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning cites this paper.

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:52.705443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:52.705443Z digest=sha256:cfea396f1ad4f847ecf92662a175c1962175ea1e4f46dd4ae088a328d75f131c

Observation 1c2f0ad6-3756-4909-bd74-6cc5f8cc0ade · inbound

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training cites this paper.

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:27.329895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:31:27.329895Z digest=sha256:ffd019b7db79e16759f9f95079f8d7fcbf1e811d35d4c529e32e2a7f532f0a17

Observation 9faaa57f-7cfc-4f53-a10a-cf3e13efeb9a · inbound

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning cites this paper.

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:48.153646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:48.153646Z digest=sha256:422341c3d81d43ac9a4f75fa4ff1e6d360411e257c90eb6fe720c5797ccddf5d

Observation 1b5c0a99-8510-4849-bb84-3a0c4d6132a3 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:29.917252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:29.917252Z digest=sha256:86fb133d0f604eb3478784e75ff61ffef6e7b64de133eaabacc616c9e353f9c5

Observation 4b2f041c-56c2-4491-8c7b-7e956d196ed5 · inbound

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding cites this paper.

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T22:01:25.954563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:01:25.954563Z digest=sha256:8cca238d76179a8587c7f2e8c63f2208a4603f959ccf025ed0b10bdef38d7e2a

Observation b82daba1-c250-4a89-b780-db46bb8c4e09 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:18.043392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:18.043392Z digest=sha256:4f7408c075bb90ef4536729122c83bf1b90d5fb72369b0c24f67550877596bdb

Observation 23cb6696-7f93-4766-a4bf-49a5c7d15e2a · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 289

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:11.305339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:11.305339Z digest=sha256:1ce66f5a72955ee2d85a4fc467609871218426869d0c9b1a827014be17bc922e

Observation 401f72b7-170c-4bda-bf6e-900197acb79d · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:08.739671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:08.739671Z digest=sha256:6075b70143add53c2ae424ddc8d0dc58555306471611690f8dc987ea13282505

Observation 5b5dc49a-a1c3-44d3-a5c9-676f2c6b03a1 · inbound

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding cites this paper.

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:16.757814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:10:16.757814Z digest=sha256:f52ef6520ac756fcbdd3fbd41f9c36e5531284f9e75936f96fbc426fdd9b6942

Observation f02cc425-c245-4bfd-a4e3-6ef5769889ee · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.966796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.966796Z digest=sha256:a7a2a7de807b73fc4142d7120a4877001ed3a6ed942b823063dd5a7054813855

Observation 0069d9b7-6e4f-4340-84fc-ae2d1be8d8c7 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 288

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:24.771758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:5092946dfdd651ffc2df16e5145efc4d318385ec178acfdf1902ff0169a76d0f

Observation e23bcd9e-c412-4c1e-8920-f9b1cf979d0b · inbound

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? cites this paper.

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:41:37.838936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T16:40:16.630602Z digest=sha256:02c9133aebead2b1b31b561e75517fa31726a9dd8fde693e0fd7ff1976ede4b2

Observation 239f54c4-f565-4212-99fa-7e479d963c5c · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:39.319465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:39.319465Z digest=sha256:f2b3b10bac888437401d190b0e9e586929c95dedab3800c58042267682ba8979

Observation 0de31263-5f81-451a-b9a4-92328d93d3b8 · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:08.076869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:08.076869Z digest=sha256:9aea6409f85d37314fc55a64e528423bde0581de6a7426c22a1e34f94bab6ce2

Observation 19ba47a9-2a76-445d-bb0d-80d45cbb1acc · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:25:22.662856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:c96295e8243e8e1df500636a615d989078ba2d637ac06508c9ebf11fe81f09dd

Observation c7eb89a7-6f99-44a6-b40e-10bb3b0d5418 · inbound

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models cites this paper.

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:59:08.539859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T05:55:11.495430Z digest=sha256:cfca3c584cc8297d36b475bc597f2b9519559ca37419c2dadf719dc8c895e30a

Observation 278caf4c-6ebd-4563-beb2-ea12b21b56cc · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:04.532245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:04.532245Z digest=sha256:d57a6ac5458dbb38fccecc0bc41e076c657056038e1c0cffaf410becaea1cd0c

Observation ddf0c330-017f-4c29-bea1-aa1305376dbb · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:31:32.080338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:dbe6eb0f9a17d801aecb1063ac81e0872f818bb3f574afb1e4e60b6f285b2e3e

Observation 1cb2b924-c12a-4ca9-96ca-0fca3369f606 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:11:26.483635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:2576a4cafeac64fa8c712563ad8d7d4b04f6fa39d3929afd646a279296ebe205

Observation 272ecf85-534d-471f-9d58-5b4347c00ece · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.299074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:6e1aa9c3a1a81f7f74e4f23579231e3544bfc380d290304e9c715285f1c5a5b2

Observation ebd662ae-d3aa-4f92-8dd3-71ab58479374 · inbound

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning cites this paper.

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:44:15.958395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T16:43:11.995960Z digest=sha256:685521fcaf2470f170a909c4de5973d92355537e88d969a3c5ea483a305a48fd

Observation 53febf27-0cf2-4e48-8eeb-49912ad75579 · inbound

AdaTooler-V: Adaptive Tool-Use for Images and Videos cites this paper.

AdaTooler-V: Adaptive Tool-Use for Images and Videos VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:28:34.402101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:23:33.598026Z digest=sha256:a6ddcaa8a743c3d8a2ee27d642c4c22613f64c79c062a6483664bb3a482f0e22

Observation 9d184e8b-ca35-425d-bf8e-48ef6c188f4b · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:02:42.495434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:340d6b81920166c69d502ab752bdec4d44de7f14f322839b8d10cdf0a7aed372

Observation 7b18476e-6052-409e-920d-660b4714a53b · inbound

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking cites this paper.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:ede0e7adc38547b779ccc72a0bb145b23dee7ac0e22ee64c235885a5e607ecba

Observation e9ed3b7d-bb1c-48e3-8590-3c8bb8682315 · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:04.026500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:04.026500Z digest=sha256:3120ce2ecff0ace211c05d2cba89fb87fb520b13e6bc4d62885b09d4504800a6

Observation f029ec41-1522-4cfe-94a9-50598acb9512 · inbound

Motion-o: Trajectory-Grounded Video Reasoning cites this paper.

Motion-o: Trajectory-Grounded Video Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T08:46:42.967714Z digest=sha256:68bdfe7aed70d2c3ebad06744a09331666c0948b839272aa921f1263b25156cc

Observation 4d382b24-8045-4d30-b89e-405acbb2ae25 · inbound

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering cites this paper.

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T21:12:29.596207Z digest=sha256:accc60153389b7929a193aa9d15571c62380d6ad559fd6269f79d823390b2097

Observation 34f00026-90a2-4eca-97ae-c5b49e4a8cf3 · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:c1a931b13e40fd1f9cb27fe8ee0efd7ef3768d77de62e6050d4c3853456ffce4

Observation 77af5656-4bdb-4564-b432-21db53b904c7 · inbound

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding cites this paper.

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:47:32.778695Z digest=sha256:e454f7cf1fd3001fad036da8ebe224b8e3e4a2de7cd8efb4b292f8904bc7c499

Observation 3cc11797-6e60-48af-ac23-67d680bc715b · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:bde160a755739bac5090f36bec56fa412f899153f3a33b581dda8d2953b79487

Observation c9098f84-ab77-4cf7-9a18-93b7b1ab0635 · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:0821fc143e221cb80dee401c7e7941475cd56ff6c92c94f7d7b68870843cf9e0

Observation c8075d91-dd5f-4480-a807-187cccc0518f · inbound

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt cites this paper.

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T12:37:17.843365Z digest=sha256:810ef9af0161fc834f6a607f13866be3dcb0f9211e9b2ce05a95d59c979e7fbe

Observation e545c590-f325-49ad-884f-785c2cdb5d62 · inbound

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding cites this paper.

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T12:03:09.408019Z digest=sha256:579b5baf997e6493cebf2f649af1352d19ce580e39ab62bfcdd764a7fcaf4205

Observation 90a97d87-3ad0-480b-8686-b9de36f8e569 · inbound

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding cites this paper.

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:37:41.747803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T17:34:10.344111Z digest=sha256:bc98daaa8eb850459bb2ce3923c7db6a06d96583c0dcc5126d3addcbd381e283

Observation f95fa5a6-1185-4170-9278-7b5cf2a48a97 · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:11633051aa180e2e7fff419ab99e792e6e43f92c5048072bc80bd71a0541c7fa

Observation ae99d3ab-70ef-4d78-9908-1c8ea09b7610 · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:7a1e0d1fe48857c331277508bb0fd7bb1208b16af5c1631bdc2106de3f86cffe

Observation 35e59ea5-74f7-4ecf-9c2c-1018e7a83f2c · inbound

Co-Evolving Policy Distillation cites this paper.

Co-Evolving Policy Distillation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T08:23:41.819485Z digest=sha256:f3212bcb94c9a69758cbcb76ed441df5865983e5533e255c9ecfc366267644f3

Observation 298cc5d4-2e2a-4be0-a24c-22bba090a0d8 · inbound

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs cites this paper.

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T14:30:56.297653Z digest=sha256:607bc7b5fa42f1427f259b88013181e3a8d2e6c268404b5b7722d6e3c3f63760

Observation 11afa599-ce7f-4970-a977-cc8a890fa03d · inbound

From Priors to Perception: Grounding Video-LLMs in Physical Reality cites this paper.

From Priors to Perception: Grounding Video-LLMs in Physical Reality VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T17:41:23.233366Z digest=sha256:e8fbdb6ad3c0ea52e5cac112fda5395c567304e80999497b65f33777dbe1dda9

Observation d487b742-4795-40b3-b9b9-98fd84932553 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T14:06:27.953376Z digest=sha256:92dac2bf4f6896df6a396d2992d86ea5b55c1ec6fe285783c10ac108f99f0f64

Observation d2798d8d-f70f-49da-a34e-c84744afd231 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:49:41.654207Z digest=sha256:a61ac680017721f7ec1df651c28fa1dcf803cd25361b36f352ee332a5f9a6aaa

Observation 25dcefc6-2205-4ac3-8f2c-0e71b1f10f61 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:35:59.553683Z digest=sha256:4b4d6b6a310082afcab902fbf79ea48ebe93700d5ad6bc2e921b42f213bea233

Observation 6dfabf31-5cd1-42df-85d8-5c72b2de8dc1 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:10:24.035073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T06:08:19.956833Z digest=sha256:6d66074a3783e43d2a413899716778e50db65fdcf3ff373b6aadc49f64fc6681

Observation 3d089cef-4ebb-49da-9599-7c23c46f5dd5 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:ff23248a76e1f08973b445d44292c6375c9eef526c8e7669a36e609a42618717

Observation 7438fca0-9a81-4291-bb54-64b8093f02db · inbound

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection cites this paper.

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:07:29.463188Z digest=sha256:a28575b212c2fdbb2acc243beef09f7a8ab264c3d5ae379e1fae7358d6ca8c41

Observation 1e8af7f8-6af5-480f-b5ef-b5f13d186f03 · inbound

AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding cites this paper.

AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:43:29.123615Z digest=sha256:25d64ea7a67bcf90741cdccfb90f92e9d98344bc6294811abe54b5ba03d0e3bb

Observation 59026968-ea35-4332-85db-263a2d6e1a6b · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.931779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:ab8cc0b0baeff436b1ae541f69e6e21a4f50d1f185db1c1ea73199a14b3e31f1

Observation b090bbc6-a4e5-40b9-95dd-473201a05b45 · inbound

Video-Zero: Self-Evolution Video Understanding cites this paper.

Video-Zero: Self-Evolution Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.414548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T21:32:16.939563Z digest=sha256:7998254c91721186764d26dbfc6d0d68421a82939bc4a453c8b1642570984c6b

Observation fc5a4d8f-d159-4562-a050-5776a51eb635 · inbound

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation cites this paper.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.110678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:1a078bd85a3b1d56c43152d062f1079edeadba7a316390f4198b410e215f79ba

Observation 5cf5bd21-1f17-4c31-a076-f9069932bc32 · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:28:04.421488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:ff49736fecc235569405690f3f57e3f467f35d9b5d6f29b073d78f26cf642a8a

Observation d0f36932-0877-4d83-8c1f-840b20b47bd8 · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:34:02.634001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T07:32:12.180233Z digest=sha256:0578ecb5d9244a5ba0350d24d535c1e21b6fb6d627426493f94c642dab5a07be

Observation dcd26375-f579-4fbb-a128-30fe1406fc17 · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.480880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:efb909b469dd042f2ab8efaaea1816649efb90abc732b8919bafaa67eba42ebd

Observation f405f0e9-49e6-405a-99a9-31353a6283af · inbound

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models cites this paper.

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:21:12.916335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T07:19:30.508843Z digest=sha256:89f54968ba1edc919716fa372ded514609816bf414ee8106bcb861cda6bd14c2

Observation 5d32eb89-951c-4883-8c22-8b87bfdc0195 · inbound

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues cites this paper.

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:14:42.493677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T07:13:43.716510Z digest=sha256:a98ac91009445ef31716e5c27a8492d8a2b41f4b89b4fe26157b191752196dd4

Observation 2fe30dba-9dfb-4ce8-be48-5cfac79a904c · inbound

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning cites this paper.

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:46:14.920464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T07:45:56.473188Z digest=sha256:4b107307bdd49886f442959af4ad44b9d7304fa9d97baa36b8a17711ec7c38a3

Observation 6753f6e8-e03f-4a03-bfc0-89edfc8d4f62 · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:55:24.906324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:f5c3e3723b70b33641aa71ff8ea0baa7b5bb22913402b06ae780f7b29339e3b1

Observation 18ec9111-7679-4de4-a9be-4565ad8735d6 · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:55:23.105864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T04:54:23.077914Z digest=sha256:8fa8a397f2b352c65241be98c24019b5fddea69392cbf213050c6999be30f78f

Observation e65074ca-46b3-42d8-a862-1e983ea558ac · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.441865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-04T00:41:02.284215Z digest=sha256:7d6f903748c75db9c00abdb5649e1a2d384794523b3b8553081e1f8932d1ff4a

Observation ec21962f-3b64-4db1-883d-9015dceab48c · inbound

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding cites this paper.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.109477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:76a402e731a74b21fdd8d0e55bd48e0654f1286f2dce308eaa195829627e6abf

Observation dc1994c4-aaa5-456c-8f28-57128a758db1 · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:23:28.488984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:a4fbd5ed68cb376ebc59d31f777a10f18915bcb7c7adfc3e0d602027ae3901ea

Observation f36497b5-429f-4978-bb95-23677da03585 · inbound

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events cites this paper.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.898675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:0eae69e24960d642e9f3e7183d85f77aaae382fd84d2ccf0b9a4b49e94e2de9b

Observation 556d55e7-8c26-413f-a7a5-2789fbf16a06 · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:56:55.479185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:2f3f7afb24523df951992647d38af028e40ed57e40ced03c509324d216ba4935

Observation 94a384c8-21c9-42f5-a36b-10c117f93309 · inbound

Towards One-to-Many Temporal Grounding cites this paper.

Towards One-to-Many Temporal Grounding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:16:57.732756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T02:11:48.455492Z digest=sha256:e02ff389ecb9393d9de3bb1f98cc6f7e370bcee848db90398c29b41f11e55925

Observation 7f614dd5-8a08-48aa-ba7d-a4f47eecdcf0 · inbound

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA cites this paper.

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:57:30.184032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T16:55:35.743040Z digest=sha256:2314b07cc2ca23a9433a89df484b83c740cbd1ed7ed904e3a7ec0041263aa7d7

Observation b9d5fb77-a970-443d-956a-c03bc3b0ed74 · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:58:57.736106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:b17857487a8333dedaf31ac84560741b6d799ea11e58ffa0db01af6f6fbeaecf

Observation 14397b1a-ba52-4726-8172-2f7b22fca07f · inbound

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding cites this paper.

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:39:58.305025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T00:19:26.153682Z digest=sha256:5d9608ffd61aea4f47476118b4bfd1b435c5b5df966835c01163bfcee2dd2e6e

Observation 4a016ead-b161-4f75-9c21-69231b25c231 · inbound

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards cites this paper.

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:49:57.207524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T00:16:58.697810Z digest=sha256:c3da86a15883caa53585bead0f0c467551297003130795912c426c8089d56b48

Observation 914c7baa-e9a1-4992-99d0-62cdf3a31be1 · inbound

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence cites this paper.

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:57.436227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T00:15:02.499635Z digest=sha256:1b9f1bc6de7bf738fae94a79279ca76d218390ac6d1a9b95fa5f31eaba73ec87

Observation 30d2d600-87d3-489b-a7f9-dd3126e21177 · inbound

Confidence-Aware Tool Orchestration for Robust Video Understanding cites this paper.

Confidence-Aware Tool Orchestration for Robust Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:19:50.412449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T05:23:06.035263Z digest=sha256:61c4efaaead3ba3e84707200e4d710b244ae8d5c5ff1568647e00d579fdd2ca9

Observation 681b6f61-0b12-4792-bef6-8b8f45d5e7e0 · inbound

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs cites this paper.

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:34:19.518606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T06:25:38.593423Z digest=sha256:1348498cabbb2c0a9bda8363182f2776de08636633967a094c2c9535c02a6eff

Observation 91b7455a-3319-4643-9018-ee038b7dfbc1 · inbound

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context cites this paper.

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:04:21.278200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T06:01:49.904752Z digest=sha256:de53ab8cda9ab3981ee2f30745c8e28410b4be22b6926886acfc7d4a1648d59f

Observation 00312180-ed44-4d12-944c-208c8c81fd53 · inbound

Incentivizing Vision Language Models to Search for Long Video Question Answering cites this paper.

Incentivizing Vision Language Models to Search for Long Video Question Answering VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T05:50:16.895740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:50:16.895740Z digest=sha256:32c7027521aea715fe507d4fd8bf7c1054f412c0a0f27fc1f3d7816505afbc64

Observation aee1a0d5-b869-43bb-9627-ed8ab4b7ad00 · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:c91a75982bd2007cf0d028f9d80ba305c7086f3c78dd3eaf09596cba79faebc0

Observation 768652a0-16ca-45b3-a088-c0df4cacf6f0 · inbound

SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection cites this paper.

SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T05:07:18.364787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:07:18.364787Z digest=sha256:2a0dfdc8f095f4b92a59606e25166dcf33cfcf84e3952fca8a15804e400ed1a6

Observation e262efdc-a516-4bb8-8808-a1c2677d9af5 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:ccab29e7dc9f70ca74feb11b1fa77e131923e1bb641445013ffc594aa1493606

Observation 17d442bd-d18f-433d-950b-b6bd4e6db1e1 · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:50.850439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:50.850439Z digest=sha256:6c6b11af443a6d208bb01015489c90f893c24c67b5dab0dcf40cfb73a0b88158

Observation 2d96c88b-f6e2-46a0-8165-73e347507928 · inbound

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA cites this paper.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:43.463242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:43.463242Z digest=sha256:36536a54f3a630c7726ff51f0102c5eb0c694bf56aa88f9ceb73a996e84d56c7