Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T20:56:07.247122Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 100 inbound Pith citation observations for arXiv:2504.06958.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T20:56:07.247122Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:41.669333Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T16:49:57.434938Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fa300db1-6c7e-4328-b28a-2e9b467ab154 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3be6774c-fefb-404f-8bcc-7cad0151f320 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2c866938-5ec6-4e64-b313-5723a5276481 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8a4e1c4c-5e79-476a-99f5-59a5023b75de · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 565583f5-7cc3-4243-877d-b51e82a91195 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b1ed0ac3-259d-40bb-88a1-bd76ffd01727 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 04fc8b7a-aa16-4854-b090-bd69702cb522 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Tall: Temporal activity localization via language query
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49cb3267-5edb-4867-b5b5-dc8809246fac · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Saliency-guided detr for moment retrieval and highlight detection
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0949eb1c-3dcc-4606-8923-aed3099d416c · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ace0399-f7f9-47d5-8d00-07ca26340e47 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Got-10k: A large high-diversity benchmark for generic object tracking in the wild.IEEE transactions on pattern analysis and machine intelligence, 43(5): 1562–1577
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8c402674-59cd-4bb8-8c71-e557319a38ff · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Online Video Understanding: OVBench and VideoChat-Online
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b2d7e38-b69b-412e-852b-4cdb9c56169f · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning OpenAI o1 System Card
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7365848c-3171-4ddf-af3c-d18e296cf342 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Dense-captioning events in videos
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 660889a1-6a46-4982-a916-8f1ef2e247ea · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoChat: Chat-Centric Video Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 38592194-23db-41de-bebe-606b597e70fa · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d74d232f-e4f9-40be-b7c5-f9d0342000ba · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 304f2a4f-3440-41b9-a78a-5c5e63b912ff · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b1f451b9-5f8d-4ee7-9797-66ffb66a0549 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ec234bcd-674c-4608-888f-c610fe3f33f1 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f5ee1669-04c0-4714-a61c-7722cfde4a38 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 552140d7-f2db-42cc-b8c9-085333b56e8d · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36: 42748–42761
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 60953261-4d9c-40fd-a447-9f79c74b7298 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 829f1d38-6a9c-4983-ac44-fafe008a183c · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c409a1c8-e932-48b3-aa22-b95800169138 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ba35c61-2595-4c3e-81cc-798898ead23d · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2887b8fb-5f10-4006-aa6c-aac2821363c2 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8470cac6-2907-4aee-90c7-e6e76052aa6a · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 08e619e6-26c9-4e84-aa76-f1dfd6212508 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Internvideo2: Scaling foundation models for multimodal video understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a0ccc2c1-0ac2-406c-8c64-f23499a54bad · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd8ec7a5-ec15-40f6-b610-69e30b6d4c39 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37: 28828–28857
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 04a206cf-250a-48e2-ae51-bf2eb5a759b9 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Can i trust your answer? visually grounded video question answering
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 262fe47c-c479-434d-ad18-cf9fbb39a787 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 38640935-a25c-466a-b623-b70ebbdb310d · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e3e3247f-29a4-4ae2-a1fb-2f6d60552615 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Qwen2.5 Technical Report
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ab973397-268d-4593-8393-2e207b42a010 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4eb0ca53-8467-4220-afe8-7fa34cd6c1a5 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Merlin: Empowering multimodal llms with foresight minds
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 159f87b7-00c8-4b9a-ac7a-9274e28fa6f9 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dd411d1f-0a5f-44d6-b5df-ac076a45f325 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b1cca0fc-c2a3-455a-86fd-09981ad80142 · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3caa9af6-8bcb-4af9-a24a-5ae85ccc00ba · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 01cb1b88-b7a3-439c-8d34-65dd1858338a · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning.arXiv e-prints, pages arXiv–2503
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bbcc9512-14cb-4621-bf6e-37c7f8efb5be · outbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 34216a65-fbda-4284-9c4c-1b6854591650 · inbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c0dda25-23c1-44e5-9915-fb6ca1751178 · inbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda6905f-c5b1-4793-87b6-7c120fbe169d · inbound
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f889743e-cbc9-441d-a88a-fbc561be813c · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 553478bb-aace-4b35-baf5-b767d1820f7c · inbound
Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f7eead2-aa70-42dd-8e9d-189cd6d4bd3a · inbound
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2411361e-59bf-47fc-a2d9-edc8ed537e64 · inbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation be058646-4150-404a-a6f0-1aa8b1fe742d · inbound
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0a9a00ae-3fd0-46c2-be4e-a04520f0240e · inbound
Reinforced Reasoning for Embodied Planning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed9d90f-79aa-4cec-9f55-8a61297044ca · inbound
VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd845bc2-083f-42e8-b64c-5e2e92bbb9f5 · inbound
Grounded Reinforcement Learning for Visual Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 059e7857-0160-4348-a977-ab5a86d5b89e · inbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c4361f0-b713-4b46-a04a-8ee3ac1a028f · inbound
Reinforcing Video Reasoning with Focused Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c600c9c-ab3a-40f5-b0e5-da9bb6a2016c · inbound
SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00491f7a-f816-44ec-89d2-873dd68e7f75 · inbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e302ca-b92a-4cb5-9391-af5c5397bfa4 · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e2c5b33-baa2-47ee-8ec1-9c08b265e0d4 · inbound
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9da0316d-eead-492c-9689-d5a407928dfb · inbound
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56be8d9d-0980-44d9-8a08-da5276953aa3 · inbound
Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af25d190-2708-432b-b466-11303fba6d59 · inbound
CyberV: Cybernetics for Test-time Scaling in Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37320b77-74e9-4433-921e-1f4164d9d5ac · inbound
Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 85e6105e-b847-47d6-90f3-150fec9a4228 · inbound
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a7f6a6-d097-497d-887f-4378e5386efb · inbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4bc7e62-d089-4692-9099-0801bb3d77e5 · inbound
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2f0ad6-3756-4909-bd74-6cc5f8cc0ade · inbound
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9faaa57f-7cfc-4f53-a10a-cf3e13efeb9a · inbound
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b5c0a99-8510-4849-bb84-3a0c4d6132a3 · inbound
Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2f041c-56c2-4491-8c7b-7e956d196ed5 · inbound
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b82daba1-c250-4a89-b780-db46bb8c4e09 · inbound
A Survey on Video Temporal Grounding with Multimodal Large Language Model VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 127
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23cb6696-7f93-4766-a4bf-49a5c7d15e2a · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 289
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 401f72b7-170c-4bda-bf6e-900197acb79d · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b5dc49a-a1c3-44d3-a5c9-676f2c6b03a1 · inbound
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02cc425-c245-4bfd-a4e3-6ef5769889ee · inbound
Reinforced Visual Perception with Tools VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0069d9b7-6e4f-4340-84fc-ae2d1be8d8c7 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 288
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e23bcd9e-c412-4c1e-8920-f9b1cf979d0b · inbound
TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies? VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 239f54c4-f565-4212-99fa-7e479d963c5c · inbound
RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0de31263-5f81-451a-b9a4-92328d93d3b8 · inbound
Video Reasoning without Training VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ba47a9-2a76-445d-bb0d-80d45cbb1acc · inbound
VIDEOP2R: Video Understanding from Perception to Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c7eb89a7-6f99-44a6-b40e-10bb3b0d5418 · inbound
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 278caf4c-6ebd-4563-beb2-ea12b21b56cc · inbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf0c330-017f-4c29-bea1-aa1305376dbb · inbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1cb2b924-c12a-4ca9-96ca-0fca3369f606 · inbound
OneThinker: All-in-one Reasoning Model for Image and Video VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 272ecf85-534d-471f-9d58-5b4347c00ece · inbound
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ebd662ae-d3aa-4f92-8dd3-71ab58479374 · inbound
Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 53febf27-0cf2-4e48-8eeb-49912ad75579 · inbound
AdaTooler-V: Adaptive Tool-Use for Images and Videos VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9d184e8b-ca35-425d-bf8e-48ef6c188f4b · inbound
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7b18476e-6052-409e-920d-660b4714a53b · inbound
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e9ed3b7d-bb1c-48e3-8590-3c8bb8682315 · inbound
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f029ec41-1522-4cfe-94a9-50598acb9512 · inbound
Motion-o: Trajectory-Grounded Video Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4d382b24-8045-4d30-b89e-405acbb2ae25 · inbound
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 34f00026-90a2-4eca-97ae-c5b49e4a8cf3 · inbound
Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 77af5656-4bdb-4564-b432-21db53b904c7 · inbound
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3cc11797-6e60-48af-ac23-67d680bc715b · inbound
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c9098f84-ab77-4cf7-9a18-93b7b1ab0635 · inbound
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8075d91-dd5f-4480-a807-187cccc0518f · inbound
Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e545c590-f325-49ad-884f-785c2cdb5d62 · inbound
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 90a97d87-3ad0-480b-8686-b9de36f8e569 · inbound
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f95fa5a6-1185-4170-9278-7b5cf2a48a97 · inbound
EasyVideoR1: Easier RL for Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ae99d3ab-70ef-4d78-9908-1c8ea09b7610 · inbound
Video-ToC: Video Tree-of-Cue Reasoning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 35e59ea5-74f7-4ecf-9c2c-1018e7a83f2c · inbound
Co-Evolving Policy Distillation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 298cc5d4-2e2a-4be0-a24c-22bba090a0d8 · inbound
Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 11afa599-ce7f-4970-a977-cc8a890fa03d · inbound
From Priors to Perception: Grounding Video-LLMs in Physical Reality VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d487b742-4795-40b3-b9b9-98fd84932553 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d2798d8d-f70f-49da-a34e-c84744afd231 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 25dcefc6-2205-4ac3-8f2c-0e71b1f10f61 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6dfabf31-5cd1-42df-85d8-5c72b2de8dc1 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3d089cef-4ebb-49da-9599-7c23c46f5dd5 · inbound
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7438fca0-9a81-4291-bb54-64b8093f02db · inbound
MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1e8af7f8-6af5-480f-b5ef-b5f13d186f03 · inbound
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 59026968-ea35-4332-85db-263a2d6e1a6b · inbound
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b090bbc6-a4e5-40b9-95dd-473201a05b45 · inbound
Video-Zero: Self-Evolution Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fc5a4d8f-d159-4562-a050-5776a51eb635 · inbound
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5cf5bd21-1f17-4c31-a076-f9069932bc32 · inbound
CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d0f36932-0877-4d83-8c1f-840b20b47bd8 · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dcd26375-f579-4fbb-a128-30fe1406fc17 · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f405f0e9-49e6-405a-99a9-31353a6283af · inbound
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d32eb89-951c-4883-8c22-8b87bfdc0195 · inbound
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2fe30dba-9dfb-4ce8-be48-5cfac79a904c · inbound
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6753f6e8-e03f-4a03-bfc0-89edfc8d4f62 · inbound
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 18ec9111-7679-4de4-a9be-4565ad8735d6 · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e65074ca-46b3-42d8-a862-1e983ea558ac · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ec21962f-3b64-4db1-883d-9015dceab48c · inbound
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dc1994c4-aaa5-456c-8f28-57128a758db1 · inbound
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f36497b5-429f-4978-bb95-23677da03585 · inbound
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 556d55e7-8c26-413f-a7a5-2789fbf16a06 · inbound
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 94a384c8-21c9-42f5-a36b-10c117f93309 · inbound
Towards One-to-Many Temporal Grounding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7f614dd5-8a08-48aa-ba7d-a4f47eecdcf0 · inbound
Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9d5fb77-a970-443d-956a-c03bc3b0ed74 · inbound
Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 14397b1a-ba52-4726-8172-2f7b22fca07f · inbound
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4a016ead-b161-4f75-9c21-69231b25c231 · inbound
SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 914c7baa-e9a1-4992-99d0-62cdf3a31be1 · inbound
EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 30d2d600-87d3-489b-a7f9-dd3126e21177 · inbound
Confidence-Aware Tool Orchestration for Robust Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 681b6f61-0b12-4792-bef6-8b8f45d5e7e0 · inbound
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 91b7455a-3319-4643-9018-ee038b7dfbc1 · inbound
VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 00312180-ed44-4d12-944c-208c8c81fd53 · inbound
Incentivizing Vision Language Models to Search for Long Video Question Answering VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aee1a0d5-b869-43bb-9627-ed8ab4b7ad00 · inbound
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768652a0-16ca-45b3-a088-c0df4cacf6f0 · inbound
SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e262efdc-a516-4bb8-8808-a1c2677d9af5 · inbound
TimeThink: Reasoning with Time for Video LLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17d442bd-d18f-433d-950b-b6bd4e6db1e1 · inbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d96c88b-f6e2-46a0-8165-73e347507928 · inbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.