Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T12:26:35.347190Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 23 inbound Pith citation observations for arXiv:2511.20785.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T12:26:35.347190Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T04:44:30.446320Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T23:57:29.009867Z
78 of 78 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1803a682-99dd-4f8a-8d36-d60618b5508c · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6a8aaab0-b1f6-4fde-bec6-dbbb05df58b9 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 33645a7d-b6ce-4652-9c36-c75bc0cbcb9f · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Eliciting good teach- ing from humans for machine learners.Artificial Intelli- gence, 217:198–215
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2d58c8e1-1d83-48dd-b329-731707038f6e · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Scaling rl to long videos
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a31ee9d4-276b-4c1e-ae4a-d15a60da42fc · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7375c2a6-00eb-42f6-949c-aa42fe7a2fce · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation defd5cf4-3f2d-4a55-a75d-5f90f4a6b999 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Grit: Teaching mllms to think with images
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3370bf28-1857-433f-94df-791457f86533 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99a61f56-340b-4303-ac3e-1e475c8eafbd · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4a9e2d1e-a724-4175-a50e-7ecfea661c8f · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Tall: Temporal activity localization via language query
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1860182d-18fe-4106-a32e-c950dc30d4d5 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d56bec6f-774d-4c95-9ee0-4f1f46841545 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T17:38:13.144769+00:00.
Observation e7c55f0d-303b-4a64-9a88-1f4fcf08e9cf · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation db538753-134e-4d52-bcdc-7e22f7d36f1c · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Multimodal Pretraining for Dense Video Captioning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4b039082-438e-4e4f-b285-182a2e3320e8 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 75ed6b6b-8a84-4e41-a107-081a0296402e · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling GPT-4o System Card
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e013938b-4d46-4e6b-9aef-0ee10b22a32b · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling OpenAI o1 System Card
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a40fa3c1-a31f-40e4-bdab-6b820b5dcc20 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Dense-captioning events in videos
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3d5e9c8b-9e50-406b-9a2a-1182e14119b3 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Efficient memory management for large language model serving with pagedattention
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e162c062-dcd5-40dc-8676-d4dfd24e22ec · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Mmr1: Enhancing multimodal reasoning with variance-aware sampling and open resources
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8f292c1d-1083-4035-b4a8-109329d1dec2 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ecf324f7-8933-45cc-be95-8205f0c92abd · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Getting more juice out of the sft data: Reward learning from human demonstration im- proves sft for llm alignment
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 50c9cc20-5e77-411b-9639-32ae51d09f64 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ddf0c330-017f-4c29-bea1-aa1305376dbb · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e5e01f37-b4c6-474a-ad85-93bcaf1b25e0 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Improving LLM Video Understanding with 16 Frames Per Second
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6dcd0194-5192-4bb7-a88b-9347000e7bcf · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling TempCompass: Do Video LLMs Really Understand Videos?
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eb9ad06b-178c-4d09-a0f6-cdb573499716 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a812e6fe-f005-4658-a14e-be7ad891c435 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Visual- rft: Visual reinforcement fine-tuning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9610c8f8-e10f-4c9f-9e40-f77919ce2090 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Lmms engine: A simple, unified multimodal framework for pretraining and finetuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 287c10d8-49d4-499c-8759-9ad328db39b3 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Decoupled weight de- cay regularization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e830ca06-9952-4c82-8162-2386324ae9e8 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16d2845b-c608-46ab-a7b2-e6df01955518 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Multi-agent tool-integrated policy optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5703c378-62c0-4dbb-92fc-9865edcc08b0 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 019fc509-e74e-4ac3-a632-740d465dc066 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e850eea-5ab7-47b2-8dad-97170513fb76 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation df2d44da-ba2d-454b-b5ea-d86e7af558d9 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6c832576-f115-479a-9c46-b7c4754a6207 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Hybridflow: A flexible and efficient rlhf frame- work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10f7f488-b341-44b6-9381-09780649fa15 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Moviechat: From dense token to sparse memory for long video understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e796a7eb-080e-45f2-a817-76cd249db505 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5c9dc37b-4062-4914-bf2e-5acfc1754921 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8b062510-547a-4e97-8dfa-59e07ae1c9ca · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a03504dc-a5b3-4e47-ad21-f04005391ee2 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Introducing gpt-5.https://openai
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e94abd91-47ea-4c4a-99f6-47dffcb31e31 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Thinking with images.https : / / openai
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 01a1aca1-5554-4d88-99b2-06a043325214 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Qwen3-vl: Sharper vision, deeper thought, broader action.https : / / qwen
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 25424899-9b6f-4a06-8cc5-23aa0f89dd3c · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3c1eae4c-1337-4205-8701-1199a1732dfd · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3d31a6b-3bf4-423a-8d39-35b9d2f45f75 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Video-thinker: Sparking” thinking with videos” via reinforcement learning.arXiv preprint arXiv:2510.23473
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 19aa4829-85a0-4bfe-ad7d-51b8bda8b69b · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling LVBench: An Extreme Long Video Understanding Benchmark
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e0cb23fb-6be2-490f-aad7-b67241350fb9 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a526b2e7-b85f-4423-9f46-b3b7a97e27a2 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 171533f5-08aa-41cc-a06a-e0d09c4a442f · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 883dd067-7b04-4678-bc66-34470e1a4def · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 08680bff-c31c-449f-8452-ea23d07eb820 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Llava-cot: Let vision language models reason step-by-step
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 74e7bc5f-527d-4477-b757-a666d5559e89 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Vidchapters-7m: Video chapters at scale
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10228c5a-7805-42a1-95d5-5b0c4b9a0011 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Qwen3 Technical Report
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9865d19b-5c60-41d9-b3bb-e105a21e42ff · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Mermaid: Multi-perspective self-reflective agents with generative augmentation for emotion recogni- tion
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c4e0aa3-df5e-4ede-81a3-84bc2f0d80c2 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Timeexpert: An expert-guided video llm for video temporal grounding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 746dd75b-1364-4a2a-9759-2f5224962600 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation decedb93-683f-46c0-ade1-39bc7d6a6c37 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 83322ef6-a248-48a8-8aac-fb5d57a6bb21 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Lmms-eval: Re- ality check on the evaluation of large multimodal models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4fc2f709-6320-443d-8391-1827576d54bf · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Open- mmreasoner: Pushing the frontiers for multimodal rea- soning with an open and general recipe
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 12dcbb22-01b4-4ab3-adb3-bf8dfcb4a39b · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Sglang: Efficient execution of structured language model programs
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e825e151-6870-42c7-859f-199a1db069bc · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9a349702-83c9-405d-ae4c-4abb8a12968e · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dcbe1b05-3af5-4d01-abf3-0068c9992cbf · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Thinking with Long Videos
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 63cd8ef0-b944-4d84-b09b-6b7da544ddbf · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling global- to-local
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c0c672d9-2f1b-4278-b36a-6728a0236cb1 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling black-box
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61b5ff77-1fa6-4418-87c8-73736968a10a · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 002d6096-c4c9-4e69-b0f6-5bce9fd7ae37 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling For a sequence of to- kensx= (x 1, x2
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5ad3a10-90e2-40e9-bf04-fbcf1d36130c · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling segment
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b6f0680c-c77f-4eb0-9f8f-7dc3203ae0e9 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling of Training Steps 3000 160 1600 No
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ab04516-4e0d-4549-bcbf-e862db263788 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling To optimize training throughput and mini- mize memory overhead, we employ an online stream pack- ing strategy on iterable datasets
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 38bf6c0c-35b3-47bc-93f7-c592a6dd817b · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling blindly rephrasing
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8a7f044a-4386-4568-936a-957187704e86 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Figure 8 shows the RL prompt template, while Figure 9 presents the evaluation prompts used in LLM-as-a-Judge [55] for measuring an- swer’s accuracy during RL
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4ff94d2-b6dd-45af-a3d6-42b9ee544eb0 · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling which video-game device
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8e9e3b7e-3719-488e-87ee-e72e6c5cee4d · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Manager Agent
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 96ba2574-6902-491f-97d7-db468b8068ce · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation edd5039e-2c14-44f5-aa43-f6f811f9e65b · outbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling type\": \
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8fefce4c-9f16-436d-8401-14a1fc6533f7 · inbound
Better Call Grep: Evaluating and Improving Grep-Like Lexical Retrieval for Repository-Level Code Completion LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1853d597-68bd-479a-9d74-f4b5a5de2089 · inbound
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6fdfb335-2546-416f-9041-2116e76163e5 · inbound
Towards Temporal Compositional Reasoning in Long-Form Sports Videos LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c82f0a56-b982-41c8-abe8-02f30c77477d · inbound
Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9c9cd432-51a9-4eff-b391-3f0a94580521 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 87b5d782-aa15-4247-9fbb-5a5496436de7 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d7afae0f-6d56-4e21-b48f-558fcaf230c3 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 27f57a54-0925-4793-9327-0f5a7c295afc · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5f715dfd-16b0-4121-bb77-847e8cd17f85 · inbound
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2d7aa7b6-1fb2-4c32-aab1-d57edb0175ce · inbound
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f89b4fde-71dc-4784-8cb6-0e114f505ec9 · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d494c51-66e2-4e41-8b30-5e2ce9880069 · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f2a19d64-227e-4966-b2e5-461947f50484 · inbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f51a4f7d-3e4e-4c37-8c49-fe4634c4e81d · inbound
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61dc3136-4bae-4c4c-b7e3-41a7414eda5d · inbound
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 72dbfd35-0cd6-4c54-a29b-e77cf02b29ce · inbound
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7fb84d69-481c-4343-82a3-1c1197003617 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 225
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 01eed325-0932-4037-8540-251c083a5513 · inbound
See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 33581240-59c3-4c43-9c98-6cadb8054af6 · inbound
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 76827ce1-ff53-47e5-bcc8-1c7bf879aab3 · inbound
VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ffb45fa-d594-474d-aa85-e13fd2a50c35 · inbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f57e935-e817-4ec8-8ad7-79753866fcaa · inbound
Thinking in Video: Can Video Generators Really Reason About the Real World? LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fba83264-40dc-4253-a0ee-d9b25e798be5 · inbound
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.