Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T02:40:06.454859Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 67 inbound Pith citation observations for arXiv:2503.13377.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T02:40:06.454859Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.215349Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T09:49:44.696023Z
87 of 87 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 096c1a33-4843-4c20-92da-e0ce2affb49c · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e61632f0-3b9a-4f38-ba2c-f454ec0df487 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Ht- step: Aligning instructional articles with how-to videos
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e2ef44aa-49b9-40a7-8309-e4a21faae837 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Localizing moments in video with natural language
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee4c2169-a8cb-459f-a967-4ed71a2d93ad · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c5ea3db8-702b-44c1-afe5-c9b46f2a8aaf · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Activitynet: A large-scale video benchmark for human activity understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 23757ce9-7cf5-4e21-acd0-efa2698440a5 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Quo vadis, action recognition? a new model and the kinetics dataset
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6a476057-7e8f-4ff9-bc39-04119bafc60f · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding R1-v: Reinforcing super generalization ability in vision-language models with less than $3
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9a5d427b-7514-4874-add1-6467a324e001 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f43f72c3-cd01-4541-ba18-6a13ad6c8519 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Space-time gestures
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2c13ca72-7eb0-4508-9ce8-1c02e190ced2 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Gemini 2.5: Our most intelligent ai model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3223d133-1d40-49f7-9453-fb96ce97b930 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fa54be3e-b58e-4f18-aded-7e569bad7d7b · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding BERT: pre-training of deep bidirectional transformers for language understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9d6d75e6-3091-490a-930b-8337c7f39a16 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5e55fd23-8794-4d2e-bc3a-51d5f22a30cb · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Temporal localization of actions with actoms
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 43375593-cc1d-44d5-a288-6f65c76d8cc3 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Tall: Temporal activity localization via language query
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 08d09f77-dc32-4adf-a481-0249d88cf950 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Ego4d: Around the world in 3,000 hours of egocentric video
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 12ebfb1b-4d40-4fd9-9038-7786c0f35795 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 19caab2e-e0db-403e-ac63-bf1a537f359b · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding TRACE: Temporal Grounding Video LLM via Causal Event Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6a400c74-73e5-4ff1-a38f-e7599b192f3b · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Revisionllm: Recursive vision-language model for temporal grounding in hour-long videos
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0bd322c1-83a4-4854-9fd3-4f13e6bcb140 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Lora: Low-rank adaptation of large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1f3d2a80-e189-43d4-a033-07cd3e570de5 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Vtimellm: Empower llm to grasp video moments
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b579597b-975f-4620-bc20-fd659085394a · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Knowing where to focus: Event-aware transformer for video grounding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16c6a277-cee3-4627-93bc-e820a671a932 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Gonzalez, Hao Zhang, and Ion Stoica
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 83ecb689-6839-4cf0-8971-0b82530f2539 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Retrieving actions in movies
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0b25d990-ef0e-44c3-b91b-f7586e9a8437 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding iMOVE: Instance-Motion-Aware Video Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 916126f3-5e6c-4fe1-b6c8-47c8909da601 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 59801fb1-50fb-4317-8a32-1acef7e4adef · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cfa03b87-9845-4c98-a800-14a7f564b40f · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Improved Visual-Spatial Reasoning via R1-Zero-Like Training
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a6fe9708-c2a2-4120-aee6-de805018e3d4 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Egocentric video-language pretraining
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5ead59bf-2a9d-41f1-a14d-1875523d9552 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Univtg: Towards unified video-language temporal grounding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2c8fb89d-1652-47a9-ad68-f5abce5fe348 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding TempCompass: Do Video LLMs Really Understand Videos?
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 229a11f8-6ec8-4c1e-877a-89efe0cfa7ca · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2121138f-f80c-4055-8ca4-e9f064f465f5 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 37a1727a-3f48-47d6-8f2d-c82e3bd5dbf2 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Walk these ways: Tuning robot control for generalization with multiplicity of behavior
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eb2ce270-e81c-495b-b521-a160656a4591 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ae26ceb-0ea6-4106-989c-ca8b99c7e736 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 855c0589-bc9e-4713-9a19-73979d0ec93c · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Snag: Scalable and accurate video grounding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8507b6bd-8911-47d2-94cb-bfacd6a2b27d · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Queryd: A video dataset with high-quality text and audio narrations
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 31298708-6ed3-4cf9-9d6e-ece7d411a55e · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Openai o1
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f1181f51-d28e-4a5e-a059-f36ff1822e89 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Training language models to follow instructions with human feedback
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f2a44993-321b-4f5b-82b5-965363200b0f · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Chatvtg: Video temporal grounding via chat with video dialogue large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc1a7ef4-e5f6-47b5-9a6a-5cab08db7920 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Learning transferable visual models from natural language supervision
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1247a069-af12-456f-a12e-87990e9d1861 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Grounding action descriptions in videos
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c93386c5-7836-4f5d-a4e2-70c2854afbb8 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 776c28a4-9841-49f9-b982-87d8e324bae8 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be34329a-5ba4-4f07-8c99-23204cf7a19b · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Hollywood in homes: Crowdsourcing data collection for activity understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bec547c2-e15e-4994-9ade-b935f7dcc8b1 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7b00d6fe-61bc-4d53-a4ea-d4b73102bea5 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Reason-rft: Reinforcement fine-tuning for visual reasoning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af0bf21b-6be7-4169-b4fe-d7b0c07e597d · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a6aba1f1-2c0c-449a-a5bc-5ba2a720944e · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Hawkeye: Training video-text llms for grounding text in videos
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a188824f-9424-494d-8c8c-37732a744538 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7afd85b1-2212-4dd9-962f-4510d4c0470a · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Number it: Temporal grounding videos like flipping manga
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 47042016-d8ce-4724-81b4-50e670085c40 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dcfcf2b1-ca8f-49cb-9773-83344c729dbb · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4c540c59-2304-492d-9cf2-6de6259cbc06 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Egolife: Towards egocentric life assistant
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 88fb637f-2325-443c-82ff-053f31c17ad3 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71abcb1e-a5fb-48ea-9a87-c8f7c08a6df7 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 770d5252-9b0f-4a9e-a152-8dbe3e983746 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding A closer look at temporal sentence grounding in videos: Dataset and metric
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a30f710b-2cf7-4679-b3d3-eb390e898793 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Hierarchical video-moment retrieval and step-captioning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc3e341c-d304-41c2-a4cc-9a7d84a37ee4 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Timesuite: Improving MLLMs for long video understanding via grounded tuning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aeb2920f-7f7b-4e2c-b315-2eb2ca6f9ae9 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Temporal sentence grounding in videos: A survey and future directions
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3473930b-d8a9-4e1a-930d-9d2f910109a6 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Multi-scale 2d temporal adjacency networks for moment localization with natural language
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 214a3c37-a574-4ee9-bb33-fa83f116273e · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Learning 2d temporal adjacent networks for moment localization with natural language
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c6903ccf-ac39-4c50-b14e-b931bf6a40e2 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d49be883-0dc4-4ac5-91cc-2ff3762cb4ee · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ecea9b7-6894-4e80-94fe-aeae55a5af41 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding goes back to the pink bucket
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16676725-fb9c-4b0f-80c1-f5b102fc4186 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d3dec532-6633-4b06-a093-61664ada7bda · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2fda06f6-df62-4bed-9597-0fcfa4e3d564 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d3f4fd69-5c3c-49fa-afab-5d40bde1f321 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 50fe21c4-4fd3-49bb-bd5a-fc8cd43834d6 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51f5e0d6-26d6-42a5-b103-e092533179f4 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Given this analysis, the pineapple is indeed being pushed forward by a person
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d51a9da3-bd7d-4d81-a399-ea01118c2b79 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding This is the first major action
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d5823636-cde7-441f-ba7d-ad0cd4062dc0 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1ff8aebe-eebf-4020-b0be-fa2116c3ae52 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8772ad1-80f8-4ead-abaf-819868dd0505 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Now, let's evaluate the options: (A) C folds the dress, places it on the ironing board, and then hangs it up
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0a1e88d2-239a-4441-b0ca-8298de99277a · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding - Examples: - person opens a book over their head
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 43495fd6-4926-4859-bb8b-d2a7b69a253c · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding - Examples: - He is talking while several people are using rowing machines
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f0c8edbd-42ad-496d-bfd3-359c92b7a16a · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding contains multiple actions, each with a clear start a nd end
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 555dd73e-5504-4304-b496-108255165969 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Posture descriptors, positional prepositions - Examples: - Several other people are in the background working out on the equipment
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 673f6430-6d99-4bc7-9e00-dfb4238aa5e0 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Simple location prepositions
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6f409486-3f7e-456f-b787-0e685e4f1111 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding after/before [action]
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1148f522-c872-4dfa-8472-11532f8d8ef4 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Property descriptors (color/size/material) - Examples: - what material did I pick from the shelf? - what color is the toilet bin?
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cf68276c-5f57-4f5b-bfc6-cee39bbe9229 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Numeric quantifiers, plural objects - Examples: - how many tissue paper were on the floor? - how many rolls are in the tray
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4fd9afe7-62c0-419f-84c5-d4ed3374740c · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Transformation verbs, completion checks - Examples: - The bulb is broken apart
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 46e892a4-f1b2-45a4-8c8c-707f0dff21c6 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Transient elements, overlay content - Examples: - video ends with clothes/captions scrolling down
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3e0f3466-77cb-472c-bf94-0ac28d781105 · outbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9df5ebc1-7b43-423b-8179-047756f94512 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 294
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8470cac6-2907-4aee-90c7-e6e76052aa6a · inbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d6cae0dc-e243-4e0e-a681-2f9ff39a6224 · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e45504bb-270e-4351-b01d-cf44a4d26041 · inbound
Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4738e270-1e50-44f5-8425-a2d3eae29a93 · inbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 567138e8-a2ed-4853-a95f-46264c02480d · inbound
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c1befbf9-e216-4ae7-bb19-6a771e49d146 · inbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12058bba-b8d3-4e0e-bbc5-eab30b9b953e · inbound
MiMo-VL Technical Report Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c5df9a-3d2a-4ad6-bcb9-dd537cb0aaaa · inbound
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78530adf-24b4-4652-b6b5-b232e77c5410 · inbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db7d921-f0b6-4b81-9228-fda0ec850c77 · inbound
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 181b8286-f0e3-43cf-8069-371c4827b030 · inbound
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb440322-a61d-4a06-8915-060afcbbc754 · inbound
Advancing Visual Large Language Model for Multi-granular Versatile Perception Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67e962c9-f402-42cc-accd-73d4f0f5d21b · inbound
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9add315c-0754-41b9-ba02-1dd5eb712c19 · inbound
Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa93c31f-f6eb-454e-9388-ac956e1afd3f · inbound
Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aeda5a9-ec24-4ea2-bb08-123d214dacf4 · inbound
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b55908-738b-43e8-aa27-bbc80aecf7be · inbound
A Survey on Video Temporal Grounding with Multimodal Large Language Model Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 125
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c68a39ef-5cef-4f18-a352-8e87bf843846 · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 288
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092dee13-de5c-4c06-89c5-20e794ea9026 · inbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f14d52f-3aa6-4dab-b8e7-339eb77843a5 · inbound
VIDEOP2R: Video Understanding from Perception to Reasoning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 43b4d501-e9a5-4dd1-b1f0-1aa7af507935 · inbound
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e0cb23fb-6be2-490f-aad7-b67241350fb9 · inbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9b187b0f-a38d-45a9-a9db-c82b763254f3 · inbound
OneThinker: All-in-one Reasoning Model for Image and Video Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fb327704-a237-408e-84ee-7a14cf52e717 · inbound
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 092753b1-aab3-4aee-806c-3130dd836dc9 · inbound
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3eb714c-6453-48d2-84e9-d25fe74e9893 · inbound
AdaTooler-V: Adaptive Tool-Use for Images and Videos Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6bc99a11-4a63-499f-ab31-df0d8b8969e7 · inbound
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4495bbd2-de89-44cc-a83e-b0b39d756171 · inbound
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b1878c71-7d1b-42b3-a3cf-d3016bbb1f29 · inbound
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d177492b-4fd6-4df0-9a25-d1f3229a1487 · inbound
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 556f6a39-46e1-4030-bf4d-23a5f96c1067 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5e8ebaf9-3410-4ae5-bd62-0a945189ece1 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3062b4-9aed-4595-859c-7a7a414ca15e · inbound
Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d80fae7a-6546-44ee-8e6a-ea9528528f23 · inbound
APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 87db41aa-b680-4d07-b2e5-55283a5d870f · inbound
Video-ToC: Video Tree-of-Cue Reasoning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3dc60eae-4c0e-466a-8773-d25ec3fb2e34 · inbound
Towards Temporal Compositional Reasoning in Long-Form Sports Videos Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e5200291-cb6f-42a9-b04c-0cc105e58cb3 · inbound
Towards Temporal Compositional Reasoning in Long-Form Sports Videos Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5ed8ec9-9d40-43d4-a267-4c05d366684a · inbound
AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 24bee16c-f63c-4f5a-8081-e71523b2584f · inbound
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9cf10bfe-4981-4ecf-9d99-11d5955202f0 · inbound
Co-Evolving Policy Distillation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0e3d2f1d-0933-43a4-ae2d-945ea35d1fa3 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation df3f111c-e38b-4cf7-9430-3a8e53b84360 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ba62a7b2-e7d0-4945-ae60-f46f4bbe7630 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4a13e00c-9ba8-46ef-b63d-ffc85dda9839 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4faa07dd-96e9-45f5-943b-e21ef325d1f5 · inbound
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eb64bcdd-6264-4ce8-85f0-d4df3dd37934 · inbound
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4ceac38a-888f-45e7-b5ec-08545880b531 · inbound
Video-Zero: Self-Evolution Video Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 87698de1-9b46-4c55-aa86-242a60097b5f · inbound
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c317675e-1d6a-495e-a30b-121de9680c68 · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e7c0e7d3-1838-4170-b85c-c3ce2f1d2511 · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 35d05aec-0652-4be4-b667-3eb96cb60d69 · inbound
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c6773e2-b13a-4afd-86f4-be124fc01250 · inbound
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e89650d7-ca77-497b-858f-285831a96273 · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 598c87d5-05b6-469c-bbe2-d13a9179001c · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9b9223fb-c6e2-4540-b335-32e90639ff5b · inbound
Towards One-to-Many Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cf68d152-da0c-4486-9e6e-bc9893b58257 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b1e67d90-a1a5-45cd-a365-3bf83f3ce2d4 · inbound
NEST: Narrative Event Structures in Time for Long Video Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 94b3e95d-b8e4-46ef-ad21-ef862a41684a · inbound
VideoLatent: Video-Language Learning via Latent Self-Forcing Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1e7841c4-0434-4cea-b3ac-ea001b0d250f · inbound
DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3ebdf02f-ac69-48be-bd01-caf85a2ca558 · inbound
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 983ab6fb-dc9e-4efa-8b06-00dce13acfc7 · inbound
TimeThink: Reasoning with Time for Video LLMs Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4169a6b-2f5d-4b69-910a-9db16aedc76f · inbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0841f998-78fb-4532-98b2-2e9a50486d46 · inbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c6cb9f-a569-4123-b64e-02ba7a841a4d · inbound
TimePLE: Rethinking Temporal Representation for Video Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7d5b55-0444-40c4-aacb-a628a9681f73 · inbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c7eacd8-50a3-4af6-b457-574f29e6d56f · inbound
I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.