Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T13:13:40.485342Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 6 inbound Pith citation observations for arXiv:2505.20715.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T13:13:40.485342Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T21:32:16.939563Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T17:27:15.832913Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6dd261fe-b1eb-40c4-b967-4b2a66177feb · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding E.T. Bench: Towards Open-Ended Event-Level Video- Language Understanding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6fedff7f-aaec-4ec1-8767-dea3b12c325d · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 70e69085-bb97-406a-9398-6610082473d3 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94989262-7ce9-4daf-be32-0c9da045630a · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TALL: Tem- poral Activity Localization via Language Query
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8f47666e-7b6c-4623-bfbb-7765955299fa · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8fb212dc-ce3a-4e6a-8a33-7209fc3388d6 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Can I Trust Your Answer? Visually Grounded Video Question An- swering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f7631895-8de3-448c-bd4d-403b6684c574 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding GPT-4o System Card
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aef4662d-eab8-409d-aec9-2cc0e487018c · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Gemini: A Family of Highly Capable Multimodal Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f6b746be-397e-4fd5-89b8-4d5a90703ae2 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Qwen2.5-VL Technical Report
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 744e55f6-1aad-43b6-a785-dd0a29cf7ff1 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TempCompass: Do Video LLMs ReallyUnderstandVideos?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1f9d51a0-8931-4a44-990f-d86b12465850 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8994b48c-70a8-44b5-8d48-9514de05ad00 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f35a0220-7d5f-46b0-b1f9-d9eef2fbc11d · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f4fa31d3-aefd-42f6-bdb9-25d25e8d84be · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2411361e-59bf-47fc-a2d9-edc8ed537e64 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4738e270-1e50-44f5-8425-a2d3eae29a93 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3203da8e-6f23-4b56-81fa-9d2d711bda85 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 88cddf4d-0279-493e-99fe-e38e602db18a · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding ViViT: A Video Vision Transformer
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 76fda725-8870-4c5f-8b49-153e8ef93ea6 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding CLIP4Clip: AnEmpiricalStudyofCLIPforEnd toEndVideoClipRetrieval
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3d42ee2d-0399-4b83-a985-107023163e17 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Video Swin Transformer
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 078136f0-ee82-4530-8a4e-662dfd2f8952 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Un- derstanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ebf26064-cb9d-4720-8361-058a97b11899 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Spatio-temporal interaction graph parsing networks for human-object interaction recognition
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 993b126d-19ce-48c5-babc-25fbeff00541 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Learning Streaming Video Representation via Multitask Training
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 09f7888c-e262-44d6-97a5-d6b748194f20 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Chain-of-Thought Prompting Elicits Reasoning in Large Language Mod- els
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 358b7850-828e-4e44-905f-3a296da34640 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation efbd6b76-da6c-4f41-b6ab-8d99e4d6e141 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Online Video Understanding: OVBench and VideoChat-Online
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 96dac2a4-83a7-4538-9fea-5902b26b70e3 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 249be1f3-43de-4813-ae3c-4aace30d42df · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Training language models to follow instructions with human feedback
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7abc83f5-7dc0-4db5-9125-79761991f997 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Proximal Policy Optimization Algorithms
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16b70194-ddcc-46b7-a55c-1a94ec77c2fa · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d580f246-b69a-4ce7-9e75-ff8a42927870 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Reinforcing Video Reasoning with Focused Thinking
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9a006b0f-30f2-434a-95b3-625296a5cfd1 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4650fff8-3c9b-476a-8024-f79880618b71 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding VersaVid-R1: A Versatile Video Understanding and Reasoning Model from Question Answering to Captioning Tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c2ebbb9-6cd8-4aca-b5aa-dcd69f3dfad5 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 80f074d6-c1de-4c12-9940-7631671e56f9 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b062386c-c594-4c2d-8f49-f356164b8325 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Qwen3 Technical Report
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a801cf4-fe91-4b49-90ec-259806dcacab · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a88e529c-9124-4ad6-9b94-dab205df6007 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Unhackable Temporal Rewarding for Scalable Video MLLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a5627b68-e500-4c6f-99b4-a727ef0e9a70 · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding The THUMOS challenge onactionrecognitionforvideos“inthewild
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f6605d93-f11a-490b-a9c6-9549b2a5a8dc · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding PerceptionTest: ADiagnosticBench- markforMultimodalVideoModels
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 355866a6-f768-4440-b238-001bd72317ba · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding MVBench: A Compre- hensive Multi-modal Video Understanding Benchmark
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a871f567-3a2b-4ed5-8e3f-1a771549acaf · outbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a810bf84-4b8e-4662-a534-5e19cae1e3c7 · inbound
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6b13fddb-5b7b-4527-826c-4a725636c01f · inbound
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3973e7e9-69de-4291-be4f-254c269ec2fd · inbound
Video-Zero: Self-Evolution Video Understanding MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e6c3920c-a017-4167-b8a0-f0a04637f133 · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3333968e-5d03-4a75-8f9f-4b7c2843ad0f · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0cba4df5-6108-4d84-b38a-8c1e04efdca5 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.