Pith. sign in

Paper Citation Record · LEDGER

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

As of 5 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 6 inbound Pith citation observations for arXiv:2505.20715.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20715 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T13:13:40.485342Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:32:16.939563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T17:27:15.832913Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact26
  • verified fuzzy16
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6dd261fe-b1eb-40c4-b967-4b2a66177feb · outbound

This paper cites E.T. Bench: Towards Open-Ended Event-Level Video- Language Understanding.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding E.T. Bench: Towards Open-Ended Event-Level Video- Language Understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.428798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:9b90baf94d7707f9fcdab4f9acf2246e04e17841b0fa7a1ee5324cff4aff35f8

Observation 6fedff7f-aaec-4ec1-8767-dea3b12c325d · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.536691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:70125af3908046b41ef70cdc01b125def9923f2d571f81183c2afa1af8bf0a3a

Observation 70e69085-bb97-406a-9398-6610082473d3 · outbound

This paper cites V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.593246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:e8fed4f7b4a8008551ddd993ef3358f31110ec638d4bc12c448d899e50e090ae

Observation 94989262-7ce9-4daf-be32-0c9da045630a · outbound

This paper cites TALL: Tem- poral Activity Localization via Language Query.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TALL: Tem- poral Activity Localization via Language Query

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.435759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:abecd876872017918e1661191a7b9f2065d5483009a6a466c32301c634a24138

Observation 8f47666e-7b6c-4623-bfbb-7765955299fa · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.490044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:6b43f288d112eea7774b4122833508ad67857fac5678ff771a9089820ba17b4f

Observation 8fb212dc-ce3a-4e6a-8a33-7209fc3388d6 · outbound

This paper cites Can I Trust Your Answer? Visually Grounded Video Question An- swering.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Can I Trust Your Answer? Visually Grounded Video Question An- swering

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.418967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:70a65ae3b2ce11e5ba1e29e68382c21e9664c1638182b4676d73d8be63e6d795

Observation f7631895-8de3-448c-bd4d-403b6684c574 · outbound

This paper cites GPT-4o System Card.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding GPT-4o System Card

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.493689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:cdc151959ec105f986894aad4b7dd5992d97d6a0ab49b43122ac83840bfb4c20

Observation aef4662d-eab8-409d-aec9-2cc0e487018c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Gemini: A Family of Highly Capable Multimodal Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.549783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:13df019fea64423d265300d64dc5d3ef4707e1815e629bbdc143dd7c3668685e

Observation f6b746be-397e-4fd5-89b8-4d5a90703ae2 · outbound

This paper cites Qwen2.5-VL Technical Report.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Qwen2.5-VL Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.566980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:3521064061c57a1b13788ac5facb639b55d246ac6bd9b90f78c89eb0a9463c77

Observation 744e55f6-1aad-43b6-a785-dd0a29cf7ff1 · outbound

This paper cites TempCompass: Do Video LLMs ReallyUnderstandVideos?.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TempCompass: Do Video LLMs ReallyUnderstandVideos?

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.410181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:80b62fc45394c306197e241b0dbda68ba10fba3ddfd5e146a28af82968ec258b

Observation 1f9d51a0-8931-4a44-990f-d86b12465850 · outbound

This paper cites Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.571134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:106df3bd63bcbc408d64cd765fc058c6ab8b520a0203f1355a401e8813955367

Observation 8994b48c-70a8-44b5-8d48-9514de05ad00 · outbound

This paper cites LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.526285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:da277ee6d6931d8e7ca8cc63a6a69aee04fc9d67a2c7b5249436b91e6c7e5348

Observation f35a0220-7d5f-46b0-b1f9-d9eef2fbc11d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.597230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:7e996294a759207d9bcd4c3f35f66b697f36ccb45565d7bf0d9b87abbdd49afd

Observation f4fa31d3-aefd-42f6-bdb9-25d25e8d84be · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.601457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:1eba8e0fac65e5c09fb14c71eed050dd92ff927ede188c132e23e41cd3e8ea0a

Observation 2411361e-59bf-47fc-a2d9-edc8ed537e64 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.558866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:1a5906642e30326b1f66ce7987658499b961b9bdd7649c08c9a043b5377e87dd

Observation 4738e270-1e50-44f5-8425-a2d3eae29a93 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.579232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:880862cd0e2519079037eef849a432f438ba6f1bd2cfd32062d971a08dee0fd5

Observation 3203da8e-6f23-4b56-81fa-9d2d711bda85 · outbound

This paper cites TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.563294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:630066e75196fe452b6eac9a3cf7816dce24e4683085f0de266bdca1c889a812

Observation 88cddf4d-0279-493e-99fe-e38e602db18a · outbound

This paper cites ViViT: A Video Vision Transformer.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding ViViT: A Video Vision Transformer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.407273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:929c94f9e8e3a6a36e2d5d7d1473a36fe5bd969802a4036b2052508875709e89

Observation 76fda725-8870-4c5f-8b49-153e8ef93ea6 · outbound

This paper cites CLIP4Clip: AnEmpiricalStudyofCLIPforEnd toEndVideoClipRetrieval.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding CLIP4Clip: AnEmpiricalStudyofCLIPforEnd toEndVideoClipRetrieval

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.404493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:bd55772404a11719168ff64dc13118e0df1bf6cdafa6a9b754e4df1379a19022

Observation 3d42ee2d-0399-4b83-a985-107023163e17 · outbound

This paper cites Video Swin Transformer.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Video Swin Transformer

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.401743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:9a1af0b727ebc992fc43056a5f51951b19a8fadabdc87a8eb7f9eaa4cdc3353f

Observation 078136f0-ee82-4530-8a4e-662dfd2f8952 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Un- derstanding.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Un- derstanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.399060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:914db66d95edea587e5b4f0fb22a297d46914f4bd1e867b96d453789455696b3

Observation ebf26064-cb9d-4720-8361-058a97b11899 · outbound

This paper cites Spatio-temporal interaction graph parsing networks for human-object interaction recognition.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Spatio-temporal interaction graph parsing networks for human-object interaction recognition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.396056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:b02c30717891118e52a90ab3966b6dc5be7ac97c4a15486ac5b5335bec2deb2c

Observation 993b126d-19ce-48c5-babc-25fbeff00541 · outbound

This paper cites Learning Streaming Video Representation via Multitask Training.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Learning Streaming Video Representation via Multitask Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.512710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:65bcd7a6eb671cf5ce06ec5e45b9a029e806bb7f749434604224c9ca507d3d35

Observation 09f7888c-e262-44d6-97a5-d6b748194f20 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Mod- els.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Chain-of-Thought Prompting Elicits Reasoning in Large Language Mod- els

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.390190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:1e18ae815b29960e0c5cbe31e1a3da8e217c04420c2a32ed6878c069b617b10c

Observation 358b7850-828e-4e44-905f-3a296da34640 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.545824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:871ec502eaa1587144d269d3564a074fa8154ae20303967863ca5e3eb6881a1d

Observation efbd6b76-da6c-4f41-b6ab-8d99e4d6e141 · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Online Video Understanding: OVBench and VideoChat-Online

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.583606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:77196f3ed4f5b346bbcec40813e73a297a80e57cc4b3c3657967814f07c05d0d

Observation 96dac2a4-83a7-4538-9fea-5902b26b70e3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.507544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:25e46b6e8bc002946042eb9ecc4e763735fcb3dd8dd7ea5edb20dbe8bbc6b5d0

Observation 249be1f3-43de-4813-ae3c-4aace30d42df · outbound

This paper cites Training language models to follow instructions with human feedback.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Training language models to follow instructions with human feedback

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.392850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:655cc8a997cc8648af2a03869c808b81512cd9cd9a7b72d572d6f5611f42f660

Observation 7abc83f5-7dc0-4db5-9125-79761991f997 · outbound

This paper cites Proximal Policy Optimization Algorithms.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Proximal Policy Optimization Algorithms

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.554055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:d84419ef93443e256f20b4c32a8dc24a54ca712be2a16ef5f7737ddc5cde4c1c

Observation 16b70194-ddcc-46b7-a55c-1a94ec77c2fa · outbound

This paper cites Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.517310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:f342432af0fcd53e4763e7b1ffcf712972ada04a450e4b37e61d0243aaecae5c

Observation d580f246-b69a-4ce7-9e75-ff8a42927870 · outbound

This paper cites Reinforcing Video Reasoning with Focused Thinking.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Reinforcing Video Reasoning with Focused Thinking

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.498226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:31bed6e8074bea438a12db4687a4c79156fea637916f7746c51f96ba46c20b1b

Observation 9a006b0f-30f2-434a-95b3-625296a5cfd1 · outbound

This paper cites Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.503478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:87bf4aa9fee3611316fa688bb985ce5b661c5a0c22c4e9f3cd4ee02d027f3f7d

Observation 4650fff8-3c9b-476a-8024-f79880618b71 · outbound

This paper cites VersaVid-R1: A Versatile Video Understanding and Reasoning Model from Question Answering to Captioning Tasks.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding VersaVid-R1: A Versatile Video Understanding and Reasoning Model from Question Answering to Captioning Tasks

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.541000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:cb7102f3688852f4265bce05edc1e5024c283caed8fd6c5b7056eb49865e02bb

Observation 2c2ebbb9-6cd8-4aca-b5aa-dcd69f3dfad5 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.530952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:024bf09e4bd2aaa6a80d3a33b7fcd5ded85762e105c9e8b6f1751fdce4aa6f66

Observation 80f074d6-c1de-4c12-9940-7631671e56f9 · outbound

This paper cites TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.587949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:39eb5c85240c1d8f38551a99327d192c0d89abfa395d03909e301f80291c0ec5

Observation b062386c-c594-4c2d-8f49-f356164b8325 · outbound

This paper cites Qwen3 Technical Report.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Qwen3 Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.521743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:7bb61ab15138c4d53f761543d5a8df645b70301793d5b70389b62ec5615328ac

Observation 8a801cf4-fe91-4b49-90ec-259806dcacab · outbound

This paper cites Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.432593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:8c5c7859eef73e7d13a974214538c0a89aaba99343c7cf565ab4e43ba8495a4a

Observation a88e529c-9124-4ad6-9b94-dab205df6007 · outbound

This paper cites Unhackable Temporal Rewarding for Scalable Video MLLMs.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Unhackable Temporal Rewarding for Scalable Video MLLMs

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.575156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:fd0368afdc548ae5f8b8778ded5d3587b546b2a7779c71ddcd6568d316762f0f

Observation a5627b68-e500-4c6f-99b4-a727ef0e9a70 · outbound

This paper cites The THUMOS challenge onactionrecognitionforvideos“inthewild.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding The THUMOS challenge onactionrecognitionforvideos“inthewild

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.425531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:630ef5d6326e13b37a36a5cf46f6a66f6bf0cb899e7156179a3be3fead4b4962

Observation f6605d93-f11a-490b-a9c6-9549b2a5a8dc · outbound

This paper cites PerceptionTest: ADiagnosticBench- markforMultimodalVideoModels.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding PerceptionTest: ADiagnosticBench- markforMultimodalVideoModels

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.422367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:aed9041d672c943e28d524ff56a9ac7db164c1bfe86add471e7aebf54db33435

Observation 355866a6-f768-4440-b238-001bd72317ba · outbound

This paper cites MVBench: A Compre- hensive Multi-modal Video Understanding Benchmark.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding MVBench: A Compre- hensive Multi-modal Video Understanding Benchmark

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.416222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:7650d59fdd3f2864625a9759d9503861a2e67fed8eb3738d028143185783a2a6

Observation a871f567-3a2b-4ed5-8e3f-1a771549acaf · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:17:19.413273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:24fd7973df50b552b7ba1e9d39a30f16287968b47de2909ab1085fbd0eb370d6

Pith citing papers

Observation a810bf84-4b8e-4662-a534-5e19cae1e3c7 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.382073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:bab209a192b6068bbafa9b7006e1e30dc872f84ad72d8f81638b453d22fb0925

Observation 6b13fddb-5b7b-4527-826c-4a725636c01f · inbound

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking cites this paper.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.314324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:3d9f1b9df0e56fb8e1bb465b9caf05bc3b28cf2e79533e9a91a6849d65d13054

Observation 3973e7e9-69de-4291-be4f-254c269ec2fd · inbound

Video-Zero: Self-Evolution Video Understanding cites this paper.

Video-Zero: Self-Evolution Video Understanding MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.468590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T21:32:16.939563Z digest=sha256:7cbd97db81aee57e9f8224b05111c2800343438a063464526ef1388711bca951

Observation e6c3920c-a017-4167-b8a0-f0a04637f133 · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:34:02.601054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T07:32:12.180233Z digest=sha256:0c200e7b672a0ebabb3a7738e18b78e8ffc7ce453ccb07da2da3bb3f54769e50

Observation 3333968e-5d03-4a75-8f9f-4b7c2843ad0f · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.306955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:b63e4bf10009ff5960315454c477f88f8509d2afa567ee90e60ecd3096b0b95f

Observation 0cba4df5-6108-4d84-b38a-8c1e04efdca5 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.834103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d37e2202520f69d46c369431f9c79d2c8f4e5b9bd679b226e35f21dcd7770da4