Pith. sign in

Paper Citation Record · LEDGER

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

As of 12 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 67 inbound Pith citation observations for arXiv:2503.13377.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.13377 v3

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T02:40:06.454859Z

measured 154 of 154 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 67 of 67 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.215349Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T09:49:44.696023Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact16
  • verified fuzzy60
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 096c1a33-4843-4c20-92da-e0ce2affb49c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:40:06.573742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:8c3df39d06aa233860888761d159e0c6d19d8df25d0a2a6df917dce3c4854db0

Observation e61632f0-3b9a-4f38-ba2c-f454ec0df487 · outbound

This paper cites Ht- step: Aligning instructional articles with how-to videos.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Ht- step: Aligning instructional articles with how-to videos

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.726653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:dc246e8dc7c340e396715dfd22443328ac2249be2cfa6607c856f5480d6badf0

Observation e2ef44aa-49b9-40a7-8309-e4a21faae837 · outbound

This paper cites Localizing moments in video with natural language.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Localizing moments in video with natural language

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.728902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:b4a4edbd7cebcf46c74a9babb4742310bcbd62ef7fa888c025e855ae83eb96f4

Observation ee4c2169-a8cb-459f-a967-4ed71a2d93ad · outbound

This paper cites Qwen2.5-VL Technical Report.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Qwen2.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.507723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:525b6e0287c7f804aec08624d3e93fff0969dbd4181d7230027d670cdf5e80d5

Observation c5ea3db8-702b-44c1-afe5-c9b46f2a8aaf · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Activitynet: A large-scale video benchmark for human activity understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.731397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:3e0de72ff971f3af79d883df544982f2bfa06692273ed4478d7c9ef0ec31963f

Observation 23757ce9-7cf5-4e21-acd0-efa2698440a5 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Quo vadis, action recognition? a new model and the kinetics dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.733491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:4e9cbf5db83af4900a2c6691e27bc3127a7919b123a6700d171e38e4ff1bbc4b

Observation 6a476057-7e8f-4ff9-bc39-04119bafc60f · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.735812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:e5dad304e6c78f1c4eb5fcbfd915fe79d7c7dbf781a9deab515e5bec538e56b2

Observation 9a5d427b-7514-4874-add1-6467a324e001 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.738373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:8c4ad0c7830e71c3413043ee2ab2d75fa040d94d1a3c0d69407a4c98a34e7961

Observation f43f72c3-cd01-4541-ba18-6a13ad6c8519 · outbound

This paper cites Space-time gestures.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Space-time gestures

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.740776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:e9ffc4f9d29fe8790820aa9b84a9cd6df057547b004587ec5077a11737060357

Observation 2c13ca72-7eb0-4508-9ce8-1c02e190ced2 · outbound

This paper cites Gemini 2.5: Our most intelligent ai model.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Gemini 2.5: Our most intelligent ai model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.743001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:e83d74bdb8d79a83238e8426d7030ca2c589232f1fc17191d29e682259156877

Observation 3223d133-1d40-49f7-9453-fb96ce97b930 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.554694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:0dde2eae7ea0d5cfd045e982a735f445a997c4315752b188d8afa792171f4b1a

Observation fa54be3e-b58e-4f18-aded-7e569bad7d7b · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding BERT: pre-training of deep bidirectional transformers for language understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.745512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:722ac545480f75cd3f766399c45797a95671b1644fc73dfb29c114a32afbb1e1

Observation 9d6d75e6-3091-490a-930b-8337c7f39a16 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.747757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:4a106848d78b7011e9ddce9520552310023032a11b2e6c02a8b1bce893fe3336

Observation 5e55fd23-8794-4d2e-bc3a-51d5f22a30cb · outbound

This paper cites Temporal localization of actions with actoms.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Temporal localization of actions with actoms

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.750429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:fbe22321cbeb9c455b49a50eb4af399fde08f79d7b3358bd854e8881fd69178c

Observation 43375593-cc1d-44d5-a288-6f65c76d8cc3 · outbound

This paper cites Tall: Temporal activity localization via language query.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Tall: Temporal activity localization via language query

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.753043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:6b43b775ec0dc80abc92d7f5366159967ec205b432cff640493a36431971d7b4

Observation 08d09f77-dc32-4adf-a481-0249d88cf950 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Ego4d: Around the world in 3,000 hours of egocentric video

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.755326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:444bc9fd43c54affe3f3782fe622f5e8683b1bd7cd6e6e3b7e6ed42f35d65eb5

Observation 12ebfb1b-4d40-4fd9-9038-7786c0f35795 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.757551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:e89fd194367605bb09d8204d9d865bd38c519a8ab22f7c548038c83df3e6d565

Observation 19caab2e-e0db-403e-ac63-bf1a537f359b · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.559123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:1d3c10ccf69f891e0696686fd5a6573de652c25c4a0eaaceb7bd8361aa295517

Observation 6a400c74-73e5-4ff1-a38f-e7599b192f3b · outbound

This paper cites Revisionllm: Recursive vision-language model for temporal grounding in hour-long videos.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Revisionllm: Recursive vision-language model for temporal grounding in hour-long videos

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.760082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:6564b217aeece0447f3a389e5209e9071362b6f0d00c91919118e21ddbff9dc6

Observation 0bd322c1-83a4-4854-9fd3-4f13e6bcb140 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Lora: Low-rank adaptation of large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.762465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:adbb3f1e2969c4828947d4c6b5d7f68a7f2802afe6f25edfcf622ac07e12598c

Observation 1f3d2a80-e189-43d4-a033-07cd3e570de5 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Vtimellm: Empower llm to grasp video moments

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.577628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:b94b3e628a8e78beaeb26a0e6f72e397348cdd98fbe1abd4826ae63782cc2a67

Observation b579597b-975f-4620-bc20-fd659085394a · outbound

This paper cites Knowing where to focus: Event-aware transformer for video grounding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Knowing where to focus: Event-aware transformer for video grounding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.581626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:47ce19c1f2b4b60a1f570bb7b8c79df81b1a9ea0479c5e5c1cf9bdccf010ccc9

Observation 16c6a277-cee3-4627-93bc-e820a671a932 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Gonzalez, Hao Zhang, and Ion Stoica

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.585180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:bf12f784121669136a2866c4316188da9d1e109ff1d6249c33e496d16230d82e

Observation 83ecb689-6839-4cf0-8971-0b82530f2539 · outbound

This paper cites Retrieving actions in movies.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Retrieving actions in movies

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.588828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:5690591e94f9a07a9b7d4e887e3a2ac06edcf086ba86ce0fc67d6fd9458ce1ff

Observation 0b25d990-ef0e-44c3-b91b-f7586e9a8437 · outbound

This paper cites iMOVE: Instance-Motion-Aware Video Understanding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding iMOVE: Instance-Motion-Aware Video Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.493518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:ecad8fe2ab08ff1a4a13070ea054a4a12dfaf7543b3b7ce11c59c4792d18d9e5

Observation 916126f3-5e6c-4fe1-b6c8-47c8909da601 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.592471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:0a937739f02bf089e34132aa1a55342b34f43395471e5df51104f01f586b9374

Observation 59801fb1-50fb-4317-8a32-1acef7e4adef · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.915683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:9da0765667b758c67c10ceb66d997af57b6ee98b79707daedbbc97920a57cf74

Observation cfa03b87-9845-4c98-a800-14a7f564b40f · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.550790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:195df5f45f3755992b551669b830f4b3605476dcecdc811c4741bfebda5e1415

Observation a6fe9708-c2a2-4120-aee6-de805018e3d4 · outbound

This paper cites Egocentric video-language pretraining.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Egocentric video-language pretraining

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.595983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:f926a20f5e2d1d4ea8e7582abce33efff1c412a55ea590f77da7b8dd0cafd20f

Observation 5ead59bf-2a9d-41f1-a14d-1875523d9552 · outbound

This paper cites Univtg: Towards unified video-language temporal grounding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Univtg: Towards unified video-language temporal grounding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.599563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:7f65869e3235255d4316a0a3c9eda178d4657c03c79bcc1e35ca6bbc59e7123d

Observation 2c8fb89d-1652-47a9-ad68-f5abce5fe348 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding TempCompass: Do Video LLMs Really Understand Videos?

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:17.144047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:502e447425e95c060c3b8bacea9904c1ce7b6e965c32cb12f1f52fd68aad62f2

Observation 229a11f8-6ec8-4c1e-877a-89efe0cfa7ca · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.569222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:be300223289f68345463a1afbdadb134edd0af0f0b28a7c0989b0fcdd40a3724

Observation 2121138f-f80c-4055-8ca4-e9f064f465f5 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.602907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:cdcd6c39bc9ee34b96f0f3b519882bce00a002aa3651602dfbb3f4f2ca843f74

Observation 37a1727a-3f48-47d6-8f2d-c82e3bd5dbf2 · outbound

This paper cites Walk these ways: Tuning robot control for generalization with multiplicity of behavior.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Walk these ways: Tuning robot control for generalization with multiplicity of behavior

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.606458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:8554193f0c03032b7f6dc7a531b237238e66290f183e76a93e93cb42dcf3ed9b

Observation eb2ce270-e81c-495b-b521-a160656a4591 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.512877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:5dc8b9d26dc4649e8d8220e1c437dfde405c9b5e70ea50b2c5d55ac95bbe605f

Observation 6ae26ceb-0ea6-4106-989c-ca8b99c7e736 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.610238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:a6b94e97e7316c0f7e521b19cd7b7418a5eb8d7585eefa70fdd2da16725875c8

Observation 855c0589-bc9e-4713-9a19-73979d0ec93c · outbound

This paper cites Snag: Scalable and accurate video grounding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Snag: Scalable and accurate video grounding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.613081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:cf50745b084b29e4d2cbfdec3decdc082b770236eb6f0fa257e31a9b009fe41c

Observation 8507b6bd-8911-47d2-94cb-bfacd6a2b27d · outbound

This paper cites Queryd: A video dataset with high-quality text and audio narrations.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Queryd: A video dataset with high-quality text and audio narrations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.616428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:b4cd1e566d3f6539dc639928ddd85ac83dd275947c8e1a186e2e3a5a39318fa8

Observation 31298708-6ed3-4cf9-9d6e-ece7d411a55e · outbound

This paper cites Openai o1.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Openai o1

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.619036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:497a12b6331984e950fedadd192798fe8b6a947e64d22a8982b469d4d64e98db

Observation f1181f51-d28e-4a5e-a059-f36ff1822e89 · outbound

This paper cites Training language models to follow instructions with human feedback.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Training language models to follow instructions with human feedback

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.621670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:d23a23908f368a42de3e59362225681681c0ae52090b4bbed407941364c2fcf4

Observation f2a44993-321b-4f5b-82b5-965363200b0f · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.624221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:4fff93d01d885a92c0d5100c2009e69b473be4b7ba793a069ac09c95c7dd5191

Observation cc1a7ef4-e5f6-47b5-9a6a-5cab08db7920 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Learning transferable visual models from natural language supervision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.626902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:216fb90b16ecc3269ff8989c525df5cb8070010f095c028488f02e5ea56eb2ae

Observation 1247a069-af12-456f-a12e-87990e9d1861 · outbound

This paper cites Grounding action descriptions in videos.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Grounding action descriptions in videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.629529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:80d0b6ecc2e65c5e307ad5739a0a05631af87c3cf75d82f53733d4a453de28da

Observation c93386c5-7836-4f5d-a4e2-70c2854afbb8 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.632096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:be53bb8ec8f8d61da744942fda4329873eba813ced653edea98debd306e61b5a

Observation 776c28a4-9841-49f9-b982-87d8e324bae8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.517789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:cbf8ee4ed38d29468806897a7d0c43624fda73551b729e1a487b220e4030c503

Observation be34329a-5ba4-4f07-8c99-23204cf7a19b · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Hollywood in homes: Crowdsourcing data collection for activity understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.635021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:f7166dc2c29eb18c309cb54fced43d7e8ff74fea91cb4db82779990ccf3026e5

Observation bec547c2-e15e-4994-9ade-b935f7dcc8b1 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.528100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:6450b4c02c05e9df5166a037861020d191576d62033ece2af2175fda6c974f2b

Observation 7b00d6fe-61bc-4d53-a4ea-d4b73102bea5 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Reason-rft: Reinforcement fine-tuning for visual reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.533183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:49e7e2a2667b17d1d14b48bff688ed5a70f15e219466c9e2f831caa7387a7c0f

Observation af0bf21b-6be7-4169-b4fe-d7b0c07e597d · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.538109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:39c6df487363a5f3a68525a14c75afd6fa117e38e4c917977b8fb7e80526d908

Observation a6aba1f1-2c0c-449a-a5bc-5ba2a720944e · outbound

This paper cites Hawkeye: Training video-text llms for grounding text in videos.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Hawkeye: Training video-text llms for grounding text in videos

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.637627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:ca61ec1802e33bcfa1209c6cb67af23a4611fa4cf8c1642a33bedd2f49c19151

Observation a188824f-9424-494d-8c8c-37732a744538 · outbound

This paper cites an unresolved cited work.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:40:06.640284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:e983c7a874817ca71e6a1d8d544e8c0066ba84303aeda6c1d66f995b74c27235

Observation 7afd85b1-2212-4dd9-962f-4510d4c0470a · outbound

This paper cites Number it: Temporal grounding videos like flipping manga.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Number it: Temporal grounding videos like flipping manga

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.642826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:7b5cc9eacfe9c20ad6954940f3e89ce2bdf6cad3d76bba8f1c525892d2ffc76a

Observation 47042016-d8ce-4724-81b4-50e670085c40 · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Vid2seq: Large-scale pretraining of a visual language model for dense video captioning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.645489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:44f11d3c96ec65985527c41d459b8be5a2d556bed7fbe8a08729156dcd2d71a5

Observation dcfcf2b1-ca8f-49cb-9773-83344c729dbb · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Vid2seq: Large-scale pretraining of a visual language model for dense video captioning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.647930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:92774a0a3c6b00fa5021fae7a76c088287fde89de768c7e154e6f1d59cce22bf

Observation 4c540c59-2304-492d-9cf2-6de6259cbc06 · outbound

This paper cites Egolife: Towards egocentric life assistant.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Egolife: Towards egocentric life assistant

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.650286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:50c101f69835d84a8ab0f68318aae5421904b42b492bf1a43d346bb538e1633f

Observation 88fb637f-2325-443c-82ff-053f31c17ad3 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:40:06.501096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:8a4dd1ecf081f4d10ce3daa18e85e67db99a6c5d181bf5bf280c571680ac03bb

Observation 71abcb1e-a5fb-48ea-9a87-c8f7c08a6df7 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.652696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:3639062fc3fa914068193ae3422e7b32e260b91edd267bf4c702cd710ab76e5f

Observation 770d5252-9b0f-4a9e-a152-8dbe3e983746 · outbound

This paper cites A closer look at temporal sentence grounding in videos: Dataset and metric.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding A closer look at temporal sentence grounding in videos: Dataset and metric

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.655502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:889c7bde062e4ede1576f5482876fd8b9ab21090ec66f6bbfce731b9d770621a

Observation a30f710b-2cf7-4679-b3d3-eb390e898793 · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Hierarchical video-moment retrieval and step-captioning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.657989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:53a16aed6404bb47d7e266c0a95d082aa7bfbe6c2f1d78e6440dbfede55b8ae1

Observation cc3e341c-d304-41c2-a4cc-9a7d84a37ee4 · outbound

This paper cites Timesuite: Improving MLLMs for long video understanding via grounded tuning.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Timesuite: Improving MLLMs for long video understanding via grounded tuning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.660390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:db3aa67ae97c10f8d01d77e4e0bd33fe336a9f7dd7b13f046b86a8da3439fed9

Observation aeb2920f-7f7b-4e2c-b315-2eb2ca6f9ae9 · outbound

This paper cites Temporal sentence grounding in videos: A survey and future directions.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Temporal sentence grounding in videos: A survey and future directions

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.662592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:f2d71d7e5a6eced60658b704115d80c5c0552441847277f9ded48d9fce6acce3

Observation 3473930b-d8a9-4e1a-930d-9d2f910109a6 · outbound

This paper cites Multi-scale 2d temporal adjacency networks for moment localization with natural language.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Multi-scale 2d temporal adjacency networks for moment localization with natural language

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.665055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:8f963d40d3ad7dc0ab192a67301a921f1470ba20a544672adf5a579a9e3c1e79

Observation 214a3c37-a574-4ee9-bb33-fa83f116273e · outbound

This paper cites Learning 2d temporal adjacent networks for moment localization with natural language.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Learning 2d temporal adjacent networks for moment localization with natural language

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.667570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:072d1d853dada61ada0e168aa84001f613d9c2b65157a829d6dc407bbb04d276

Observation c6903ccf-ac39-4c50-b14e-b931bf6a40e2 · outbound

This paper cites TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.542292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:41dd25eb47cbeaee58f95a6a4ca5977dca36e07c35b024f2c7214d6ec68c9e67

Observation d49be883-0dc4-4ac5-91cc-2ff3762cb4ee · outbound

This paper cites VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.546326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:5b9c1aa75fc279c075582cddea32fec005823df0296a9e851325afb71608036c

Observation 6ecea9b7-6894-4e80-94fe-aeae55a5af41 · outbound

This paper cites goes back to the pink bucket.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding goes back to the pink bucket

Reference 66

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T02:40:06.669897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:1657a6f33daf1c6add116274b122d23ad83cc2102b7e80cf4af6ecf6e4745095

Observation 16676725-fb9c-4b0f-80c1-f5b102fc4186 · outbound

This paper cites an unresolved cited work.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:40:06.672868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:f081f2fff7f66b8d239a974f575c10ab00d0aec36505774619fea416f220ad18

Observation d3dec532-6633-4b06-a093-61664ada7bda · outbound

This paper cites an unresolved cited work.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:40:06.675761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:c8d2312831262f5f396a35ee82e76e729bd2410173490841a020406ea8924f09

Observation 2fda06f6-df62-4bed-9597-0fcfa4e3d564 · outbound

This paper cites an unresolved cited work.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:40:06.678391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:b5fa4bed51b3a457d221032135af76e7e3db7048d424c4bd6f2dcae00e15b5d5

Observation d3f4fd69-5c3c-49fa-afab-5d40bde1f321 · outbound

This paper cites an unresolved cited work.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:40:06.680904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:b70c1a2e258fae8bcc0fc4be646e15b723a92ccc09c89b6ee45b3c759245e7d2

Observation 50fe21c4-4fd3-49bb-bd5a-fc8cd43834d6 · outbound

This paper cites an unresolved cited work.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:40:06.683328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:00becfabc53bb988bf8b5d62a90161c79401b51c732df59d3c55a1556e2eb2be

Observation 51f5e0d6-26d6-42a5-b103-e092533179f4 · outbound

This paper cites Given this analysis, the pineapple is indeed being pushed forward by a person.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Given this analysis, the pineapple is indeed being pushed forward by a person

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.685762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:4d8661cc3c4615b6af2fd7359bbb08766808dd99bc1381474787b16dec3e6310

Observation d51a9da3-bd7d-4d81-a399-ea01118c2b79 · outbound

This paper cites This is the first major action.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding This is the first major action

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.688251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:fea47ac1af61e17e650602daf9e8331db21fe1a485092c3064f4a4e433e5a705

Observation d5823636-cde7-441f-ba7d-ad0cd4062dc0 · outbound

This paper cites an unresolved cited work.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:40:06.690563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:5f96703e95720f279f784c11182d7af50a4618f3818e3614394e914aef648606

Observation 1ff8aebe-eebf-4020-b0be-fa2116c3ae52 · outbound

This paper cites an unresolved cited work.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:40:06.692989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:cdf6c65117a9493fc297cf8b78a0ecaacb12284c7abbc67ef42fc23d54d451aa

Observation c8772ad1-80f8-4ead-abaf-819868dd0505 · outbound

This paper cites Now, let's evaluate the options: (A) C folds the dress, places it on the ironing board, and then hangs it up.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Now, let's evaluate the options: (A) C folds the dress, places it on the ironing board, and then hangs it up

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.695767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:239e10e9f4d01a1fdcb2e429114264776704f90720d0d561b8aa7af6a0c79f71

Observation 0a1e88d2-239a-4441-b0ca-8298de99277a · outbound

This paper cites - Examples: - person opens a book over their head.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding - Examples: - person opens a book over their head

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.698397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:341c589f11a899bc71d48e3d2873aad03191e6d8e62f6cc9431e9bcc48eda861

Observation 43495fd6-4926-4859-bb8b-d2a7b69a253c · outbound

This paper cites - Examples: - He is talking while several people are using rowing machines.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding - Examples: - He is talking while several people are using rowing machines

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.700777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:ed37c348a3994d59b32fd01a3c52da8a415ab857cb9b73eed82d2c4b5d22200e

Observation f0c8edbd-42ad-496d-bfd3-359c92b7a16a · outbound

This paper cites contains multiple actions, each with a clear start a nd end.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding contains multiple actions, each with a clear start a nd end

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.703203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:f1cd35e2dea68faad797b399e0fb8a4bcdc5848c6151e3b983b98c8d8afddfb0

Observation 555dd73e-5504-4304-b496-108255165969 · outbound

This paper cites Posture descriptors, positional prepositions - Examples: - Several other people are in the background working out on the equipment.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Posture descriptors, positional prepositions - Examples: - Several other people are in the background working out on the equipment

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.705831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:6fb3ad623bee9619ec9ec8b664e334092c67b04bf533afd6c7c47ba7c65900e8

Observation 673f6430-6d99-4bc7-9e00-dfb4238aa5e0 · outbound

This paper cites Simple location prepositions.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Simple location prepositions

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.708144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:46bb78dc223a25387f8407313b84f2e7a709349031ad95e44977a291a4e7841d

Observation 6f409486-3f7e-456f-b787-0e685e4f1111 · outbound

This paper cites after/before [action].

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding after/before [action]

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.710693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:ad87781fc9835dd3bce62cb9e93af69e4663699e3ceb7b2273dc48620c6f5f94

Observation 1148f522-c872-4dfa-8472-11532f8d8ef4 · outbound

This paper cites Property descriptors (color/size/material) - Examples: - what material did I pick from the shelf? - what color is the toilet bin?.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Property descriptors (color/size/material) - Examples: - what material did I pick from the shelf? - what color is the toilet bin?

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.713130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:c3c587f5c529112829eba7b89096c0bddb6a3806b89356e088e09c9c7cb8db21

Observation cf68276c-5f57-4f5b-bfc6-cee39bbe9229 · outbound

This paper cites Numeric quantifiers, plural objects - Examples: - how many tissue paper were on the floor? - how many rolls are in the tray.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Numeric quantifiers, plural objects - Examples: - how many tissue paper were on the floor? - how many rolls are in the tray

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.715764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:319a2ce8d21fbcf1d9ad3481187af042fea0a8ddb42a62c5c1ff933d5f382968

Observation 4fd9afe7-62c0-419f-84c5-d4ed3374740c · outbound

This paper cites Transformation verbs, completion checks - Examples: - The bulb is broken apart.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Transformation verbs, completion checks - Examples: - The bulb is broken apart

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.718304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:f20781dc1b200c431be563781673c825453f936000934f989307304e922e382a

Observation 46e892a4-f1b2-45a4-8c8c-707f0dff21c6 · outbound

This paper cites Transient elements, overlay content - Examples: - video ends with clothes/captions scrolling down.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Transient elements, overlay content - Examples: - video ends with clothes/captions scrolling down

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:40:06.721157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:c1388eaa03c88fdcbe67ac76e068c62db5504e5e09706e9dbd494b3f9a6f156d

Observation 3e0f3466-77cb-472c-bf94-0ac28d781105 · outbound

This paper cites an unresolved cited work.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:40:06.723947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:438fb3f4ab8a2a6e9cea0f2ea0fb3aeef16af0c075fc3fdbcd81d0f7d71a13b6

Pith citing papers

Observation 9df5ebc1-7b43-423b-8179-047756f94512 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 294

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:47704124dfd56809e54a86abbcbf8d9179c5e22ff3e0677945b9d8c33a3b2668

Observation 8470cac6-2907-4aee-90c7-e6e76052aa6a · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:b3d90f5ba6e35dc8fef916456a23afd50ad9cf6e1ade450caa9ea6dab4bd1b73

Observation d6cae0dc-e243-4e0e-a681-2f9ff39a6224 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:18.006760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:18.006760Z digest=sha256:4efe3346b4afd8a28d79e3a2d74036cd8de98a4e5bbdd58d10c7e85e988a3b7e

Observation e45504bb-270e-4351-b01d-cf44a4d26041 · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:13.173718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:13.173718Z digest=sha256:11c533065e97defd626bfb0b13b18fa556e2424b0aba139d9a8416b68589226b

Observation 4738e270-1e50-44f5-8425-a2d3eae29a93 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:17:18.579232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:01ebc65d5d7410ccfe41664b72f6d22f63dc18c751f8ac3660db2c9c083bb8e4

Observation 567138e8-a2ed-4853-a95f-46264c02480d · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:40:56.053231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:a89d1336f8336bbd2d7298bbb332edd97cec0180afdad2e1c9bc1f654877af82

Observation c1befbf9-e216-4ae7-bb19-6a771e49d146 · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.023507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.023507Z digest=sha256:c43a089a78f93597ddd213d2f556d29a2411f86eedbd3ea5a2f53282d1b7b50d

Observation 12058bba-b8d3-4e0e-bbc5-eab30b9b953e · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.234976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.234976Z digest=sha256:e791e64ffd7623908d4dca74d5d22c8cc661b61a19170b289c3536cdbd56ba07

Observation a6c5df9a-3d2a-4ad6-bcb9-dd537cb0aaaa · inbound

Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward cites this paper.

Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:50.420151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:50.420151Z digest=sha256:ef3f206ff0908470bd981aef481b475e13d23088937085e88abdbe30e8060baf

Observation 78530adf-24b4-4652-b6b5-b232e77c5410 · inbound

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations cites this paper.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.624919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.624919Z digest=sha256:67dd41a841b747a8964bd5d7130fdda858f3cf5524b7ae12e746199a87dcc6e1

Observation 6db7d921-f0b6-4b81-9228-fda0ec850c77 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:52:08.036603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:3c7596aa9b77f063735432b90205168099d267ff353bac86349ea475a09dc0a8

Observation 181b8286-f0e3-43cf-8069-371c4827b030 · inbound

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning cites this paper.

Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:52.737642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:52.737642Z digest=sha256:9f6a01908e5d03a7303371a1cad29a42756539d864d7225409a766f5f6108233

Observation cb440322-a61d-4a06-8915-060afcbbc754 · inbound

Advancing Visual Large Language Model for Multi-granular Versatile Perception cites this paper.

Advancing Visual Large Language Model for Multi-granular Versatile Perception Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T15:25:03.137410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:25:03.137410Z digest=sha256:1dfe200fa55e277cb3e812dd2a8536da9ade113a3d02f10e855e85c41c160a39

Observation 67e962c9-f402-42cc-accd-73d4f0f5d21b · inbound

Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning cites this paper.

Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:28.039862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:44:28.039862Z digest=sha256:b0e7652e2c966c74499363f6e64eadbe4881e5953040df0714b8f6e3c443a0af

Observation 9add315c-0754-41b9-ba02-1dd5eb712c19 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:33.307499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:33.307499Z digest=sha256:6a804d982855d50e7aedc7814126cada0df5930ad75e10a8dbdaab1b420aab20

Observation fa93c31f-f6eb-454e-9388-ac956e1afd3f · inbound

Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding cites this paper.

Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:32.447112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:54:32.447112Z digest=sha256:b05ffbc7c67563424a7937b3e054e50a9e383122d02155c371da5aa6408e3bf9

Observation 6aeda5a9-ec24-4ea2-bb08-123d214dacf4 · inbound

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding cites this paper.

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T22:01:26.003333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:01:26.003333Z digest=sha256:3072a33b232d3349e71b821948511f37781b628a730c3d558e7314d0167667a7

Observation b7b55908-738b-43e8-aa27-bbc80aecf7be · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:18.032059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:18.032059Z digest=sha256:163291170d55766495d45cba5c53a9773b0b547c46dac0f968add74cd1400799

Observation c68a39ef-5cef-4f18-a352-8e87bf843846 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 288

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:11.193788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:11.193788Z digest=sha256:042f4414d1f47a51322d92ef535b0fa882adb91c6c50ad132e416e4073b33f25

Observation 092dee13-de5c-4c06-89c5-20e794ea9026 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.108809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.108809Z digest=sha256:5cf1c0c80b4ccee812f1643635016cead99487a80248208d3df13a30be18efb2

Observation 5f14d52f-3aa6-4dab-b8e7-339eb77843a5 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:25:22.540751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:50f7e4e54ab94162cb0c34f4a1a586e71e37c243e66e86d0a8578a79fc2855f9

Observation 43b4d501-e9a5-4dd1-b1f0-1aa7af507935 · inbound

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding cites this paper.

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:20:22.937278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T22:19:36.366837Z digest=sha256:0e670c233e4824748067b608b909dbea5f168906834735067f5d325a1db20902

Observation e0cb23fb-6be2-490f-aad7-b67241350fb9 · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:31:32.196643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:82409bc82ade60f1a9fdc568d879347fe55b375c0138e2a7a16cc4d4463d496b

Observation 9b187b0f-a38d-45a9-a9db-c82b763254f3 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:35ce8dcead2ef4125465b819f2e20590de8cea609b4ee2bd22eda2d2552a9223

Observation fb327704-a237-408e-84ee-7a14cf52e717 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:1056ce12f0f90e0469f18825781cefb13da84eb007163fc70a344125c4dfc342

Observation 092753b1-aab3-4aee-806c-3130dd836dc9 · inbound

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding cites this paper.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.236016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.236016Z digest=sha256:9a9aad6387036a20f790bee345a0e1ad9c04f3b9d0805d91dc8eff9002174440

Observation c3eb714c-6453-48d2-84e9-d25fe74e9893 · inbound

AdaTooler-V: Adaptive Tool-Use for Images and Videos cites this paper.

AdaTooler-V: Adaptive Tool-Use for Images and Videos Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T21:23:33.598026Z digest=sha256:3c59494c0b40512a9fc848683ab88b263d9083ac5d9061ce92ef55e481197e0c

Observation 6bc99a11-4a63-499f-ab31-df0d8b8969e7 · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:f9d2c0c637c51c3f65cb6f4b07bc979b328c3b189baa545d38c53b8c53e1a638

Observation 4495bbd2-de89-44cc-a83e-b0b39d756171 · inbound

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking cites this paper.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:05bbcb04196ecf29af85040c9dc3a95b5caeec56b2943f9820a8ada2efe2b037

Observation b1878c71-7d1b-42b3-a3cf-d3016bbb1f29 · inbound

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation cites this paper.

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T00:15:23.806078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T00:15:23.806078Z digest=sha256:ec84519358fa8e41b50b414fd4b9ed55a10a409265750e2e8039eba11d0c8a1a

Observation d177492b-4fd6-4df0-9a25-d1f3229a1487 · inbound

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering cites this paper.

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T21:12:29.596207Z digest=sha256:40865ca22ad409a101591954f8600eab2e6f65ea422c16762b59a559bcbfc95c

Observation 556f6a39-46e1-4030-bf4d-23a5f96c1067 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:23f62cab46ef85357958fc6c332671a42d9cec5b8683478d611544a41b912256

Observation 5e8ebaf9-3410-4ae5-bd62-0a945189ece1 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:ec7a929e14b98aa57c188a8a42aa56e50b1306a9d3a9fd4eab52f64769f8a330

Observation 1f3062b4-9aed-4595-859c-7a7a414ca15e · inbound

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt cites this paper.

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T12:37:17.843365Z digest=sha256:b8c1e911519d5ece1a4c60215a656851973a9c8eedde28cf8e4fc0371c482085

Observation d80fae7a-6546-44ee-8e6a-ea9528528f23 · inbound

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track cites this paper.

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T03:29:12.783993Z digest=sha256:d072d0c6d6655e28ffb4f2874176998445701978f4b14e8848de39d011a3b2ef

Observation 87db41aa-b680-4d07-b2e5-55283a5d870f · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:ba310130fb6f362f5ba066a97c53814bd87ac26cd2e595e59d63008a4dc76d01

Observation 3dc60eae-4c0e-466a-8773-d25ec3fb2e34 · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T12:45:50.422679Z digest=sha256:ea3bbc57c06f478005bf10a7bcb1d28b5f34fb6bf2cecc3b212309f34572538d

Observation e5200291-cb6f-42a9-b04c-0cc105e58cb3 · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:fcab189585e2b6b50f855d16f112189ff3faea82b3978d9099318f5bb143b2df

Observation d5ed8ec9-9d40-43d4-a267-4c05d366684a · inbound

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method cites this paper.

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T05:14:25.423302Z digest=sha256:7050c9df6b9b527bb69199bb602680a222cef8de074480b66082b3aa05e96787

Observation 24bee16c-f63c-4f5a-8081-e71523b2584f · inbound

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding cites this paper.

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T16:51:43.783133Z digest=sha256:499a5c6b8bab94df114dc7f3d1c5273ae0ce9809d8bdc390a7635db226b796ed

Observation 9cf10bfe-4981-4ecf-9d99-11d5955202f0 · inbound

Co-Evolving Policy Distillation cites this paper.

Co-Evolving Policy Distillation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T08:23:41.819485Z digest=sha256:d5a3755200ef539a6b7f44c98aa02df7b0f2531d41ef7f6578d3657c3642f5f5

Observation 0e3d2f1d-0933-43a4-ae2d-945ea35d1fa3 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T14:06:27.953376Z digest=sha256:3bef2a8ac04398b5688288d1b9c53a53814e8b796fa2a38f6d3ff4d87e27e738

Observation df3f111c-e38b-4cf7-9430-3a8e53b84360 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T01:49:41.654207Z digest=sha256:01a717bf108c2f21f17537126e07c24b3bcae4fe5864a76b50f86a4c5abd2a58

Observation ba62a7b2-e7d0-4945-ae60-f46f4bbe7630 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T03:35:59.553683Z digest=sha256:764c2f32703de1432f9e4aaa0f94cf338fd2ccb04ec81405fff93aaed530afd7

Observation 4a13e00c-9ba8-46ef-b63d-ffc85dda9839 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:10:23.838149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T06:08:19.956833Z digest=sha256:4d75c9ab392dbc24d085e45e4ec55f7070920313220840c8039e5feee0e7d97a

Observation 4faa07dd-96e9-45f5-943b-e21ef325d1f5 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:1c5fbe4123ff2e29ae5ca30502663cba4c04759e31f7614c4a8a20579f496fff

Observation eb64bcdd-6264-4ce8-85f0-d4df3dd37934 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:15b9b7feb941fe9c392cea227f73b6d47fa5eb1fc755ba6616542e5ebed2301d

Observation 4ceac38a-888f-45e7-b5ec-08545880b531 · inbound

Video-Zero: Self-Evolution Video Understanding cites this paper.

Video-Zero: Self-Evolution Video Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.398647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T21:32:16.939563Z digest=sha256:86b2f1c2cb26aa621b58559f402db006ebca08ff0c2d5fa1f27ea5ee2297f3eb

Observation 87698de1-9b46-4c55-aa86-242a60097b5f · inbound

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation cites this paper.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.133728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:fde22733097fccaf25bb25bde88668787a3dd8dfe6423572b379fe4ac67a2baf

Observation c317675e-1d6a-495e-a30b-121de9680c68 · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:34:02.631324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T07:32:12.180233Z digest=sha256:5828e53342e193a974674f3d0bec8b8e3c881a1e14e271be1b86e00530cba351

Observation e7c0e7d3-1838-4170-b85c-c3ce2f1d2511 · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.294041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:a7c0ade0a61c47a8650369482b3909564b793b9a421de3e44f8fad84fd8d3ada

Observation 35d05aec-0652-4be4-b667-3eb96cb60d69 · inbound

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models cites this paper.

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:21:12.922592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T07:19:30.508843Z digest=sha256:e83ee5642485d5c85d00adb0930b2c69e78838d4641a92b5de620fe82c52d7af

Observation 8c6773e2-b13a-4afd-86f4-be124fc01250 · inbound

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues cites this paper.

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:14:42.423672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T07:13:43.716510Z digest=sha256:753eb368b4743ccd08f1bd8200f0623cda5ba45fc27e9ba6e52b3391852226d3

Observation e89650d7-ca77-497b-858f-285831a96273 · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:55:23.205843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T04:54:23.077914Z digest=sha256:c511c671b3975c593866263525fbc5fec360e028e1e4dd173e26db8c90d20952

Observation 598c87d5-05b6-469c-bbe2-d13a9179001c · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.488864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-04T00:41:02.284215Z digest=sha256:079e9d9ec7fca986780261c9af7a5df98ce471bdd32616280135d51fd9e7603a

Observation 9b9223fb-c6e2-4540-b335-32e90639ff5b · inbound

Towards One-to-Many Temporal Grounding cites this paper.

Towards One-to-Many Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:16:57.661656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T02:11:48.455492Z digest=sha256:b2db1a193847e796aa64e7399251c6feb92c03a44dbd60f5cd881be48b14e9e7

Observation cf68d152-da0c-4486-9e6e-bc9893b58257 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.836506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:356dab0d3298205edfb6720440c685fcb9685beb498544a057ef34b21a26b0ec

Observation b1e67d90-a1a5-45cd-a365-3bf83f3ce2d4 · inbound

NEST: Narrative Event Structures in Time for Long Video Understanding cites this paper.

NEST: Narrative Event Structures in Time for Long Video Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:31.224782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T17:57:55.366051Z digest=sha256:5707ab1e83693d3b49cdf6da9d289dd156eba7ad8328e869796ee94cd87ba841

Observation 94b3e95d-b8e4-46ef-ad21-ef862a41684a · inbound

VideoLatent: Video-Language Learning via Latent Self-Forcing cites this paper.

VideoLatent: Video-Language Learning via Latent Self-Forcing Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:49:44.697249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T09:25:14.441526Z digest=sha256:22c917ebcb70f749f45e521917bac4136d66eb7a570da191aa982ac29accfba9

Observation 1e7841c4-0434-4cea-b3ac-ea001b0d250f · inbound

DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding cites this paper.

DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T14:37:03.232673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-02T14:27:15.302565Z digest=sha256:0c00920186f6ce641f84af6884e53e4f74a0bdc55a6c98633aab9d5424fdaffc

Observation 3ebdf02f-ac69-48be-bd01-caf85a2ca558 · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:96ba55be9d2d3b00a16b95a7c5abb2bc9c38c7c894693a1da81efd2822fae439

Observation 983ab6fb-dc9e-4efa-8b06-00dce13acfc7 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:e391e95cb0c8af2bdf052714720f526f8f7320c47df8b9e074376dccba385c19

Observation c4169a6b-2f5d-4b69-910a-9db16aedc76f · inbound

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA cites this paper.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.070624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.070624Z digest=sha256:30135653295827001b47d7f4f71d8691b623a3d3c8bac3117b0949408f7cce2f

Observation 0841f998-78fb-4532-98b2-2e9a50486d46 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.679249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.679249Z digest=sha256:c498020fe4949862d3e16368537d2be686539b89ce4f48d04904a93967df6edc

Observation 40c6cb9f-a569-4123-b64e-02ba7a841a4d · inbound

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding cites this paper.

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:54.162233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:32:54.162233Z digest=sha256:485831853ffb5b8602e3868826643809e27e17f62226f6239b35d8ad1e66e356

Observation cf7d5b55-0444-40c4-aacb-a628a9681f73 · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.215349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.215349Z digest=sha256:516e166d45a9995e438670e883f4c0ab48c318535d97103f2638695c8f8886e5

Observation 1c7eacd8-50a3-4af6-b457-574f29e6d56f · inbound

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning cites this paper.

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:59:13.736545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:59:13.736545Z digest=sha256:3f9940b75aaab4f4b0fc013d5fc764bd68c10d05f68ff851c858c49564d66e01