Pith. sign in

Paper Citation Record · LEDGER

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

As of 4 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 3 inbound Pith citation observations for arXiv:2512.03963.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.03963 v3

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T02:18:21.718091Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:17:27.916261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

72 of 72 outbound references displayed

  • verified exact35
  • verified fuzzy36
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9e924d67-296f-4bc0-97c4-95c6e387773f · outbound

This paper cites GPT-4 Technical Report.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.240459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:711333fc90b29873d4817eb4b46a0f8987353e7005e7d3f736e085f8a1375fca

Observation ca102fe4-0294-4bc1-9fbe-3246bf0dcebe · outbound

This paper cites Localizing mo- ments in video with natural language.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Localizing mo- ments in video with natural language

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.929711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:502e141c962ad9aaa659c67bb3cb84bd9c886b789249442da420bf56acb5c8af

Observation c2da755f-03bf-48c4-8594-4bb41a80e636 · outbound

This paper cites Qwen2.5-VL Technical Report.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.236564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:2982788ccc0f317892224ba2288bf153559a61bfbe16eb5a298ddcdb887e6c00

Observation 2d14289e-c51c-4387-99ec-71096cad233b · outbound

This paper cites UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.266369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:a4797560b4a2a08fbf8db3cf80ff26e5d317cd3463e7e80b4e02413685da798d

Observation 23aaa48a-bfdf-4790-98c1-42b278245487 · outbound

This paper cites Dense events grounding in video.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Dense events grounding in video

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.927612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:becb94e7d3d694cb39f6a95fe5d8aafff546f0d4bb460aed0d74ef55ae82c098

Observation c6fec625-18a3-4d76-9307-8443d9984044 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Activitynet: A large-scale video benchmark for human activity understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.925448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:393bf3b7af0a8d5b5d734017a3c1c15f69c65d87d303ca01a534c91248c6b6c1

Observation 6a65651b-f8ad-4a84-80c0-e377032b98ea · outbound

This paper cites Flashvtg: Feature layering and adaptive score handling network for video temporal grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Flashvtg: Feature layering and adaptive score handling network for video temporal grounding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.923441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:8ae5200676432b93ea11364b5de38a4acbda10508670d1f8d182f9168fa56a78

Observation cd5c16b1-0bb6-4f72-b4ca-8e53382e30c7 · outbound

This paper cites Scaling rl to long videos.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Scaling rl to long videos

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.378019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:538e00d4f56921f15d4f86bc92e5195295dbdbd8f6dc7ab58321ce371141a1f9

Observation 42e182a2-3342-41fe-9686-0d1cf0474f6d · outbound

This paper cites VisRL: Intention-Driven Visual Perception via Reinforced Reasoning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.374154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:d8c952f7da95c028b38d191d77863d58197a4e45daa31e2a0a30d7e494ee6898

Observation 1b69e8a0-90a9-4dce-8cd0-419b334bec29 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.279116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:66fd32a362d8de738e11a77a7c0c7a7b34904074636b9646ca64db17051c2c7a

Observation ebccfd11-141e-4040-af8c-0d4fabdf6973 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.921377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:947402992b0e70976a87494ae6e475517d54474201336a6d318ad57175b0c08b

Observation 68b7ba79-714e-4aeb-92c4-ee1a6901d77f · outbound

This paper cites Tall: Temporal activity localization via language query.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Tall: Temporal activity localization via language query

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.919212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:f1c6bfb5de57a7de3f18e45efab3d1026e2a4a0d52fb7e91189ea12b06d0c044

Observation ecea151b-2de8-4a09-91a9-e993788542b7 · outbound

This paper cites TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:17:07.011358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:5c6aa6c18b3af263f9685ac799cd4583de9e92857d7187da908aeae022eb25b9

Observation 520d7b46-2d7f-4254-a6a2-34761e1c5191 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.257349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:5d316fc52246fa2009e6b37aeec2351b05bfb6eb3054993bdf7560579f0a2d6e

Observation 169e689b-ea81-47b2-ae35-2c795d170f3c · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.385855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:64320cacf674d442f740855d215c5714a84bdbf019f2fa49ac69c08347a5f59e

Observation c545c391-4de8-4bf9-9af6-5814e52614dd · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Vtimellm: Empower llm to grasp video moments

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.917145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:90cd337b8f2c589a2a4cd91c51f59ea4efcd1bf92ff6bdc6b6680a9e4aab9cb1

Observation a01b7b66-7d29-4ad0-94ec-e21dfb402f7b · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.289328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:32548dba1d890ca65460082538ba0b833b0a2f13b3e79d2627a17e0a66574567

Observation 9807d542-83f1-4aec-9b6c-a81ac8618563 · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Online Video Understanding: OVBench and VideoChat-Online

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.274928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:d07716b2cdb39bc32986c84988c107da1da632807b703aefa55670bd68262db1

Observation 4623874d-a6f0-4582-b3d5-c36b164e0407 · outbound

This paper cites in the wild.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning in the wild

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.914937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:ddea6b37c5f22753000d773c949ed420e22bc72a359afc42ccca612761226874

Observation f2974a46-ed9a-4530-a305-09ee1de41461 · outbound

This paper cites OpenAI o1 System Card.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning OpenAI o1 System Card

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.365201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:cdddf96cea0768d5cab896a1c75c924a8c657672bb0db8f4cbcf7e09ec3e1f42

Observation a1bfd47c-1aea-4da4-a8d9-9ffb80f8ef83 · outbound

This paper cites Knowing where to focus: Event-aware transformer for video grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Knowing where to focus: Event-aware transformer for video grounding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.912742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:ee4830015052ac604a5b448e95ae6742533f548c63d4f891188e0f6e1e5a1752

Observation de5e88f1-43a6-4c72-9611-8fbe04d39583 · outbound

This paper cites Dense-captioning events in videos.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Dense-captioning events in videos

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.910580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:3e174ca57950dd2ff8463b6290ddd570426877109fc25c10bb5c392b628b1c8b

Observation 6c2f1939-f828-48b5-941a-67ca5bde11a3 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Detecting mo- ments and highlights in videos via natural language queries

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.907868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:95fbc0e3331fd2a0bf6ce149f369a2276c18f846ce81bade15996da8cb8c49cb

Observation c12dd039-c2f9-4ce8-818e-28bf6ffea39e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.294965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:8a68ae3843d8ec8f0241805e142d9b29731cc913f4c77bdd88073832b9164e12

Observation b52f970c-bf0a-4400-9c52-e9db42d95dfe · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Unmasked teacher: Towards training-efficient video foundation models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.905506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:41ae51bac1ca836790265d9bb4d417869eeb8f474e6902938681df44a3e54207

Observation 01f84cc9-2158-465c-85cd-b87dd00ccf88 · outbound

This paper cites Mo- mentdiff: Generative video moment retrieval from random to real.Advances in neural information processing systems, 36:65948–65966.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Mo- mentdiff: Generative video moment retrieval from random to real.Advances in neural information processing systems, 36:65948–65966

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.902868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:5f9848f8e68c68893d0a56f30ac615042e760e26de34e4d626f46bbcba276d25

Observation 272ecf85-534d-471f-9d58-5b4347c00ece · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.299074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:1d30e866fec37a18be91972df678c6c7684f415e4d061a53411b32e825a2bd12

Observation 93ac2896-e900-4888-a00a-4a25f5a5458d · outbound

This paper cites Ground- inggpt: Language enhanced multi-modal grounding model.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Ground- inggpt: Language enhanced multi-modal grounding model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.900629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:a61b1e69451abc377d414a0de598537664d15439585c3d5b420e2e5dc19d3039

Observation d7c17fae-3323-4fa8-bc27-438b4c46233a · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.841392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:c37c0d779c748181e344fe2d95d6ff5fa403f3b508590237f562548aebc1186c

Observation e56712fc-fcc8-4a34-96fa-f3204fd3b467 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Univtg: Towards unified video- language temporal grounding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.898220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:4b74e5033b85703b0983be7a96d56de50d5b8f22e5ebe9455e0f60013d5ee9d8

Observation 681f3c37-9bc1-4d10-bcc4-050b33aa5877 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.896063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:486db2159b27b41623039e4411c3633221e0a68548a79a5d46bc04db4c65b161

Observation b8c9c9b0-1b9b-45fe-920c-371565cd023a · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.893610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:8d3b01a49b897aa26af69ab9b4e1d14d119bbb28e12b3cd6bae020d75476aa2a

Observation 672cd8b6-56a2-4a32-a61b-78fd726cba21 · outbound

This paper cites r 2-tuning: Ef- ficient image-to-video transfer learning for video temporal grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning r 2-tuning: Ef- ficient image-to-video transfer learning for video temporal grounding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.890767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:ba1449105802dbd0f7dba669af795a29d74a175dc5dd8c9c00a509af12d6048f

Observation 28b801d3-bcaa-4514-b818-0c7cdf0e3eab · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TempCompass: Do Video LLMs Really Understand Videos?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:17.144047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:6594e9b0ae6d93015f00e04db7f58d70f11a6fd23993430274e6ac59a6852a03

Observation 442ec364-5fa3-4d06-918a-45a016735ab8 · outbound

This paper cites an unresolved cited work.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:18:52.888371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:3ee3070e2c43d58188ea08d642c753187f5bbdd4a9d513adce36220f3783fd1b

Observation 4f370744-bb59-4f16-9eb1-72d9333d5167 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.360675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:c84137c41f0b032d25b105b4be280b683e85147c1bd4fece2cf0f367141f801a

Observation a810bf84-4b8e-4662-a534-5e19cae1e3c7 · outbound

This paper cites MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.382073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:68cf2cc66ee47305987c5cb8889bb57bc4ab3c12f31ab3ec7a24a83febcdbff0

Observation ce32cea4-58d5-429f-acab-add86dd88f39 · outbound

This paper cites Correlation-guided query-dependency calibration in video representation learning for temporal grounding.CoRR.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Correlation-guided query-dependency calibration in video representation learning for temporal grounding.CoRR

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.885789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:4d2ba15adc60ddfe39891f34bde4ebe1d428e194b35b381b12f9859724f3c66c

Observation a2f9eeaa-84bd-46d6-ac96-24c143f27db4 · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.883621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:ee0f361cea11b4f7ee8352e1cf498a10aa2116f9cf4d152a0b3faec5f40a54bc

Observation 9dc16043-078a-4a3c-8d20-3d1239b663b5 · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.244560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:f137d5697b9470f08bf8cdd3c595ed1d4eedd5b361657be8bd629dc47d8aa53b

Observation 1a5daa42-d2f4-4fb3-af9d-0cf2f6326335 · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Sys- tems, 36:42748–42761.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Sys- tems, 36:42748–42761

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.881302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:f4b5188fab9bad51cc5de9f999fd43c3426810334a099f8eec8a361e8ca6cde4

Observation 11d2dfa6-32f8-47de-a608-34136f45a500 · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.878949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:3c4044eaf6e934b37c1fa6d0269f655ba04c41aff406f229b4752e497092a625

Observation 59f2e683-7596-40d4-8c8c-769b76484e69 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Learning transferable visual models from natural language supervi- sion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.876556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:6b9b2be9183dfd8eff469723c4c87f84d7aea480a285351b119780014cca38d4

Observation 6dbe3818-759c-4abe-ae7f-eb186a401164 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.873515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:a5188b9414d216051f92f31b7c7ff9a48abc1642096706b7d7c22286e41f8742

Observation 4f603076-e551-4285-867f-d759b3d35f27 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.248653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:3fbe296484580877cb953abfbd496be71fcb03396ce7c61e34ab74f4d4cdaf3f

Observation 6d828343-faca-4796-998e-a3172586f1a0 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.253010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:07a7aef1b34917f48501ef7dee29d966481832df002076b898c9d5f1ac5d5f1e

Observation e4818a9e-30d2-4592-9db8-3dd509e5649c · outbound

This paper cites End-to-end dense video grounding via parallel regression.Computer Vi- sion and Image Understanding, 242:103980.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning End-to-end dense video grounding via parallel regression.Computer Vi- sion and Image Understanding, 242:103980

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.870909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:448a610d8466f8f0da66086f67947a57c1ffb3ac981d3144607d71cdd4660030

Observation 1459d225-e80c-4a46-bab7-4e39c11f7548 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Moviechat: From dense token to sparse memory for long video understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.868466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:7e0eb7d2b880bbcd93b9d22c1b45147dc2f1abbed0bd6c5595f763b42d399ffa

Observation 7e7cee7c-7670-4ba0-9e21-e339cd5aec3f · outbound

This paper cites Tr- detr: Task-reciprocal transformer for joint moment retrieval and highlight detection.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Tr- detr: Task-reciprocal transformer for joint moment retrieval and highlight detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.865957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:a9575cf75ec43302c57f2b4ac4f6d5f332367d9fffb7d918c030906d929ae949

Observation f233cbca-53fe-4135-bf8b-a2ad4156c04e · outbound

This paper cites Hierarchical semantic correspondence net- works for video paragraph grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Hierarchical semantic correspondence net- works for video paragraph grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.863440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:c8843c8fdc803d2bb8da6c50fc7766d036e4e9fd6f4927a22cceb63b1f80b067

Observation 15dd6e88-c57e-4c4c-837a-008464009a0f · outbound

This paper cites Hierarchical semantic correspondence net- works for video paragraph grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Hierarchical semantic correspondence net- works for video paragraph grounding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.860479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:94af558ff7bd6cfc8d37abc249741b86f2ef32e7304fe3c84fa5e0b0809b4b21

Observation dc84d14d-1295-446c-b300-38fb62d65e0f · outbound

This paper cites Tspo: Temporal sampling policy optimization for long- form video language understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Tspo: Temporal sampling policy optimization for long- form video language understanding

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.356084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:7860a20f4d149d33c36eb5448a994dff872e8eb00d606dacf80057c82317f67a

Observation 862cdbb2-b580-43cb-b6df-242a8ec5a795 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.302760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:91c785cd798c834de9b9c371ea66b59891c4cd91e1310c53166394d78b37f5f7

Observation c6333379-37db-4fd4-85b4-7502732d183e · outbound

This paper cites Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.351204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:fe087bcf070f9783a4ca0d2d6e036a9c79dd50e05da118bcb97f5ccd5e12cfd9

Observation a8c5d2f2-9c0e-49b3-b0e0-b42e0379d4fd · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.306944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:e5bcde075b9382aeeb94108322f2e62bc11a8eff3de0cce70d359037e03af61d

Observation f8404e5c-5d49-4305-b484-126752c65f39 · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.857516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:fb80f4101c317561a49c8382775780f54dbe06681dbc440071ffa4c9c47323d5

Observation c2633b38-8492-4fae-b0b6-62c18d264932 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.343793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:68910aec8afb8f372f752092746d1716a3a9dcce53bcc1ef1ea1b9595d446c94

Observation 5de38293-abb8-4a48-8d1d-41bbe9413f47 · outbound

This paper cites Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.339400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:09907d2d96edf39aa0bd9c29cc07b02be8b3c7e422d1f3725ce11270ab79d0a0

Observation fb327704-a237-408e-84ee-7a14cf52e717 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:14ca406037c2e36fe66586a79101277d6a2a570fbcf4fe6c061abeb9551d3378

Observation 1a187403-cbb5-4c2f-8eba-367d2645778b · outbound

This paper cites Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.284171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:3d048384de26154d33cb1a54c3c21e281830e4e05c966d3d901a3ce83893d2c0

Observation 6a123d00-0d84-4dba-92cd-0ef2c475cd52 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Can i trust your answer? visually grounded video question answering

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.854622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:6f49b63253bd12bdc3165842c707459894fe5ff6d4314540eacd1508884d1b39

Observation 1b539996-056c-4907-9e4c-50a5e54ba3c3 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.851888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:bafdb9dff63e5690de93624d7da5d69ed79ad44879e2722901ec750367ceb81e

Observation a30d6193-0b77-4a33-8a28-fdcf02fa13f9 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.335434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:612f221881bb4d1c25e29fe20b5063da1bdb5ac761f6a3f5e6085a887d55f963

Observation dbb608f1-cb63-4ed3-ad90-f84c4bacb0c0 · outbound

This paper cites Videochat-r1.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Videochat-r1

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.314177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:113588083aec827f38eb3e435a9d2b07776a776f754d44c5e320573c686d819f

Observation 067c528d-c020-458e-8a3a-3237b3e1d735 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.328163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:a34f4023cfc885f781dd1b74af17959f1b2ef0a22cc4ff8632d2b02bec1a7f71

Observation e0fb155f-e300-4ee5-a867-e80b8a875f35 · outbound

This paper cites TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.310729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:634c37d61dad136bbc2c013d76cc3966281c3d82fab9119172f00924afd1e0a6

Observation 3b27a866-be10-47e8-90f6-c954341edf38 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.324847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:aa76607033a6dd01b6b69f15059185407930dabb05f1affb563f664d89221dee

Observation be23abb0-5110-4e32-841d-bccc4487b546 · outbound

This paper cites Sc-captioner: Improving image captioning with self- correction by reinforcement learning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Sc-captioner: Improving image captioning with self- correction by reinforcement learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.849268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:aa79071d2104b5ffaff8f429ec98ad13cc78a6a10961e63b7704419922c12b8e

Observation 7e940b54-6378-4f98-b9df-28dc33c57448 · outbound

This paper cites TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.321509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:dd20ef5db90d97262730909129694fa05c34c2d18f22ee214e0331bb90af930c

Observation 6ff04040-ff6f-49d2-9d95-6d3bd7fc1f47 · outbound

This paper cites Hacs: Human action clips and segments dataset for recognition and temporal localization.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Hacs: Human action clips and segments dataset for recognition and temporal localization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.846707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:fd6c94bf86bfe35585dc9079340a20829af1bce2f65e2f247c5489019c712352

Observation b5e099b8-c0de-49f0-bddb-704e8764274b · outbound

This paper cites Group Sequence Policy Optimization.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Group Sequence Policy Optimization

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.317842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:e90c1306dc56997fd852537073744c28d4d30b1d080c6cf25520158ec542543b

Observation 72669cb1-ea7a-4f24-8e64-5a80eb90dc0f · outbound

This paper cites Rethinking the video sampling and reasoning strategies for temporal sentence grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Rethinking the video sampling and reasoning strategies for temporal sentence grounding

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.843915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:85ed448cbee9581e7c803e904582d8ad5a59427ae6c3df81d633a6f5524cb494

Pith citing papers

Observation 38fb59fc-0b9b-4067-9d2c-c447183f7ddc · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.916261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.916261Z digest=sha256:a65322a7af20f88afb54a55b0b7c0e15b89430fc679d4fc92324ad0cf9fe1e2e

Observation ce54db9b-5438-4540-9a1b-cbffd719bea6 · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:de09fe056cb8d8be7de8c8d0ea777c4c1f693361f7267f0f30936f42df372f52

Observation 31e83b57-04d5-45c2-93b1-b4bffa41c296 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.823032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.823032Z digest=sha256:ab5ef288effbcee38b57d27d8d5578ecaf4c19cde570df61edc0fe81ca3815ad