Pith. sign in

Paper Citation Record · LEDGER

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

As of 11 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2605.20342.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20342 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T08:59:28.405218Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact39
  • verified fuzzy14
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11d64e63-8697-451f-ab17-adb889e689be · outbound

This paper cites Qwen3-VL Technical Report.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.357015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:d8ee612e79cd183dfeb599df795bac7df145a76916c2baf0ad405d28053da5c3

Observation b9ab5b36-e97a-4777-a6cb-7a72e05e156b · outbound

This paper cites Videochat-m1: Collaborative policy planning for video understanding via multi-agent reinforcement learning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Videochat-m1: Collaborative policy planning for video understanding via multi-agent reinforcement learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.363227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:055e6bd2972a98948f57412394601f6a24b08b8bfbba673b7f80abef9d011c26

Observation a07535a1-ea6f-42e5-93cc-bd3e67c0ec95 · outbound

This paper cites Lvagent: Long video understanding by multi-round dynamical collaboration of mllm agents.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Lvagent: Long video understanding by multi-round dynamical collaboration of mllm agents

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.930201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:d5497b3e043d2aed36ee4b69fd8eb35f5a33d9b99fb2cd92c42c142b6fd0002d

Observation 9c4da7a0-ffc7-4c00-80cf-4061d5bc8e99 · outbound

This paper cites Scaling rl to long videos.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Scaling rl to long videos

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.345470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:b0866aac050e7d420b2eb3b71e90e7283648850c07eb7941d9b83a890e0a58df

Observation 4bda5484-570b-42a7-a569-1bde97498c72 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.429746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:3273a2e3a41bf0f6fc65cf260750d31d92c2ac8c4a15f19c76de592a9e20e418

Observation 5d373ee7-5850-47a2-bc2d-27ed854fd62b · outbound

This paper cites Videozoomer: Reinforcement-learned temporal focusing for long video reasoning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Videozoomer: Reinforcement-learned temporal focusing for long video reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.393906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:57a5b4ff7918a4dc5c0c4eda0579214101da299b39eda2942d9dc3c65587ea88

Observation 10e7a9e4-510a-4ae6-bbf7-f0072826bd41 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.493287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:e0239455e8059df89b417d5adc133ed950e4d578e2158d0e0c92deaf530b4856

Observation 43c23bfe-62e2-4bed-9bf6-18f65201e346 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.993140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:916efbc7f6d5045c55d85a8771177021e0a4c537f390e6063200994b18253a77

Observation 9e9eaa7f-0e39-4bfc-9791-380ff2587424 · outbound

This paper cites Love- r1: Advancing long video understanding with an adaptive zoom-in mechanism via multi-step reasoning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Love- r1: Advancing long video understanding with an adaptive zoom-in mechanism via multi-step reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.313489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:f26a2b76fb3ea7458d9e0df504f44e66f8a79d4c969c4a0499c5f0d0274d7ab3

Observation 15a3adf2-8211-4f33-84b2-6124d127428b · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.399223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:19bdc5f4127d269ac758d66ca0fa284f4970a88e7ae53783dc528125906bad05

Observation 01b19d05-c715-48f4-bcee-fcf5a7706839 · outbound

This paper cites Tall: Temporal activity localization via language query.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Tall: Temporal activity localization via language query

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.974857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:cdce4bb56ccb9b3cab2a87f87bf12007b505b773931fca0e6cf3bfc9ecb58712

Observation 64069ec1-39cf-441b-bdc2-66a380f07e77 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.333120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:fce648609397a09afa45f0ae68fe0ad19fd28d29b641d0ae3bf686bf216d20e7

Observation 4491bede-b39f-4dfe-be41-aa9c2f3a23c6 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.452182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:c643222de026b63c55c3bb561e62e2f9b6979debb265747fd25141f9f47cfba2

Observation 835da0e0-7969-4c58-b50b-0d44511cf5fd · outbound

This paper cites GPT-4o System Card.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning GPT-4o System Card

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.463016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:d35ab1b52c54f731b5d43ceff25ea7735978a3d9c1243953308c5994258a166f

Observation 5c96d7f0-6cf8-4193-b55f-be3596f6bcd2 · outbound

This paper cites Sage: Training smart any-horizon agents for long video reasoning with reinforcement learning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Sage: Training smart any-horizon agents for long video reasoning with reinforcement learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.505342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:3ff7b8765500c2c8f47e120d6f19defc9b83a9564a03af6f766723e3c9979746

Observation dcd26375-f579-4fbb-a128-30fe1406fc17 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.480880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:611f477b1dadc6deab84fb2cd351893b0079f7244a2c7f1d182093c8d6e4afbb

Observation 2a996420-c494-44da-807e-579e2af75a27 · outbound

This paper cites Longvideoagent: Multi-agent reasoning with long videos.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Longvideoagent: Multi-agent reasoning with long videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.499479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:6675d8ceded16e8255d900f7ce28c97bdb88a1240d0a91d88afcd409e3ef65e9

Observation 3333968e-5d03-4a75-8f9f-4b7c2843ad0f · outbound

This paper cites MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.306955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:78247d8fc00439c3f30cf204b83431709c6d1a2c6d47a438c8288d5bc2754c7a

Observation d8dd16e2-9b44-41e6-8459-ed5c2c9ff29c · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.375108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:97a3fe30e357b3698180bd1766e6fc1838222d0aae388c2153ac0bb7e6905095

Observation 66263d3d-1e07-46c2-9bf7-313456d8880d · outbound

This paper cites Sparse but critical: A token-level analysis of distributional shifts in rlvr fine-tuning of llms.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Sparse but critical: A token-level analysis of distributional shifts in rlvr fine-tuning of llms

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.418117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:a70cf93f57d96eca9d93742b2488fc7c5c374a6a77879fa7e4cd9b279a1dcaef

Observation 63265969-52a1-418b-911e-bf4039e3173f · outbound

This paper cites Conan: Progressive learning to reason like a detective over multi-scale visual evidence.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Conan: Progressive learning to reason like a detective over multi-scale visual evidence

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.288368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:403b88a678f07ca373bd9094b8e57fd84cc561d22a8b1da33fc4e6f05b699e94

Observation 4dceac67-6f16-4c5c-b430-8424f0cdd054 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.446979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:68ed90668b4eb90954b3210397c1807dfba12be47c749d00228baaa3574ab9fe

Observation 0c86dbcc-5cd4-4e1c-85ed-99cd3007d537 · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.405588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:828ee322a7decbc7ee4049bdc1e3784c7bbb43d95e871d995fccca323d4d513a

Observation 9577281b-0530-42f5-b15d-1405b49175e6 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.468570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:592a192e2a6b21305a7425642e066dc63019a32c8dcdc1f71afeab5c0058e91d

Observation 9fec77ae-da37-46ec-8e5f-eb70100f53c7 · outbound

This paper cites Qwen2.5-VL Technical Report.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Qwen2.5-VL Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.320041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:294d020a1c4c2d822b914398a1834183ec8ad69482cc04988b940eebfcfb499f

Observation 318600d3-154c-4be7-b6a7-e1ddb14efdd7 · outbound

This paper cites Revisiting the Superficial Alignment Hypothesis.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Revisiting the Superficial Alignment Hypothesis

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.369543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:7c6dabbe06e45e2dd9426d13c6076277072777177739239d290ad50426d1f457

Observation 34c13d4f-32c9-49db-a094-eb0077c25ab4 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.Advances in neural information processing systems, 36: 68539–68551.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Toolformer: Language models can teach themselves to use tools.Advances in neural information processing systems, 36: 68539–68551

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.984137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:abdf482ab5caa73f164aeb2161320690f857d02c059561a95d765728c9f65fb6

Observation 93b694e0-e2b5-4647-bc9f-35bd577131b2 · outbound

This paper cites Zoom-zero: Reinforced coarse-to-fine video understanding via temporal zoom-in.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Zoom-zero: Reinforced coarse-to-fine video understanding via temporal zoom-in

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.339818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:9057788a38bc607a1f380715a889a05cbca5f343226e8693a8df0defd4a5d9d3

Observation 6335e4ab-42b2-4114-be63-130c9bad0569 · outbound

This paper cites Defining and characterizing reward gaming.Advances in Neural Information Processing Systems, 35:9460– 9471.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Defining and characterizing reward gaming.Advances in Neural Information Processing Systems, 35:9460– 9471

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.978251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:33cfca8d840886e70758fbc3641407a2231bf86a658b6d33f352c457bbd2fa40

Observation 0e369321-cba3-46a9-bba5-bdc3c1ffcb3b · outbound

This paper cites Enhancing agentic rl with progressive reward shaping and value-based sampling policy optimization.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Enhancing agentic rl with progressive reward shaping and value-based sampling policy optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.282064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:ffeda6c12fc0b7de898ef453928a08d0a9cf593f01b7c01f9f113fb39d0cbdf1

Observation 426f22e1-0f62-4fb6-989e-b07d4be97c33 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.326982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:2ad6f153d6a4d4e736bedfc38cbce76112f85c949e48eea1f7704c6bfe35f024

Observation cc5ffa93-4aa7-4fbb-8afa-8c7c0ad5032c · outbound

This paper cites Ignore the kl penalty! boosting exploration on critical tokens to enhance rl fine-tuning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Ignore the kl penalty! boosting exploration on critical tokens to enhance rl fine-tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.987834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:fafa5b089cf8ac68e832fdf6ee28f143f7905e23d84b123824ac476f944c3a9b

Observation f56c6107-f6a2-4284-9e06-839afd8ae116 · outbound

This paper cites Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.412131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:fbe9630da6135592c1c9d6dfe24cc8ea44633fa6f44c1a967d778ca740585312

Observation 447ca47e-b469-4201-ae53-6c35e39e9d54 · outbound

This paper cites Video-thinker: Sparking” thinking with videos” via reinforcement learning.arXiv preprint arXiv:2510.23473.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Video-thinker: Sparking” thinking with videos” via reinforcement learning.arXiv preprint arXiv:2510.23473

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:01:19.423879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:56cc06bb249fc0d3b7307f091f5e8ff73f62e7f965e68120eff760013fadcec5

Observation ee9f4ea0-3d77-455a-9db0-c087dd67995c · outbound

This paper cites Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.381541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:54d3af7c4929c8bd97b3b2f74232f881dc8ce873ca2d2f5d298f258f2af62ab1

Observation 4a6146b1-1afc-48b8-aecf-5823935a520f · outbound

This paper cites Lvbench: An extreme long video understanding benchmark.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Lvbench: An extreme long video understanding benchmark

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.968664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:77116c0e6668b10b5f6a8110d008fb24d9d276435c5646aaea9019d5a0632148

Observation e7c0e7d3-1838-4170-b85c-c3ce2f1d2511 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.294041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:7f279f42c814625eef1c921fca0b673f0b42621b1e44b230083098df6ca788e9

Observation 16e6aac2-8b40-4e69-92b0-fcd993528ea5 · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.922182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:a0c86f7a8b96c3d83b50e427b81c309e7dedbdbcca6ad7784414086849dc340b

Observation dfdd3cfb-8742-4a43-94a0-7c5b9e6e09e9 · outbound

This paper cites SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.300469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:de2d141ac77fa1207d1e41e9891d19fc238d5cbf1fb50f03203bec652494876d

Observation caa806a2-4ea6-4fe6-8ce5-3067612f734c · outbound

This paper cites Inex: Hallucination mitigation via introspection and cross-modal multi-agent collaboration.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Inex: Hallucination mitigation via introspection and cross-modal multi-agent collaboration

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.934646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:a74b20ceac1b2f93b996eccfa68d76493585b7f382091b82ec99962ff395b848

Observation 9d494c51-66e2-4e41-8b30-5e2ce9880069 · outbound

This paper cites LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.351659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:308e83d66c5b0a43a34ca8716d991b014810fe3ca342067c5fb4617bd0622a2e

Observation a9c212c8-6ef0-4c3a-b18f-328a8d6ed33f · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.387738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:5b7f1fd60c9a4949d8f464a4275d61f8176c11619b8b92cf72cfacf0f139c8d5

Observation 32069add-ff4e-47a2-87aa-6ed8d5967fb7 · outbound

This paper cites Re-thinking temporal search for long-form video understanding.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Re-thinking temporal search for long-form video understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.916561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:ff5f2c197756e941c71974513a23d5f55d6583e064e673c3211eb0e425f165de

Observation 3085af64-b957-48ed-8cce-23e773a66606 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.441375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:b86494e9befb3a97eb8d84005a53f0688db4acb5961f03e2494145c098046302

Observation bd90550d-0b58-44b7-9105-d3fe25a305a5 · outbound

This paper cites Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.474877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:b835e7ae3ddb9380fe46cb99e762b3b315822e6effb8c3ec58ef88b52bd620db

Observation 5b30b0c5-9fc2-4cc5-8af5-5cda085f1ec7 · outbound

This paper cites Rewatch-r1: Boosting complex video reasoning in large vision-language models through agentic data synthesis.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Rewatch-r1: Boosting complex video reasoning in large vision-language models through agentic data synthesis

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.511483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:9ceaa60092a2b3fbdd0de17d9197b816b6ccfd68d6ce571e4300df7ff04268c8

Observation bb64a801-46c1-4a1a-96ba-865f08b1e2d4 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.487650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:efe5e491673dd9a01790d8aad7651f86bacae352abf847b47c5ffe6a40eca2ad

Observation 14135592-ea68-402b-900d-2c6c8c6161e5 · outbound

This paper cites Open- mmreasoner: Pushing the frontiers for multimodal rea- soning with an open and general recipe.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Open- mmreasoner: Pushing the frontiers for multimodal rea- soning with an open and general recipe

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.435943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:5a8218afb36c4eac15bd0405a9f8e28434360bbe57eaf2d753324c4cce7fbc7f

Observation f33abe2f-cb93-4886-89c8-3f7ba45f623c · outbound

This paper cites Deep video discovery: Agentic search with tool use for long-form video understanding.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Deep video discovery: Agentic search with tool use for long-form video understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.517616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:0caeea72e4886c1fb902860ecffaf63dc9b7bf4ba0bd52548329568cc997feac

Observation 8ceb8f83-95d5-4e8c-83b8-9d9cccdc2844 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.457893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:ae616e4ef12df468eda9dc1b961997f9423ee60ca0e62b11fae15ea26a404fa0

Observation 73b7b9ae-961d-4a95-a61b-aab26c351ad9 · outbound

This paper cites Mmvu: Measuring expert-level multi-discipline video understanding.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Mmvu: Measuring expert-level multi-discipline video understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:21.002747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:1f67d0e1e1dff2304c8a7cae19f6eca9533e0b8b90da46cbc096980de5941717

Observation b7f2ff73-8b85-4aff-9276-6c7d28623b5f · outbound

This paper cites Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.960210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:958e7c5c0db26f8432f385ffbb6f712d8e156a22bf8b76814229cdb32eb83b65

Observation cb1244fc-d923-4232-8cb4-750fef90a4c6 · outbound

This paper cites inspect 00:30–00:50.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning inspect 00:30–00:50

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.952117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:e5b74e9970e0fbfd8a47f7c02e242d66a14136a0ee2239815340cb101e9b6a3b

Observation 4be40d03-7c59-480e-b679-19a571246cfb · outbound

This paper cites an unresolved cited work.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-22T09:01:20.945995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:7151c9cbf66b457f50d956cba75563337718d15b17308d10f15e2515efc7fc13

Observation 1b5ec251-eabc-421c-a2cd-5e2391d8222a · outbound

This paper cites You may issue multiple <tool_call> blocks in one turn to inspect different temporal windows in parallel.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning You may issue multiple <tool_call> blocks in one turn to inspect different temporal windows in parallel

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:01:20.939425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:dd2a01372a99c35dcf5ccf38ac126b174d9b981a844397e8198a8dd4f13bab76

Observation 2ea4707c-0ac8-4e46-a62d-54cea6b15466 · outbound

This paper cites name": "crop_video.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning name": "crop_video

Reference 56

Resolution
malformed identifier
raw_fallback, observed 2026-05-22T09:01:20.997201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:6e46e2710311ec786c25c46523ffade08b2ae18332478fa511c6862c0343dd49

Pith citing papers

No inbound Pith citation observations are available.