Pith. sign in

Paper Citation Record · LEDGER

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation

As of 13 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2605.16079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.16079 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T19:37:09.244578Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T19:56:09.820812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T21:07:23.929020Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact36
  • verified fuzzy4
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9fcace9-5fcb-4838-934a-6dced20c9da7 · outbound

This paper cites Qwen3-VL Technical Report.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.049715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:aa30e0d80cad326061e361bfec6d53f1f323675f83043702b733b00f4c8f8d3a

Observation 42db01d1-b377-4212-af33-2e67655d431c · outbound

This paper cites SAM 3: Segment Anything with Concepts.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation SAM 3: Segment Anything with Concepts

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.043858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:5e32a0fca927869485e516bbc1d9481cadec7506a647f1448db9bb2a2a2665c4

Observation dab4ccb6-f27a-449f-87aa-0db2c25d18af · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.036772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:636f90eef7941e9003d68c7feff20a19fa6666db50163508a09ce2b6a6b6e715

Observation 4080fef3-fcd3-4b96-ac0b-02bb11bc753a · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.822686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:9a7a62b1c0f8da131c0d1e4983bedaee2e0e78be76de77980ded63355a0ee3d8

Observation 50e8a264-80a0-4048-b984-0c90ab22ea70 · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.825932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:ac372142b63a2dac55c9d57901c49bd84cb41b7945721c043c570fcdeef61995

Observation 3e0da6a4-9c59-4c5e-b96e-034b676df229 · outbound

This paper cites Scaling rl to long videos.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Scaling rl to long videos

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.030979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:85c40cfef5d4052d877be2b3bb4cf173e6b85392c60d93f523568ad2b2fb320f

Observation 133add86-adfe-404f-af4b-8518ce6262c6 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.023943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:be291bf15ce14afef178bb71764fe40e4ea60bd4f4fa825459e16f4a1341c490

Observation fb1dcc8d-2260-4d4b-9ddf-679b459cb2cc · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.160444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:a4dcbc79fbf9cc484a1edc456a10465f7cec1ff616a47a214f0ba41b4a779f49

Observation bef46cfa-f24a-48c4-a295-3b866e1e7c1a · outbound

This paper cites Deitke, C.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Deitke, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:38:56.841511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:66cd96d5450cfe05d00e836faf3656556b65de9440547bbeb93bb49f2c226dd3

Observation 954bc2f1-15de-4634-b0a3-6da3899382b3 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.101572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:ec04315ae00922a608accfc0f4591e1aa2f7b09d32d274d7c31a68099911605f

Observation 7240c959-ae34-4a33-9109-438ac891ddf7 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.077048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:052abd0d348866fe36d362a60c74da0a6082614fb77b0ae7070482177b620830

Observation 956c3ab7-0b4a-413b-9f99-ddc874cae06d · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.835091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:693f3a53f445a97d7eb35d6210c541b078a673467e53f8b55a3cbfa73406b12c

Observation b84f70d0-e1dd-49bd-8aba-a00038a72761 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.095685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:c4f52a5c5394d7426aa31bf6657e7b575a32471425f29a995b8454853cfdf756

Observation 6667f4f9-fa11-4a7c-b2f3-5486177912bb · outbound

This paper cites DeepEyesV2: Toward Agentic Multimodal Model.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation DeepEyesV2: Toward Agentic Multimodal Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.194938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:7586c9f66a0ebee1179d8ba3e6d262ea46e51a8e803ba0d0f03e897f103838b4

Observation 724b1f36-3924-465b-8c8d-8c3e27935671 · outbound

This paper cites GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.204131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:0c087b190f964ee13dfb96af3034578aa4a332b0312e29e0ea7a75c756cb3418

Observation 033e7999-356e-44da-9eb9-f5ec98b5dee9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.115052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:1f853a40a1ddb052d77f6ebb01ae6515436bbf6d2b82de6247681430e9b8a5fe

Observation 36c84333-657a-4de6-9f99-c8c2a3e817e1 · outbound

This paper cites GPT-4o System Card.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation GPT-4o System Card

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.137975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:9ee484a356a201e93ee472c807ace2e7da9a4f0ed40d190bf47634953e5a0fd5

Observation 4c1b3d22-9c60-4068-b415-1ae35841e91f · outbound

This paper cites OpenAI o1 System Card.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation OpenAI o1 System Card

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.214119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:002fa072f55d4c81e09bb89fd0023860cf4e587bd1e56ae338d1d3577ad625ef

Observation ece79d5d-84b9-4c61-b10d-db5e7d654379 · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.850465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:784ccaca8e5934dcf4ab5af1b29435c99b5936a49ef7c48a363c1cf32da8c7b6

Observation fc5a4d8f-d159-4562-a050-5776a51eb635 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.110678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:afb42ddb8135f922248cbbb35bf9fa6cde880d6e3e7dfde1f0b06dc2a53588e1

Observation aef889cd-b2f6-4c31-af80-19783c661372 · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.862369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:d28d55e6a30c93e4b16562fe954da1acb6f012a403597ec8efc6ad2b0b6dc574

Observation 93ae7cbe-87a3-460f-8006-6d501706aac9 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.081767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:c381260eab5b5a05b16a307716f640cd75ac93e1bfc403a5405f76bb97df2ecd

Observation d8b2736d-495a-4c49-935c-a4edf52757d8 · outbound

This paper cites VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.199595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:779dec91fef41092b6179e54c8115182ecbd490259e469cd897db971e966c3ec

Observation 0aa0c5cf-ce6b-4c0d-88bd-5e6a763ca483 · outbound

This paper cites Rafailov, A.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Rafailov, A

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:38:56.831681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:3787e52270d4bf5c383d9a9ebbbb2f98f23c9bd4b3d9a4221b57b14da52dd2f9

Observation 1cdf0074-dc20-4286-9008-cd13f20818b3 · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.865253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:1e873621e9f061ed829137a29a7a720893c788b29ccea08794821d8a39ae06c6

Observation 4a691e5d-624a-4824-a9d7-58de885ad914 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.119699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:91a1128625e559a53cc8c690c22b9a90e9e02d736fdb26ccb7ce8723d9861424

Observation 390cf3a5-ba9a-4ac8-bb4e-a3208d65369c · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.071277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:338ded1f301c4dbd0c33eecf06f94d7173ed9b5fd48d9a519af781170dfaf1a1

Observation 75c0588e-3c48-4560-ab2e-ba78f482dc68 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation HybridFlow: A Flexible and Efficient RLHF Framework

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.060197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:d79f078a92cf8b2b80f326e7d113e2e01bfa6acc6dd2a62aff2a52d7dbd0984d

Observation 8aa51fc0-5987-48ad-87c3-a7f570d73e63 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Kimi K2.5: Visual Agentic Intelligence

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.054884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:ad1cc07f92e72908d45f31c006e827abe980aa0488cd565b1ab0349d8b2edcb8

Observation e8e268aa-037c-43b6-a53f-1d66bda2faf7 · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.828669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:25af0c62d67ea5b59a0b1ef3b6efe2491c837fad77ad7132dd06628a2803c71e

Observation 109b9733-637a-465b-a045-29273b932d53 · outbound

This paper cites Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.129604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:504f22ee151bd423b28a0c93f9b862fde504fe51ceae55bab62dc9daf6e6e19d

Observation 2929d93c-b76b-4d33-b6d6-7102987ee55a · outbound

This paper cites AdaTooler-V: Adaptive Tool-Use for Images and Videos.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation AdaTooler-V: Adaptive Tool-Use for Images and Videos

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.147802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:184fd5e2282aa0c077843d46cb0d7312b37a32e9ff298addbe6a57a4439872ff

Observation 5556ed8a-a58e-4756-b793-846d66f0c2cb · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.065492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:a2c22782bc7f80023a1a15ad7475e5755aa00fb07a3a7f2a65c5df0e598eb264

Observation c2dbb5f8-e64c-4eeb-8cd1-be369d07c850 · outbound

This paper cites Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.186027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:e79830393991af3d5ba7e8b27a9ceb75680955c72512e4e540f368514d488366

Observation 58b44ea4-88a4-4ac5-a044-eb752caec526 · outbound

This paper cites Video-thinker: Sparking” thinking with videos” via reinforcement learning.arXiv preprint arXiv:2510.23473.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Video-thinker: Sparking” thinking with videos” via reinforcement learning.arXiv preprint arXiv:2510.23473

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:38:56.086589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:654c63ddbf17f1dd13dae16e53ae4d228783147ba1bb2ca91523bd1b1cad5dcc

Observation c8fd74a3-a013-429c-8099-de10dffb0c1a · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.175051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:0f5b17acd82d9100a1e59f495805ade945297c5caeb0a7e923afb0a49cac9e6b

Observation 87698de1-9b46-4c55-aa86-242a60097b5f · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.133728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:5ad7f4e70d6897b73e3af272d4058019082293d99bb358fe5f5e868d585e7032

Observation 9614726e-eb0b-4fda-bd34-84be4af8c00c · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.859353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:9a91e334f9eeb8404264c81ada38521c8823d577b81091b18836cce11597b954

Observation bcbe7025-38af-495a-a854-55d0d44d8e57 · outbound

This paper cites Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.190292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:54e2691581a46645e65cb2e200df36ab8257766d33c781ca994739fa91b0db90

Observation b6c88fac-f9bf-4ae4-a27f-f742260f2aa3 · outbound

This paper cites 2509.22647 , archivePrefix=.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation 2509.22647 , archivePrefix=

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.179679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:5870d623b80d1e8d3f0e88102cf247ccbac7ff813da0a2fc943686c70dd75180

Observation da80de83-4719-4fac-b2ca-1930ce3bec7b · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.847454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:c98aed858e0a5c8b3a4482d2b919d86e3540fe03fa69c2ae0b0b2f83a56a6ca9

Observation 0f5c71be-3f89-4238-9bb5-5e86c8acfa21 · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.853507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:cd923bf9d17f21b69974f0c16a8001d82a665c68e0b4d6a51d63f4b640caf1d5

Observation 2d7aa7b6-1fb2-4c32-aab1-d57edb0175ce · outbound

This paper cites LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:52.219011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:dbea90d1b2467546622267bb154b29144dea092a2ed1ac277ca91b221685a325

Observation 048032d9-4f39-43d3-b943-7d2ca08caca8 · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.151767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:0de8db867ffde6cb77376fdbb59152f6fa413cc0097186bc6f19018abe25e88e

Observation ff4f1e5b-159b-451e-af00-bc5c58d53205 · outbound

This paper cites arXiv preprint arXiv:2510.01304 , year=.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation arXiv preprint arXiv:2510.01304 , year=

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.106220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:707354cdc06ad7f9e8cfad888e53eebb849ca3b9eed29cb356df2ad1ed713014

Observation 76a0227e-420e-4729-9094-3c4b979242f9 · outbound

This paper cites an unresolved cited work.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:38:56.844594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:ba99d1cc93b62fc57bd0293b5ee2b850af14d77e090cddf04ba35e2ea83804e9

Observation dc24a0bc-2ecd-4ac9-a3cb-69a18fe83560 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.091241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:a133a5865ede8bffb0f2927fc77632677ec714b45e659b60daadc1a3510d407f

Observation 448d1f33-452d-4cf6-a735-0979ea1fcaf1 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.164285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:dbcf1a4812871eaea41762cf05c1a759a4149e1326bfab3fe438ed82bffacb06

Observation c30a972a-050c-4254-b6a1-64422ca970f8 · outbound

This paper cites PyVision: Agentic Vision with Dynamic Tooling.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation PyVision: Agentic Vision with Dynamic Tooling

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.208850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:359eb10ede9f785356f58e1c9f159f67f00c2937c49a700b9d3243b1a738f844

Observation 2559cda7-ba7b-4b36-aee5-c2a2bd123b0d · outbound

This paper cites arXiv preprint arXiv:2503.17736 , year=.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation arXiv preprint arXiv:2503.17736 , year=

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.123950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:dc65de4bbb665972fcaad7531d812c83efc7d54280f2106073c22a615f3a1e0e

Observation 6be654b1-6c48-4f26-8947-ee1dd3361bf8 · outbound

This paper cites Zheng, R.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Zheng, R

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:38:56.838197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:5a2b97422d50746e0d8224078a77ddc8bd8921658faa271fa12320f2fe5a1f29

Observation 3d3cbc2d-c047-4a1a-95b9-48328891898d · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 52

Resolution
malformed identifier
local_arxiv, observed 2026-05-20T19:38:56.143658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:c3fe5dac9c02f6a127f54d3854820ca126ddcb8dc1febeeb22e46bb19f741007

Observation c6949fc3-ed6d-4f23-930b-9d79029ca9a7 · outbound

This paper cites Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:38:56.856749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:9a4bc998eae7176873e046ea825ed3eadcce27f431cb837d0d5573fab62c26de

Pith citing papers

Observation ffb93de8-b762-45ef-97df-7016bb2954df · inbound

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning cites this paper.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:07:23.930391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:1673954ae5624baef98e3421626a3ae2a2b8c4158b5b400276fa9275fcfdb318