Pith. sign in

Paper Citation Record · LEDGER

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 35 inbound Pith citation observations for arXiv:2508.04416.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04416 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:01:18.518663Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:58:12.505711Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9bf4b689-edaa-40a0-b169-434f072761e6 · outbound

This paper cites As illustrated in Fig.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning As illustrated in Fig

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:01:20.272317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:17.925091Z digest=sha256:00ffdf270de2e257c4f18ec855623d4174dd113144e19548bdf3d2f9b09af905

Observation c06b0847-a345-4f76-b3db-c840703af3f8 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:01:17.576049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:01:17.576049Z digest=sha256:ba057e5ca4fea87b7b7798410a603cf4f5b13e223e5b6c1eb61b265e8161c9c1

Observation 2e44997b-56e4-4344-b202-d0a780dbe678 · outbound

This paper cites an unresolved cited work.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:01:19.185951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:18.292055Z digest=sha256:3b48645c2de1cbf780398380f609a9346e4d95043a90ef61723119a8688e6e31

Observation 4e95ea3e-1340-477f-b457-ce21aecebbb9 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:01:17.748181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:01:17.748181Z digest=sha256:825069e83fe94a29884b11e8eb699476e259c0c09ae05116ca912eb41e0311e5

Observation f31ef202-e132-4529-a38c-ec59d6d81794 · outbound

This paper cites an unresolved cited work.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:01:18.805831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:18.518663Z digest=sha256:607db0c9e36a468bd5c1afabfd1a2cde92c254ea902b536d74a8333a6d536d52

Observation 88c03a28-2d84-465a-8aa7-c480a48e061a · outbound

This paper cites an unresolved cited work.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:01:20.091176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:18.008566Z digest=sha256:14bb068b4c1f2a7fd5222ee8397db6f242cd64d51c4fa9a52e07b9ef28531938

Observation bc806f4a-465d-4b52-8987-6886d89fc911 · outbound

This paper cites Do not alter the given ref in any way.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Do not alter the given ref in any way

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:01:19.917996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:18.075481Z digest=sha256:8302b14b4ba5f615fa3910764a7e5d2c5bd448fb22fdafbd531097932269f0d7

Observation 8994c051-27e3-48d7-8da8-aeed4dde6785 · outbound

This paper cites If the object is unique in the scene, a slightly more detailed description than <s object> is sufficient.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning If the object is unique in the scene, a slightly more detailed description than <s object> is sufficient

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:01:19.760711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:18.141793Z digest=sha256:ddd12af982c4424e408f9375f1a7e3873955bf01f1fb06711453ae14a9171ba8

Observation 74af6b7c-0cb2-419a-8ad4-cab413254677 · outbound

This paper cites This complex ref must unambiguously and accurately refer to the exact target object uniquely identified by the uid and its associated pixel mask in the video.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning This complex ref must unambiguously and accurately refer to the exact target object uniquely identified by the uid and its associated pixel mask in the video

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:01:19.567475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:18.230666Z digest=sha256:614b9b1a8e429d99248f7907ad8e723f03325b1d4c7d5dbc42ec69a508889f93

Observation 698c8b69-8512-454f-aaaf-0a5c9c398e2d · outbound

This paper cites Crucially: avoid any simple, direct descrip- tions.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Crucially: avoid any simple, direct descrip- tions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:01:19.353791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:18.285228Z digest=sha256:054a43b55bdd95633d40ad53d1ebec92d6d338158385b3f5c3dd59621f18ff34

Observation 0f5df691-9c33-4820-9a98-26ce02fc306c · outbound

This paper cites an unresolved cited work.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:01:18.990770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:18.406402Z digest=sha256:a0c0834aec5b543487e1555b30cde1d68c025cd5262cd06c73015c791675dce2

Observation b932b4fc-e7c5-45e0-84a1-be76a92f7c6a · outbound

This paper cites We observe that the majority of CLIPScore values fall within the range of [0.8, 1], indicating strong se- mantic similarity.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning We observe that the majority of CLIPScore values fall within the range of [0.8, 1], indicating strong se- mantic similarity

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:01:20.431128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:17.866317Z digest=sha256:e47df26a0f822a16b96aaaf84fcd10a9638db9739c7a54b1a754e584732cdff8

Observation acaddf52-4655-4fb4-9ffb-019b7e26124e · outbound

This paper cites In CVPR, 19108–19118.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning In CVPR, 19108–19118

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:01:20.618275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:01:17.632701Z digest=sha256:99720bc51f37c60d46402f0871cb81dd88e6f051981385627e49ec6916e802fc

Observation b08d946e-b2e5-4dc2-ace7-5b1c264ee31b · outbound

This paper cites Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T00:01:17.810853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:01:17.810853Z digest=sha256:da1d5d5e95a63796c29849cd537454f5a79800fdd399c58bf55c471e3c890699

Observation 0fc7d1e3-6c86-42ca-a4ee-6686936fc6ae · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T00:01:17.525381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:01:17.525381Z digest=sha256:262c5e013ab74a69b81551189ebe81cf947cdbc861bb1e27b7b0a9a4c1f957f2

Pith citing papers

Observation 2c2ebbb9-6cd8-4aca-b5aa-dcd69f3dfad5 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.530952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:51f2c0c9a6d05b5af4bc9282b751bf383108d8f0755a01333a95df5fa6f07ad9

Observation decedb93-683f-46c0-ade1-39bc7d6a6c37 · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:31:32.058146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:2818632d299b01859db91a9125751f78dd401db5b23fa65abdb69c0b40fce1ff

Observation 8e8a7963-12fb-423e-b531-20ddac204a8f · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:26.519397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:25eaddda6ed6cd920c10c67e1034492937943f2addb0cb842dba4a313048e898

Observation c8449ed7-602c-437b-b133-fa54015da88b · inbound

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation cites this paper.

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T14:32:00.400123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:32:00.400123Z digest=sha256:4f87b14f098f2fc18b2649fa01ae4120ccea788098464eb70d04682e5a4a67dd

Observation fdaf31ee-5295-4b34-94be-02fcc51a19b7 · inbound

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning cites this paper.

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:17:51.863189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:17:42.135851Z digest=sha256:7fc9fa745e1a130702f63ed435406922d5e01d808b9a38d87f8cbc74e0896371

Observation 8ef12da1-0c3f-4e51-aeab-d8f38fddd796 · inbound

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking cites this paper.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.308558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:4d3ab09561469b6a50e2fc8ccca3c4448b32799795e8d3d262e6f904b957eae8

Observation 2ab52996-c550-47cc-9227-f4c788ad34e3 · inbound

EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models cites this paper.

EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T22:25:04.080629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:25:04.080629Z digest=sha256:5b723495bf877467ea9d47386a4a4b175ab149729c8106e77e6af48e4ff27889

Observation 590a387c-64df-4182-874a-c6d345964da9 · inbound

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration cites this paper.

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T16:50:50.552104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:50:50.552104Z digest=sha256:d855ca825525d626fb59086832a10ab842f75a931e6391e2bc9dfc7eeabe7437

Observation 3b3a6c40-0dfa-4452-a3ea-117cff48f0b4 · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:51:20.956951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:4af0183f4a8d64b3485eefad911151aec6ca42fc8a0c1541c80807af751f2dcd

Observation 70c70560-a7fa-4a35-906f-cfabf5963d6e · inbound

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping cites this paper.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:35:32.730359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:6f9e2ca23146bf46525afd0b403a5dbfa8579156e8dbf488b9b86b2aa1043acd

Observation 65bb68ac-ab07-47ad-b95d-ca38c9340817 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:03.762912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:d5e40a7589976ce10e41787ee78934d67616a61ead579512072527ee6d9c455a

Observation 4bd62dca-4a09-499b-bc07-6e22a784805f · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:c6d15e7be61ea98090862ae3c6a51a410d59f28220509d0986aeae8d448e3672

Observation 45b638aa-bc33-4a04-9004-bec1c65f6b51 · inbound

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding cites this paper.

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:09.890845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T09:12:17.734060Z digest=sha256:572a9064a4c23439fc225dc2413fbc35c11bc0a50671504fa95715963ae86546

Observation 77e937b3-811d-4d06-8936-40a7ddb33a2b · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:08.768431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:06:27.953376Z digest=sha256:68113c1264d0ec1dcdc217181f49ac945ad8d913f13adcfcdfbb26e56ba211df

Observation b94cf810-3ece-406a-9d8e-4ce5a796e368 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.232728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:49:41.654207Z digest=sha256:ceaac6b36aba24dccd3634f73ac98167c4aa226db47d14a0e0e302d29887fdb3

Observation adb14a59-39b4-47f6-b8bf-bd693545d279 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:36:20.134373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:35:59.553683Z digest=sha256:15688abb334aebee56d958d644a9c898006e705808504bdd0d8d6e748bc505d2

Observation 7bf99d19-731b-4993-bbad-f79538464a1e · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:10:24.077403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T06:08:19.956833Z digest=sha256:ef1e31aa8e1dc27c1844ee4a9a36cb88e0bb5068989ff6002538170a3e499a1a

Observation 76a612d9-b5dd-4dd9-87ac-42cfb4cbbc39 · inbound

ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding cites this paper.

ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:26.487783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:08:13.214803Z digest=sha256:d39ea8212850b5732df5696136bd0c22b40054a83780f741f66c9b8f0c70a51d

Observation dc24a0bc-2ecd-4ac9-a3cb-69a18fe83560 · inbound

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation cites this paper.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:38:56.091241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:a370dbaf01f83dd93b6ee049358e65bb7aa25412ef85bb6841aa90d4c074cfe8

Observation db95fa81-b137-4ef0-98d9-dad96a32170f · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:34:02.608812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:32:12.180233Z digest=sha256:15be09561e00b55b56187e8c1656d3f75ed5bbfbb27630f7a30f71675e10add5

Observation bb64a801-46c1-4a1a-96ba-865f08b1e2d4 · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.487650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:610b4361a2299232a20447a633928c31cb159e6bd4e7c25ddeba3a47b6d56861

Observation a015ea5d-be13-4fe4-88d9-85fbbb21645c · inbound

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning cites this paper.

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:36:10.511284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T06:34:57.483234Z digest=sha256:ee158e490d45be3632e82044f04745c2b1a4f69c428be843cc03428aaf045824

Observation c0ed52b8-4eab-4acb-b59e-023b3e58e831 · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:23.142526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:54:23.077914Z digest=sha256:51ccefcd4abe3eb2f79c6734c408aa4a7b44a7f7b6816dc1bf144d8bbc31b420

Observation 9e4ab436-1ca0-42e6-9997-ffa36ee1fb70 · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.535083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-04T00:41:02.284215Z digest=sha256:6127a13ad49c312d59dbb7a34e0165726de269cf0a5b1bd51bb0aaa9d4487802

Observation c00334ef-161f-491f-be05-6bc80315fd53 · inbound

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding cites this paper.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.123574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:c240b62a3f2d7e16f027cd352e3939a778735a182c5d43f504bbf651e6bab5e7

Observation 4094a22e-e0d7-47df-a43f-674766ef90e5 · inbound

SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence cites this paper.

SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:02:46.722292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:57:38.477065Z digest=sha256:0323b4ecce3365eff9b789ac9737d5d2f68d0edd87073efac9ccaf3a91d55345

Observation 368d2bae-e360-4f63-87f2-626cf7a6c461 · inbound

SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence cites this paper.

SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:57:25.649868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T22:53:35.213874Z digest=sha256:81462719f6bffbe42a1662dac3c3d5c880938e14ba7b1c1c15f16772c806ffed

Observation 26674536-e164-4883-88af-f7ed786e4eae · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.802070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:c993d5991eddaaebd80cee0fd69334d3ca1bbbc19e5aae12e658ae32207aa2ac

Observation 6fb61329-55b6-4ac2-96a2-70ba13836a0f · inbound

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning cites this paper.

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:23.078516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T20:08:29.208550Z digest=sha256:4f73437d2c807c27ee1120fff322724493e3769fb1a720c52c61fb812a874cdd

Observation 944ab3b2-5312-479c-b9e6-dd6e3d6c2e17 · inbound

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding cites this paper.

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:29.038040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:35:07.667016Z digest=sha256:7b50b2425339f5126b9087f61e5d68139d2b748f3f3dc3ffcea99b38571e5943

Observation 683ccd4f-4c63-4a4a-ac9b-5b1c2838605e · inbound

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning cites this paper.

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T06:04:16.637378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:04:16.637378Z digest=sha256:67e6c3cf339d07d343cc4ed830be623f1b86d19d15116c832e5f49d10fe62374

Observation 2d72fa15-8089-477a-bab5-02b19d16776e · inbound

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory cites this paper.

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T06:35:35.951554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T06:35:35.951554Z digest=sha256:aeb48575d8b854b291155d8b3e2a95e68c5700d9c4e2772ddc185115e4859d14

Observation 1a0e3e4e-fe69-4804-b274-743ef21e3497 · inbound

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA cites this paper.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.802892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.802892Z digest=sha256:6dd04f71dfde8da97995cd81379314b0d35d44bfd7f8d22019c77b1283882caa

Observation b943b3f8-bd6f-475f-b17a-530aae2b75e3 · inbound

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent cites this paper.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:30.525201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:30.525201Z digest=sha256:e838fa172f7215e4615beefbabca58e01eb4571cd886fe75e6e6edef0f64d54e

Observation 22419d38-e134-4199-bea9-b3855609dab1 · inbound

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs cites this paper.

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:58:12.505711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:58:12.505711Z digest=sha256:65ed80433921004b3e01796bdf6b22c12e5b8a6e12fe7849f404113b39dcac42