Pith. sign in

Paper Citation Record · LEDGER

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding

As of 10 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2605.26680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26680 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T17:50:00.740770Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact33
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd60909a-2a0e-43f3-8e2c-d3ed7f1c64cf · outbound

This paper cites Qwen3-VL Technical Report.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Qwen3-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T17:53:47.110244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:c22ecb70910e7be67c2e12235de9016d5dac79e43845c7bbf72be6a24fb4d396

Observation a2dca62f-22ee-4a28-bac8-b4589b824905 · outbound

This paper cites Qwen2.5-VL Technical Report.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Qwen2.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.126670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:9e9324e35a9f7a256c19ae919a4595bc3a4acb8894a4101bd68b9d5fddf48b42

Observation e2916a80-bca9-4063-b06e-8d2eedb67c33 · outbound

This paper cites ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.143553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:25ebab488e295a8bc6e37bc2a6ddb818b6928b347c2faab9f47175d2e8844889

Observation 9a1382c3-4616-4011-9027-4ced28821d41 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.054431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:ff4d4e4abe6c18851b643a28c0080f0e290fd244cc221a98fc9aa8381473eeef

Observation e6b53214-d894-4308-ad19-8923e8722b6b · outbound

This paper cites Videozoomer: Reinforcement-learned temporal focusing for long video reasoning.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Videozoomer: Reinforcement-learned temporal focusing for long video reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.149003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:d59b7616936899f274dc93c4e0cb120724373dc1fac5321ee70eb4b88d007d32

Observation b83c7ede-0965-4cb3-bd97-00d5d0cf9956 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.146062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:68f7e4c299addfd19b0f89aacb100c90c49e0d0a6fdce34d821fa5c45853192f

Observation 7f4a8675-abfe-4692-a958-19d83c573a17 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:abc2478aac9ac75b113fbf7b74a8fc6c5fd96f57dbc3732ff2d24baefa279fd6

Observation c78ab457-45da-4bcc-a101-5a727d13b26f · outbound

This paper cites Love- r1: Advancing long video understanding with an adaptive zoom-in mechanism via multi-step reasoning.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Love- r1: Advancing long video understanding with an adaptive zoom-in mechanism via multi-step reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.061147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:0b5ab8c1c58d172a1c5b7ba2efd6a57655bdd12be20d3e43259f6b926f9e60f6

Observation fb944604-a3fd-49d6-a8d4-57514e97112e · outbound

This paper cites Tall: Temporal activity localization via language query.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Tall: Temporal activity localization via language query

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:fb4480db2bef19d8dc88a34b52a38a9a5bcdaceca8ce3f3ea45244e238ab6e40

Observation 394b3f26-cdff-4683-9944-e0aa643365c6 · outbound

This paper cites Gemini 3 Pro Model Card.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Gemini 3 Pro Model Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:f4e439f6870b802a2ad26972af66731ad6d0d64dabc686369109fdfe3347132f

Observation 72b1b12a-568c-471e-b8e7-cbc03326b264 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.063996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:34f3b394827050959cd16a7d1dd79c7e9df12570ad612e7974a30b4c5e107082

Observation 0d0bc326-c0a6-4b3d-8765-bd9697045bbf · outbound

This paper cites Framethinker: Learning to think with long videos via multi-turn frame spotlighting.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Framethinker: Learning to think with long videos via multi-turn frame spotlighting

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.055731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:d9473534d7755332f87944bf9af479163ee018884f207b78195e67e150a69080

Observation c17be9e3-4f0d-4877-8f56-8a2df063a1b4 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.066375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:b43bc2f629eeeae6080e55030b07efce5b2a855949865b47349e77ac46af5088

Observation ef8e1294-98b4-4ddf-bb3f-eeee7d438180 · outbound

This paper cites GPT-4o System Card.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding GPT-4o System Card

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.115262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:bcfa1e91aa7270063cb9498a59aff96ac5f496120285032e150b108c34358d00

Observation 8054942e-8613-4f4b-bc66-225174eb4613 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:91c972cbb47466bba258f2a0d54e3dcf579479a8832cb6045a4d2d1e69e38ec3

Observation e0eea99c-922c-43ce-a610-f15bafc3d2aa · outbound

This paper cites Kimi-VL Technical Report.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Kimi-VL Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.117940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:1b2a7a014f688b100dfdd183016fbfa0b1a8e211793fb98fceba21286709a947

Observation 15a68bf9-b97b-4b4a-8ea3-abfaaf930ed2 · outbound

This paper cites Dense-captioning events in videos.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Dense-captioning events in videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:ce51470d64a701136f74f2e75baf334dcf57e61fcc7f48a6364b1cf8412b73ae

Observation 1951d29b-8c83-4961-a868-06c5ff0cfc84 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.132160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:9dce1c7894504a4356ef5e2843e32cbf844c3f69cbd04a85b8409f408e0f0824

Observation 5f329e02-4e84-42af-94a8-3b9914d312ba · outbound

This paper cites Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.135590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:c56bd6604ff69ea23a7fcf63bfa388c917559b71fe314d60c2fdb3377a98ba49

Observation ec21962f-3b64-4db1-883d-9015dceab48c · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.109477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:bd2abad3eaf7895c060ddab9938012ccafc1c4e0bbf3175170f4f5437f36a91c

Observation a3c115e8-2e73-4b43-85f0-1d82a03280be · outbound

This paper cites VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.138226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:3cf2d5dae0687b5d6ef55ac464e3e57e6c7d677efe3f081e8c9cf1da2f4d023a

Observation 107391d0-6271-4530-93eb-512804475c8d · outbound

This paper cites Video-rts: Rethinking reinforcement learning and test-time scaling for efficient and enhanced video reasoning.arXiv preprint arXiv:2507.06485, 2025.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Video-rts: Rethinking reinforcement learning and test-time scaling for efficient and enhanced video reasoning.arXiv preprint arXiv:2507.06485, 2025

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.124073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:77c77a708d1e6a90dfdb4e8f16e707869f5d16651ea7a13bf1acaf9617e9d8ba

Observation 0c2ea2c2-2550-4613-8d19-f7645ee5f965 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.094370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:bd6ed20c56edeab29ca721b481f0181c11362531e6da49f58cb44aee9e5b1823

Observation ff8b3412-52c2-45bb-a696-002569f432df · outbound

This paper cites Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.087447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:fd7cb1b6862d3e9b1fc571483c1b7508316981a0f22f62efdd4e7200216935bc

Observation 1644c19d-11a6-47f4-8706-7495d3d7a138 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:b55536f3e7ab0abe5be1a80f35befb4d7b51a76952fe2041c2fa6423cce259c0

Observation a7303139-5c10-4923-a01b-e1d20d2ad53c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.103700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:a62f029fecc348e4be93e7a2a8901c9c3f931f09b90ddc6dfcfef28a9b28a055

Observation 0e945c59-2149-4293-8372-a9166d2c4bb0 · outbound

This paper cites Temporal Grounding of Activities using Multimodal Large Language Models.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Temporal Grounding of Activities using Multimodal Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.085857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:27a1524675160063f49ea76005adcc68bd8874d4639c712766415b8cb49c27b0

Observation 4a0e107d-5a71-460e-9250-32ae08b5c9d0 · outbound

This paper cites Adaptive keyframe sampling for long video understanding.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Adaptive keyframe sampling for long video understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:2002ef0cf0465c84a27e9050903ebbd2fdf5a52120da1627a48a88353a1b152f

Observation 9b56a2e3-91fe-42c4-a10f-3c2b1b66a312 · outbound

This paper cites GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.092273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:e0b10bfab75605440a692a4c93ac5b2d0efaeb7ba0bb23ded3251ca944a06dde

Observation 5fbac1b6-d5cb-4ca0-b42d-7dab98014b6b · outbound

This paper cites Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.101742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:c92995ccf555ea2064ba05375601d24091552143fa3c78cf08369dcad2aa30ab

Observation 1d80d592-0544-49b3-8c17-2ea82b75c705 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.131542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:4053aced7a8ae3c9d8c702bf04a677047cb017fe417e7d06d9308f4171759f9a

Observation 92f3acad-7a0e-4a6e-85a3-d4edb8fd8e77 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.115143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:a63b5d650b74adbed79dcb49ee884eeeea21de22662070d227cb8d00d4bd2113

Observation 0391f5a4-5d2d-453f-829f-e644916d4f07 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems, 35:24824–24837, 2022.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems, 35:24824–24837, 2022

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:8ddb575480b6c4ee221deae7ccacc77f6d9a8742e97e0bd2118fcf2c96429b82

Observation 36096b60-118f-48c3-8771-3ba85442b49e · outbound

This paper cites Can I trust your answer? Visually grounded video question answering.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Can I trust your answer? Visually grounded video question answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:86f67c1e2367dc54acec28483498872c9d1bc62dc56c32c59e07957a3b8a33bc

Observation 3d8fc918-8948-46c9-893b-e1e52b821ff8 · outbound

This paper cites Videochat-r1.5: Visual test-time scaling to reinforce multimodal reasoning by iterative perception.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Videochat-r1.5: Visual test-time scaling to reinforce multimodal reasoning by iterative perception

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:fc208b4757072326aab16d6badfee99bceb7fbb037e8384f8e2cc20fc23b6242

Observation f51a4f7d-3e4e-4c37-8c49-fe4634c4e81d · outbound

This paper cites LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.134508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:8002ddb905c8d27c1923100773852f4156bd2433bb924dc39ebc2101127a8d74

Observation 0029ac56-f577-430e-b1a3-1fbcf02a33d1 · outbound

This paper cites Focus: Efficient keyframe selection for long video understanding.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Focus: Efficient keyframe selection for long video understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.137255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:ccd484005fde1f3f9bff3ff51f98db8122cf4ce86d6c221e13458a6ee7aaaf21

Observation 838367aa-8689-4197-9e11-14311c95be71 · outbound

This paper cites Frame-voyager: Learning to query frames for video large language models.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Frame-voyager: Learning to query frames for video large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:30eecee10328c3492eb61eb6bc95845b75e7f439061f53259962649b6c4388f8

Observation f1224666-04e5-4bcb-be0a-0ba2d4c7743c · outbound

This paper cites Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.121262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:eb8b2618fe65932f45e8ed69e1365a237be193b380a7cea45a156113bb729e43

Observation 585e08e9-4107-4b7b-b8a9-e61d522f013e · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:008f2c2e8ec87ddcb2aad68add20fe91e860f222c5787f82fb103b4dbb4dc646

Observation c00334ef-161f-491f-be05-6bc80315fd53 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.123574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:c240b62a3f2d7e16f027cd352e3939a778735a182c5d43f504bbf651e6bab5e7

Observation dbed8afa-7fa0-4bc1-af93-846a9da60741 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.126233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:3b6f74ccce444e56804371165b6a06ce8cdaeaa4af7c9c7ec7edb63c6d7fb17d

Observation aff93130-9265-4d0a-b03c-6ca38dc78df7 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.139711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:27a6dd954f5121c34b5a672ec7fb03937fa139ce6d0d509886a7869abfd031d0

Observation 8f0e0d7e-22b1-4299-a4e3-4b78ef8ff049 · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.140730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:d0ee816119a855b2f81172c86c45f5043e2b603f74fa646669b9ade0e35f64c7

Observation 4c05d73f-2aa9-4308-9574-8d453e49aa30 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.080480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:b2d9c9067a68a47e080fb94584446183c9cf58e602a67d87f07843e8862be10e

Observation 6afaf91e-c776-440d-8e98-62407c20a90d · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.129478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:733b3d41381254f022078b14122c684d1094884469e3634a66edfbeaffd2c332

Observation 4931a756-9708-46db-88ca-18a2999ec714 · outbound

This paper cites an unresolved cited work.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:8125f1d8daf19520eb52d554b4e3cc7b79669e4679dcc4893c60d29a846b2304

Observation 1ec86a77-a6e5-43c9-86e3-c15b10774c4c · outbound

This paper cites Explain why this portion is critical.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Explain why this portion is critical

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:4659dfa1c1b180c34b6d21e72b88630bd975acbfe0bfdabe379158a019304f78

Observation 4262fb4f-cbfe-4cb2-957e-1f27ade60ac6 · outbound

This paper cites an unresolved cited work.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:feaaeb54161494dabb31b71c669344806380f7fbd878107b904aba0ad9874ffa

Observation b5c64ec9-779e-47a3-9855-b09c6bfb9178 · outbound

This paper cites zoom_in_cot.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding zoom_in_cot

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:bd9c585a2f905ee37869e688e95a02692eb1d7c2ce073d9c4db0322b2a86023b

Observation 0e959838-c1dd-4c85-9232-371c320b6787 · outbound

This paper cites Using the visual information in the segment, reason step-by-step to reach the answer.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Using the visual information in the segment, reason step-by-step to reach the answer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:e513acf3caa181007ee85a9443b439a82c57e91af127f9538ec98d3f6ba625bb

Observation 26696af3-b55b-4b4c-9e8a-e38be90fb1cf · outbound

This paper cites answer_cot.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding answer_cot

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:33797b3d19aa78dd4c6787bde97e5a621ac4a88959fe4fbe2939ca6f55b0e159

Observation abf719dc-f887-4c9f-abdb-568734f7f3b6 · outbound

This paper cites an unresolved cited work.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:10a3a810c2dce11f3c5e3c16f55bcca68400d25cf9ce76ae56b97d05b6485e36

Observation e5c16279-b35f-4827-8bc8-6787148f2b7a · outbound

This paper cites an unresolved cited work.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:f86db7f14e80786930b583b300a30a359481d16605d219cbaa89e10bb256789c

Observation e0db0caf-164f-44a7-9846-382b920b2d3f · outbound

This paper cites Keep the reasoning concise and non-repetitive.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Keep the reasoning concise and non-repetitive

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:5bb4f8f22a8ce1954e4d6772afda8e9d1e69f731333eb2d4d024fbc5ebcddc2f

Observation 0219894b-ceeb-4d60-8b6c-e4be36b5bfef · outbound

This paper cites an unresolved cited work.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:d1df4b5b2e190d294003cb7c1dc99566acc399f9674e3b59160fe070c9b39919

Observation 84ab49aa-8ad4-4259-8c7b-1a2b2a60c609 · outbound

This paper cites Max retrieval / injection.

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Max retrieval / injection

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-29T17:50:00.740770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T17:50:00.740770Z digest=sha256:7e4885d799d8160deb05157024f72d1dcc1f753e3189fa745ab83878019ad50c

Pith citing papers

No inbound Pith citation observations are available.