Pith. sign in

Paper Citation Record · LEDGER

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

As of 22 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 11 inbound Pith citation observations for arXiv:2512.05774.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.05774 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:22:33.231784Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:35:26.623428Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T15:28:33.947770Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved84
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab07ad85-2248-4752-af36-3559ed0356f4 · outbound

This paper cites Psychology Press,.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Psychology Press,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:23.527845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:23.527845Z digest=sha256:469793d1f0e19f3fd72a1aa18e834082ea8985414ea65e354d02cdf1ed62281e

Observation 0979b839-73ae-4a52-8a19-cc714b8676c9 · outbound

This paper cites Temporal chain of thought: Long-video understanding by thinking in frames, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Temporal chain of thought: Long-video understanding by thinking in frames, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:23.597949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:23.597949Z digest=sha256:950f193b57ddc9e83fa70548ac64330c855c72bec96b3605347c2f3a932b565c

Observation 4f03292c-2032-4565-b079-2752b514e104 · outbound

This paper cites Qwen2.5-VL Technical Report.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:23.657740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:23.657740Z digest=sha256:96bce7f397a96ac9b9e7cc37776b41ece381e6cd4ad0c21467da84016e4b6aef

Observation d99e710b-87bc-46e2-b2bb-08b3c58d2bcd · outbound

This paper cites Active perception.Proceedings of the IEEE, 76(8):966–1005, 1988.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Active perception.Proceedings of the IEEE, 76(8):966–1005, 1988

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:23.725387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:23.725387Z digest=sha256:22cd0455d46c6550fc3062eba1c02dae30de12f612963c2e7c1e7cb9adf944ad

Observation 8a5e3cfe-1e3b-4420-b5a3-e6927aeb36be · outbound

This paper cites an unresolved cited work.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:23.836886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:23.836886Z digest=sha256:fd69813910016f0d07893018d5921c343de0a83aa39f587c283eaea3e45273c1

Observation 8550eae5-9bca-4c99-8114-ea05c2beb871 · outbound

This paper cites Revisiting the “Video” in Video-Language Understanding.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Revisiting the “Video” in Video-Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:23.954077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:23.954077Z digest=sha256:e138739479e170a956d8a9a66b2213e824a98975af19a1fa45e50f2c2ad63245

Observation e5b78b50-cced-4924-b0c2-b7cbc42108e1 · outbound

This paper cites Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:24.071228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:24.071228Z digest=sha256:4a0c4835a4b1a2a68e35421e1580409ec20f9fb420083a8569d164f8dbffc6e2

Observation ee6c8446-6848-47c6-9f25-d3ddc059263b · outbound

This paper cites Lvagent: Long video under- standing by multi-round dynamical collaboration of mllm agents.arXiv preprint arXiv:2503.10200, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Lvagent: Long video under- standing by multi-round dynamical collaboration of mllm agents.arXiv preprint arXiv:2503.10200, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:24.284882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:24.284882Z digest=sha256:6841594d145e85a34cc18c6c688233cf55f0ff86a22f04902d5ae373cd7ebd68

Observation deaf634e-0f3c-48a4-949b-3bffd6c46c17 · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos,.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Longvila: Scaling long-context visual language models for long videos,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:24.500550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:24.500550Z digest=sha256:6aad20cd5e981e5728ee20a3d80febf333c62bdaf8dddcc6114e4d3d7dc54fd1

Observation 2ab6de7d-87f0-45bd-942a-646db4b28c65 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capabil- ity in llms via reinforcement learning, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Deepseek-r1: Incentivizing reasoning capabil- ity in llms via reinforcement learning, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:24.635886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:24.635886Z digest=sha256:eae0eb30860444ead20a083cef9f96141bd71dc5d0023ac8f52d1868bd5e55bf

Observation bb4b9c68-1c53-4556-b642-dc18befe91d5 · outbound

This paper cites See what you need: Query-aware visual intelligence through reasoning-perception loops, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding See what you need: Query-aware visual intelligence through reasoning-perception loops, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:24.740213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:24.740213Z digest=sha256:147c2358b59e2fec025375c4803bc047757c4fd723f586633db6dadd5c74e1f0

Observation a8c5a29f-c674-4f3f-9044-c2efc0b0b7d9 · outbound

This paper cites Agentic keyframe search for video question answering, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Agentic keyframe search for video question answering, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:24.903019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:24.903019Z digest=sha256:6b28e2c43b2836da8f7333b01ff0c500884f90811739f4196819413cd228fd00

Observation 1a9f88bc-a7c9-455e-9640-acddcc9f2a6c · outbound

This paper cites Videoagent: A memory-augmented multi- modal agent for video understanding, 2024.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Videoagent: A memory-augmented multi- modal agent for video understanding, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:25.067335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:25.067335Z digest=sha256:70a6b9bf08279f6c58ea9902f5c69672185231cd2afed92eaea90604e76efe16

Observation 39d9c632-a053-4ba3-ae37-12a08dfcbd05 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:25.203479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:25.203479Z digest=sha256:a2eb75df649ddcef4d1f4e229250eb519c91c7240aa93939e3725d6a1d82452e

Observation 5f3990bf-f693-454b-9f2b-671870d8e6a7 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2024.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:25.296291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:25.296291Z digest=sha256:c7705feeb44bb01078b82e8fb25d972bedf996fddef6a2c76b95b32ecc4759a1

Observation c0038f4c-31d6-4a46-993d-fff38859e842 · outbound

This paper cites Love-r1: Advancing long video understanding with an adaptive zoom-in mechanism via multi- step reasoning, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Love-r1: Advancing long video understanding with an adaptive zoom-in mechanism via multi- step reasoning, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:25.458253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:25.458253Z digest=sha256:41c4f8adf2c4f6ac65e10d4634f5db93507dfb3a4a9d2f286920c6a5d2e9291c

Observation 766d3b5e-c281-4990-b36b-46392ae6fc2e · outbound

This paper cites Framemind: Frame-interleaved video reasoning via reinforcement learning, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Framemind: Frame-interleaved video reasoning via reinforcement learning, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:25.562342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:25.562342Z digest=sha256:47642006908516abff6afae1001041cde8b8494680345ae125adaa9d7b07f6a4

Observation 5f2603e6-7635-4503-9759-a25615faac7a · outbound

This paper cites Framethinker: Learning to think with long videos via multi-turn frame spotlighting, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Framethinker: Learning to think with long videos via multi-turn frame spotlighting, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:25.668330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:25.668330Z digest=sha256:b2ee2736995c8defc81d00be3ed4097950a77e8fc6c00d97be785e4da0417cd1

Observation 382a94c5-c17e-4d1b-8216-c14d2a5dc0ae · outbound

This paper cites Adaptive video understanding agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning, 2024.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Adaptive video understanding agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:25.818730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:25.818730Z digest=sha256:9da8f428bdad506488ad3a3f448f941eb894e120b57adf8a4e92ac55d5519dbd

Observation 890b45a3-048d-4226-b78d-f82a5079b549 · outbound

This paper cites Language repository for long video understanding.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Language repository for long video understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:25.912949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:25.912949Z digest=sha256:99497aed50ac1abba8fca0d28ddc3a84d557d114a4040befaa6a1fb2aad6bc51

Observation a609a22e-1532-431d-874e-840a93ff94d5 · outbound

This paper cites Videomultiagents: A multi-agent framework for video question answering, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Videomultiagents: A multi-agent framework for video question answering, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:25.998750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:25.998750Z digest=sha256:dcce786dc0d381b4e378f64850d119b22cd7d62c445cef67bb6a8121e54e77a6

Observation 262ee400-c61d-4a27-bdf0-8b6e971c0587 · outbound

This paper cites Aria: An open multimodal native mixture-of- experts model, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Aria: An open multimodal native mixture-of- experts model, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:26.168260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:26.168260Z digest=sha256:69493f032778380e53134b74d71e8ee3f22900039f05461dc61f52ab94494619

Observation 0443557b-3779-4956-8287-b4338890baad · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding BLIP: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:26.390300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:26.390300Z digest=sha256:c65f4c7f96175fd4b992b0efcd308b8f806e4b666b45fb1b874fea1811d84e7b

Observation f5e69584-5200-4fa3-af81-4031da93e4d1 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:26.553542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:26.553542Z digest=sha256:43f5e302612b79d2a12ffaa4c85edad3275b15fdd6f4763c61a637e2c6e7a4a1

Observation 001a86da-c99c-44df-80e1-33be32c08153 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:26.641945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:26.641945Z digest=sha256:7531a0be9cd48d181c62ef53f0cb3012df36b5114dfa7b72f728c60750e6cab6

Observation 5f810102-3f8f-484f-9f81-e56ce2ae8592 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:26.762749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:26.762749Z digest=sha256:81dba92066d44d2f34265ee7f76829ebc99847982dcb8642da657cbd710dbb77

Observation 26962856-9399-4286-ac5d-1ca9184f4e80 · outbound

This paper cites Videomind: A chain-of-lora agent for long video reasoning.arXiv preprint arXiv:2503.13444, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Videomind: A chain-of-lora agent for long video reasoning.arXiv preprint arXiv:2503.13444, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:26.967561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:26.967561Z digest=sha256:7352e29ff01f019dc57039fb30ef644f0fa1ec4ae90e14e3969072722c6b5d38

Observation 7c8a033b-4daa-4509-a97c-29ee15866f44 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension, 2024.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Video-rag: Visually-aligned retrieval-augmented long video comprehension, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.084218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.084218Z digest=sha256:d98f48e5fe4eefe835d22c4aec5556814ca6a6db9646373c6d140aae1a5f6691

Observation 1a637fc8-5cf3-487b-a4c9-cf1a96e1e427 · outbound

This paper cites Video active perception: Efficient inference-time long-form video understanding with vision-language models.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Video active perception: Efficient inference-time long-form video understanding with vision-language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.138082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.138082Z digest=sha256:79437567b67da9d33b5b96a83169504a887386736367f763abe70b998fa675fe

Observation 2cee4a0c-a239-4bad-b23a-d498d99d529c · outbound

This paper cites Drvideo: Document retrieval based long video understanding, 2024.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Drvideo: Document retrieval based long video understanding, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.224491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.224491Z digest=sha256:8d480159a481bfe6a6e38d5b5d28c990d9facb84656b0a5c8753bc5808649d15

Observation 05599ce7-37fc-452c-abec-68d460b0c9e9 · outbound

This paper cites Caviar: Critic- augmented video agentic reasoning, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Caviar: Critic- augmented video agentic reasoning, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.307201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.307201Z digest=sha256:2c6310577b857e64736e9a7f62444a266a1e746aed160f8faaba1410e8ebf67b

Observation 49631c8e-88ed-4461-a3c3-7fc26d48f1c2 · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Morevqa: Exploring modular reasoning models for video question answering, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.377043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.377043Z digest=sha256:3a744f280ca9d51047b64679b1bdb76e74e31a99748a4d93ec4ae1f042ef85cd

Observation a6f4cec4-ea96-4fd3-9f37-7fdb3c242dbd · outbound

This paper cites Minerva: Evaluating complex video reasoning, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Minerva: Evaluating complex video reasoning, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.424656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.424656Z digest=sha256:c13378e0eb3075b3f39865ed08d34cb7b101dda128a389da191fa042ccabe3b1

Observation 2ff1b48f-623d-4978-845b-eb7449aa3c95 · outbound

This paper cites Gpt-4o system card, 2024.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Gpt-4o system card, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.486611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.486611Z digest=sha256:c8366aaee5b2841512066bb2178c35ec2deefac62b274fbacd0017320f440484

Observation 132af810-d857-4e6f-a1eb-be56ab1e266f · outbound

This paper cites Gpt-5 system card.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Gpt-5 system card

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.558715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.558715Z digest=sha256:003ab164c4e35f5c7667038cd722337c6bc94486c07c0a15d81c844dc28c2ed0

Observation 99181d7a-f852-4af0-8980-2471e20ddb0a · outbound

This paper cites Introducing gpt-4.1 in the api.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Introducing gpt-4.1 in the api

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.648141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.648141Z digest=sha256:07197ba7faa5e8eb82e72dcdc41a2fc7a0aab7cf77b4fcdf07593c9776ed7335

Observation 911c8666-d36a-4a65-b883-971d8ce3b972 · outbound

This paper cites Openai o3 and o4-mini system card.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Openai o3 and o4-mini system card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.738975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.738975Z digest=sha256:2612639f8343e034da068e9064eae3066e9d00ebb51e4d173499c5b67d003bbc

Observation 2dd44134-331c-4ae9-af8e-7cf9453fe103 · outbound

This paper cites Conan: Progressive learning to reason like a detective over multi-scale visual evidence, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Conan: Progressive learning to reason like a detective over multi-scale visual evidence, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:27.892470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:27.892470Z digest=sha256:3f3e31d0dcb49fb926f44c268567e234b8f089bd6bd1092aacfb3167935e1900

Observation e77f76e3-50af-4a2a-8762-2e1d9b573389 · outbound

This paper cites mapreduce.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding mapreduce

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:28.061767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:28.061767Z digest=sha256:58e022c1473cab8c2264822d33e6e8dbe293882eebf2502311308d6a3a854fcc

Observation af23dfd3-3c97-4ace-a30a-6c7158542e7e · outbound

This paper cites an unresolved cited work.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:28.145393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:28.145393Z digest=sha256:65873641de86f56e2b8d865e6073720dc299573def21636d6fb1ce6e3c45e669

Observation 619ddfdf-c83b-45bd-a8c6-fc89ff2b5cdf · outbound

This paper cites Understanding long videos in one multimodal language model pass.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Understanding long videos in one multimodal language model pass

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:28.284376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:28.284376Z digest=sha256:c17503096b1c79ea8383eb9bb4d51b681c2bbaa049f5788a899008c3a9acb657

Observation 12dd4acb-da3c-436d-b321-44e1a56ecbb5 · outbound

This paper cites Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Jongwoo Park, Kanchana Ranasinghe, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin, Le Xue, Manli Shu, Jongwoo Park, Kanchana Ranasinghe, Silvio Savarese, Ran Xu, Caiming Xiong, and Juan Carlos Niebles

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:28.409448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:28.409448Z digest=sha256:078d61d8d8b8d44cc9c8b6021abfa1c0177a9ff7009a6fa906c336cae41af2c7

Observation 087a47bf-97b4-4a42-aae8-c0b9f794f8db · outbound

This paper cites an unresolved cited work.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:28.572270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:28.572270Z digest=sha256:18f44d0bbf97d8ca729360aa60e7b165224feaa293ecf2fe06fa09554da397cc

Observation e21123ba-21f9-4d4f-a740-e8356176de58 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:28.724700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:28.724700Z digest=sha256:da86b4294adc882100f5e06a2c90f35f27f90f1af9706d9bffb3bdd48c2fb51e

Observation 490429ce-0cf9-43c4-be67-7f09c0650180 · outbound

This paper cites Vgent: Graph-based retrieval-reasoning-augmented generation for long video understanding, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Vgent: Graph-based retrieval-reasoning-augmented generation for long video understanding, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:28.882214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:28.882214Z digest=sha256:bba15048c3aff3c521e1b68f0c9492da16e2d2b28c766938840c0c3225f7870a

Observation f7ddc866-bff3-4019-abdb-451f732f8fd9 · outbound

This paper cites Enhanc- ing video-llm reasoning via agent-of-thoughts distillation,.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Enhanc- ing video-llm reasoning via agent-of-thoughts distillation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:28.991847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:28.991847Z digest=sha256:9409c2388a4fbe7fc4ab8c6bc23616fb0456a4ebea2361036cd161fd92bf9c4d

Observation b0b2081b-5f53-4d21-8b42-08d2cc83303b · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:29.150095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:29.150095Z digest=sha256:6a455f093524c27091c2d751c6b25aef1a6e361fe38d98d56a3aa4f4d17046c0

Observation 1e234f49-5701-4ddd-9fb5-01a86e983313 · outbound

This paper cites Scene exploration by vision-language models,.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Scene exploration by vision-language models,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:29.221334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:29.221334Z digest=sha256:4824d23fa4131ce555cd1a1d425a27754305663e58d864b04505e2fe3bd4b2b5

Observation fa5167e0-e596-42fc-a9fd-4e1682a1dbce · outbound

This paper cites Adaptive keyframe sampling for long video understanding, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Adaptive keyframe sampling for long video understanding, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:29.372562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:29.372562Z digest=sha256:3d28caeffff910db19ced6ea8214cad9f64a2c96c8a0d744fe78df889fe75b3f

Observation 2d5d6bbf-c1e9-4f25-9bfb-b5f651670046 · outbound

This paper cites Moss-chatv: Reinforcement learning with process reasoning reward for video temporal reasoning, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Moss-chatv: Reinforcement learning with process reasoning reward for video temporal reasoning, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:29.477922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:29.477922Z digest=sha256:5bffead4ce2e3b556d1d71d51ee113a9e3e5374fb05337fa85116b5f1ffe7c0f

Observation 87eb7486-4da8-47a0-b224-99d022bbcb82 · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:29.633102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:29.633102Z digest=sha256:a90e59c0642fde7389ce01881686cb9bde95d03a1c246add0e6e9ea8c3a59f7a

Observation 18b20e16-d5a0-4109-86bd-9bc1c76dc5e9 · outbound

This paper cites Qwen3-vl: A general vision-language model.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Qwen3-vl: A general vision-language model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:29.838794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:29.838794Z digest=sha256:68d6310909b4841150676553d9af6144cedc2fad59cabdc90a1154b4a6c17b75

Observation e776a9fe-6365-4d5a-9a4f-f9700f38a5e4 · outbound

This paper cites Seed1.5-vl technical report, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Seed1.5-vl technical report, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:30.027703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:30.027703Z digest=sha256:5d982168bc52bdbba5e4a8992c3267ab15c4373d76c417349e9cd3dd4cd3991d

Observation 38cae0a9-08a6-4b36-b17d-45cc3b815ccb · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models, 2024.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Tarsier: Recipes for training and evaluating large video description models, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:30.241779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:30.241779Z digest=sha256:539689c2c1f612ed238c13d2fbe1e957a74ea8401309f52cd37da0c2b79e1ace

Observation f36d90b2-6847-4ac8-beb3-1267664dc549 · outbound

This paper cites Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:30.363995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:30.363995Z digest=sha256:aef3f5b641117d284b5b4eba833d1faf239c350e916cd6f9fe9a0716d100b837

Observation d1f8c0b5-cc6a-4f5d-af24-f97a3fe5713f · outbound

This paper cites Vamos: Versatile action models for video understanding, 2023.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Vamos: Versatile action models for video understanding, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:30.550128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:30.550128Z digest=sha256:120db8842062cea54deaf13253e6899c26366cd65c1f55492b1fccc55036b5d3

Observation 6265b3ef-4462-417f-9743-97636983d1a9 · outbound

This paper cites Alvarez, Lei Zhang, and Zhiding Yu.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Alvarez, Lei Zhang, and Zhiding Yu

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:30.715885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:30.715885Z digest=sha256:c2c80aa4a3f6dcfd9bdc205e208c64075ed40d06e917d991c4de7d7681490439

Observation ff1aabae-64f4-4165-a500-5d40bcea4c2f · outbound

This paper cites thinking with videos.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding thinking with videos

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:30.823378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:30.823378Z digest=sha256:c4db22f92a61a3a2bd425bef30b954ef2ea42ecef6401d18262d997562383d6c

Observation 71804053-5966-4169-a2a3-7532df86e9e2 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:30.883136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:30.883136Z digest=sha256:eb7d7f265bb202faeb521edc3d813e084888e3179b60db776786b634980d1347

Observation b39fdc55-995d-479d-8887-7be32c9ce189 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent, 2024.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Videoagent: Long-form video understanding with large language model as agent, 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.006785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.006785Z digest=sha256:130b263485f8e30ecbc5abb2ad53c41ae2ff1c6ef9950fd7e01f8ea85cf4e1e7

Observation cf36da34-5af5-4ead-bb12-041df948da16 · outbound

This paper cites Adaretake: Adaptive redundancy reduction to perceive longer for video-language understanding, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Adaretake: Adaptive redundancy reduction to perceive longer for video-language understanding, 2025

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.129405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.129405Z digest=sha256:31530a61c272c935551ccdca93ef98938eec8b2fef25e42ba594a97b93e3c620

Observation 092753b1-aab3-4aee-806c-3130dd836dc9 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.236016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.236016Z digest=sha256:0f67d85dec02f105aef6ac98e6145064e1296c92e28777f3df2ad4ecc86ca2bd

Observation ecc2491c-45b2-4cab-bf74-e1e0057f03fa · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.365284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.365284Z digest=sha256:1bd65a80a9f6478f30d049070bc35cc93b0a2aabeb0cca05da240fef4e00b8d3

Observation b369d48e-fe31-46df-bd24-a9df4d21cbb2 · outbound

This paper cites Videochat-a1: Thinking with long videos by chain-of-shot reasoning, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Videochat-a1: Thinking with long videos by chain-of-shot reasoning, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.498639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.498639Z digest=sha256:7c50edd8d02b959f1405f6e6eb801af556297b3ca87d23a4d5a64fecd5d68a1e

Observation 89b3596d-ed6b-4f1a-9adb-0972a57588e4 · outbound

This paper cites Video-RTS: Re- thinking reinforcement learning and test-time scaling for ef- ficient and enhanced video reasoning.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Video-RTS: Re- thinking reinforcement learning and test-time scaling for ef- ficient and enhanced video reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.638487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.638487Z digest=sha256:b9c758418dc29c4ee38af05d7dfcbe8bdf9d9484ccb50abee8538d83d6711efe

Observation fce5a10f-891c-46f8-8bad-75a845ac952b · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Informa- tion Processing Systems, 37:28828–28857, 2024.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Informa- tion Processing Systems, 37:28828–28857, 2024

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.768007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.768007Z digest=sha256:4e99d0657147ee31b67e1594452fd77e1f898106d10896ae211ff8ee6685ebaf

Observation 149c61bb-fc16-4a96-9749-ad062f8d1013 · outbound

This paper cites Vision in action: Learning active perception from human demonstrations, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Vision in action: Learning active perception from human demonstrations, 2025

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.863417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.863417Z digest=sha256:b1aa38736a210aff90332291f1f60086eaa2e48707fe3c30dbe92eccbb503c40

Observation c8f813f9-4627-43ba-a7b0-2668afa96c4c · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.984503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.984503Z digest=sha256:41a8a9d1fcbc787f5c7c9e80402d13ae2c2da7d6cca4357fdf5c81f409153e9a

Observation 86733491-3b8a-4b3c-809b-1c36a5df3da9 · outbound

This paper cites Ava: To- wards agentic video analytics with vision language models,.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Ava: To- wards agentic video analytics with vision language models,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.088897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.088897Z digest=sha256:0e41600e8a3c99c5fa6890f9f920ef02690bcd64d2cc7400a313f1e1d4a75e86

Observation 2c3b0acc-dca0-43c3-8a5e-8d78aea8b37e · outbound

This paper cites Generative Frame Sampler for Long Video Understanding.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Generative Frame Sampler for Long Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.132710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.132710Z digest=sha256:07c151676d8ba7c8241f5885d1701672cb816e1f680d1aac9d2ada8d77b483dc

Observation 38331be7-93c8-4c26-bff3-c5e9b0faef54 · outbound

This paper cites Re-thinking temporal search for long- form video understanding, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Re-thinking temporal search for long- form video understanding, 2025

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.183016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.183016Z digest=sha256:b27e6764b7da416602570821fb7908fddf15f95782f2ecbc8483429c6e5d3354

Observation f8569c7d-3f0f-4825-8be2-47c14e8f7893 · outbound

This paper cites End-to-end learning of action detection from frame glimpses in videos.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding End-to-end learning of action detection from frame glimpses in videos

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.259304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.259304Z digest=sha256:f4228672c4cced376d4b66e2b10beeda7e8902a5018b10f7a4d2d74103a438c6

Observation 2aac6a72-52bf-480a-951c-0ed00dfe12b4 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Self-chained image-language model for video localization and question answering

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.320679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.320679Z digest=sha256:d2980e878b9d188138b600a050880a78d7735d43bce14c9da985a452485e5b1f

Observation 0414c75e-1cd8-4249-b2e6-2879148ed658 · outbound

This paper cites Videoex- plorer: Think with videos for agentic long-video understand- ing, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Videoex- plorer: Think with videos for agentic long-video understand- ing, 2025

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.363429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.363429Z digest=sha256:ec7b9aff762197af3e2eb6957b3dd39562848895418559cdbafe2be9f6f715ee

Observation 0a4a2750-876e-4854-b316-4c81608b8cc8 · outbound

This paper cites A simple llm framework for long-range video question-answering, 2023.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding A simple llm framework for long-range video question-answering, 2023

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.419785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.419785Z digest=sha256:0abe5b41c7731614db8be7fe60c7d0fe7d62c8ccf5b58937d2c94aee433525c7

Observation 980d2234-069b-411b-86f1-5cef7ff19c02 · outbound

This paper cites Silvr: A simple language-based video rea- soning framework, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Silvr: A simple language-based video rea- soning framework, 2025

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.473110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.473110Z digest=sha256:24cb06b4bdb59531848e9f424e4ad0d5b7fe48ec5e5430b20b97e57d1300ad40

Observation ab704eb6-7243-428e-b602-cab8b96c728c · outbound

This paper cites Thinking with videos: Multimodal tool- augmented reinforcement learning for long video reasoning,.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Thinking with videos: Multimodal tool- augmented reinforcement learning for long video reasoning,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.545083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.545083Z digest=sha256:494e1fed9d58a81ebc456daf77241b71c21ce899e81f27b2201760fd7d33cf68

Observation 82e609f1-d849-4451-89ca-dde64e8fbb9e · outbound

This paper cites Long Context Transfer from Language to Vision.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Long Context Transfer from Language to Vision

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.646227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.646227Z digest=sha256:29b4c67593163c667325a2ab822d5921dea0274ca303ebe82c4790bfeebba388

Observation 769280f2-ed68-4882-8a21-12a21bf414ea · outbound

This paper cites Deep video discovery: Agentic search with tool use for long-form video understanding.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Deep video discovery: Agentic search with tool use for long-form video understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.730480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.730480Z digest=sha256:19e617e434705043b08a0a9c1fc2d20265a12425e90d9d65f4bd0322e72a1f09

Observation 67934772-bf1f-4c7c-b7cf-dd600c3842b5 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.845533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.845533Z digest=sha256:dd9e159a9afeb6df5b0e9fbba31d99f7e0ac44fdf5294201d272b5e8cf07dd74

Observation f027c173-a06d-429c-a644-f5f35ad74c61 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.900214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.900214Z digest=sha256:70bb0c16bf8859f2206ae15325544d3753277dcffbd11c2b8fe7303efef9c806

Observation 200105ba-b457-4315-ac7d-5d04921ef0d5 · outbound

This paper cites Reagent-v: A reward-driven multi-agent framework for video understand- ing, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Reagent-v: A reward-driven multi-agent framework for video understand- ing, 2025

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:32.953045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:32.953045Z digest=sha256:cf5ba01e200c5925a0f0ded0ebad1a56c889127e690703ff70275cd90fa79bb9

Observation 73c19dd3-b2f3-42dc-a70c-1abfcae64bf0 · outbound

This paper cites Active-o3: Empowering multimodal large language models with active perception via grpo, 2025.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Active-o3: Empowering multimodal large language models with active perception via grpo, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:33.016012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:33.016012Z digest=sha256:2ab34f93b9d14459630fe4910a2be0a84328fb18edee9da712b7d55d1bfc2cc0

Observation c8804358-9b18-4108-8265-4f6569830d65 · outbound

This paper cites A.i.r.: Enabling adaptive, iterative, and reasoning-based frame selection for video question answering,.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding A.i.r.: Enabling adaptive, iterative, and reasoning-based frame selection for video question answering,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:33.147794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:33.147794Z digest=sha256:0b3993a22fe7206ff4418a8e22aaae138e4ddc990b72ee9a311af08cd321adb1

Observation 17bc61a3-56c6-4fe4-ad3b-890354a3163c · outbound

This paper cites usually range from 4 to 5 feet in length.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding usually range from 4 to 5 feet in length

Reference 85

Resolution
malformed identifier
no resolver link, observed 2026-08-03T18:22:33.231784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:33.231784Z digest=sha256:b55494a1b31f25dfe2c2983eb4d54bfcb07c87ad0955d7fb885f8f3e0cf19bf5

Pith citing papers

Observation 1912ee73-8485-4f2d-9144-2047a4d59586 · inbound

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding cites this paper.

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-06-05T02:15:31.456438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T04:37:46.090777Z digest=sha256:ae3d8b156543724ad8b712c7f303f217566561f3b2f77a7dfd0753d49406a080

Observation 86164427-32ac-4232-b257-fb932e53e883 · inbound

Personal Visual Context Learning in Large Multimodal Models cites this paper.

Personal Visual Context Learning in Large Multimodal Models Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-06-05T02:15:31.456438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:42:15.402131Z digest=sha256:b7fc3943c0363ae4b1abe1fc2997eea78f02eddf45f3b41ca02cb97db0459fbc

Observation 734bca2d-c3c2-43d1-bad0-151018abff92 · inbound

Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval cites this paper.

Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-05T02:15:31.456438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T04:18:58.470652Z digest=sha256:74f5125a54d646b3305ecc2538614d0393f318410c73db35b554420907ef4030

Observation f96c7881-8649-401f-973e-f129e244a4c6 · inbound

Benchmarking Visual State Tracking in Multimodal Video Understanding cites this paper.

Benchmarking Visual State Tracking in Multimodal Video Understanding Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:28.019866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T10:43:06.811228Z digest=sha256:a0883a6c4f7632bff105707f1a03d1cfe9cf5319fbe55fd3a968527745bf22ce

Observation b81f57c6-b73c-4239-8309-38574c2ad9a9 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 230

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T17:27:15.077171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:328df44bcc03a93969e4bd75792d6b59bcc069e1f7f094809ed5adab4019828a

Observation dfab930b-09ce-410f-b217-5ca8bc44e275 · inbound

Rethinking RAG in Long Videos: What to Retrieve and How to Use It? cites this paper.

Rethinking RAG in Long Videos: What to Retrieve and How to Use It? Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T15:28:33.949814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T06:30:33.428489Z digest=sha256:cacd76c905ea78e0710841252973563cd3448797c687eed50da087512f272566

Observation b2160f28-107f-4777-aa54-4fcf8fa980b3 · inbound

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA cites this paper.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.205921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.205921Z digest=sha256:2d3f1a195b55d80a532d4e9373ea860d2e6936474094a65c9ec816e34c7b442d

Observation 46a95a9a-82d5-4324-8595-74302ef3e2b5 · inbound

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding cites this paper.

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T02:59:25.213497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:59:25.213497Z digest=sha256:96922ce1344091089eb3b7ac021b83fc8a658162753f9945434df84604042b45

Observation 0dbe0cc8-4ad6-44f9-a539-8402412a25ad · inbound

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding cites this paper.

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T01:46:07.582823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:46:07.582823Z digest=sha256:040bb47105f25e480ac7b61c6c7b23c5f8f8e6ced1a3ed66c3fccbfc7a2c9aaf

Observation 73d96fa0-a39a-4bc8-b68c-deef7daed697 · inbound

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning cites this paper.

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:43:38.391136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:43:38.391136Z digest=sha256:fd210d5332d2d3405b7a9f275672ef94e41c7ba69aa134fc590b77e6891c4ced

Observation 7f4451c8-e1d5-4fe0-8868-2e4a348865b4 · inbound

REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering cites this paper.

REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:26.623428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:35:26.623428Z digest=sha256:89a2f8768e3e99142beef96c11e691259fce7a749f814717474c2be136e07f41