Pith. sign in

Paper Citation Record · LEDGER

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models

As of 4 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2605.01662.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.01662 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T19:17:20.745425Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:16:51.254681Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact21
  • verified fuzzy47
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce2e797f-d450-4d35-94f5-cb930986bfd6 · outbound

This paper cites GPT-4 Technical Report.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-09T05:55:31.666045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:235171db0c072eee3b7f5532896a3fb262615228bd13dbd9f77f834138184a0b

Observation cf96cc15-4114-47fa-a1f5-b32fa019df95 · outbound

This paper cites Psychology Press.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Psychology Press

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.080570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:9525551bfdda9b94b7313831d3755009686587c8c8aa56eda4436a6a39709116

Observation dd7e063a-af38-48bb-9927-890f171fb4c1 · outbound

This paper cites Active perception.Proceedings of the IEEE, 76(8):966–1005.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active perception.Proceedings of the IEEE, 76(8):966–1005

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.037518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:a34f8f319ecdaa2aaddb07a9d65a8ea50fd954497df6057916a7f1b6ca3b1350

Observation 709ce9dd-cfc9-4231-840a-3029cb23cd1f · outbound

This paper cites Revisiting active perception.Autonomous Robots, 42:177– 196.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Revisiting active perception.Autonomous Robots, 42:177– 196

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.960694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:99a3d65acedd4e4cc8e4e8b05ba50421db5ba5572585d654d9444406e3f658da

Observation d5e2eb7f-f225-4aeb-baff-79bb96e3df92 · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Activitynet: A large-scale video bench- mark for human activity understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.083827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:c149bd47c46c4c84d62bd041392fcb6e91655967a90f0cb1b46f3321885089c5

Observation 9a3f3610-568d-4b6c-b370-3290297c9f52 · outbound

This paper cites The perception-behavior expressway: Auto- maticeffectsofsocialperceptiononsocialbehavior.Advances in Experimental Social Psychology/Academic Press.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models The perception-behavior expressway: Auto- maticeffectsofsocialperceptiononsocialbehavior.Advances in Experimental Social Psychology/Academic Press

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.071575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:41db5e643198d51fd5dbc4433994a21128a7cbcd675e2e57e399f1d0e06657b5

Observation 4c8a43a9-a33f-41e1-bdc1-6dc4586347d7 · outbound

This paper cites VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.659733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:ced99de96234081ffedde42bf3a8724f725e955198dc79d009ba390f5bb58d44

Observation 06bb5e40-0395-4daa-8df7-db205946974d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:58:42.298802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:474b335870b5b966fa41289704b7a80def1c5650259c0b4ccf911fc7ddfb9e3c

Observation 13146a99-53d0-40b6-8d45-5b85be2eb64f · outbound

This paper cites an unresolved cited work.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-26T03:11:35.942275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:15194865652ce67d6a6e2cd6c236b858dcec2d6cd2df98c0be7713254c704f19

Observation 1bc0fe3e-f829-48f8-b39a-2714959ba37e · outbound

This paper cites Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.141895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:643d79a05374436ff5275ca6a3767985a7abd6169f583bc6ac102eb99c832939

Observation 8a5109d8-fdbd-4bdb-909f-60461c6da5da · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:24:30.435137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:53497fa7a91cf77538e54a4b8e737207424f13ef2532e4546feaf8ff9a55324b

Observation 45567afd-c411-48d6-aafd-5be69a8bfeea · outbound

This paper cites Real-time intermediate flow estimation for video frame interpolation.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Real-time intermediate flow estimation for video frame interpolation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.990837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:b9ac3b8c6c76ea4e5145adfec654b47b13cc96655ebc9454f21fa7b9633d4c02

Observation a8868eec-f52f-467c-b308-bbfa81ddf0bc · outbound

This paper cites Building a mind palace: Structuring environment- grounded semantic graphs for effective long video analysis with llms.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Building a mind palace: Structuring environment- grounded semantic graphs for effective long video analysis with llms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.032537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:1155029fc682d073cff72f8160f27cf4611c0403695dd870b8f0dce93a877273

Observation 9b2f6a14-ebc1-4b1a-b8c5-7995a2857a84 · outbound

This paper cites Adaptivevideounderstanding agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Adaptivevideounderstanding agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.046062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:2d44bc61ba6c4736eb9bc655b5169d240dd9ce103831e903e4277b178dbf5ae0

Observation 5d7fa9f5-5917-45cc-af1b-e1899bf07194 · outbound

This paper cites Cost-sensitive feature acquisi- tionandclassification.PatternRecognition,40(5):1474–1485.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Cost-sensitive feature acquisi- tionandclassification.PatternRecognition,40(5):1474–1485

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.917041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:09662985ecee9f0c084ef8c75f15952d3f4f93d5396cf5604c1dffa10d382343

Observation 0d7fc838-0364-463c-8dcd-219658f94108 · outbound

This paper cites Egoschema leaderboard.Kaggle Leaderboard.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Egoschema leaderboard.Kaggle Leaderboard

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.067395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:e1bd04af024979cbc205112929bf6db93cffd75dbb52eb57c4f3fe7141ae35c1

Observation 71b9b300-e04e-4249-bfa0-b5a8df08b993 · outbound

This paper cites An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.716021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:e5d989cf33f7c31f44a06c260ff4abf0088e3b2d260a54e7cdecde3d06481392

Observation 9037d6a7-2e45-4b58-9f13-29c4cd34cda2 · outbound

This paper cites Active Acquisition for Multimodal Temporal Data: A Challenging Decision-Making Task.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active Acquisition for Multimodal Temporal Data: A Challenging Decision-Making Task

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.670146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:ab47e5089da3b0943b011e57c4e5ef34e30f70068bf7021aace8e0b2027ad815

Observation 7d72b43c-e4d0-475d-9d19-20b2d94a7185 · outbound

This paper cites Active Data Acquisition in Autonomous Driving Simulation.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active Data Acquisition in Autonomous Driving Simulation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.718988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:360c6af300a493c18f287f603c583d35600a94e36ca72e97c01418214398f66e

Observation b58129b5-a744-45f4-a04a-d389436fbe33 · outbound

This paper cites Accurateimputationandefficientdataacquisitionwith transformer-based vaes.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Accurateimputationandefficientdataacquisitionwith transformer-based vaes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.964705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:06648033184b09f9135cad2827085d607bf8a9606fa6dffeace94700916977a4

Observation e0ceafe7-78bc-42a5-8464-a0a72126412a · outbound

This paper cites Lmms-eval: Accelerating the development of large multimoal models.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Lmms-eval: Accelerating the development of large multimoal models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.117369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:d0fcceb225e784a450c6036efe507650b2a222aa0084a91a2347fb595ca6d481

Observation 075bbe11-cfd6-45f2-8a1e-3aad7457b462 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.663235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:00bcbe07db32a6ae1111b97e61079869c73585473c46d39282b73a02be653898

Observation 14800086-7ba6-4d6b-bcfb-7a988035f45e · outbound

This paper cites Intentqa: Context-aware video intent reasoning.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Intentqa: Context-aware video intent reasoning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.003594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:18f419fd7292504388a8e3fb2061f36098b5e8b62488b91adfe49ccdefbe1c57

Observation d28fd3c5-02a7-4686-9a77-795a12fef6aa · outbound

This paper cites Mvbench: Acomprehensivemulti-modalvideounderstanding benchmark.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Mvbench: Acomprehensivemulti-modalvideounderstanding benchmark

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.938574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:8e39fc8c9349c69ee4fc435bd22c94a8e520fec54635f77d6ea16009291637ac

Observation f5d3bc76-2c7a-4f85-a395-2d6392101f73 · outbound

This paper cites Foundations & trends in multimodal machine learning: Prin- ciples, challenges, and open questions.ACM Computing Surveys, 56(10):1–42.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Foundations & trends in multimodal machine learning: Prin- ciples, challenges, and open questions.ACM Computing Surveys, 56(10):1–42

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.076925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:bfac98fdec540677351465a9b9a6a674b18bafb2be3e9b548d4d2839a4f987e9

Observation 918d1069-ca0e-42e7-8737-9b6fb13fc9c0 · outbound

This paper cites Videoinsta: Zero-shot long video understanding via informative spatial- temporal reasoning with llms.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Videoinsta: Zero-shot long video understanding via informative spatial- temporal reasoning with llms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.950383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:ee78ee4e1b1cddc3523914c2179cca0dfd7181c36e17bf724de00b0a12a61ec7

Observation 2744caf3-5bfd-4538-841a-9d57203b1203 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Video-rag: Visually-aligned retrieval-augmented long video comprehension

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.092283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:a02154ed87825df25845a1f63cac7f2929ce9d23afa6626b0bac0f5132c2e7a9

Observation 9f332021-f86c-44eb-933a-166e2130c0ff · outbound

This paper cites Drvideo: Document retrieval based long video understanding.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Drvideo: Document retrieval based long video understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.946090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:c3bee68d4bc843ba9760ee54d23b9ecba05195624a0832815d0387f3475ffb12

Observation 347b50f8-096c-4186-bdad-cf71a0de69e1 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:36:18.685087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:a454ea7a81b8538bbc042a4950e0973a279c1b910e09a29522eb27efac627206

Observation ee9374e1-f65d-4d21-80a6-25b3a9ee1123 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural Information Processing Systems, 36.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural Information Processing Systems, 36

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.058546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:fa6e71611386ecbef571144a62e21a2d7c8b744ad47a3733231aae5ffb1766e1

Observation 0d5845ac-fee2-461b-ac51-b1e55371dbd6 · outbound

This paper cites an unresolved cited work.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-26T03:11:35.981872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:b52d9368fa2817ae7892c9ba5eabc39c3ebf4a0d53c1f5951b38bde9eef28559

Observation ff96e788-ecd0-420a-a400-de6a8e9219bb · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Morevqa: Exploring modular reasoning models for video question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.054444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:401130b8d8b25ce260f40392f747c1c10ca8f734128949a49a5bafe1717dfd3b

Observation cf4c8499-f989-4bec-b0d6-6c90de6f7152 · outbound

This paper cites Gpt-4o blog: Hello gpt-4o.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Gpt-4o blog: Hello gpt-4o

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.012441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:b484855b290036823b1c60151c2ed13820f2f2105992525f1670d0f5ae6d5905

Observation 50b021f1-5f33-43a3-ae69-91eb1321c503 · outbound

This paper cites Too many frames, not all useful: Efficient strategies for long-form video qa.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Too many frames, not all useful: Efficient strategies for long-form video qa

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.736782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:855d085fc0f3e6a8f8b988c8b12ba8dc0be1729909f5b441c7863bd6f5925bc3

Observation 50fcf6c8-8ed9-4a60-a74a-026e58887ad9 · outbound

This paper cites Scalable diffusion models with transformers.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Scalable diffusion models with transformers

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.956408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:8fafdfcfe03afc6f78033fd9dc8a0e53b1872169be8ad89a5c78cf8fe572772a

Observation 14004e7a-64c2-44c6-88eb-3c91636cf148 · outbound

This paper cites Active percep- tion: sensorimotor circuits as a cortical basis for language.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active percep- tion: sensorimotor circuits as a cortical basis for language

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.973168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:147e9eca3543f0e95729d694040e93f24b52c64b1df9dd8aad1639e3ad583552

Observation 949a800e-5209-4b33-b563-6e0aa5f4eb68 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.062882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:5bfc95fdbf80fa00dfb69fdf72de1837ef6e119ef942ae322f1274b421c809f2

Observation 16fce5a7-3ba6-4cd1-9247-f9ab32c70935 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:35:23.646633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:e01cb71514c17475c9ab1f09fcee56aa631507294f8514587406cba6a681b31b

Observation 45deb8c5-0c72-44eb-b45f-dbc578ce3ed2 · outbound

This paper cites Active feature-value acquisition.Management Science, 55(4): 664–684.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active feature-value acquisition.Management Science, 55(4): 664–684

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.088324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:aa02f4a5a77313aa4e47d4259e4e04489706ae4cb7d612db50b7f0e7bfa7c5fd

Observation 2b954211-0a04-4145-995b-8d0b29936c70 · outbound

This paper cites An ecological approach to personality: Psycholog- ical traits as drivers and consequences of active perception.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models An ecological approach to personality: Psycholog- ical traits as drivers and consequences of active perception

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.986520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:4e071e08b60284417ed8843ac6b98329c924767f76ee057c2bc7b7f0082d57f6

Observation ddd73909-1e2c-4225-a0de-19fe7adfd299 · outbound

This paper cites Maximizing Information Gain in Partially Observable Environments via Prediction Reward.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Maximizing Information Gain in Partially Observable Environments via Prediction Reward

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.712709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:a31eff1463c71cfccdbea4491c8483012fb91c3c2d2f8ba4099a2c584a9f83e9

Observation f0c0b2b5-76d3-4261-9bc9-be0d93d8cf86 · outbound

This paper cites Traveler: A modular multi-lmm agent framework for video question-answering.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Traveler: A modular multi-lmm agent framework for video question-answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.977258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:adb58abe93a8cd2b5b339a4e7a65615b18000cacf52d0bee4073745e5917eb29

Observation 8e21caea-f888-4a0d-8adb-79d7baf086a0 · outbound

This paper cites Joint active feature acquisition and classification with variable-size set encoding.Advancesinneuralinformationprocessingsystems, 31.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Joint active feature acquisition and classification with variable-size set encoding.Advancesinneuralinformationprocessingsystems, 31

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.041817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:7bbaef204bcaa37a6115bd7cf385de166823c88b71297fcf3ed021ef574cb161

Observation aa3a7a44-82f4-4d86-bc40-dfd3cefffa77 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.923878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:36725636ee730c0e2dff01e91a96a47c8106cebc41934d7b4526c1b3f27b4a5d

Observation 41eef350-d07b-450c-b82c-9b8e3edc65d7 · outbound

This paper cites Stanford University.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Stanford University

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.120944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:591de079d9eccd32068aa3b46b1e10d91c099f99574c19d284d6bd9ae3719918

Observation bd15e77c-f160-475d-99b8-5cb61f3dc4b4 · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.693563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:33379e214e97545d389e099cd6a734a0d062eab9c90059e0c9b43cbc0c80b8bb

Observation 01fe7a43-bdd3-4e21-a48d-92919b0efde8 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.687909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:d68f759ae3be6ef3d27c78b478c14a5f3657d34e2a44f1291da559e823618a0f

Observation 89895639-7621-42d4-9830-a2cde7fdf635 · outbound

This paper cites Network mechanisms of ongoing brain activity’s influence on conscious visual perception.Nature Communications, 15(1): 5720.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Network mechanisms of ongoing brain activity’s influence on conscious visual perception.Nature Communications, 15(1): 5720

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.129658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:e2cf97be1d1b0dcb8d09f2bbd98d1a5c767777e5ca6060174978b47ede38253c

Observation ceae1ef3-5e0f-4ade-bc64-11b529f55834 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporalactions.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Next-qa: Next phase of question-answering to explaining temporalactions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.097657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:a0cc324d860cdf5c6975e83f1837e07113106ac111e667becf8752b239b5ec44

Observation 526113ce-1e40-46e2-af1c-e6c8b50c2d24 · outbound

This paper cites Slowfast-llava: A strong training-free baseline for video large language models.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Slowfast-llava: A strong training-free baseline for video large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.105435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:c3522b1c4e9d6a0e990c6c26b4ba158424bdd2ba69bfe215d5ec146c4d94cef0

Observation 3fc1a1a1-25ca-4c17-be13-44cfab9930e2 · outbound

This paper cites Evalai: Towards better evalua- tion systems for ai agents.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Evalai: Towards better evalua- tion systems for ai agents

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.113039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:22a2deb6248eda3aa0abcb2c4e19789d92cae7e09763c322d42dfcc476051f70

Observation 78fe61ae-b9e5-4928-b71c-7613f598bfca · outbound

This paper cites Active sensing in the categorization of visual patterns.Elife, 5:e12215.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active sensing in the categorization of visual patterns.Elife, 5:e12215

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.137850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:8e7c6b303c4f41826e0a37f0cdb55e8c116b16983fb0652671ee2cab2991c215

Observation e9154b44-e78d-459e-b865-96d65919d340 · outbound

This paper cites Vca: Video curious agent for long video understanding.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Vca: Video curious agent for long video understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.101504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:02b70f1c164fae471d4a91a36b282c6d07ea4aaeac00d7b206684b7e025a9373

Observation d0ace8d5-ce54-4560-86bd-8b84845abce8 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:26:22.534732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:4d74e4e7a64045cc4a9f5fdd2723c08e37b12b6528f6b053616d2d38f9b10f8c

Observation 69b7c51e-3217-4686-86c7-6b892ff83e4b · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:01:50.618167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:67a28ef993958faa4d8aa4d90dfc0b6621e54a39967b9d35467d1577148747c0

Observation 7df06415-8597-4de5-b8da-a03ff47487a1 · outbound

This paper cites Reinforcement Learning with Efficient Active Feature Acquisition.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Reinforcement Learning with Efficient Active Feature Acquisition

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.684974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:af210f23cccd1a824129fb6e4aa82ca1616c723f1f4f06bde5a0c772ddb7f4c3

Observation 1edb73f1-2eec-447e-914b-f89d4cf37902 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:2e0263978207fffee38ae52a8975503573d7a1e8e4ba2d68a3c8a5a6fc8f0974

Observation 1065da24-55b8-496d-98b2-8391869bce17 · outbound

This paper cites Active sensing.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active sensing

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.994976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:dc49931818f4a07492d0a29bcf1159d5ad10b8cebd646d99e56540d8b805d75e

Observation e1df9adc-b2d5-457f-a911-213fce65abab · outbound

This paper cites Activitynet-qa: A dataset for understandingcomplexwebvideosviaquestionanswering.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Activitynet-qa: A dataset for understandingcomplexwebvideosviaquestionanswering

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.929678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:70035d3eca1467ddbdee7a59dc8545de3e750b1d2cf30e8504c17de357d6a978

Observation 2fef03a1-0e95-4b77-949b-833037f68e2a · outbound

This paper cites Active Perception and Representation for Robotic Manipulation.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active Perception and Representation for Robotic Manipulation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.646806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:9c0d862c13afda12cd6eeb9c2afc06d0e4cfdbc96df75cd4a2399dc15091dcbe

Observation e15f4043-62b2-46b3-9d2a-6ec710a8141d · outbound

This paper cites A simple llmframeworkforlong-rangevideoquestion-answering.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models A simple llmframeworkforlong-rangevideoquestion-answering

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.124998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:633c737c2513672f01fcaee1ef09935a88c5735643e3de8d5a2af0a522f3cad1

Observation 7b06da63-8fe3-4fc4-8bd2-69d84e1a9f45 · outbound

This paper cites Safe Occlusion-aware Autonomous Driving via Game-Theoretic Active Perception.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Safe Occlusion-aware Autonomous Driving via Game-Theoretic Active Perception

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.681923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:d1b6571a0d0bb08d62688a32283888963fed0549db65faaf0bfa0c96c056ff2d

Observation 5e93fbef-b51c-4c4a-a725-6f99990aac5d · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.650724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:e020d80b2eab5a04204b87707b2eb571a46a4cb52fe196d22a458cb7e8832dda

Observation 1eee0626-5379-4ccf-b776-4d3934a28137 · outbound

This paper cites Weprovideadetailedpromptwithexamplestoleverage the video generation model’s capacity as much as possi- ble.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Weprovideadetailedpromptwithexamplestoleverage the video generation model’s capacity as much as possi- ble

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.023935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:c9a789c869bf5996918f6a2fc810997ee8f30771f26e2fa281d0e30a9a06f313

Observation a41c8e9c-96fd-4a29-a2b4-522e594ebfc1 · outbound

This paper cites 1b)Extract visual cues that indicate the environment (e.g., indoor, outdoor, time of day) and participants (e.g., people, animals, objects).

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models 1b)Extract visual cues that indicate the environment (e.g., indoor, outdoor, time of day) and participants (e.g., people, animals, objects)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.050545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:39fa7d89da8d3e99ee696934e5b2966d58a28bec08a82e61362c3cfa2c0a030b

Observation 28658935-4fb8-454c-9ba1-91088bd5029a · outbound

This paper cites 2b) Consider each possible answer to understand different potential outcomes or scenarios.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models 2b) Consider each possible answer to understand different potential outcomes or scenarios

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.934272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:354cd6c5381ae6eb809c24a82694099ac640f66b4343164e98bfc0dcc3aae446

Observation eea959ae-ae5d-4947-80e5-b8c646c7928c · outbound

This paper cites 3b) Focus on generating dynamics that would lead to scenarios described in the possible answers.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models 3b) Focus on generating dynamics that would lead to scenarios described in the possible answers

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.008002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:0dfe05e9e7a7b2f9d8468233df377c5909bf265db38bd981e7e75202f3ac3d7c

Observation 0b002e87-c3c0-435b-9953-97047572d18b · outbound

This paper cites an unresolved cited work.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-26T03:11:36.133711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:ae80757dc4a534069636c11f662e04fb54371455f4d7192e33f83a65e3d1c564

Observation 88229131-0c9c-424a-85a2-c7c015f61943 · outbound

This paper cites an unresolved cited work.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-26T03:11:36.109028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:3bf17ca119de5af03e974da1a8ac364abf28970d71dfc6dabbdbf79015b7b1ab

Observation f27a14a8-eeb1-49a7-a4d7-cc7fd7fba306 · outbound

This paper cites an unresolved cited work.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-26T03:11:35.968728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:36879a084183a8f3a7bae5c08a893e798cda86fc6da2dc5833fe83f3abd3f7fe

Observation 34b523ee-c76f-439d-8963-ab4721446afc · outbound

This paper cites 1b) Incorporate common sense and logical reasoning to predict what is likely to happen next.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models 1b) Incorporate common sense and logical reasoning to predict what is likely to happen next

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.028199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:1e14802031494805afe0825e05407f1e88dcd3daf5fa42f4e134c79a7efa2eb7

Observation 2b2eb67a-7416-4987-bcde-7a30a63eb63b · outbound

This paper cites 2b) Highlight events or actions that would help distinguish between the different answers.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models 2b) Highlight events or actions that would help distinguish between the different answers

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:35.999526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:0cfdab8b0dd9452c4d5c2ece8b6a11dd88e19882e29bcaf44dff62c537bc0178

Observation 4a3581e0-c858-4a56-ab44-17e19e2f2cd3 · outbound

This paper cites What does the person do after the light turns green?.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models What does the person do after the light turns green?

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T03:11:36.019864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:2ed8a4b13bd21513a49d37ebc3b4bfaca44b50a6a35bc03dc06d9321aede3710

Pith citing papers

Observation b977cab4-e1ab-4819-829d-695c14c02af6 · inbound

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? cites this paper.

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T16:16:51.254681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:16:51.254681Z digest=sha256:a6227c01f566ebf320dc1df19d968920ec47b987bf6387dce758ae434bc41f7f