Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T19:17:20.745425Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2605.01662.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T19:17:20.745425Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:16:51.254681Z
A source-named dated measurement, never combined with another source.
Source: cited_works
73 of 73 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ce2e797f-d450-4d35-94f5-cb930986bfd6 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cf96cc15-4114-47fa-a1f5-b32fa019df95 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Psychology Press
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dd7e063a-af38-48bb-9927-890f171fb4c1 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active perception.Proceedings of the IEEE, 76(8):966–1005
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 709ce9dd-cfc9-4231-840a-3029cb23cd1f · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Revisiting active perception.Autonomous Robots, 42:177– 196
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5e2eb7f-f225-4aeb-baff-79bb96e3df92 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Activitynet: A large-scale video bench- mark for human activity understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9a3f3610-568d-4b6c-b370-3290297c9f52 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models The perception-behavior expressway: Auto- maticeffectsofsocialperceptiononsocialbehavior.Advances in Experimental Social Psychology/Academic Press
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4c8a43a9-a33f-41e1-bdc1-6dc4586347d7 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 06bb5e40-0395-4daa-8df7-db205946974d · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 13146a99-53d0-40b6-8d45-5b85be2eb64f · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1bc0fe3e-f829-48f8-b39a-2714959ba37e · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Video diffusion models.Advances in Neural Information Processing Systems, 35:8633–8646
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a5109d8-fdbd-4bdb-909f-60461c6da5da · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 45567afd-c411-48d6-aafd-5be69a8bfeea · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Real-time intermediate flow estimation for video frame interpolation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a8868eec-f52f-467c-b308-bbfa81ddf0bc · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Building a mind palace: Structuring environment- grounded semantic graphs for effective long video analysis with llms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9b2f6a14-ebc1-4b1a-b8c5-7995a2857a84 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Adaptivevideounderstanding agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5d7fa9f5-5917-45cc-af1b-e1899bf07194 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Cost-sensitive feature acquisi- tionandclassification.PatternRecognition,40(5):1474–1485
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d7fc838-0364-463c-8dcd-219658f94108 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Egoschema leaderboard.Kaggle Leaderboard
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71b9b300-e04e-4249-bfa0-b5a8df08b993 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9037d6a7-2e45-4b58-9f13-29c4cd34cda2 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active Acquisition for Multimodal Temporal Data: A Challenging Decision-Making Task
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d72b43c-e4d0-475d-9d19-20b2d94a7185 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active Data Acquisition in Autonomous Driving Simulation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b58129b5-a744-45f4-a04a-d389436fbe33 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Accurateimputationandefficientdataacquisitionwith transformer-based vaes
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0ceafe7-78bc-42a5-8464-a0a72126412a · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Lmms-eval: Accelerating the development of large multimoal models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 075bbe11-cfd6-45f2-8a1e-3aad7457b462 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 14800086-7ba6-4d6b-bcfb-7a988035f45e · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Intentqa: Context-aware video intent reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d28fd3c5-02a7-4686-9a77-795a12fef6aa · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Mvbench: Acomprehensivemulti-modalvideounderstanding benchmark
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f5d3bc76-2c7a-4f85-a395-2d6392101f73 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Foundations & trends in multimodal machine learning: Prin- ciples, challenges, and open questions.ACM Computing Surveys, 56(10):1–42
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 918d1069-ca0e-42e7-8737-9b6fb13fc9c0 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Videoinsta: Zero-shot long video understanding via informative spatial- temporal reasoning with llms
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2744caf3-5bfd-4538-841a-9d57203b1203 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Video-rag: Visually-aligned retrieval-augmented long video comprehension
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9f332021-f86c-44eb-933a-166e2130c0ff · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Drvideo: Document retrieval based long video understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 347b50f8-096c-4186-bdad-cf71a0de69e1 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee9374e1-f65d-4d21-80a6-25b3a9ee1123 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural Information Processing Systems, 36
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d5845ac-fee2-461b-ac51-b1e55371dbd6 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff96e788-ecd0-420a-a400-de6a8e9219bb · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Morevqa: Exploring modular reasoning models for video question answering
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cf4c8499-f989-4bec-b0d6-6c90de6f7152 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Gpt-4o blog: Hello gpt-4o
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 50b021f1-5f33-43a3-ae69-91eb1321c503 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Too many frames, not all useful: Efficient strategies for long-form video qa
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 50fcf6c8-8ed9-4a60-a74a-026e58887ad9 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Scalable diffusion models with transformers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 14004e7a-64c2-44c6-88eb-3c91636cf148 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active percep- tion: sensorimotor circuits as a cortical basis for language
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 949a800e-5209-4b33-b563-6e0aa5f4eb68 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16fce5a7-3ba6-4cd1-9247-f9ab32c70935 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 45deb8c5-0c72-44eb-b45f-dbc578ce3ed2 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active feature-value acquisition.Management Science, 55(4): 664–684
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2b954211-0a04-4145-995b-8d0b29936c70 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models An ecological approach to personality: Psycholog- ical traits as drivers and consequences of active perception
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ddd73909-1e2c-4225-a0de-19fe7adfd299 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Maximizing Information Gain in Partially Observable Environments via Prediction Reward
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f0c0b2b5-76d3-4261-9bc9-be0d93d8cf86 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Traveler: A modular multi-lmm agent framework for video question-answering
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8e21caea-f888-4a0d-8adb-79d7baf086a0 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Joint active feature acquisition and classification with variable-size set encoding.Advancesinneuralinformationprocessingsystems, 31
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aa3a7a44-82f4-4d86-bc40-dfd3cefffa77 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 41eef350-d07b-450c-b82c-9b8e3edc65d7 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Stanford University
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd15e77c-f160-475d-99b8-5cb61f3dc4b4 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 01fe7a43-bdd3-4e21-a48d-92919b0efde8 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 89895639-7621-42d4-9830-a2cde7fdf635 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Network mechanisms of ongoing brain activity’s influence on conscious visual perception.Nature Communications, 15(1): 5720
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ceae1ef3-5e0f-4ade-bc64-11b529f55834 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Next-qa: Next phase of question-answering to explaining temporalactions
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 526113ce-1e40-46e2-af1c-e6c8b50c2d24 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Slowfast-llava: A strong training-free baseline for video large language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3fc1a1a1-25ca-4c17-be13-44cfab9930e2 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Evalai: Towards better evalua- tion systems for ai agents
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 78fe61ae-b9e5-4928-b71c-7613f598bfca · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active sensing in the categorization of visual patterns.Elife, 5:e12215
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e9154b44-e78d-459e-b865-96d65919d340 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Vca: Video curious agent for long video understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0ace8d5-ce54-4560-86bd-8b84845abce8 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69b7c51e-3217-4686-86c7-6b892ff83e4b · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7df06415-8597-4de5-b8da-a03ff47487a1 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Reinforcement Learning with Efficient Active Feature Acquisition
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1edb73f1-2eec-447e-914b-f89d4cf37902 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1065da24-55b8-496d-98b2-8391869bce17 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active sensing
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e1df9adc-b2d5-457f-a911-213fce65abab · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Activitynet-qa: A dataset for understandingcomplexwebvideosviaquestionanswering
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2fef03a1-0e95-4b77-949b-833037f68e2a · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Active Perception and Representation for Robotic Manipulation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e15f4043-62b2-46b3-9d2a-6ec710a8141d · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models A simple llmframeworkforlong-rangevideoquestion-answering
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b06da63-8fe3-4fc4-8bd2-69d84e1a9f45 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Safe Occlusion-aware Autonomous Driving via Game-Theoretic Active Perception
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e93fbef-b51c-4c4a-a725-6f99990aac5d · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models MLVU: Benchmarking Multi-task Long Video Understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1eee0626-5379-4ccf-b776-4d3934a28137 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Weprovideadetailedpromptwithexamplestoleverage the video generation model’s capacity as much as possi- ble
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a41c8e9c-96fd-4a29-a2b4-522e594ebfc1 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models 1b)Extract visual cues that indicate the environment (e.g., indoor, outdoor, time of day) and participants (e.g., people, animals, objects)
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 28658935-4fb8-454c-9ba1-91088bd5029a · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models 2b) Consider each possible answer to understand different potential outcomes or scenarios
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eea959ae-ae5d-4947-80e5-b8c646c7928c · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models 3b) Focus on generating dynamics that would lead to scenarios described in the possible answers
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0b002e87-c3c0-435b-9953-97047572d18b · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 88229131-0c9c-424a-85a2-c7c015f61943 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f27a14a8-eeb1-49a7-a4d7-cc7fd7fba306 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 34b523ee-c76f-439d-8963-ab4721446afc · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models 1b) Incorporate common sense and logical reasoning to predict what is likely to happen next
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2b2eb67a-7416-4987-bcde-7a30a63eb63b · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models 2b) Highlight events or actions that would help distinguish between the different answers
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4a3581e0-c858-4a56-ab44-17e19e2f2cd3 · outbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models What does the person do after the light turns green?
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b977cab4-e1ab-4819-829d-695c14c02af6 · inbound
RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos? Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.