Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:29:51.391770Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2412.17415.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:29:51.391770Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T01:52:44.785582Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T12:46:56.808179Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 36b59a45-b65a-4919-a5fd-3d4d50acf58b · outbound
VidCtx: Context-aware Video Question Answering with Image Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0b15ddd-9712-425a-8d9b-836f8593e592 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d862ea53-59ad-4fec-a7e4-5741daf06e7c · outbound
VidCtx: Context-aware Video Question Answering with Image Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 976db303-747f-4c80-bb8d-17d02eaadee7 · outbound
VidCtx: Context-aware Video Question Answering with Image Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cec20e0-c2b2-46ca-926a-2cd54c61a4e0 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Mvbench: A comprehensive multi-modal video understanding bench- mark,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 34d2f7df-a516-4953-bf2c-2d7e264b4e14 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f5ca850-f005-485f-bcec-d2d08ccae1fd · outbound
VidCtx: Context-aware Video Question Answering with Image Models InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b73ffffa-6723-4754-8183-1b97d3d9654e · outbound
VidCtx: Context-aware Video Question Answering with Image Models Self- chained image-language model for video localization and question answering,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e782af15-4a88-435d-b6b3-b7598ae1a837 · outbound
VidCtx: Context-aware Video Question Answering with Image Models VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a4f946-5951-4242-9af4-05643838b515 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Videoagent: Long-form video understanding with large language model as agent,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 779c2b7a-300b-454d-b289-4758696ef7eb · outbound
VidCtx: Context-aware Video Question Answering with Image Models A Simple LLM Framework for Long-Range Video Question-Answering
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9a8cb9-4666-4f03-af54-b49098b2ee9b · outbound
VidCtx: Context-aware Video Question Answering with Image Models Language Repository for Long Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daf1d424-5f21-4a3c-8b3d-d98dfa29eb43 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Question-instructed visual de- scriptions for zero-shot video answering,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 84ac54d0-55df-4757-a43e-2305ddb96686 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6736c950-611b-4dab-ad9e-ced58d2443c2 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Large language models can be easily distracted by irrelevant context,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation db2d6455-f26a-4ea6-bc2a-23a4cd9236ee · outbound
VidCtx: Context-aware Video Question Answering with Image Models Learn to explain: Multimodal reasoning via thought chains for science question answering,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e862572-2f5e-4d59-b473-7d4514489060 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Multimodal Chain-of-Thought Reasoning in Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1290a7d-a931-433b-a740-3d6bc81295ca · outbound
VidCtx: Context-aware Video Question Answering with Image Models Ddcot: Duty-distinct chain-of-thought prompting for multimodal rea- soning in language models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c3947363-3c7c-4dd3-b60e-3dc3edf60c6b · outbound
VidCtx: Context-aware Video Question Answering with Image Models Enhancing multimodal sentiment analysis via learning from large language model,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd9a91da-8403-4e3e-90cb-d3d9c73af29a · outbound
VidCtx: Context-aware Video Question Answering with Image Models Video-of-thought: Step-by-step video reasoning from perception to cognition,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3e5aa05f-549f-4f40-9081-d470ee0498b7 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Vamos: Versatile Action Models for Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f5fac8-7b5a-489e-bab4-c6a35ee3a06c · outbound
VidCtx: Context-aware Video Question Answering with Image Models Large language models are zero-shot reasoners,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7e21f624-5f9d-4657-9b4b-40809b13690e · outbound
VidCtx: Context-aware Video Question Answering with Image Models Chain-of-thought prompting elicits reasoning in large language models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 09818ffb-dafd-4aeb-a383-1fd0d3bc1f0d · outbound
VidCtx: Context-aware Video Question Answering with Image Models Compositional chain-of-thought prompting for large multimodal mod- els,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fa23c586-feed-418c-91dd-c826e7f0678e · outbound
VidCtx: Context-aware Video Question Answering with Image Models Next-qa: Next phase of question-answering to explaining temporal actions,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed5e1194-f9b4-4603-a465-a88004e6276a · outbound
VidCtx: Context-aware Video Question Answering with Image Models Intentqa: Context- aware video intent reasoning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1dc03d09-aa3f-4fc7-b7d6-a78b825ce3d0 · outbound
VidCtx: Context-aware Video Question Answering with Image Models STAR: A Benchmark for Situated Reasoning in Real-World Videos
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29a74aa0-e440-4351-acaa-3a00735895d7 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Verbs in action: Improving verb understanding in video-language models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 527a5034-1726-4fac-984e-db07468774e6 · outbound
VidCtx: Context-aware Video Question Answering with Image Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 199033f3-a787-4436-aa41-9da7e36fc172 · outbound
VidCtx: Context-aware Video Question Answering with Image Models Llava-next: Improved reasoning, ocr, and world knowledge,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4dccfb87-3f24-4aae-bd55-8b45679640d5 · inbound
VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning VidCtx: Context-aware Video Question Answering with Image Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.