Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:22:35.965901Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2510.01009.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:22:35.965901Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c1fa159f-0ae2-4cd6-827d-62e3c9089184 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d28358f7-6e0a-471c-bb90-8db79045d677 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eb72ff8-da18-4084-97c8-d18d3098e792 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f771fd04-7b2d-4444-8768-5e7858c4cd8f · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Vivit: A video vision transformer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab416c63-d1fa-4e14-9e0a-437d18f683c7 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Goldfish: Vision- language understanding of arbitrarily long videos
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13fa6ef2-939e-4d02-822d-ce5f3cc4761f · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Machine Learning, pages 813–824
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc12f8f4-c088-4f66-94bc-2f9d30b69dd3 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcccc084-5db0-4ccc-bb78-657ac82e8683 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Vindlu: A recipe for ef- fective video-and-language pretraining
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8df1c74-c858-4cbb-bfa8-f46a40823712 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be4c9d6-cf02-4784-a35a-fa43352f0d1f · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Qlora: Efficient finetuning of quantized llms
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfaa2350-16c1-4ef5-8d5a-ba4f979d49a2 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Knowit vqa: Answering knowledge-based questions about videos
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13ff2d9-b68e-4c4d-9bce-c3afb6d3f4b4 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a62ce53-ac7a-4488-ad1a-10bb4b1ceaf6 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge mem- ory
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a829eb62-cfc0-4e14-8b83-8a4677c108ee · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Ku, Qian Liu, and Wenhu Chen
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33fb768c-b02c-49e9-9b1a-c66d0889da50 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Visual question answer- ing: Datasets, algorithms, and future challenges.Computer Vision and Image Understanding, 163:3–20, 2017
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1704e858-929a-4b94-aeda-2aff16fb65db · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency An image grid can be worth a video: Zero-shot video question answering using a vlm.IEEE Access, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0619942-7f50-4460-b84f-b62a7f3495d8 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency TVQA: Localized, Compositional Video Question Answering
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56dd022-f6b6-4b89-b253-71bc8c61c996 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96318be8-ddff-4ddb-88fc-bc9c5f7b8dd6 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ebd77bb-2af6-41d2-9d3b-68ddd1a43fd1 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b3eb857-445b-43f8-b62d-fb3afcfac7d1 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency HERO: Hierarchical encoder for Video+Language omni-representation pre-training
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 230df3f3-e26b-4782-94b9-2932af801295 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Video-llava: Learning united visual representation by alignment before projection.EMNLP, 2023
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776dcebe-c3dc-4c7a-8907-93b65501d099 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Common- sense video question answering through video-grounded en- tailment tree reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91238d44-23f5-45e0-9537-c046fabea754 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc2e1ca2-8b63-4f5e-add4-4a821f048960 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Question-instructed visual descriptions for zero-shot video answering
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986407c1-5c37-447b-a358-0ae5bb9f5ce9 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Streaming long video un- derstanding with large language models.Advances in Neural Information Processing Systems, 37:119336–119360, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35b22632-d826-47b1-b4fd-0e6c1e51a8f9 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Direct prefer- ence optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4624298-6b57-4ac2-a74f-d0f676cc4057 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d22ebfc7-ef6e-418b-af03-7c5fdf45f922 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Action Recognition using Visual Attention
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1c2e0aa-9720-440b-b197-e7c5170677dc · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Moviechat: From dense token to sparse memory for long video understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb01734-afaf-43e6-9f72-eb607114c636 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Modularized self-reflected video reasoner for multimodal llm with application to video ques- tion answering
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a26c2812-3d26-407c-b46a-4376dada978a · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Movieqa: Understanding stories in movies through question-answering
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b5ffba-2c35-4879-a1f6-5f492720c4b7 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65826b61-f3c3-4dc0-a3cc-ab8cf5ebf585 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Video- mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural infor- mation processing systems, 35:10078–10093, 2022
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a34a74f-3e2b-405c-9619-8b161a62fbdb · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Fastvlm: Efficient vision encoding for vision language models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b1023c-47f2-4c5e-8eae-f3c77288b4ea · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49598a56-397e-4a7c-90b2-9d5a7e7f9b95 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Vila: Efficient video- language alignment for video question answering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cab872b-e3ab-488b-b404-c942bc03d25f · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9d193d-3232-42ab-9b75-3bd67ae6e484 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Longvlm: Efficient long video understand- ing via large language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bccc7492-cf70-4dea-a402-134cf0aea9e6 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency STAR: A Benchmark for Situated Reasoning in Real-World Videos
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b384d6-847f-4ca8-9baa-7cd7870bde0d · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Adaframe: Adaptive frame selection for fast video recognition
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 286b7aeb-d09e-4542-85c9-27b73dd22711 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Next-qa: Next phase of question-answering to explaining tem- poral actions
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1228da27-da17-4897-b1a4-2f6bdd2c8883 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Zero-shot video question answering via frozen bidirectional language models.Advances in Neural Information Processing Systems, 35:124–141, 2022
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d6a8ca9-fca7-4d50-a6d2-bf099951be2a · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Qwen2.5-1M Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73be145b-6fc5-4477-a6ee-694af2d996e8 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ef35a2-c38d-4799-9ae0-68ffb6073932 · outbound
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Self-chained image-language model for video localization and question answering.Advances in Neural Information Processing Systems, 36:76749–76771, 2023
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.