Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:43:17.664668Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 7 inbound Pith citation observations for arXiv:2504.17447.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:43:17.664668Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:42.167776Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:09:57.135194Z
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1ada4143-35d9-45db-94cb-e14e07ef3683 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f89af446-a8ea-472b-903a-6491e0a3f9c1 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Memory Consolidation Enables Long-Context Video Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b633a9e-24a4-4eea-b8d9-6ed4b62d17bd · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Scene text visual question answering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e6059f84-e991-4470-9dc6-8990444182dd · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Gram: Global reasoning for multi-page vqa
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 178e8cc1-86ac-4fb2-8dca-3d66c9b14994 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce5fd533-6a47-45dd-95a7-9eeefb8dfe10 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c822df91-9e47-4ab0-ab21-1567456bcd32 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6de5e810-6df3-4d82-b3ea-bae87cf764c5 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af9e289b-df83-4efe-a58f-38d089121ef5 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 20baadda-17aa-49cc-bb4e-72a54323a851 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e97737cd-0f35-4ae0-80d8-acfea68998ad · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c9a97cb-fba8-4e27-a9d7-e688c47c9bb3 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f6767b8-faa8-49bf-b35d-290f5056b60a · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Scsampler: Sampling salient clips from video for efficient action recog- nition
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b86dcd42-eb52-46b0-8cce-528cb082ef19 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Text-Conditioned Resampler For Long Form Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36dc74f4-5747-4cbc-a24f-9190fe7a50ef · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Building and better understanding vision- language models: insights and future directions., 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 58863500-c589-487f-90ce-db3d2db49149 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 011b615d-ccd1-4494-8f3d-1ca725d6d4cb · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1217fc4-556b-436b-9b8a-84ed79f83219 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43381703-c98e-4d46-b49d-f46925d5cb91 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Vila: On pre-training for visual language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bb7b0b27-0a16-4ee6-9ad6-9cfce0a83da2 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Visual instruction tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4759bcae-9a95-4113-ac08-8a535f435ccb · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 402462e8-cc6e-4395-b2f9-27734b8d7c44 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4c7c1924-1e88-4e5e-b8fd-eb3bbd8c7894 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Docvqa: A dataset for vqa on document images
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c4baba-695d-4bbe-b9c2-6f2680e3f6f3 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Too many frames, not all useful: Efficient strategies for long- form video qa
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82fedbd9-d060-4a2a-be62-bceff9ec89ba · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Streaming Long Video Understanding with Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1eed15c-fea4-476d-bc42-759301067f22 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1857db82-5077-4b20-9c24-09374eac1b50 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Slidevqa: A dataset for document visual question answering on multiple images
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d7b5f9fb-345a-4dae-acf8-1fe0e52fb72a · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Instructdoc: A dataset for zero-shot gener- alization of visual document understanding with instructions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dfc16086-dddb-4efb-ad8e-79989b8bfdc0 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Hi- erarchical multimodal transformers for multipage docvqa
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1298d1a8-b6fb-4a92-8886-b831c7296f81 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58ae6ca0-c904-4f4e-ab8f-d3c7e6208b1d · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Lvbench: An extreme long video under- standing benchmark, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5206eac9-1b0d-441b-943d-4ca90021443f · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a83f2be9-a27a-4baf-acab-f113d7f31e9a · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b4584ac-7149-4d2c-8a41-910b8a38eea8 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LongVLM: Efficient Long Video Understanding via Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83dd7ddf-af62-4574-84b1-80f13ceefb27 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 354fad58-8c66-4189-8728-22959c02b8fa · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef481088-fd65-43c7-9e6e-30081001d4aa · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Next-qa: Next phase of question-answering to explaining temporal actions
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8c82477e-1992-47de-bbd0-7070ecd3c9bb · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d5ac8eb-d889-43ed-a7b2-f3af9282befb · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4f6d4a8-51e4-4acc-9cb6-8c8208e0830f · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Self-chained image-language model for video localization and question answering
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a1c67c44-1506-41e7-bf7e-73884a4d6389 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Frame-Voyager: Learning to Query Frames for Video Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d8254d3-2911-476a-afff-4abea8b596ce · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Sigmoid loss for language image pre-training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a7fa0d-12d6-43e3-bddb-2f8f684dfb1c · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Long Context Transfer from Language to Vision
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3a52f5-c8a9-468c-855a-24bd4205849f · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Llava- next: A strong zero-shot video understanding model, 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 272592ad-4daa-4e37-95db-74214230670d · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding MLVU: Benchmarking Multi-task Long Video Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38d4079-ca5a-428e-a6af-ee1209482b68 · outbound
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding Selection Scoring Prompt Our prompt is based on the prompt used in SeViLA [40]
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 34d43881-d5b2-4eb7-a79c-a3532b767e37 · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b25ee4a3-3132-4bde-b546-d701ac7253df · inbound
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9928cae8-7e0d-41b4-9ec5-2f63929189fc · inbound
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d481f219-c623-41d4-8a26-96118247b5cb · inbound
Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4065af9c-3783-4808-9ff1-13758e62168b · inbound
QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d598b017-5f0a-4bc8-9264-60f35e2e2fb8 · inbound
QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 917e71ba-b271-41a7-b130-9ebdc8cb9808 · inbound
Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.