Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:05.248695Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 8 inbound Pith citation observations for arXiv:2506.22139.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:15:05.248695Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T18:39:24.915547Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:09:57.149192Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a2c548b6-25d1-44d2-ac91-7ec88012bc3a · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Qwen Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea63bb34-9624-424a-a837-d4b9cafa0642 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Robust motion-guided frame sam- pler with interpretive evaluation for video action recognition
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86249126-8260-4bac-8d4e-d5b7523ba862 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Sharegpt4video: Improving video understanding and generation with better captions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ca6dea2-4ce0-4342-9bda-680746d4e6d6 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636f60ca-4311-4df5-bf28-12d428f28cbd · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55a5ff53-44d3-4273-8475-82950400326c · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260f516a-eb60-4a70-aa94-5a542647b211 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ed91ed-796f-4940-b039-802d573b9f48 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Categorical Reparameterization with Gumbel-Softmax
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1403065-aac9-4fcd-bfdc-403074e446ef · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49eacf54-21f3-4a01-b6c3-b164f5b11144 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f53e097e-80e5-4101-b645-9a7a49a929b7 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a797c863-8993-4d59-b303-c2213ace2331 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76dc95d8-f7e4-4b88-9a02-94e10a7c9b77 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs KeyVideoLLM: Towards Large-scale Video Keyframe Selection
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2dd8b82-c552-47d0-885a-c4f771f400d5 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-llava: Learning united visual representation by alignment before projection
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2475e21a-6cd2-4261-8e08-9faca5ee815c · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Vila: On pre-training for vi- sual language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5d2932f-6139-47ee-8d33-e1b0201e24e6 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Visual instruction tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73badf97-5289-4ab2-9a12-ae2e30d2d29d · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Improved baselines with visual instruction tuning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8379fdf-f320-447c-9623-29220e3249d2 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Llavanext: Improved reasoning, ocr, and world knowledge, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f33d806-78f9-4ccf-9bd3-34caa259491c · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 666db4a5-e386-4eac-b8a7-5681a16d0392 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad62d436-89cf-4000-b98e-3e65e6ed29ca · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Hello gpt-4o, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9fbe48c4-ba7c-415b-a164-a6f7af190310 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Learn- ing transferable visual models from natural language super- vision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20018ddc-e8da-4d12-8dd6-b41804ca24cd · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51b1b8f9-be18-4d35-bc41-a649916768d0 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3b6777-e92f-4325-91a5-bf8b25fc6660 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Moviechat: From dense token to sparse memory for long video understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b47ab2da-1955-449d-998b-88df156c4469 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video understanding with large language models: A survey
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4359f59-bc8b-46fe-a718-628a619f961b · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f4c3484-0531-472f-ad6f-3aa4f348d22f · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Internvideo2: Scaling foundation models for multi- modal video understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e39c4189-926f-46bb-b906-7f1542fec27a · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e73ebea1-c1cf-4f09-8858-be514963c3a3 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d916726f-0327-4842-ae41-4fcc2571c33c · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Adaframe: Adaptive frame selection for fast video recognition
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4b2f519-101d-44cc-af46-640da2b1c2df · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 665bd746-852e-4a8b-a786-607ecfbe37a9 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Frame-voyager: Learning to query frames for video large language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb4d00fa-9cc6-44eb-bf77-de94c0c3682b · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Sigmoid loss for language image pre-training
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0e680ee-9a99-4961-857a-fd21ac5307dd · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d985fbf-4b73-481f-b255-2b737e68829e · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef8de3d9-0280-4974-bb2f-47853121def6 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Long Context Transfer from Language to Vision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a1f0d10-f715-47c1-b335-f47ffddde9c3 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c307f2e-7cdd-4c08-b992-c4fc22cea03a · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc18216-6abf-4dc7-915d-f115ba96b59f · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Mgsampler: An explainable sampling strategy for video ac- tion recognition
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f71645bf-2cfe-46ad-8924-4ae73dd3d962 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs MLVU: Benchmarking Multi-task Long Video Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eef322c-e13d-4329-90c3-8916f65c57a9 · outbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Limitations Q-Frame enhances query-aware video understanding, but it depends on pre-trained models, lacks explicit temporal modeling, and operates within a fixed token budget
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b6ad907-50e4-4dc3-bfa1-ecd47740152c · inbound
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33844b79-8431-4eda-bdce-8ca1e28b38dc · inbound
Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eacdb4a2-0578-4ee1-87d6-3cc48c57148e · inbound
Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 079f6c34-65d1-42bd-aab7-82261c6a427e · inbound
PEEK: Picking Essential frames via Efficient Knowledge distillation Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9498d452-63c8-44f2-8693-0f8a72695bb3 · inbound
Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de4ae604-e237-4b7a-9565-2c18234cc729 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9824f3dc-f80b-4c09-b06c-904f35ade8c6 · inbound
Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5051ba87-dc57-4d7a-946d-cbf469c36426 · inbound
QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.