Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T06:37:33.282545Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2607.28463.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T06:37:33.282545Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b85b768a-13b4-4b7f-93f5-64695f20e3f9 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Video-ChatGPT: Towards detailed video understanding via large vision and language models,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb94bac5-d01f-4503-a2fe-744d3185aec7 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LLaV A-video: Video instruction tuning with synthetic data,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea62da92-f269-41ad-bcee-8fc920916d88 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LLaV A-onevision: Easy visual task transfer,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c38278-5b66-4c27-a4c0-7dfe9ba73f39 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dea72b5-709d-439d-9d01-f6f95e915957 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1157303-bd66-4240-ab16-66892b8deb5f · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9009200-6f30-4bb9-82d8-340c60ad6ba6 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LongVideoBench: A benchmark for long-context interleaved video-language understanding,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 021b3815-4d18-4ce1-8094-b13d61d8070c · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding MLVU: Benchmarking multi-task long video understanding,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46c9606-395b-442f-8570-7fd1f68bc4fa · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding SlowFast-LLaV A-1.5: A family of token- efficient video large language models for long-form video understand- ing,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e7a7f3a-19a0-40c9-9c1a-774f61299f3d · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Long context transfer from language to vision,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d1ad064-51f3-4478-97e7-c2463ed108ca · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Adaptive keyframe sampling for long video understanding,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0060e0c6-ac84-43e5-afbf-a2f61b50c93e · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Q-Frame: Query-aware frame selection and multi-resolution adaptation for video-LLMs,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53d1db39-be30-4927-a2ba-103836f8a624 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding BOLT: Boost large vision- language model without training for long-form video understanding,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec4378cd-1ac9-40e2-a7a0-18679656ba4e · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding MDP3: A training-free approach for list- wise frame selection in video-LLMs,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8524394-313c-4732-8c3b-06a91ccf5ab6 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding KeyVideoLLM: Towards Large-scale Video Keyframe Selection
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d408267-d69b-4a37-8902-1a9e53a87c11 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding MaxInfo: A training-free key-frame selection method using maximum volume for enhanced video understanding,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6659840f-9c13-4b3c-9844-5e3dee41d8f0 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding AdaRD-Key: Adaptive relevance- diversity keyframe sampling for long-form video understanding,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32732981-9dac-4e5e-b682-39fa7b19ab3a · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf471211-9d4a-4e9a-a4c2-231a20393fde · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1204b20-c776-4798-8fe7-6f10d97f197f · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LongVU: Spatiotemporal adaptive compression for long video-language understanding,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8088731-e9df-4580-9e07-dab1a9233c9f · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Flexible frame selection for efficient video reasoning,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b70ab4-dcb9-4bc8-99f1-2cc42e68a8b3 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Frame-V oyager: Learning to query frames for video large language models,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f811a44b-73e0-42e6-b383-f8d559a35a9e · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding M-LLM based video frame selection for efficient video understanding,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 640fca01-8386-449c-9aed-13b844b1a0ee · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Self-chained image-language model for video localization and question answering,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f0856bc-5b41-4f4a-bf23-0f2a705c0209 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Generative frame sampler for long video understanding,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78136fb-ea79-4290-acee-5f1f6fefb0fe · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding K-frames: Scene-driven any- k keyframe selection for long video understanding,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a45a8612-d07a-4653-8684-a7c7a3ea477f · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Event-anchored frame selection for effective long-video understanding,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5544f8-12ce-43b3-ac89-201b1cfa80a2 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Wavelet-based frame selection by detecting semantic boundary for long video understanding,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22fbb8da-d30c-463f-b493-f56d00a0e88d · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding The use of MMR, diversity-based rerank- ing for reordering documents and producing summaries,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 352f5502-243d-4268-9229-1230e2ffa605 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ce74f0-32f4-48b5-ad90-54912741af47 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding DINOv2: Learning robust visual features without supervision,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f27c1b9a-4834-48b7-8b62-00864158f02e · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Determinantal point processes for machine learning,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d92d894-6d94-46ab-9a1e-9e57c1fdec77 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding k-DPPs: Fixed-size determinantal point processes,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454b277f-3406-4a69-b501-b4562d78f71e · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Fast greedy MAP inference for determinantal point process to improve recommendation diversity,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a6d5c48-3ed5-460e-ab82-d798e829484d · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LMMs-Eval: Reality check on the evaluation of large multimodal models,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eba41349-aa3c-40b3-97e3-4978d4db6642 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Qwen3-VL Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 561d2acc-260f-4c6c-955c-ba1818d5b686 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ed70c2-851b-46dd-9b5a-cd8e2a53dc17 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LongVILA: Scaling long-context visual language models for long videos,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c71d7b-bc7c-46dc-bb34-a7b0fe49d5a9 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Video-XL: Extra-long vision language model for hour-scale video understanding,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdedef4e-de7a-4baf-8e13-0724e96f652b · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Learning transferable visual models from natural language supervision,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5fcc2c9-5ecb-4808-b93f-f6c5de58c6b8 · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Sigmoid loss for language image pre-training,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec06c51-aa56-44c5-803d-55c037d2c96f · outbound
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.