Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T22:32:19.740039Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2603.18558.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T22:32:19.740039Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d6738284-e9e6-4e0c-98d9-987e1d095e19 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9958e024-debf-4b83-a338-c1f63c36abe7 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae7a0ef-d2df-401f-a625-01de8a66f0c0 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering GPT-4o System Card
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a73728c2-cab4-4408-90c1-b73503e49797 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Learning transferable visual models from natural language supervision
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec097bab-a728-45fe-9453-927093126e24 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Sigmoid loss for lan- guage image pre-training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82487c28-2fca-440b-933c-74a95e60b8ea · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Bolt: Boost large vision-language model without training for long-form video understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c968043f-7be3-4680-b77c-f8119f3fb747 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Adaptive keyframe sampling for long video understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8832c725-0f15-4777-a70a-0ff31e2dfc81 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Mdp3: A training-free approach for list-wise frame selection in video-llms
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2142abf1-54ca-4439-8c8c-bc4c2c3bab69 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Videoagent: A memory-augmented multimodal agent for video understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44b23e5-5622-4de7-a18b-873e22040413 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Videoagent: Long-form video understanding with large language model as agent
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa88792c-7a4d-46b3-8c14-5fae0adfbd1a · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering LVAgent: Long video understanding by multi-round dynamical collaboration of MLLM agents
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a805de4-dfdf-49e6-a27a-88656248cfd5 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Longvideoagent: Multi-agent reasoning with long videos.arXiv preprint arXiv:2512.20618, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fb535fc-be98-4885-b1f6-98241c05d648 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering SeViLA: Self-chained video localization and answering via llm
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d591e8a-fccd-47ef-838b-c877caf6de4b · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering A.i.r.: Enabling adaptive, iterative, and reasoning-based frame selection for video question answering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ccf10fa-0efe-452b-a796-0381f79a7deb · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Video-MME: The first-ever compre- hensive evaluation benchmark of multi-modal LLMs in video analysis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3cee4a0-0c8d-46b1-8d66-3921a4d4b07b · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering LongVideoBench: A benchmark for long-context interleaved video-language understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f3aa5d2-7c05-42fd-a2c7-ae2e4072ed79 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering HERBench: A benchmark for multi-evidence integration in video question answering
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 750a6999-5e16-44f1-81ce-fbf84f95c254 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Frame-voyager: Learning to query frames for video large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fcb63c7-1a55-439b-8732-800df64e37c1 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Flexible frame selection for efficient video reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29a2740d-4d32-437a-b3cb-e982ca1d89de · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering M-LLM based video frame selection for efficient video understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f949c76f-ee69-4056-8799-6da396f15b08 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering End- to-end videoqa with frame scoring mechanisms and adaptive sampling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6b4837a-04c2-4c22-973f-12c158a2ccaf · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Re-thinking temporal search for long-form video understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4b392d7-061c-422f-952c-cd1ba676b666 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering YOLO- World: Real-time open-vocabulary object detection
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97377aef-fbd8-4043-a728-e47ec53a17af · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Logic-in-frames: Dynamic keyframe search via visual semantic-logical verification for long video understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ec8ec76-e740-47e3-aa28-e1bfb2158534 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Neus-qa: Grounding long-form video understanding in temporal logic and neuro-symbolic reasoning.arXiv preprint arXiv:2509.18041, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c31ab5-6778-49ac-9a9d-98c233cfaedf · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d4120b4-b2af-43c9-a419-e20f7f227502 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering VideoTree: Adaptive tree-based video representation for LLM reasoning on long videos
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbdb37c8-e94b-4ecc-9760-a13f1f585eb2 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Videozoomer: Reinforcement-learned temporal focusing for long video reasoning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94d446aa-3605-4c2b-94cd-b1c4a5e75e73 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, and Vikas Chandra
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e10ab2-780e-48ce-869b-371687cdf43c · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b44f30f-fc35-4945-8818-6d42b7970cf9 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering docTR: Document text recognition, 2021
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc03578b-c0ae-44cb-9aee-0c1352a0eef1 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Robust speech recognition via large-scale weak supervision
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23808f1-337e-46b7-b49b-431f687413ac · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Data filtering networks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 549a1198-f88d-47f5-b6f6-baf0c119bc92 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering VideoChat-A1: Thinking with long videos by chain-of-shot reasoning.arXiv preprint arXiv:2506.06097, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a41949c-1634-4641-89f0-d4e9a82a8f00 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Qwen2.5-VL Technical Report
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d99c68ec-9514-4525-983e-998dde4490f4 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb763ee-3ff5-4972-8ff9-7fcda555b71f · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering EasyOCR: Ready-to-use OCR with 80+ supported languages, 2020
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1947a29a-bdd4-4bc0-a9d5-acd628fd37d9 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d36972-f3f5-4954-b87a-57bf0651a701 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed63f80e-9304-436c-81bf-8eb3892c6e6a · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a163c3-2c02-47c7-aae1-136e7241d4cd · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering person
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea5ee306-92ae-473d-8226-735d3ac467fb · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Exit " ,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d494086-659f-4af9-a962-d9424631e46b · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering person sp ea kin g
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7802363-5cff-4fb9-ae35-de30c1b60986 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Add an ASR leaf with related spoken ke yw ord s a l o n g s i d e visual leaves
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e21d6f-ac7b-41cf-998a-175e972b21a1 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering do or bel l ringing
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c799bf1-e1b7-4d68-abad-8d430ca62ab1 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Never make a tree with only one expert type .] [ ELSE : Use mu lt ipl e visual experts when po ssi bl e .] [ IF_ASR ]
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41c5d913-dee5-4753-a11c-67553c717f42 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering [/ IF_ASR ]
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65aa8b31-bb53-4af4-b63d-f8bcc1f5eacc · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e02f608-dea1-4cbf-9ae9-5b57503d0957 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d2ff39d-a7e1-4033-8e21-c130568ed575 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31a53722-8804-4dd9-aa85-22d0ae6f6f60 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df2777b-14ab-43aa-bfba-16145fb8d74c · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee2a23d-d4f9-43e0-87b5-577f30c086b7 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Same " ,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3713b8e7-a3be-43bf-b0b1-6e35b97742ee · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 745758b3-123c-4406-9028-eb863821327b · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b438ece-e30c-47d2-acc5-65551b996fc6 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering [ IF_ASR ]
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df87a3ae-98a2-4c0a-aebe-53e14c050d6e · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11798827-8f08-468e-9e10-9e2750afe87a · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d0e814-6060-459f-a824-7f1670ead1b1 · outbound
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering op ": " AND
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.