Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T18:09:59.236030Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 7 inbound Pith citation observations for arXiv:2603.01455.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T18:09:59.236030Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T15:27:46.664630Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T10:58:03.480373Z
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cc156a03-de17-4bd0-8e80-bba526553996 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 664a6953-e737-4c58-a9e9-b3faf37c7571 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation de650f8b-18bc-46f8-bf4b-94dfd41637c4 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Videominer: Iteratively grounding key frames of hour-long videos via tree- based group relative policy optimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2a28a8ff-da50-4c82-87a3-e1e5b4e4fcf7 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 24ec224f-bd4c-47f2-b2fd-670a6ac9458a · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Towards large language models with human-like episodic memory.Trends in Cognitive Sciences
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 65645373-9018-4f07-9609-564a85ed5cf2 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c9b382f0-5f33-433b-b3b0-7ed2fe5477cf · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8fbf239e-910b-4428-b273-4da57db7d6ba · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Inducing high energy-latencyoflargevision-languagemodelswith verbose images
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e5a71de-842e-4006-8d89-b03d77d2a21a · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Ma-lmm: Memory-augmented large multimodal model for long-term video under- standing
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1212c5af-ecda-4b89-bcd8-ef90209d5865 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents View from the top: Hierarchies and reverse hierarchies in the visual system.Neuron, 36(5):791–804
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e2605757-1e64-4e5c-a073-994223232c67 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Lightweight and cognitive agentic memory for efficient long- term interaction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 24f5bf25-7579-4845-9bfb-692322018610 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents GPT-4o System Card
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f92fd66a-635c-463a-acec-78abb64e8fa4 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Chat-univi: Unified visual representation empowers large language models with image and video understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9c12cb24-b5e0-4cff-88cd-342167ecb798 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 247fea39-5a72-4225-bbb3-d20fd3775268 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large lan- guage models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 39938b33-4bf6-4e1a-8b85-7a39d71390ab · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Llama- vid: An image is worth 2 tokens in large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8d334c04-4072-44ba-a6e6-53b6f39117b7 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Video-llava: Learning united visual representation by alignment before projec- tion
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a4ec8e0f-80e4-48f5-aefc-54e53d20ed84 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Visual instruction tuning.Advances in neural information processing systems, 36:34892– 34916
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 72c29613-599b-4000-a140-8b60cfd7b9eb · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Seeing, listening, remembering, and reasoning: A multi- modal agent with long-term memory
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f0a909e8-7e1e-47a4-bdf8-799ecaa1a65a · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents arXiv preprint arXiv:2411.13093 , year=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a3d7dae4-f278-46bc-ad08-e2d3849bc656 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a97c1720-f839-40f2-924c-2f4121c7565e · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents M3-embedding: Multi-linguality, multi-functionality, multi-granularity text em- beddings through self-knowledge distillation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 771debdc-c122-4e95-a4a7-a4e8c4991b56 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Memgpt: Towards llms as operating systems
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49364b86-735e-4834-baf8-9e4ec88862df · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Hd-epic: A highly-detailed egocen- tric video dataset
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3bd30bc5-7c1e-4930-b98d-59144df894e2 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Streaming long video understanding with large lan- guage models.Advances in Neural Information Pro- cessing Systems, 37:119336–119360
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a39f7bc-7546-4152-80e0-510cde6ab1da · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Fuzzy- trace theory: An interim synthesis.Learning and Individual Differences, 7(1):1–75
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6935991b-0208-474b-9d5b-a6bcf70e3bbe · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Vgent: Graph-based retrieval- reasoning-augmented generation for long video understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0148f92c-dc6a-4ec8-9ff4-e9345b7b3c4c · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Moviechat: From dense token to sparse memory for long video understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b3b003ad-ebba-447e-ab18-7af2c0aad524 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cc9c67ac-c6d2-42bf-bf95-e3816e82a5b4 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents The information bottleneck method
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1f8f433f-4812-4bc4-8563-c4b538c5550d · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e6b9451-beb0-4ed6-8a1f-14600b7a3fea · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3d023001-1c3d-4ded-b914-42f940f2a27b · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Videoagent: Long-form video under- standing with large language model as agent
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 849b962e-fa03-47a8-a029-0688a4c5f44b · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7cdf5f1f-6e6f-439f-a89b-abbc27e4c6a9 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Videotree: Adaptive tree-based video representation for llm reasoning on long videos
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6530a7df-536e-4c8b-a1f4-e3f07abfd18a · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents C-pack: Packed resources for general chinese embeddings
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2a99e6f8-c524-448c-96a2-ce1119528069 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Large Multimodal Agents: A Survey
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ba2968c-dcaa-4206-a256-cb8e7eee4cc9 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents A-MEM: Agentic Memory for LLM Agents
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5cab4dc2-71de-45ad-b4fd-c21a3270fe6e · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 172309e8-90f8-42e5-8eea-cb96806110f7 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5bc8b24f-ac61-4f4b-a48e-e571ca4389d6 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1862eb9f-6e6d-4eac-9e5b-43713013f906 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Long Context Transfer from Language to Vision
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fc7ef0be-1d40-4e91-b4a3-82756e7975b4 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0abd9925-a65c-4c13-b7d5-788292e8eb20 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Memorybank: Enhancing large language models with long-term memory
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3435fd34-5232-4c8e-ad6e-a0f0247dce24 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Mlvu: Benchmarking multi-task long video understanding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1f98cdf6-5093-482e-9e24-ce409a1251b8 · outbound
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Select the best answer to the following multiple-choice question based on the video.\n
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3b0736f9-0c87-4c2f-8e15-64cb654a8117 · inbound
Audio-Oscar: A Multi-Agent System for Complex Audio Scene Generation, Orchestration, and Refinement From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ecb2deb6-c7e5-4163-9018-55b689de3521 · inbound
MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2f4d6e20-97a5-4bdd-a9ae-d16db14bc916 · inbound
Xiaomi-GUI-0 Technical Report From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cfa67b57-089a-469f-9130-3974647e9a7b · inbound
Xiaomi-GUI-0 Technical Report From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation baf01ff1-873c-48ed-839b-1b9b7c85d152 · inbound
UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17f89242-abae-4e70-a51f-d88a8812ddd7 · inbound
FOLIO: Focused Semantic Memory for Streaming Video Understanding From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f54c517-b46d-4551-8608-2ad739758dec · inbound
DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.