Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:40:23.771906Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 10 inbound Pith citation observations for arXiv:2411.15024.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:40:23.771906Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:08:09.432264Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T14:31:40.739574Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 530f79c1-4e59-4f8a-8606-9bd9662b71ca · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Token Merging: Your ViT But Faster
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cd564c9-5191-458a-aedd-52c0215a2c01 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a99899-602b-4928-917b-a694295f98bc · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df522aa3-0ead-41ab-be2a-79ddb14ba327 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95365fbd-939e-4e5a-a890-886f386ae174 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 662dffc3-0195-443e-ad69-8d197f7c7c51 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e3850772-27f2-49c4-a1e0-a3a152f19c32 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b7dab3-7667-4bb2-91bf-6d44d6c44b0b · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b194b244-12ef-4680-adde-3143b57ddb2f · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% perfor- mance
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bbf33c42-8533-40eb-b11c-1c4948d21924 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Masked autoencoders are scalable vision learners
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 56211308-d450-46b8-92cd-3d234d3530b8 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Lita: Language instructed temporal-localization assistant
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6bf8a6d5-43ad-4ec1-b6b1-ed857fabb90f · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models An empirical study of llama3 quan- tization: From llms to mllms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0d8fc5da-66ec-44b0-9c2a-62c868714354 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Phi-2: The surprising power of small language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8c8faeba-6d76-4ed4-aa97-f43027f3e992 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Effec- tiveness assessment of recent large vision-language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 890ac9fe-8c70-4b02-8a4b-12d0d1fe1eac · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Chat-univi: Unified visual representation em- powers large language models with image and video under- standing
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 80b32cc2-0d2b-4c0f-8c76-96ca4d5939a0 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Lmms-eval: Accelerating the development of large multimodal models, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444d4a0e-ab21-4579-958c-ccb442e8d808 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeb20722-42e1-4cb7-b13f-c73494ca640e · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca789b6-071b-4145-ba0d-ae457bde8cca · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cca5709e-d411-4a1c-baff-a9270c33971d · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Tp2o: Creative text pair-to-object generation using balance swap-sampling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ad9317d3-9ab5-4df5-877d-82cced281e69 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ac6aed7-a5d3-4faa-b167-7b296074fdfe · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 98ff61bf-30bd-421d-8bda-b5463da11204 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 686683fb-6050-47cf-a523-1f7a04717388 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40330da5-7f79-4122-a3dc-ed1e6da79230 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d988134-9f0d-49c4-acc3-8016818a5067 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Vila: On pre-training for visual language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a101bd37-0ac1-417c-baa6-8b59d8bce438 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80057844-2fee-4228-a585-4e8065661293 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Improved baselines with visual instruction tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f046d1-4fe9-479b-858f-621ffe97f0e4 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Visual instruction tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e7c4b62-9d33-4e80-aea7-1a1b8a3af018 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video detail caption, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c073fbc5-43d2-4f94-954e-6da67bfaee97 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75e6fbb-1254-4b8a-8ca3-76666289f524 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Gpt-4 model, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0aba2f3e-02b9-43dc-856b-01242661eed0 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Per- ception test: A diagnostic benchmark for multimodal video models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0499b5c-2f5d-412a-875a-3c0fa62174d0 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Learning transferable visual models from natural language supervision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 03bae0ac-1fbe-4342-b3e4-9420f2fece5d · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0dbfeb-3365-4a3a-a75e-4a19c42f8261 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7118a95f-454d-4822-be9c-bf4500cacae8 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models PB-LLM: Partially Binarized Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aea40c27-f9b4-47bc-8ff4-03947e4ffb37 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b10fefc-2f69-4e0c-ab17-c914a975ecb5 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a70519a-1c22-4609-9647-d7ff754b9db5 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video- mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0a677e7e-f101-400f-b5ad-a44a2a77e18c · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6822db5-7cbd-4ebd-8023-6aa5e8df054d · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3bac1f9-6d6d-4692-8fc9-d665278ec2a2 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Small Language Model Meets with Reinforced Vision Vocabulary
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation badc2f20-4a8f-473c-a9a0-017195bb0e1f · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Mamballie: Implicit retinex-aware low light enhancement with global-then-local state space
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 134df1f3-fcb4-491a-adcc-9379af1190e9 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Next-qa: Next phase of question-answering to explaining temporal actions
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d607b380-1909-462d-9095-53da1f0b70c4 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Robustmq: benchmarking robustness of quantized models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation be08aeb6-a7ee-4f92-bbee-39f6c09fb38e · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Novel object synthesis via adaptive text-image harmony
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a3ae0252-04b4-478c-b369-08b8ebd2b89b · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae70fbcb-4bca-4025-b71c-330e23fd0387 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fef99eab-0e7f-4355-80be-4bc14eb44e66 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d754c0a-2b72-492c-9eec-8d3682e7244c · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2d8b2c-e692-45ed-96e0-d939d4fe532c · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Lmms- eval: Reality check on the evaluation of large multimodal models, 2024
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 096daa0b-aece-40b1-b582-65cb437d3be2 · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc4c9fcd-7491-4696-ace4-b1513a46cc7c · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1056b463-59c5-4538-8142-48d35f621a2f · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models Llava-phi: Efficient multi-modal assistant with small language model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 542b4934-4004-4159-bff7-cf40f1ae2ebb · outbound
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models counterfactual reasoning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 07f9dd78-b7ea-4ce0-afbc-03b2027df9ca · inbound
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3f1acf08-3f3e-4d4d-bf9c-889c2570e0ce · inbound
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2792e61-63a5-472a-8d51-e05cbbedfa04 · inbound
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23f69557-35b7-43b7-9b7c-2d0d26d6a526 · inbound
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fdda23a-e379-4a88-8fda-78e5a5fba147 · inbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09cc8165-9107-402e-b976-e1c701189c49 · inbound
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2b08a16f-941a-4a64-8f19-744306d50b72 · inbound
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6c70ceb7-65d6-46d6-a34a-39eabd1dee2e · inbound
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e3d60966-5cf6-41f8-a06f-eb0f69c2434f · inbound
TTF: Temporal Token Fusion for Efficient Video-Language Model DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2dad76de-74ca-4d25-993a-6a3554dad9ed · inbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.