Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T04:01:48.625937Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2607.23265.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T04:01:48.625937Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 44e2cbf1-01e8-4686-a792-913be99b683a · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdae0b6-6e02-4163-9122-2a3b1334ee2b · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Token Merging: Your ViT But Faster
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f6a29bb-70d6-4dd3-8d81-cd39c0c5e2db · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c3485c-74b5-4288-82fd-120e407d87bc · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Flashvid: Efficient video large lan- guage models via training-free tree-based spatiotemporal to- ken merging.arXiv preprint arXiv:2602.08024, 2026
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53dee95a-1806-44a1-b376-81e4ce862e8f · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Video-mme: The first-ever 8 comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 974bb5d5-894c-4430-9d4f-4be4e26c6499 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation ToSA: Token Merging with Spatial Awareness
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18bebb59-76ba-4594-94a4-67a8a572192a · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation GPT-4o System Card
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328529b7-f3d2-4427-8bf9-350335e6e3f6 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation LLaVA-OneVision: Easy Visual Task Transfer
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad61f7e-cc04-463a-a242-3cbd1623810a · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f210869a-2640-4a7e-9a62-7ea7291c484c · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8716d82-f30e-4be0-b6f8-33615a4b0f8c · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation VideoChat: Chat-Centric Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c47054be-7fe5-4001-a7a0-3c1377f61fbf · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Llama-vid: An image is worth 2 tokens in large language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21db5e46-752e-4f57-8a64-d2ded3633464 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Video-llava: Learning united visual repre- sentation by alignment before projection
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa1ab547-1060-4554-8727-6b75540b0809 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1b8d28e-7d83-408b-a3cc-124242cba75e · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f319d61c-aaeb-465f-a9e5-ab11449f5ee5 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation St-llm: Large language models are effective tem- poral learners
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eac30308-e9d8-4b64-866c-96827b195d1f · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Quota: Query-oriented token assign- ment via cot query decouple for long video comprehension
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aaadd5e-6331-453a-a5a3-c7d534396285 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural In- formation Processing Systems, 36:46212–46244, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80e4d298-80ab-4d3f-947c-2cc9deafe599 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867c1159-ffdf-499f-994e-29a705b326b0 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Holitom: Holistic token merging for fast video large language models.arXiv preprint arXiv:2505.21334,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd2233ae-0764-4622-ba01-0bc71e99655c · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3200ee9a-b65a-4800-9985-973813bc3e5f · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Fastvid: Dynamic density pruning for fast video large language mod- els.arXiv preprint arXiv:2503.11187, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d925bb5-be95-46b0-9e07-dc9989631636 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Dycoke: Dynamic compression of tokens for fast video large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2512aef0-09e6-40ec-a128-6c78c1d71e0d · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation LOOK-M: Look-once optimization in KV cache for efficient multi- modal long-context inference
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81fd0740-e503-445f-a570-13e1ead2dbda · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3d32d06-f585-474a-ab13-67d59a8b579e · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d40669d-0f7f-4372-8775-f43c0d696663 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Lvbench: An extreme long video understanding benchmark
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6474f9-9105-4521-aed5-33b16a18475c · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Informa- tion Processing Systems, 37:28828–28857, 2024
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cf8eaf9-3adf-41a1-aa63-a04b8662a820 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30edd1d2-1f31-40a6-a0b5-1034d98c124f · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Visionzip: Longer 9 is better but not necessary in vision language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0b34fd-c262-4970-970d-62920d28128d · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Vflowopt: A token pruning frame- work for lmms with visual information flow-guided opti- mization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d1800be-f5b8-4c39-90f5-c747fdf34328 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Wave-vit: Unifying wavelet and transformers for visual representation learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 558d0d5d-1581-409f-98fb-9543daec9110 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5631c45a-2fe2-4b42-830b-5889dc58e5cb · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c940e4a4-e27a-4e34-b139-3498d8d14cc3 · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca36352-5a8f-416b-ba86-639dfcbb76bb · outbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.