Pith. sign in

Paper Citation Record · LEDGER

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2607.23265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23265 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:01:48.625937Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44e2cbf1-01e8-4686-a792-913be99b683a · outbound

This paper cites Qwen2.5-VL Technical Report.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.498396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.498396Z digest=sha256:708734de537c430fc1b5a96a0ff38668a4b610090c4d0c2bcde2f8281eced000

Observation 1fdae0b6-6e02-4163-9122-2a3b1334ee2b · outbound

This paper cites Token Merging: Your ViT But Faster.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Token Merging: Your ViT But Faster

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.503555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.503555Z digest=sha256:89f63e9fafd408212d36278a80c5168c5705233071315efa1ec082c5d76481cc

Observation 5f6a29bb-70d6-4dd3-8d81-cd39c0c5e2db · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.507558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.507558Z digest=sha256:43d2a9c1212293b493da69a1d27f73ac08d5fa89dd66fc5c9a9b90bc88b7af7c

Observation 78c3485c-74b5-4288-82fd-120e407d87bc · outbound

This paper cites Flashvid: Efficient video large lan- guage models via training-free tree-based spatiotemporal to- ken merging.arXiv preprint arXiv:2602.08024, 2026.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Flashvid: Efficient video large lan- guage models via training-free tree-based spatiotemporal to- ken merging.arXiv preprint arXiv:2602.08024, 2026

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.511212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.511212Z digest=sha256:c4fe6d35444795f890720afddbaaba919b6631e59bf52e27d7b9647d230e71fd

Observation 53dee95a-1806-44a1-b376-81e4ce862e8f · outbound

This paper cites Video-mme: The first-ever 8 comprehensive evaluation benchmark of multi-modal llms in video analysis.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Video-mme: The first-ever 8 comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.514807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.514807Z digest=sha256:b012ed408433ac43299fd9c76b2194b3f43003880325e7cd8a5195f4c0129019

Observation 974bb5d5-894c-4430-9d4f-4be4e26c6499 · outbound

This paper cites ToSA: Token Merging with Spatial Awareness.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.518353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.518353Z digest=sha256:dd72bce795847dd0eb1bafe04415d3f851619c7d3fee41c4c67beac2e4a1fc79

Observation 18bebb59-76ba-4594-94a4-67a8a572192a · outbound

This paper cites GPT-4o System Card.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation GPT-4o System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.522434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.522434Z digest=sha256:b964efc5daf254e4a705caf41b38266516b1eb4518d199eff6d1de1e728f8f8e

Observation 328529b7-f3d2-4427-8bf9-350335e6e3f6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation LLaVA-OneVision: Easy Visual Task Transfer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.526036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.526036Z digest=sha256:64b6ca0a4a142e7b6f616814c21c417169073e39d6f77585e4f65e2ee19ec0a9

Observation 0ad61f7e-cc04-463a-a242-3cbd1623810a · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.529655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.529655Z digest=sha256:202bdb09891ef8df422b1abaee491565e3e3cd1c02fd200ae7cd03d7326f7ed6

Observation f210869a-2640-4a7e-9a62-7ea7291c484c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.533350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.533350Z digest=sha256:6ce1405097f7c0263e113355e508c771dc5a9e3813b2056c81399b4b924d550c

Observation b8716d82-f30e-4be0-b6f8-33615a4b0f8c · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation VideoChat: Chat-Centric Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.536689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.536689Z digest=sha256:810b3fa2df87988f2da27ed3ed2c12a6e1f882852d0d42a8fff1585df84b1633

Observation c47054be-7fe5-4001-a7a0-3c1377f61fbf · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Llama-vid: An image is worth 2 tokens in large language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.540425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.540425Z digest=sha256:42ef7a84235139fb94ec6db2fc99c96dc154a627e509c57c9fb16641e26be984

Observation 21db5e46-752e-4f57-8a64-d2ded3633464 · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.544682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.544682Z digest=sha256:0dd6f844ee021503271882fa1a18befc8275db2533e6e15c63890aa4fe7d6b1c

Observation aa1ab547-1060-4554-8727-6b75540b0809 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.548091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.548091Z digest=sha256:c3f807071232c39ab6a2e4f195429bf69ed18f66934eb233f471ddeb04e956a1

Observation b1b8d28e-7d83-408b-a3cc-124242cba75e · outbound

This paper cites HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.552316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.552316Z digest=sha256:53ad2af3c1ea0c81f45f82daca31d2357cc17854b773670f81e151c038d965bd

Observation f319d61c-aaeb-465f-a9e5-ab11449f5ee5 · outbound

This paper cites St-llm: Large language models are effective tem- poral learners.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation St-llm: Large language models are effective tem- poral learners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.555827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.555827Z digest=sha256:9127ed3565b4c0ac269f2df3fb944d4fb07f8006636b5462d6e5a97dedf7ab64

Observation eac30308-e9d8-4b64-866c-96827b195d1f · outbound

This paper cites Quota: Query-oriented token assign- ment via cot query decouple for long video comprehension.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Quota: Query-oriented token assign- ment via cot query decouple for long video comprehension

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.559642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.559642Z digest=sha256:04393f98eeecfa08cbaca479af7538ff9d7d3e7db0de797dff06cde3b33c5e78

Observation 1aaadd5e-6331-453a-a5a3-c7d534396285 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural In- formation Processing Systems, 36:46212–46244, 2023.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural In- formation Processing Systems, 36:46212–46244, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.562981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.562981Z digest=sha256:1f50670fcbcd57a7e2d882d48055b8f0f60584a83476db8203edb04b234671a2

Observation 80e4d298-80ab-4d3f-947c-2cc9deafe599 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.566302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.566302Z digest=sha256:d1246d90ab90ab6a01546000e271e8eee67e8eb36f786ec708a3de691e247bff

Observation 867c1159-ffdf-499f-994e-29a705b326b0 · outbound

This paper cites Holitom: Holistic token merging for fast video large language models.arXiv preprint arXiv:2505.21334,.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Holitom: Holistic token merging for fast video large language models.arXiv preprint arXiv:2505.21334,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.569757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.569757Z digest=sha256:7908a20178b247e160192e92f963ac7c80f1832bebdfee944731912bd4c37068

Observation bd2233ae-0764-4622-ba01-0bc71e99655c · outbound

This paper cites TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.573191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.573191Z digest=sha256:01fa80f643b9bcd82705b9cb059480e38148bfdc41edf3d074c8619e69adbf2a

Observation 3200ee9a-b65a-4800-9985-973813bc3e5f · outbound

This paper cites Fastvid: Dynamic density pruning for fast video large language mod- els.arXiv preprint arXiv:2503.11187, 2025.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Fastvid: Dynamic density pruning for fast video large language mod- els.arXiv preprint arXiv:2503.11187, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.576697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.576697Z digest=sha256:e93058f78f997129cdf115612edec37958ed9b1132266712028ef6454ada94c5

Observation 9d925bb5-be95-46b0-9e07-dc9989631636 · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Dycoke: Dynamic compression of tokens for fast video large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.580079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.580079Z digest=sha256:42a9df89f050949d204c23113e1a0304811d7ad6adcc4b71bdbfe767d60244e0

Observation 2512aef0-09e6-40ec-a128-6c78c1d71e0d · outbound

This paper cites LOOK-M: Look-once optimization in KV cache for efficient multi- modal long-context inference.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation LOOK-M: Look-once optimization in KV cache for efficient multi- modal long-context inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.583336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.583336Z digest=sha256:3a50741241595e05e8b01beeb2d8dabb1fde465866d99c8574486e75b1e35357

Observation 81fd0740-e503-445f-a570-13e1ead2dbda · outbound

This paper cites Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.586722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.586722Z digest=sha256:e2779e50b9cdd98f81cedcf263acfdc0880c2d018a18b5f51b0bf7c4d005f3f4

Observation a3d32d06-f585-474a-ab13-67d59a8b579e · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.590655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.590655Z digest=sha256:14324510fe1f26d65730bd65302ff82c02fe0058ad4c7d2323f6cae96be07eec

Observation 7d40669d-0f7f-4372-8775-f43c0d696663 · outbound

This paper cites Lvbench: An extreme long video understanding benchmark.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Lvbench: An extreme long video understanding benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.594179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.594179Z digest=sha256:f12299d6002d4a6e56e68a11429b581f22764157583d279a0d357453a6413393

Observation cc6474f9-9105-4521-aed5-33b16a18475c · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Informa- tion Processing Systems, 37:28828–28857, 2024.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Informa- tion Processing Systems, 37:28828–28857, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.597987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.597987Z digest=sha256:02ee3cc0ad00f0327735cc779d8ca798fd83524faf1a05836439ee8bdf9ae29c

Observation 5cf8eaf9-3adf-41a1-aa63-a04b8662a820 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.601334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.601334Z digest=sha256:57bb5a631cb1871adc093242feab99947b73b8c56ce4372d2643433365afa58c

Observation 30edd1d2-1f31-40a6-a0b5-1034d98c124f · outbound

This paper cites Visionzip: Longer 9 is better but not necessary in vision language models.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Visionzip: Longer 9 is better but not necessary in vision language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.605144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.605144Z digest=sha256:2b4d875d413bcfbd83e185d9966f61f3d2da581f7f84aa5b1e12e15d69829372

Observation fe0b34fd-c262-4970-970d-62920d28128d · outbound

This paper cites Vflowopt: A token pruning frame- work for lmms with visual information flow-guided opti- mization.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Vflowopt: A token pruning frame- work for lmms with visual information flow-guided opti- mization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.608428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.608428Z digest=sha256:fa074f3c93b6f2b518f6c71de7049205d9a18ae820d8ff14b2b650da73c63014

Observation 7d1800be-f5b8-4c39-90f5-c747fdf34328 · outbound

This paper cites Wave-vit: Unifying wavelet and transformers for visual representation learning.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Wave-vit: Unifying wavelet and transformers for visual representation learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.611796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.611796Z digest=sha256:a4f1e89223ce1f45e857fbe41d58e20a84f3445fc4eb5a449eec84988507fb05

Observation 558d0d5d-1581-409f-98fb-9543daec9110 · outbound

This paper cites A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.615080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.615080Z digest=sha256:13b70f4d9697660ff934bedca16e9438b50527a65576eb5fd0413f9614164e45

Observation 5631c45a-2fe2-4b42-830b-5889dc58e5cb · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.618509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.618509Z digest=sha256:c48a7bd54109039dca931c011734d1b74dd6526e101b4d919e7bf625ecf515f2

Observation c940e4a4-e27a-4e34-b139-3498d8d14cc3 · outbound

This paper cites Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.622068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.622068Z digest=sha256:b3165297515ae27b5c8e3dd5513b03ef44a02fb82056384fc465de46051f2b5f

Observation cca36352-5a8f-416b-ba86-639dfcbb76bb · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-04T04:01:48.625937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.625937Z digest=sha256:917c4b4ac7cece5d86e5a72b7ee0b3ee1e6d355660e60990a30d830ad2b570ee

Pith citing papers

No inbound Pith citation observations are available.