Pith. sign in

Paper Citation Record · LEDGER

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 4 inbound Pith citation observations for arXiv:2507.07990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07990 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:50.801112Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.654336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T06:58:05.988167Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 246b891c-bda5-48f4-b44b-df2299f204f7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.920094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.920094Z digest=sha256:7670c493cd4a415118f3fcbe1852c3f37a42e819d646ef6442a3564cebbd7605

Observation 7d0f0eb4-1003-46b9-83f2-6d8fd01db424 · outbound

This paper cites Token merging: Your vit but faster.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Token merging: Your vit but faster

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.444166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.007491Z digest=sha256:28e1083a3e512d2229c0d63f77066e4b7b276b0c601726dcd31c4ebc7d8419bc

Observation 1d098822-e46e-43ba-98be-c0a9f7a2e67f · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.136591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.136591Z digest=sha256:69bf97b12683d1bfcf4fa0e8524898dc7f8cd58f9d6c2e18c7d705e91624e446

Observation 0670cd7b-274a-459c-8110-fa130e3564d1 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.411633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.252215Z digest=sha256:77801852fd5c446f6ac4556795dff33e5d7b8725ad94f53b4e0dce7cd9447aec

Observation 19b8f3b1-b4e2-4330-b2cc-7d25e356367c · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.392591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.386772Z digest=sha256:55fe700aca36ca8db6e4f915753fee1ffe4ecd2ed4551042cfcf544d308a4938

Observation 37f2bc3d-cab8-4313-b194-437375aacf32 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.374669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.517307Z digest=sha256:cb5d5db430356a5b64b8cf465fd8a7e93a94768b341682aa081ce0ace3954384

Observation b7c8df05-599c-413e-b4df-e5ea529487ce · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Longvila: Scaling long-context visual language models for long videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.352568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.618165Z digest=sha256:8888b664c77c65dbed455b3771619f7e9fe903e36d823360440892280622bce7

Observation e62eee1b-61e5-463a-9b3c-2e7f17199433 · outbound

This paper cites vid-tldr: Training free token merging for light-weight video transformer.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs vid-tldr: Training free token merging for light-weight video transformer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.331836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.710034Z digest=sha256:d70c66d41f1601cfeb8248a4d0546d6d03f7204bdb5146f90859ac42696643bb

Observation 5e573a74-dfd0-4e4c-b898-554117a8c69a · outbound

This paper cites FlashAttention-2: Faster attention with better par- allelism and work partitioning.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs FlashAttention-2: Faster attention with better par- allelism and work partitioning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.306539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.804392Z digest=sha256:628d268cfba27736082b1a6344d7128770456f2c59a6271b294564c7a24e6f2d

Observation be0085a3-55d6-4f57-a88d-a988eef926be · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.284441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.866227Z digest=sha256:d62d3539b466202af7ecf950904ceca21f03e3914d389e62c9b2682deb30523b

Observation ec769102-d618-4f13-a71a-79fbae6b2218 · outbound

This paper cites The Llama 3 Herd of Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.999654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.999654Z digest=sha256:be6865592a891a28b03f70be54c40ee3debc700b36fd33f343448160f4ace282

Observation 2f038acb-cd79-4778-b467-fe202eed1361 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Masked autoencoders as spatiotemporal learners

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.265252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.168771Z digest=sha256:ff6b10cb8e794017b6286370e6ad3ee2233a2118f76a08190f339c10da72996b

Observation 77307511-04a3-486b-80c8-5018115eb2da · outbound

This paper cites Quad trees a data structure for retrieval on composite keys.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Quad trees a data structure for retrieval on composite keys

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.245530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.273906Z digest=sha256:e3de27bb39507cb10d9b12d5bb174a95f7ca29326441bdba478bbfd4565d3cc6

Observation 86f94159-5ea6-41b5-a195-dcd31f95a052 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.222087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.372576Z digest=sha256:2a7731adab985fa4b955b12fc8c3e3c2676b2188743c260e9999c31034305bb4

Observation e842939a-c212-4e01-b7fd-dfc93bcc6dee · outbound

This paper cites FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.455680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.455680Z digest=sha256:158dfa175a60496279e9d59dfecffa6d9d3022299c7cd37820e37dbb038b647b

Observation f045812b-d2eb-4998-a0d9-76435809e32a · outbound

This paper cites Caching — google ai, 2024.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Caching — google ai, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.204484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.509746Z digest=sha256:0ad08b6b6262813be9471245db92c1cf522ee3bf68906c45599b3b9da9338ee3

Observation dcec1c7b-82b6-43cc-bda6-1910dc2673e5 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.642984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.642984Z digest=sha256:305d24c5a69a7f10eb8c63a59e4567312a015165a2159974eb7b1fd8f141bda5

Observation e5902d98-fde2-470b-9f0e-c750d81456e5 · outbound

This paper cites Deep residual learning for image recognition.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Deep residual learning for image recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.759667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.759667Z digest=sha256:f963bcaba0626973bf61543529bcbea2ff732172b5d68b0e0e85c6233e461e8a

Observation 0d4efc77-4d15-4bd9-85ff-deabb6392d3a · outbound

This paper cites Masked autoencoders are scalable vision learners.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Masked autoencoders are scalable vision learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.171666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.855039Z digest=sha256:5b7da5c087ca21837221c21a970ebfadff4aedc0467201ca0281cbfb88b81103

Observation d544c777-c601-45fb-a23e-d49639e58b38 · outbound

This paper cites PruneVid: Visual Token Pruning for Efficient Video Large Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.995554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.995554Z digest=sha256:16271ce07468bb16049d90139bf3703b0203a5afeeeb59cf129377b795a99121

Observation 39e5f196-45bd-42c2-a28a-cb5acb60cc83 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.148978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.097387Z digest=sha256:af10f2b2cf692b0046b7440b3fbf914d6b28354ea5194622bbe45834a348ae28

Observation 8d0fd6b9-ffdf-4105-9e21-fbf443daf535 · outbound

This paper cites Needle in a haystack – pressure testing llms.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Needle in a haystack – pressure testing llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.130345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.206402Z digest=sha256:5376cd05fc60410977066c2bab8f2b559eb651ac0b3f50daae7dcf59ad707236

Observation e7b38e68-d904-4f5a-b31d-54e76909e05b · outbound

This paper cites Handwritten digit recognition with a back- propagation network.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Handwritten digit recognition with a back- propagation network

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.113062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.301474Z digest=sha256:575f57d550c9bf13403843a02d5c0e6ecb21647fa1ca1420c727924b7d1ff36e

Observation 85685f60-2fb4-4cc9-86df-c99a101a9a0c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.367536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.367536Z digest=sha256:f8384797bb7b0784e7b0d88fa542c0202e8e6629458097cff097fe7c0c280ff4

Observation 8b0f4a29-485c-4221-ac3b-4e0064bc0f6f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.081714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.470093Z digest=sha256:3c1ab9b10d3377be9d2b0cc6b777dfcaea80d84114478c859312f0eb717fb438

Observation 3c22a84a-1f9d-4894-bd7a-ddec8a09b6be · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.575406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.575406Z digest=sha256:d728ac2f5d8a03a57e33cbc4029dd04724ad4f687abe29d7a3db1d9b2992ede5

Observation e2e21bf8-0ea7-422d-92da-9f74675e0312 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.063042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.618941Z digest=sha256:8566a09d2e65810f46e22dad4524ec6fc57d40cd35c99aa33c5b34cfe0a3d12a

Observation 77420b45-60dd-4a2a-ae06-5c850c902b2e · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.042764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.732124Z digest=sha256:f67b1c86c33804e9edbb3c6664991e7fc54853ef294025ca783a8c050780b987

Observation bf4da92a-b536-4555-9aed-2d58320ef6c8 · outbound

This paper cites Visual instruction tuning.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.017634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.844102Z digest=sha256:60a737f0da45aa0aea713a238ec1321f874209b4cbb927dc8276e640e1f487d5

Observation 1e5324cc-5866-4d60-a22a-132504cbcc9d · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.962155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.962155Z digest=sha256:b0e97433541f736166f5fa611ae427e8f0644db263119e09de9a49bd80d9183c

Observation caeac0d6-d89a-49eb-b1f9-09ba2d3d5223 · outbound

This paper cites Ring atten- tion with blockwise transformers for near-infinite context.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Ring atten- tion with blockwise transformers for near-infinite context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.989405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.093263Z digest=sha256:d1e144667c3d8e6e2068e80427cb36e407a6ba1277036966e279c95f8e024962

Observation ed03f2a1-23f5-4a19-b587-574980db0d02 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.187988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.187988Z digest=sha256:28e91054b44f88a1a4eab8b6de24927b24bc014684734bdb167c8277aeaa601b

Observation 16437856-d4e5-4593-8847-7fa463692fc7 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.947353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.260745Z digest=sha256:3a5a2c753f52dcc1b31cece578ef6250d516fffa0a231eec0333dba2f26ced44

Observation 4b8f8318-4e79-4561-9f1e-926a1142560c · outbound

This paper cites Efficiently scaling transformer inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficiently scaling transformer inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.915585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.312111Z digest=sha256:2e0ad3acfa4d7c59c3ca26048516dac0cbdfa5b1acc31497648f5423d6e54320

Observation 687cec00-dd84-4cf7-ad26-d5f4fadf6b97 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.887336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.362612Z digest=sha256:8075f0d531819d528bb4cd8a052d43db523f96d133612a028b9fdc4d71463a93

Observation 83a72fcd-4822-4217-91f4-a59a1769c279 · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.415433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.415433Z digest=sha256:3629d6c290e2495aeabb928a879ed2e7fcb3d5d45b937a58c3340d0b4a9aaf0f

Observation 88fcb5da-9be6-487d-b531-a752d5631b86 · outbound

This paper cites The quadtree and related hierarchical data structures.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs The quadtree and related hierarchical data structures

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.860677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.491626Z digest=sha256:8fa07803d29a081c2342bd517772049108825b92621ae3d48e6d635a4df6c3bc

Observation bbe2353b-5f2b-47e9-b11b-23ed357065d0 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.573421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.573421Z digest=sha256:ba8537e5581defdf634b45bb4564200a4857059f2e651665f15fcc2ae22e497b

Observation 7ee7ef49-1ddd-424c-848b-70815a92fcd9 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Moviechat: From dense token to sparse memory for long video understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.830881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.647081Z digest=sha256:284746b6f2270f7e5d8ce47cee402f3bab6b238bed2aeaa7d461fad2553bf7df

Observation afe23cbc-779d-4373-be94-3f1a48a468b9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Roformer: Enhanced transformer with rotary position embedding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.725013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.725013Z digest=sha256:1d3865f158624288acf53ae55d93e03c49d76772f3014ddd649d8d74b05e67ed

Observation 1026ae0b-c4e0-46ae-a06a-306d747fec8e · outbound

This paper cites Efficient quadtree cod- ing of images and video.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficient quadtree cod- ing of images and video

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.786169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.768904Z digest=sha256:d7acd10bea3d442fba076c20165050e6093b23edd59004e8fcd16fee1c444099

Observation 827c1b70-8999-44b3-845f-e9a464a8003c · outbound

This paper cites Overview of the high efficiency video coding (hevc) standard.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Overview of the high efficiency video coding (hevc) standard

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.766030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.839547Z digest=sha256:4ea34c8215823ed34b93b571cf3dd819b1cec9b33e9cde8debc6c42656a00294

Observation 1e0ee72e-7428-48f7-b981-87a41aa44524 · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Dycoke: Dynamic compression of tokens for fast video large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.739311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.887333Z digest=sha256:75d665edf62c860d8a2159bd45b100857d6b08299579107531bc11db01d1d22a

Observation 1928e642-4168-443d-9c79-469fbc97316f · outbound

This paper cites Efficiency of a good but not linear set union algorithm.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficiency of a good but not linear set union algorithm

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.619484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.958480Z digest=sha256:0294d543c03fd4080872ad647baac44a421acba8815b79045ae1606f77720f13

Observation b26d5fc4-c105-4f25-9614-c9a0da4e329b · outbound

This paper cites GPT-4o System Card.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs GPT-4o System Card

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.021249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.021249Z digest=sha256:ddbbbf05ce2db0bc8a3f2d3f03256d6aada51d9f85d86b0fca6b20e0a10e8a83

Observation c7ba62c1-f17b-40e1-b81c-0cf3e1fadad7 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.511623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:49.081969Z digest=sha256:10e64b6ba1bf720adc96aec5c8807295bdc2deca2c83bd8120a2cd87b112dc65

Observation f213867e-cd85-461f-836d-1cc6c5668753 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.147349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.147349Z digest=sha256:bd0dec8bebbeffbdd8957b9acdb22a8d797feca374d0d90505dd8441779cc3e5

Observation 56b383c2-f3a5-4120-99e2-96f1213e87d4 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Learning spatiotemporal features with 3d convolutional networks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.371957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:49.185062Z digest=sha256:59af0d06805b0a31342eba9be961bde35cad59bb6c0c652750f2b2e7612da8c6

Observation 50d62950-433e-436a-9662-4ce9c855f577 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.280218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.280218Z digest=sha256:7b3062a08070c4311afc2aabf0deed525003037dbbda20154fb400185dc4479d

Observation c5a0000c-cdb7-4c5a-8f89-03803715ded6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.345961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.345961Z digest=sha256:a113b29c6c5fc4334fe8c714296624c7bca96a669981e113477f69f610d456e3

Observation 0778abac-b905-4d0b-ab9a-8dd7e3b71467 · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.423515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.423515Z digest=sha256:0c87d15a38af03268c19187830cdcb268de9b814f1a16ad4cf4b9a4471fcb170

Observation 8884aa8f-17de-46c5-99ab-cd1704ee6913 · outbound

This paper cites Sullivan, Gisle Bjontegaard, and Ajay Luthra.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Sullivan, Gisle Bjontegaard, and Ajay Luthra

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.254156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:49.484414Z digest=sha256:bf9671652df37937396c249de719762f5f9c64b563dd0deccdf9bcacd5752fd1

Observation abe2f266-b145-441e-9c51-376e10e9557d · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.120020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:49.577022Z digest=sha256:617203a4987b2bbadbdbab63e92d43a5c15d6e6eaa7b294ee0861e0662e338e8

Observation 94cc60c0-1937-4263-a654-9ed27861d065 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Next-qa: Next phase of question-answering to explaining temporal actions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.636803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.636803Z digest=sha256:cf0c4cb85be6559c1b321327316cf08c4838ae8e5b93cb08cbe27e5a06a7651c

Observation 7560d55b-5148-499d-92b8-ca763da865a8 · outbound

This paper cites Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.945416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:49.707972Z digest=sha256:3bf138b696308ca41cf70e0fd0f0d0c24f5d6895c7a8da3e264cc2a8084b50c6

Observation 45a6af3d-703e-4459-8708-f0a4d375049e · outbound

This paper cites Qwen2 Technical Report.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.783006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.783006Z digest=sha256:67dd876039e832bac854a1c7daaaffc5a799a13d9184f593908c91e6bc1e26a5

Observation 5c385049-5d20-4389-bb81-fc98a8695fd0 · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.863938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.863938Z digest=sha256:45d3f376065006036447ea37543a5bd7849ca9d3b41e5346762c581fcfe9e149

Observation e182b257-b922-4262-ad32-3d431f68bda0 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.847754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:50.034259Z digest=sha256:c68dbe6e7cc8296b18c5611bda855c83e36baeb5c79cf0db753463569f542a2e

Observation 602199ed-14c7-4fbf-9815-939352dc8fa4 · outbound

This paper cites Long Context Transfer from Language to Vision.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Long Context Transfer from Language to Vision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.185274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.185274Z digest=sha256:a113979ffb2e682968aa166c4522d01b68cd25f0eca1771c08ea83d976166e8e

Observation dd7c2631-a2ca-43c3-be08-e90da971380a · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.312891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.312891Z digest=sha256:d08f9e7f299f466cf9985a76a12220c12d529adf0d5a1b2a408be4e858612885

Observation f3fd2480-95fc-42f5-8e71-53514a96b734 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.442322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.442322Z digest=sha256:3a755bf4a20b9789286f67343af6105f65089d599df2f6c287d22951c1e1a2f6

Observation a7e64a4e-366b-4623-9790-6ec1045e116f · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Mlvu: Benchmarking multi-task long video understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.646767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:50.667041Z digest=sha256:9b128d07b04a9f57c21088089531a06e844cb920fbf0ecb171702147f174f0fa

Observation d05ef8e8-af0a-4b5f-8c80-030feee8b836 · outbound

This paper cites 11 to 14 show the absolute values for the main comparison results.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs 11 to 14 show the absolute values for the main comparison results

Reference 2025

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:32:51.576968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:50.801112Z digest=sha256:0d5e9987675935a48b80b589865cd5b95681c834ed06c999e093ca497f16f3d9

Pith citing papers

Observation e99a8d45-4dcb-453b-80d4-725afbba76f0 · inbound

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs cites this paper.

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:36:44.191878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T18:35:01.328250Z digest=sha256:4129fec925a1a9d939dece8541157cb6ca67e46e4d2e4361bd3b1e1a704c046f

Observation 8cb5c6ec-fea4-4edd-8ed6-17874b0c7018 · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:13.772797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:0a42c6d0ed28fa11e9e8e1c433eb19f3ec17d947dc751a55bf93cd006e3a3392

Observation 27a489f0-2b34-46c3-b16b-76908f60e426 · inbound

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs cites this paper.

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:58:05.989905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T06:55:01.619441Z digest=sha256:91eef10375e8005ce5f773637209575f3f5287bd4ee05bd4331c97b4baca2929

Observation d890eb34-3b8c-4d7c-8973-4ab572bb59d4 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.654336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.654336Z digest=sha256:4fce134a9cd0e9a52807de2f7d6057bba8382d4a93d34abf103ad6bc22af30ea