Pith. sign in

Paper Citation Record · LEDGER

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

As of 19 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 4 inbound Pith citation observations for arXiv:2507.07990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07990 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:50.801112Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.654336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T06:58:05.988167Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 246b891c-bda5-48f4-b44b-df2299f204f7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.920094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.920094Z digest=sha256:1acfd0553e548cb84585490aaf648f4a3d74bc55cd3dd4928200f3677773bd5a

Observation 7d0f0eb4-1003-46b9-83f2-6d8fd01db424 · outbound

This paper cites Token merging: Your vit but faster.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Token merging: Your vit but faster

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.444166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:45.007491Z digest=sha256:32206b9de66d7112319f694b9d121e001767c024196d7c5013c709d4d13b3d60

Observation 1d098822-e46e-43ba-98be-c0a9f7a2e67f · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.136591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.136591Z digest=sha256:6d5b6b0d3574296ee30f26c676f98dd0d5969889f749869f4f335de7258bb1df

Observation 0670cd7b-274a-459c-8110-fa130e3564d1 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.411633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:45.252215Z digest=sha256:5484e926b5713af2c0418835d444ee4b1c3dc7267290e3c7236794fe6b424b9a

Observation 19b8f3b1-b4e2-4330-b2cc-7d25e356367c · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.392591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:45.386772Z digest=sha256:84d8131266ccf27da390f6fc38cb2e3a830a4529f49f1aed33b613b4ae563e72

Observation 37f2bc3d-cab8-4313-b194-437375aacf32 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.374669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:45.517307Z digest=sha256:d9b94b02d39fd33143734f945d0e6d7bd79eb4484819b772aefa8bbc5800e9fc

Observation b7c8df05-599c-413e-b4df-e5ea529487ce · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Longvila: Scaling long-context visual language models for long videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.352568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:45.618165Z digest=sha256:74b9935aff328e4cf605208f8d0c84eb654499bed002a198725616ff5c2e84ed

Observation e62eee1b-61e5-463a-9b3c-2e7f17199433 · outbound

This paper cites vid-tldr: Training free token merging for light-weight video transformer.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs vid-tldr: Training free token merging for light-weight video transformer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.331836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:45.710034Z digest=sha256:1eb39e67ef8f8d32ca22a835905c3fdc13b90df33948e3505cc2cf4f435478c1

Observation 5e573a74-dfd0-4e4c-b898-554117a8c69a · outbound

This paper cites FlashAttention-2: Faster attention with better par- allelism and work partitioning.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs FlashAttention-2: Faster attention with better par- allelism and work partitioning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.306539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:45.804392Z digest=sha256:73a447f0fa7d1fb493b20dd9845f2dbe1bd8cedfb6681d3d1d8c2d655f954785

Observation be0085a3-55d6-4f57-a88d-a988eef926be · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.284441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:45.866227Z digest=sha256:c66c806378c9297de80e8e84bef41261d71b1fc6e3801b73b28ce64685fc5001

Observation ec769102-d618-4f13-a71a-79fbae6b2218 · outbound

This paper cites The Llama 3 Herd of Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.999654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.999654Z digest=sha256:515d92f8593313b979303a9ff34c7d937a3c2144151225bfe92e78352457f0ff

Observation 2f038acb-cd79-4778-b467-fe202eed1361 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Masked autoencoders as spatiotemporal learners

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.265252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:46.168771Z digest=sha256:d322810f8be463c86a420a53887f7c20e61f50d2e07d6ba9fce68d7af664f71a

Observation 77307511-04a3-486b-80c8-5018115eb2da · outbound

This paper cites Quad trees a data structure for retrieval on composite keys.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Quad trees a data structure for retrieval on composite keys

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.245530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:46.273906Z digest=sha256:29c8b991c089e478623b490222c302415c3edda267b1df463bd2395435dc6255

Observation 86f94159-5ea6-41b5-a195-dcd31f95a052 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.222087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:46.372576Z digest=sha256:f8f608f1de1c7167363cff4c23e1b22b5420172462467b36ba3a0bec8bb32a29

Observation e842939a-c212-4e01-b7fd-dfc93bcc6dee · outbound

This paper cites FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.455680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.455680Z digest=sha256:cd1f814da9e647c72e5be47fc2c10d4cf6a0a03f02e78de32a5877675decd175

Observation f045812b-d2eb-4998-a0d9-76435809e32a · outbound

This paper cites Caching — google ai, 2024.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Caching — google ai, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.204484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:46.509746Z digest=sha256:b478ddfc64a94c19d3a297dfd4eeeff3811deebd4d2ec708d48f6693e8ad129f

Observation dcec1c7b-82b6-43cc-bda6-1910dc2673e5 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.642984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.642984Z digest=sha256:612bfdeee8183e8f05c0f91fe24fa92ff31b9a798e37176cfe7831a13da70124

Observation e5902d98-fde2-470b-9f0e-c750d81456e5 · outbound

This paper cites Deep residual learning for image recognition.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Deep residual learning for image recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.759667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.759667Z digest=sha256:75a53f31d348340ca74762b6df1c1d8bb2fdbd600bc6c8301a45f82f7d70548a

Observation 0d4efc77-4d15-4bd9-85ff-deabb6392d3a · outbound

This paper cites Masked autoencoders are scalable vision learners.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Masked autoencoders are scalable vision learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.171666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:46.855039Z digest=sha256:724846f21daaf24e63b9c82880a318a8fa6565e1eb9b03c98efc88d0f631b7e6

Observation d544c777-c601-45fb-a23e-d49639e58b38 · outbound

This paper cites PruneVid: Visual Token Pruning for Efficient Video Large Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.995554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.995554Z digest=sha256:f4719cde00df27d642c6bffd3908131eb9830d07955ebcd8b534cba44c368960

Observation 39e5f196-45bd-42c2-a28a-cb5acb60cc83 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.148978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:47.097387Z digest=sha256:b3ae06c2e4003984944df44defb615d3c93fc16dd94cb622d74c7c9e218dd9ab

Observation 8d0fd6b9-ffdf-4105-9e21-fbf443daf535 · outbound

This paper cites Needle in a haystack – pressure testing llms.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Needle in a haystack – pressure testing llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.130345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:47.206402Z digest=sha256:32b5cfd9807e36965004bc46c15025c99f4f17fe0e2ca90ee12772224567c321

Observation e7b38e68-d904-4f5a-b31d-54e76909e05b · outbound

This paper cites Handwritten digit recognition with a back- propagation network.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Handwritten digit recognition with a back- propagation network

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.113062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:47.301474Z digest=sha256:8513c5d2ecbecab8097134a563476abe3aaad2b28466475e38b246f58b3fa513

Observation 85685f60-2fb4-4cc9-86df-c99a101a9a0c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.367536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.367536Z digest=sha256:636220cb9215c32346e325cc9752ad31b9ea479cb05a6588eef8138630f72087

Observation 8b0f4a29-485c-4221-ac3b-4e0064bc0f6f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.081714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:47.470093Z digest=sha256:437492de8a767f158d9924ebbddb639212a8e3ea85e774a9f560f8903b6ba835

Observation 3c22a84a-1f9d-4894-bd7a-ddec8a09b6be · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.575406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.575406Z digest=sha256:ac51a4463c28108b56f6a0d4f33c28d4fe04efe413d641353fb44f68308b5f36

Observation e2e21bf8-0ea7-422d-92da-9f74675e0312 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.063042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:47.618941Z digest=sha256:e0cfb1539bc5050ae87a5c53283c9867ef2282e83997122c0cd97464103211e9

Observation 77420b45-60dd-4a2a-ae06-5c850c902b2e · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.042764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:47.732124Z digest=sha256:61404cad16a5e9a498407cbdf0cbbd68396d9c467ef4a2ced4851569bbb23370

Observation bf4da92a-b536-4555-9aed-2d58320ef6c8 · outbound

This paper cites Visual instruction tuning.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.017634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:47.844102Z digest=sha256:954cf5ca0866e597cee9d11a561d51a0d4f9f09403c31cf9430619d302b17e67

Observation 1e5324cc-5866-4d60-a22a-132504cbcc9d · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.962155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.962155Z digest=sha256:1055f7e10fbe318e5dfc273f342c4ca136e7923986bdfe256ff504c6e90240ec

Observation caeac0d6-d89a-49eb-b1f9-09ba2d3d5223 · outbound

This paper cites Ring atten- tion with blockwise transformers for near-infinite context.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Ring atten- tion with blockwise transformers for near-infinite context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.989405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:48.093263Z digest=sha256:7e5c929e864487c6ab82d07fd79defd7cfbc02bb9303a22aa31bcd9168f4a994

Observation ed03f2a1-23f5-4a19-b587-574980db0d02 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.187988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.187988Z digest=sha256:a2ed22ccaeb74309e40076ed7fdbc81c05ccd9880aea0d5e29ddbe60fffcf7f9

Observation 16437856-d4e5-4593-8847-7fa463692fc7 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.947353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:48.260745Z digest=sha256:86521f43ef62a61b4d551a804352fc389b558e8442def5ffb97c5933fdd7d05f

Observation 4b8f8318-4e79-4561-9f1e-926a1142560c · outbound

This paper cites Efficiently scaling transformer inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficiently scaling transformer inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.915585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:48.312111Z digest=sha256:850ba688194109470f8d96b1ecd2f3de3efdf9b750fa1cf6205d57de4497323e

Observation 687cec00-dd84-4cf7-ad26-d5f4fadf6b97 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.887336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:48.362612Z digest=sha256:1d55e3c5b624528e9216cf9d674802bbe6da4cb89ec7b9b3654621caed89119f

Observation 83a72fcd-4822-4217-91f4-a59a1769c279 · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.415433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.415433Z digest=sha256:d79c6da57e41f5aca9a7ff9757da87882af6f2078c0196c61bb74b36580e0e3b

Observation 88fcb5da-9be6-487d-b531-a752d5631b86 · outbound

This paper cites The quadtree and related hierarchical data structures.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs The quadtree and related hierarchical data structures

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.860677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:48.491626Z digest=sha256:7051b49a0c0e9f014f8929d3789973353e51c8d85d18d0ae2bf967098d9e4f32

Observation bbe2353b-5f2b-47e9-b11b-23ed357065d0 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.573421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.573421Z digest=sha256:fa991ac2b169d78600cb97793ea6496cfdb5c3fe607d7a64301c27321796d9cf

Observation 7ee7ef49-1ddd-424c-848b-70815a92fcd9 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Moviechat: From dense token to sparse memory for long video understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.830881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:48.647081Z digest=sha256:2367adbc9545df0329282d1823d7f6f3a409c2ce386f62dd03dbc78dda197552

Observation afe23cbc-779d-4373-be94-3f1a48a468b9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Roformer: Enhanced transformer with rotary position embedding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.725013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.725013Z digest=sha256:64a91473d9eb87cbfa9edfb2e8b91e98ff6bc0a3ea4a06737bffdcda27422278

Observation 1026ae0b-c4e0-46ae-a06a-306d747fec8e · outbound

This paper cites Efficient quadtree cod- ing of images and video.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficient quadtree cod- ing of images and video

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.786169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:48.768904Z digest=sha256:92abf0a9e4c1c49ccfb4282f3f5efb5d225ab4208e948a2a650006194ffa778b

Observation 827c1b70-8999-44b3-845f-e9a464a8003c · outbound

This paper cites Overview of the high efficiency video coding (hevc) standard.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Overview of the high efficiency video coding (hevc) standard

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.766030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:48.839547Z digest=sha256:d8c24f5455556c35cd4fda12986ddd274495df0e21476b31e43bc3e3046316eb

Observation 1e0ee72e-7428-48f7-b981-87a41aa44524 · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Dycoke: Dynamic compression of tokens for fast video large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.739311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:48.887333Z digest=sha256:f59ac4404a30ae51384d1f2f48eae6c93ae09f65a9e5b9a1b61e0b9726126bb4

Observation 1928e642-4168-443d-9c79-469fbc97316f · outbound

This paper cites Efficiency of a good but not linear set union algorithm.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficiency of a good but not linear set union algorithm

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.619484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:48.958480Z digest=sha256:bd16ded711cabbd9fa935b0f3276ef8df19c6360f695c59a160fc4c0b17490ea

Observation b26d5fc4-c105-4f25-9614-c9a0da4e329b · outbound

This paper cites GPT-4o System Card.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs GPT-4o System Card

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.021249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.021249Z digest=sha256:246d3050545529e19de64533cef51ee271173761006e13e93688adec48f07988

Observation c7ba62c1-f17b-40e1-b81c-0cf3e1fadad7 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.511623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:49.081969Z digest=sha256:03a7ce32c475fd7a3ef0b767a9fa10f4db75a651df0f1f5d805ef4cbb771f22b

Observation f213867e-cd85-461f-836d-1cc6c5668753 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.147349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.147349Z digest=sha256:62b7b4cfefc73d4ef5a29037d9985681bad03165c0b3206f35f10e58ef6e29a2

Observation 56b383c2-f3a5-4120-99e2-96f1213e87d4 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Learning spatiotemporal features with 3d convolutional networks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.371957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:49.185062Z digest=sha256:c9ce233444e6029a3c21da4004d04d351c1de91e0865ed1e843dbf0d921d045d

Observation 50d62950-433e-436a-9662-4ce9c855f577 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.280218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.280218Z digest=sha256:e2bb6e0fee30463b2ff51f86d3679c8f698549167a90db3b9fef491f26383937

Observation c5a0000c-cdb7-4c5a-8f89-03803715ded6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.345961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.345961Z digest=sha256:357ff06a18801f6c76386f05ce01bb336e5ca32a3971180e1509f38fad959fe1

Observation 0778abac-b905-4d0b-ab9a-8dd7e3b71467 · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.423515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.423515Z digest=sha256:1984630ccd016d82133e0e966d816c51cd6e138cb3245c4345ad4a14ccaf04e2

Observation 8884aa8f-17de-46c5-99ab-cd1704ee6913 · outbound

This paper cites Sullivan, Gisle Bjontegaard, and Ajay Luthra.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Sullivan, Gisle Bjontegaard, and Ajay Luthra

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.254156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:49.484414Z digest=sha256:6bbf156f67fab358853224fc496e56500d4a900fd40177557181d4ac03cc3ab7

Observation abe2f266-b145-441e-9c51-376e10e9557d · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.120020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:49.577022Z digest=sha256:40b6fe7e4624d2353cbd5bdc45d6686da3e4affa47d2efbee59c3d557745941e

Observation 94cc60c0-1937-4263-a654-9ed27861d065 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Next-qa: Next phase of question-answering to explaining temporal actions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.636803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.636803Z digest=sha256:a9167323c5710f81c42ec9d362a43ba06b3a0097dd3ca4a74b017116bc656b84

Observation 7560d55b-5148-499d-92b8-ca763da865a8 · outbound

This paper cites Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.945416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:49.707972Z digest=sha256:4b6661dbf85a940a58f21c55b7142b9c79fbc344a7e599d9e4be3568adb26c5a

Observation 45a6af3d-703e-4459-8708-f0a4d375049e · outbound

This paper cites Qwen2 Technical Report.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.783006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.783006Z digest=sha256:7d09d824e66d94bfa45d6284de2e7cb836f0e189cc114c22d753aad70b56058d

Observation 5c385049-5d20-4389-bb81-fc98a8695fd0 · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.863938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.863938Z digest=sha256:17f088c6c5c4480cc5d638603954a083d8f07aac620c9d7d101d803cbe0c9eea

Observation e182b257-b922-4262-ad32-3d431f68bda0 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.847754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:50.034259Z digest=sha256:756cc9f676c87aa27b80b18b2c56877fe970ab515cd6eaed4b3ca79bac16d31e

Observation 602199ed-14c7-4fbf-9815-939352dc8fa4 · outbound

This paper cites Long Context Transfer from Language to Vision.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Long Context Transfer from Language to Vision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.185274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.185274Z digest=sha256:5d86925cbfecc48d16f54ce074d293b2111f56bb71c3312fff2838484d1c3db2

Observation dd7c2631-a2ca-43c3-be08-e90da971380a · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.312891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.312891Z digest=sha256:d9e43728c7fc712fa04e2b29905d656fd4b535897ceedecd9ad6c5f801b21329

Observation f3fd2480-95fc-42f5-8e71-53514a96b734 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.442322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.442322Z digest=sha256:c47f3c9ace5ad0cddfcf17282958f8a3c7196feb90bc3cb181cc46d132fa7cf9

Observation a7e64a4e-366b-4623-9790-6ec1045e116f · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Mlvu: Benchmarking multi-task long video understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.646767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:50.667041Z digest=sha256:12084ada5761ee789f5c6bd1b307250f57ce19f4029fea07fa0701233d6cb32e

Observation d05ef8e8-af0a-4b5f-8c80-030feee8b836 · outbound

This paper cites 11 to 14 show the absolute values for the main comparison results.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs 11 to 14 show the absolute values for the main comparison results

Reference 2025

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:32:51.576968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T18:32:50.801112Z digest=sha256:92146eb1d7631be2ffed29aaef8d126c1509ea9354fc3033be6b6274db91b5b7

Pith citing papers

Observation e99a8d45-4dcb-453b-80d4-725afbba76f0 · inbound

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs cites this paper.

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:36:44.191878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T18:35:01.328250Z digest=sha256:13385dd62e61e4cb86a55695a0fe69f69905b1079f365171da2621402ee5b379

Observation 8cb5c6ec-fea4-4edd-8ed6-17874b0c7018 · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:13.772797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:c7a501a78d1bc5051968c3f81f25c207ca2728aabd0a74832488af72d871858d

Observation 27a489f0-2b34-46c3-b16b-76908f60e426 · inbound

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs cites this paper.

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:58:05.989905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T06:55:01.619441Z digest=sha256:a5d6a897f0055df5a9cb6d0c0b0d567e54fb6d536a9a60b2b3c35f743e2d4295

Observation d890eb34-3b8c-4d7c-8973-4ab572bb59d4 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.654336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.654336Z digest=sha256:85bd960cf7a2a8102e97d1463864283f6b594e83e58b16f758daa1083786df05