Pith. sign in

Paper Citation Record · LEDGER

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 4 inbound Pith citation observations for arXiv:2507.07990.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07990 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:50.801112Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.654336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T06:58:05.988167Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 246b891c-bda5-48f4-b44b-df2299f204f7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.920094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.920094Z digest=sha256:416afa4ce534e0c1b605cf36bbab3eeb63d5be788907903c8531ed39e36dde96

Observation 7d0f0eb4-1003-46b9-83f2-6d8fd01db424 · outbound

This paper cites Token merging: Your vit but faster.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Token merging: Your vit but faster

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.444166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.007491Z digest=sha256:8b9b609605ae532418c57ee3d400880a98dbc1bd0a26572b3f205b521c744ab4

Observation 1d098822-e46e-43ba-98be-c0a9f7a2e67f · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.136591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.136591Z digest=sha256:684cd0387934e9f9e57ff993db2cbf1ef1cb50a08e292fea9a217789f7dcbb06

Observation 0670cd7b-274a-459c-8110-fa130e3564d1 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.411633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.252215Z digest=sha256:9db2f455346a589ed8193b43eb79bd6b01b56bb66dc21e83637d133cf189a57d

Observation 19b8f3b1-b4e2-4330-b2cc-7d25e356367c · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.392591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.386772Z digest=sha256:3d590e47c10929ee561c59b4f223925fbae1eaf9646fd73d72a78f9fadf30815

Observation 37f2bc3d-cab8-4313-b194-437375aacf32 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.374669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.517307Z digest=sha256:50803e8a73041a56c554bb67d58d8af5cfafce3fd5e8ef81fd06841e98528674

Observation b7c8df05-599c-413e-b4df-e5ea529487ce · outbound

This paper cites Longvila: Scaling long-context visual language models for long videos.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Longvila: Scaling long-context visual language models for long videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.352568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.618165Z digest=sha256:c25a95e90c9d2a4af4fe3ae6f0e6c21cc409d6c62cd6a2978f6873c8ccd2070d

Observation e62eee1b-61e5-463a-9b3c-2e7f17199433 · outbound

This paper cites vid-tldr: Training free token merging for light-weight video transformer.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs vid-tldr: Training free token merging for light-weight video transformer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.331836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.710034Z digest=sha256:5bab520eb23f72c4d903aceef5144a26e1f8af19ac59f41049fcf39579bd5d3a

Observation 5e573a74-dfd0-4e4c-b898-554117a8c69a · outbound

This paper cites FlashAttention-2: Faster attention with better par- allelism and work partitioning.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs FlashAttention-2: Faster attention with better par- allelism and work partitioning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.306539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.804392Z digest=sha256:8942e5c598a3ddfae91a244abf7d888f91d14f61b44c1cd9bbab7dd5f2645d68

Observation be0085a3-55d6-4f57-a88d-a988eef926be · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.284441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:45.866227Z digest=sha256:c991d3329e790910c6360e6d2d8a661d7876d0a43eb72c6c3382e7ca392ad63b

Observation ec769102-d618-4f13-a71a-79fbae6b2218 · outbound

This paper cites The Llama 3 Herd of Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.999654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.999654Z digest=sha256:e827f835a4b2aa3f73c9bdf93cd93d0c5d0f63c395edbe03250b74a490d87e9b

Observation 2f038acb-cd79-4778-b467-fe202eed1361 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Masked autoencoders as spatiotemporal learners

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.265252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.168771Z digest=sha256:52e9341ce63dbff32481d424952f6588ae9d183432bcb4deb95d209c12a86dcd

Observation 77307511-04a3-486b-80c8-5018115eb2da · outbound

This paper cites Quad trees a data structure for retrieval on composite keys.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Quad trees a data structure for retrieval on composite keys

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.245530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.273906Z digest=sha256:5030ec302e538ecaca77596b15d013aff64b06e68fea1def9837841ff5de9bcc

Observation 86f94159-5ea6-41b5-a195-dcd31f95a052 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.222087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.372576Z digest=sha256:8950a03d7f72e78d5526f2388600722d7acd2217690b76203cabf3ea08179f2b

Observation e842939a-c212-4e01-b7fd-dfc93bcc6dee · outbound

This paper cites FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.455680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.455680Z digest=sha256:3cb8de6be185905968beb7e687ee6f607339020c9cf470fef714408b91f5fbb2

Observation f045812b-d2eb-4998-a0d9-76435809e32a · outbound

This paper cites Caching — google ai, 2024.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Caching — google ai, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.204484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.509746Z digest=sha256:1813ba629f30bb6ac119a57730cb3cc290c67dca115e29abb73157a8e4bc8922

Observation dcec1c7b-82b6-43cc-bda6-1910dc2673e5 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.642984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.642984Z digest=sha256:59bf3005924707d30d5867a80b39718c9c449a0c70b4425827c61f313e17a2de

Observation e5902d98-fde2-470b-9f0e-c750d81456e5 · outbound

This paper cites Deep residual learning for image recognition.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Deep residual learning for image recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.759667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.759667Z digest=sha256:91340bd637e44b4686cde9290083510d56fcf4ff5879342a5f86677618e8fbcc

Observation 0d4efc77-4d15-4bd9-85ff-deabb6392d3a · outbound

This paper cites Masked autoencoders are scalable vision learners.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Masked autoencoders are scalable vision learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.171666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:46.855039Z digest=sha256:6e3aabe28a7019bc0c44927ae3f64f278fb651f16280c38507ae0e351a8be689

Observation d544c777-c601-45fb-a23e-d49639e58b38 · outbound

This paper cites PruneVid: Visual Token Pruning for Efficient Video Large Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.995554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.995554Z digest=sha256:d7330d61b56e8ad4d64a0f0280e67c00c499527428a99c41b3a222495fff4090

Observation 39e5f196-45bd-42c2-a28a-cb5acb60cc83 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.148978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.097387Z digest=sha256:4e3ca611a8b702edf5663193c4bc3fa3138c923bb67b108c8451c59a5272fd91

Observation 8d0fd6b9-ffdf-4105-9e21-fbf443daf535 · outbound

This paper cites Needle in a haystack – pressure testing llms.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Needle in a haystack – pressure testing llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.130345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.206402Z digest=sha256:c7eb0c970c1b5bfa81eec1eb6ddbcec393072dcc962c93c0ca1482709390ac06

Observation e7b38e68-d904-4f5a-b31d-54e76909e05b · outbound

This paper cites Handwritten digit recognition with a back- propagation network.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Handwritten digit recognition with a back- propagation network

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.113062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.301474Z digest=sha256:fd32823c80f27f82bc7da28b0576a7d64845740b5c68817616e2eb0ac0f3f023

Observation 85685f60-2fb4-4cc9-86df-c99a101a9a0c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.367536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.367536Z digest=sha256:eb8758d3bf848a8e93d5ecad5afd8f5365a0c4beccbc3e80932d1af4f2bdd0d9

Observation 8b0f4a29-485c-4221-ac3b-4e0064bc0f6f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.081714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.470093Z digest=sha256:38e9a24d54718ac7d29c230577e969c81119f129c8f9ad61df8fa171bbe7de41

Observation 3c22a84a-1f9d-4894-bd7a-ddec8a09b6be · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs VideoChat: Chat-Centric Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.575406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.575406Z digest=sha256:7576d3ad0b0e1fb49038066eec657ee64a58c4645e035573603cb3c3cdee7091

Observation e2e21bf8-0ea7-422d-92da-9f74675e0312 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.063042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.618941Z digest=sha256:94cadd01bd99f8656f09811086243e304c6ce335621d4fe71b7d75dc377cf14e

Observation 77420b45-60dd-4a2a-ae06-5c850c902b2e · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.042764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.732124Z digest=sha256:b9b373fbba8844bc6cee802bc89da31c027db35f1fa6cb188c6c24d4bc58ea7a

Observation bf4da92a-b536-4555-9aed-2d58320ef6c8 · outbound

This paper cites Visual instruction tuning.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:53.017634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:47.844102Z digest=sha256:9448a2c0b4788adca1587fc87f98b574fcb74dc1596aaecec8d76bd0ddf07cd4

Observation 1e5324cc-5866-4d60-a22a-132504cbcc9d · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.962155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.962155Z digest=sha256:53b9d439c53819cd1307e13e62fbb5d17c5bb0b675d547673e0d98ef32b30477

Observation caeac0d6-d89a-49eb-b1f9-09ba2d3d5223 · outbound

This paper cites Ring atten- tion with blockwise transformers for near-infinite context.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Ring atten- tion with blockwise transformers for near-infinite context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.989405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.093263Z digest=sha256:53f2afb389040fbaf78eae6599d273b901898d6ba2a1d622ad1c6d2c4cb64daa

Observation ed03f2a1-23f5-4a19-b587-574980db0d02 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.187988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.187988Z digest=sha256:494a7f18470300c48ad691348b0cb1803698c026231b82469a00c82f3a22ca99

Observation 16437856-d4e5-4593-8847-7fa463692fc7 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.947353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.260745Z digest=sha256:78ee03abd8eb3921c6d4066c6804fd0e9bbf1a702a3c4029d5ff562fb664fbeb

Observation 4b8f8318-4e79-4561-9f1e-926a1142560c · outbound

This paper cites Efficiently scaling transformer inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficiently scaling transformer inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.915585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.312111Z digest=sha256:b1806f39869c453543b0feca6e53f10411a490ea8c8cddbe61317e41f95637ac

Observation 687cec00-dd84-4cf7-ad26-d5f4fadf6b97 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.887336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.362612Z digest=sha256:fbe3d6cbc3775ceb87c8e575eb3229743045614dcefd35d2b70337946f53616e

Observation 83a72fcd-4822-4217-91f4-a59a1769c279 · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.415433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.415433Z digest=sha256:9c4236b97b98ac52d410ebf3b1b0163482389bde63bea31897a5a4463ea8a41e

Observation 88fcb5da-9be6-487d-b531-a752d5631b86 · outbound

This paper cites The quadtree and related hierarchical data structures.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs The quadtree and related hierarchical data structures

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.860677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.491626Z digest=sha256:7d4bd97b5d33d7014ea19629d3fe1c4c362de6af71ea0a9979eec491388f557d

Observation bbe2353b-5f2b-47e9-b11b-23ed357065d0 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.573421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.573421Z digest=sha256:615fc62d3046312c3d8e80202fb0f2769482a49ff9d38a0ae8e7a4e1ce34643e

Observation 7ee7ef49-1ddd-424c-848b-70815a92fcd9 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Moviechat: From dense token to sparse memory for long video understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.830881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.647081Z digest=sha256:26c3b2b270b50819e0f99a525b5513aaa7ddf31abfc032198636724e0223f627

Observation afe23cbc-779d-4373-be94-3f1a48a468b9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Roformer: Enhanced transformer with rotary position embedding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.725013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.725013Z digest=sha256:dbfe6fd4e864f1fcb1cac4bfda6361458796f212c25fef9d8bfb5c198fcda588

Observation 1026ae0b-c4e0-46ae-a06a-306d747fec8e · outbound

This paper cites Efficient quadtree cod- ing of images and video.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficient quadtree cod- ing of images and video

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.786169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.768904Z digest=sha256:189e0dacc2055cc84b74b58199916d5d077ecae62ef48edf9188cdd0ae7d596d

Observation 827c1b70-8999-44b3-845f-e9a464a8003c · outbound

This paper cites Overview of the high efficiency video coding (hevc) standard.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Overview of the high efficiency video coding (hevc) standard

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.766030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.839547Z digest=sha256:04255c8ac86966e1934185c8e47d0d9f6e4fc63d9348c75be682c067b9d36e79

Observation 1e0ee72e-7428-48f7-b981-87a41aa44524 · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Dycoke: Dynamic compression of tokens for fast video large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.739311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.887333Z digest=sha256:51167eab9cd72e6bac7231c405dcc60081ab5dbdd937431c4e475f8fa8b9095d

Observation 1928e642-4168-443d-9c79-469fbc97316f · outbound

This paper cites Efficiency of a good but not linear set union algorithm.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Efficiency of a good but not linear set union algorithm

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.619484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:48.958480Z digest=sha256:550985189fdf4469985e01a5823d83f35787b6beb21c640e9a5930ac716ef3b8

Observation b26d5fc4-c105-4f25-9614-c9a0da4e329b · outbound

This paper cites GPT-4o System Card.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs GPT-4o System Card

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.021249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.021249Z digest=sha256:7b210146d6ee8a5b805f020ac049595a28608be8486c174151521669d80ccbc5

Observation c7ba62c1-f17b-40e1-b81c-0cf3e1fadad7 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.511623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:49.081969Z digest=sha256:6157367cc67321fba7882a23a35a9ab2497954bd9a0f039bf0bebb5cd71d0d39

Observation f213867e-cd85-461f-836d-1cc6c5668753 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.147349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.147349Z digest=sha256:704f3b82e6689ee3e0255beb0b02c1c0e706d863138a1bf04ca7bff81d260adf

Observation 56b383c2-f3a5-4120-99e2-96f1213e87d4 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Learning spatiotemporal features with 3d convolutional networks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.371957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:49.185062Z digest=sha256:d0a790db837e5e6ba8f812ce66d1597db0d7bdfb29dd4f2acdd3312e04c88540

Observation 50d62950-433e-436a-9662-4ce9c855f577 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.280218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.280218Z digest=sha256:0dfc0009d872fab559d3056b4d2a69be5f36bb1b4e830d705c812b0ef0b61754

Observation c5a0000c-cdb7-4c5a-8f89-03803715ded6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.345961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.345961Z digest=sha256:634ac0bf59c7b9620300da825e0379c17993998190a092ecbfecd2f1e2d1cc65

Observation 0778abac-b905-4d0b-ab9a-8dd7e3b71467 · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.423515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.423515Z digest=sha256:11a942f8e5f323398868c424942ee566c4dd2a5fdbea5558c76205aea106eed7

Observation 8884aa8f-17de-46c5-99ab-cd1704ee6913 · outbound

This paper cites Sullivan, Gisle Bjontegaard, and Ajay Luthra.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Sullivan, Gisle Bjontegaard, and Ajay Luthra

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.254156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:49.484414Z digest=sha256:b1a48c3b58f0b07c1926ceaa72fed17ae201c74f184c5024c1d22ad57ea7f29d

Observation abe2f266-b145-441e-9c51-376e10e9557d · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.120020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:49.577022Z digest=sha256:5edcd140a04e348755bd1ac4fb7a8d3f16303eccc6e4f7ed09b34c3e853c97c0

Observation 94cc60c0-1937-4263-a654-9ed27861d065 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Next-qa: Next phase of question-answering to explaining temporal actions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.636803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.636803Z digest=sha256:39648b2e1f466b355900feb7aaca080261556098a6c2dd1359835d195574517a

Observation 7560d55b-5148-499d-92b8-ca763da865a8 · outbound

This paper cites Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.945416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:49.707972Z digest=sha256:528b8ae1319db2aee88509ab76a1c139949fd1bcd5f76afbbc97ea17d75410d9

Observation 45a6af3d-703e-4459-8708-f0a4d375049e · outbound

This paper cites Qwen2 Technical Report.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.783006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.783006Z digest=sha256:8e0c23079472ea39ab66265de78d30085ab951e18b6d62e9a773400b46121ea3

Observation 5c385049-5d20-4389-bb81-fc98a8695fd0 · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:49.863938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:49.863938Z digest=sha256:c1e1383bd3b6a1dbfe95a7153205ebd810f4c49c74155d781b4367a00655fc56

Observation e182b257-b922-4262-ad32-3d431f68bda0 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.847754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:50.034259Z digest=sha256:f1e9a6a03659ddb3aee0005d790935db0e3ea75ff69456d442f812521c174566

Observation 602199ed-14c7-4fbf-9815-939352dc8fa4 · outbound

This paper cites Long Context Transfer from Language to Vision.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Long Context Transfer from Language to Vision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.185274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.185274Z digest=sha256:044aca5492270ed8c7b1b3ffc24aeed2eed29033255234ef1157d0aa49cb8a1e

Observation dd7c2631-a2ca-43c3-be08-e90da971380a · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.312891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.312891Z digest=sha256:43116cbbb5f8b27f3aeeb76e4068e9d6aae762978acfe390464c6fc42d43f259

Observation f3fd2480-95fc-42f5-8e71-53514a96b734 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:50.442322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:50.442322Z digest=sha256:da3a18b8c7fd9836ac10962437ebb02e10df2fa0fed1268ec6e94d6ce50fab4d

Observation a7e64a4e-366b-4623-9790-6ec1045e116f · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs Mlvu: Benchmarking multi-task long video understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.646767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:50.667041Z digest=sha256:e4480afad8a1d529c626cdaf63045ba98df200d3a803ddb338761882596a1f03

Observation d05ef8e8-af0a-4b5f-8c80-030feee8b836 · outbound

This paper cites 11 to 14 show the absolute values for the main comparison results.

Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs 11 to 14 show the absolute values for the main comparison results

Reference 2025

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:32:51.576968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:32:50.801112Z digest=sha256:be8fc977e1f29200ae55c0ab6de8f55151c6af638639418a289a63cf2ba6a891

Pith citing papers

Observation e99a8d45-4dcb-453b-80d4-725afbba76f0 · inbound

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs cites this paper.

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:36:44.191878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T18:35:01.328250Z digest=sha256:eb6dc87d978380b02860654565c9f6be6c08f42ae8392a2fa2f4635a79e0e669

Observation 8cb5c6ec-fea4-4edd-8ed6-17874b0c7018 · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:13.772797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:d606f0a6475d00c5c2fb1bb2ca282e73932b2cb5442ba36fd18bc97dbd61d56b

Observation 27a489f0-2b34-46c3-b16b-76908f60e426 · inbound

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs cites this paper.

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:58:05.989905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T06:55:01.619441Z digest=sha256:da9bf5fa0f303eb947c13192b378e958531f3a264012a04cc22d5ae66457d16a

Observation d890eb34-3b8c-4d7c-8973-4ab572bb59d4 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.654336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.654336Z digest=sha256:21d95a6f5075fd2b04ee26461beecaa7df01b4c08633a8c758a34bd81ac7026f