Pith. sign in

Paper Citation Record · LEDGER

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 3 inbound Pith citation observations for arXiv:2506.05709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05709 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:29:09.281677Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:40:36.389460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T13:20:26.673617Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0888780a-745f-43c7-a35a-fbecf752f310 · outbound

This paper cites Token Merging: Your ViT But Faster.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Token Merging: Your ViT But Faster

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.281431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.281431Z digest=sha256:b48639bf4ad42592ef460534efad8474ddfa5f18ccc9b48acdcf136fa8b834a1

Observation 5cab2b06-3244-4ce3-831a-2e70ad67fb4a · outbound

This paper cites Crossvit: Cross-attention multi-scale vision transformer for image classification.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Crossvit: Cross-attention multi-scale vision transformer for image classification

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.388730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.388730Z digest=sha256:a52ea73811f626efb29bce513fc34ff5e0483ff19addfe1aa224f768ea7d50d6

Observation 0b1abe6f-8215-4308-be29-9f31d4057c79 · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration The cityscapes dataset for semantic urban scene understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.502992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.502992Z digest=sha256:706e9cb742841aba67c344f730c9d0b59971db4b1005b6b38d31445fbd90ec2e

Observation 4fb8e953-7fe6-4d69-868f-042896b35c86 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Imagenet: A large-scale hierarchical image database

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.652087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.652087Z digest=sha256:40d9d8086e373e7ba12f01cd47b3f3d56b7627872560b5d7a12c2f03c6a8c1c3

Observation e9372569-a820-4c8b-b1ed-e90341c7b665 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.745969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.745969Z digest=sha256:f7f37c23880ba5d42b1b449731e5617de0f5662a38b64cd8ed904ea5b5ef5107

Observation f5b48f0c-07db-4424-98a3-57a965da433f · outbound

This paper cites Adaptive token sampling for efficient vision transformers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Adaptive token sampling for efficient vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.711924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:04.879291Z digest=sha256:83bd7933e020f4a83edcae13a6720d8836da03b056af0647f3af9a067756f321

Observation a64d3003-f149-462d-8504-445086b060fb · outbound

This paper cites Dynamic Channel Pruning: Feature Boosting and Suppression.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Dynamic Channel Pruning: Feature Boosting and Suppression

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:04.957294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:04.957294Z digest=sha256:51b4c3f072f7c222d8eb11d386cdb0d9cc5fed5a3ecd6c6cfdbb5a56923cc6ec

Observation ef0f1ead-73e8-4bac-b6ee-714e548b1e29 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:05.047412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:05.047412Z digest=sha256:1c517447b6e9467d2fa926186dfb1821c2ac68fd65d6041a8dd9286897e5f5ba

Observation b28c7a65-1c10-4b2c-8014-f04eef38c082 · outbound

This paper cites Dynamic neural networks: A sur- vey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7436–7456, 2021.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Dynamic neural networks: A sur- vey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7436–7456, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.681292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:05.135319Z digest=sha256:53420621e9566b1f98388473b600eaba8d27f07d4eee4d375e91939efe067ab7

Observation 46435033-fb1a-4eb3-ad6e-fba69d2779bf · outbound

This paper cites Masked autoencoders are scalable vision learners.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Masked autoencoders are scalable vision learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:05.210527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:05.210527Z digest=sha256:b91c0432cdbd2a8a6a2d9459c629d0243133f542d41b4f625dedcbd15e6b1e42

Observation b56a5213-dfa2-4563-9664-36861b6d4660 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Distilling the Knowledge in a Neural Network

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:05.342081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:05.342081Z digest=sha256:17978c84f8732fdcc5912a1e15501bb153333da709be4fe2b7df31fa3e44d27b

Observation bc203aa9-11cd-4c25-9936-ab7f1167166e · outbound

This paper cites All tokens matter: Token labeling for training better vision transform- ers.Advances in neural information processing systems, 34: 18590–18602, 2021.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration All tokens matter: Token labeling for training better vision transform- ers.Advances in neural information processing systems, 34: 18590–18602, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.651394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:05.430042Z digest=sha256:b1933beba6bad07f1a33b5699049f29e43d43bef56b263eaea1a08b502ed8bb2

Observation 0db0ec9d-8003-4d08-9186-e582ca2adfaa · outbound

This paper cites Spvit: Enabling faster vision transformers via latency-aware soft token pruning.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Spvit: Enabling faster vision transformers via latency-aware soft token pruning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.624831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:05.506428Z digest=sha256:25f9cabf8c7fea08338e453a722916565ea7c42bc558a4488e8d6a1b24251e00

Observation a38b61e1-ebe5-4e71-a9c9-425a8547639a · outbound

This paper cites Peeling the onion: Hierarchical reduction of data redundancy for efficient vision transformer training.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Peeling the onion: Hierarchical reduction of data redundancy for efficient vision transformer training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.606625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:05.585257Z digest=sha256:9b9dc8c744801800dd0d5ea87872f7256123348e208c856c19c6df62956f6615

Observation 45dc1c5f-6e8c-4412-9a90-13b27f30a900 · outbound

This paper cites Token reduction should go beyond effi- ciency in generative models–from vision, language to mul- timodality.arXiv preprint arXiv:2505.18227, 2025.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Token reduction should go beyond effi- ciency in generative models–from vision, language to mul- timodality.arXiv preprint arXiv:2505.18227, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:05.660989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:05.660989Z digest=sha256:97d95bfada654ec81e105550ced8c915e96c80ba4e8a67647171beeb6e46ca5b

Observation 7745153f-5938-43a9-9bcb-ddd347d7c914 · outbound

This paper cites Mpvit: Multi-path vision transformer for dense prediction.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Mpvit: Multi-path vision transformer for dense prediction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.580185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:05.744711Z digest=sha256:1e547e9e8e33607c979a2c6fec0449afab4a760894b7d112c7e1391c9d34ceaa

Observation ebd511db-f48c-457c-a81e-9910295f8dfb · outbound

This paper cites Vidtome: Video token merging for zero-shot video editing.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Vidtome: Video token merging for zero-shot video editing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.554998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:05.823630Z digest=sha256:6ced2b1ca3763dbb1be2a043b17dbe799d2630fb26926cc0bfbf846250a3f113

Observation 35781d02-da4d-4d47-b89e-e82f19139409 · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Exploring plain vision transformer backbones for object de- tection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.527097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:05.934640Z digest=sha256:a72990115d2bb823449542cda9328f85f3e5a836a7f819472a6cd90855813253

Observation 6f97ca6a-6c2a-441f-bf9b-056cfbe1cec5 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Evaluating object hallucination in large vision-language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.499169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:06.014456Z digest=sha256:9ca2ef4f4af8b2d1a4cb5ff985faad178f088c838e438b05fd616937146bca26

Observation 3187dbc6-bf1c-4d47-b63b-1d61bb3e8225 · outbound

This paper cites Expediting large-scale vision transformer for dense predic- tion without fine-tuning.Advances in Neural Information Processing Systems, 35:35462–35477, 2022.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Expediting large-scale vision transformer for dense predic- tion without fine-tuning.Advances in Neural Information Processing Systems, 35:35462–35477, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.476808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:06.111149Z digest=sha256:71b5e212534ccde0411b32340fbd0b83c821d2b987b6732163f8b7e11c576537

Observation 4ea22053-9204-4db7-8171-e2f7f4ebb9c3 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.229026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.229026Z digest=sha256:49f7010d62d695f5b52f53846e22d69495f5b1abb11b10148d04b939f704efef

Observation 12b87abb-f5f7-4210-b8ce-fdacd94047db · outbound

This paper cites Microsoft coco: Common objects in context.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Microsoft coco: Common objects in context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.305116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.305116Z digest=sha256:7de3076bca9015c213d19f55afe415e91ff0841f3f0a90b43de9510a4f0d5e81

Observation 43770335-d3c0-455d-ad52-55898dd40e3d · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.392937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.392937Z digest=sha256:c2eaa68e7457ff7fe1b37639c945dc96298198f3545b0c85a9bfd6591ecf2b58

Observation 860ec1a7-aee1-403f-8680-eb4b3b6ebb18 · outbound

This paper cites Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-Attention.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-Attention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.482067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.482067Z digest=sha256:983bf7204b2485db8f6d992b09dc32176f8e13823f9e3c0c20777e70b374e8c3

Observation 33436fba-e7b2-4490-ada3-601545fe07d2 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Swin transformer: Hierarchical vision transformer using shifted windows

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.552259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.552259Z digest=sha256:4be2578ff9f945d8f22a853e608ad10f82fd1668b8224fc1fc8f30335b7006ef

Observation 979a44d9-6da2-46b2-9c21-3b0c8a1f665a · outbound

This paper cites Beyond attentive tokens: Incorporating to- ken importance and diversity for efficient vision transform- ers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Beyond attentive tokens: Incorporating to- ken importance and diversity for efficient vision transform- ers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.428964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:06.636831Z digest=sha256:31dbac51cbfc6d8c9d117a700dee9f4febfba771612fd84a95819a21f6e481ad

Observation 8009a827-c332-4af3-857d-f6ca9329f6b4 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.746497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.746497Z digest=sha256:c7ab9f41066b05f6e24543e5920f16553b55e2f5db7a30c480dc658ae05ed74b

Observation 2c5cf6f9-377d-48bf-9ce6-902d88ad763e · outbound

This paper cites Importance estimation for neural net- work pruning.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Importance estimation for neural net- work pruning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.391781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:06.834295Z digest=sha256:a7ceac063f837f470a37fd0063170255837d1dad6b728d4bf06a0f0c147b5cc0

Observation 389c33d8-09c4-43d5-8a3d-b274e8ee555d · outbound

This paper cites Improving language understanding by gen- erative pre-training.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Improving language understanding by gen- erative pre-training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.910131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.910131Z digest=sha256:2503285bcab55222ae7a0c68b386090469f02707fbd9355b00f831a50f46fd1f

Observation be14cf72-3f22-4df1-aa32-85e1491bf8af · outbound

This paper cites Vi- sion transformers for dense prediction.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Vi- sion transformers for dense prediction

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:43.130693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:06.981966Z digest=sha256:35dfcaf4dd9b95ed81eb173189ea5f5cf5489ac3c9ec8ee91fed3f5775a1618a

Observation bd498842-e345-4559-80c2-2c70f62f00f0 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.068786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.068786Z digest=sha256:9e91903bc8e79be182478fd392897839ce774f813e19f1e79c89ee48dcded144

Observation c4914461-08cc-4b41-aa12-1558c457f3b5 · outbound

This paper cites Beyond fixa- tion: Dynamic window visual transformer.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Beyond fixa- tion: Dynamic window visual transformer

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:42.989923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:07.153073Z digest=sha256:b99d547ec06bf2758271c91910c041004f2da493d8b50323aa9c922df39ae993

Observation c15a2e28-18e9-45f3-b269-8df503a92712 · outbound

This paper cites Tokenlearner: Adaptive space-time tokenization for videos.Advances in Neural In- formation Processing Systems, 34:12786–12797, 2021.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Tokenlearner: Adaptive space-time tokenization for videos.Advances in Neural In- formation Processing Systems, 34:12786–12797, 2021

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:42.845527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:07.225649Z digest=sha256:194377800dca14227ae728e35e413fead3729fba1d9f4eacaff7a6049544e72a

Observation 599cf5bd-dd11-4724-8683-bef66c8ba766 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Llava-prumerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.294589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.294589Z digest=sha256:000778ac29aa5517d8ed85322f7c5c2e800fdc6e36131b325d6fcab632b014c6

Observation 0c617dbe-ca68-4420-8bfa-56e394ba68d8 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Indoor segmentation and support inference from rgbd images

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.373418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.373418Z digest=sha256:cfc1b99fb69767353e6336c67788310ba8835a025213220500568dbefa3460d8

Observation f38efc45-34f6-4c89-9451-0a85d54ac8b4 · outbound

This paper cites Towards vqa models that can read.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Towards vqa models that can read

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.438068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.438068Z digest=sha256:0a80529e90bc15d67381c9f7fec0eb728134bf39531e44469956974f20fffd4f

Observation dc0e95be-ffb6-4161-bdbf-bf987a02c641 · outbound

This paper cites How to train your vit? data, augmentation, and regularization in vision transformers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration How to train your vit? data, augmentation, and regularization in vision transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:42.738363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:07.536108Z digest=sha256:6db7c5c3d25b27334271bcc72528bde6699a31af207a22e51867c1a931720ae4

Observation 180647af-62ad-4a1c-a0c8-5fd678508398 · outbound

This paper cites Segmenter: Transformer for semantic segmenta- tion.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Segmenter: Transformer for semantic segmenta- tion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.626801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.626801Z digest=sha256:cf139813222aa459a1c435a0ebf27622d0b45d9650a9f448e4cc94cf1e69e56e

Observation 3bf620c2-ed9e-41c9-9bb4-b3d7605a99f3 · outbound

This paper cites Chip: Channel independence- based pruning for compact neural networks.Advances in Neural Information Processing Systems, 34:24604–24616,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Chip: Channel independence- based pruning for compact neural networks.Advances in Neural Information Processing Systems, 34:24604–24616,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.694796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.694796Z digest=sha256:3ba97bfe1969a2da6dc1e1fca4ac31089aa6230f2596e48c8d303a80d124b8df

Observation 665c0120-0de5-4f2f-b165-53a8a84b86ed · outbound

This paper cites Revisiting unreasonable effectiveness of data in deep learning era.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Revisiting unreasonable effectiveness of data in deep learning era

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:42.560510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:07.784186Z digest=sha256:c9f014242f555248b2b2fdacb961f577b0973e5475c41c05726f0d08882fc327

Observation 7642238b-05ad-4b54-95bd-5a92d8873de0 · outbound

This paper cites Patch slimming for ef- ficient vision transformers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Patch slimming for ef- ficient vision transformers

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:42.468333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:07.843474Z digest=sha256:cb4ab7fd178d3058b149ecc2f17c5373afacf0642340249995503aaa044c0dcf

Observation 63b54ada-a354-421e-8569-7910f6534099 · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Training data-efficient image transformers & distillation through at- tention

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:11.049320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:07.913513Z digest=sha256:6e3528058d854e8b748531a14a28faff349768a299713f0ba537bc88c3d9a0d1

Observation e790d5cb-ccf1-4974-9c54-65424d8265f7 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.974529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.974529Z digest=sha256:21859366154cff8e0eaf04fcabdd2d8dddeff38dee798b1ea2834a4fe5df943f

Observation 6265c65b-848b-463c-846e-00f2bf5bfa08 · outbound

This paper cites Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition.Advances in Neural Information Processing Systems, 34:11960–11973,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition.Advances in Neural Information Processing Systems, 34:11960–11973,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.912369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:08.053653Z digest=sha256:3622d4020e9a3b970bfe16c75d6e4704346d2775f51990258b7af5af62dbf066

Observation 040418a6-dfae-4977-97f0-68ce49ba74f6 · outbound

This paper cites Qsfm: Model pruning based on quantified similarity between feature maps for ai on edge.IEEE Internet of Things Journal, 9(23):24506–24515,.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Qsfm: Model pruning based on quantified similarity between feature maps for ai on edge.IEEE Internet of Things Journal, 9(23):24506–24515,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.741796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:08.136865Z digest=sha256:1e716d1df7a6044538e2a7ea1995407c79a8a9ebb2ea55749689adf222dd3be3

Observation 63559539-8457-4376-a2e8-b9030fb3dafa · outbound

This paper cites Joint token pruning and squeezing towards more aggressive compression of vision transformers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Joint token pruning and squeezing towards more aggressive compression of vision transformers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.586451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:08.203119Z digest=sha256:fff2f0e65289ff855612d1cb4da03759f8e497d16db24a3cba9abc02df9f38ff

Observation 32e2e7f0-0d15-4941-a2d5-1f0e6db7e66d · outbound

This paper cites PPT: Token Pruning and Pooling for Efficient Vision Transformers.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration PPT: Token Pruning and Pooling for Efficient Vision Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:08.280794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:08.280794Z digest=sha256:3c1579188e9a6eed6934c32e828bab85c0ee444c642f9d9c9785a720f69eac9a

Observation eca5658d-1943-4dc2-8b1a-8f62dab531a8 · outbound

This paper cites Evo-vit: Slow-fast token evolution for dynamic vision transformer.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Evo-vit: Slow-fast token evolution for dynamic vision transformer

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.423172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:08.344854Z digest=sha256:07f4c7cfc586ecb0d309501d8d35d52961e4e70fdfadf9ac17c0105078027e6b

Observation 72664a1a-4b25-4c08-bda6-24766bf60a3e · outbound

This paper cites Global vision transformer pruning with hessian-aware saliency.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Global vision transformer pruning with hessian-aware saliency

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.254561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:08.423711Z digest=sha256:3075b7fa016b8a4f92124af864528776654bde409ff7a8cff6ecc3ee3392ab35

Observation 7c47a347-b9f8-4c80-99c0-4ebe1b3fd104 · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration A-vit: Adaptive tokens for efficient vision transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.129384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:08.504518Z digest=sha256:9cc98813bf3fd41d7847e7c647c583ab5e36285f5e73c7f1c11ccf4cd011a793

Observation 63afa598-10e5-4f57-88f8-492bde008644 · outbound

This paper cites Tokens-to-token vit: Training vision transformers from scratch on imagenet.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Tokens-to-token vit: Training vision transformers from scratch on imagenet

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:10.005623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:08.575250Z digest=sha256:bbdf05b6153b9bce79622653805caeba17769f27c3127da9ab3555d4cda48402

Observation 2e68c206-74c2-4a0f-8819-3c2255aeadc6 · outbound

This paper cites Vision trans- former with progressive sampling.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Vision trans- former with progressive sampling

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:08.661650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:08.661650Z digest=sha256:970be4710c62eee29a3c6ad41ad5298bb510be015faef247889dc7d1ccb22bd5

Observation a6d25efe-71d6-421f-a3f2-91d162418698 · outbound

This paper cites M2m-tag: Training-free many- to-many token aggregation for vision transformer accelera- tion.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration M2m-tag: Training-free many- to-many token aggregation for vision transformer accelera- tion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:09.878326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:08.730983Z digest=sha256:182966b1adf1eac33eda10b40362a7197aa5f1d004d059a3a1cf9d8eac875f5a

Observation 26b79f09-84bc-4d47-93d4-63977eb608fd · outbound

This paper cites Parameter efficient merging for multimodal large language models with complementary parameter adaptation.arXiv preprint arXiv:2502.17159, 2025.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Parameter efficient merging for multimodal large language models with complementary parameter adaptation.arXiv preprint arXiv:2502.17159, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:08.831348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:08.831348Z digest=sha256:f39833d14f2d315487cb26c2962b04d2fbae44334004c66056f1d643e8b2f380

Observation 826c9cc8-af54-414b-8c2e-909cd64c8e08 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:09.058019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:09.058019Z digest=sha256:a80f8542e9bc4a7cdb9631d1e3aeb03d88a8cd31ccf3dd248e29b916663de46a

Observation 9079dc22-0f25-4fa9-87b7-94dd8acfa084 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.International Journal of Computer Vision, 127:302–321, 2019.

Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration Semantic under- standing of scenes through the ade20k dataset.International Journal of Computer Vision, 127:302–321, 2019

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:09.750518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:29:09.281677Z digest=sha256:0c47fffdfb4d4bb88eabf72239aac7a41eca38b4ffef84b6db8807cf34238749

Pith citing papers

Observation cd619d8b-e160-4609-ba92-0965376c5715 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration

Reference 243

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:36.389460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:36.389460Z digest=sha256:7a21e593eaf2e8ee2881aa937492d5c99b2e76a7e9ebd680a0c48b078289d50b

Observation 176cd46b-7d57-4bd3-b913-244b12844db4 · inbound

MaMe & MaRe: Matrix-Based Token Merging and Restoration for Efficient Visual Perception and Synthesis cites this paper.

MaMe & MaRe: Matrix-Based Token Merging and Restoration for Efficient Visual Perception and Synthesis Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:26.675894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:06:09.392876Z digest=sha256:e0d97a7ba44455ce3b0965569caeacee151312e5330ea4331ce48adfc30f2430

Observation 3d6a8040-3b49-477c-ba34-4d78edc3078a · inbound

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors cites this paper.

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.966902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:31:52.126027Z digest=sha256:8e073dfe4f0d7b387e2b3ab428e925da702d9e4afb9b3cbba677b8dfcab6adb1