Pith. sign in

Paper Citation Record · LEDGER

Learning to Merge Tokens in Vision Transformers

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2202.12015.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.12015 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:23.804525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:09:36.498887Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c65e594d-c154-445b-a67a-def22ddd5241 · inbound

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation cites this paper.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Learning to Merge Tokens in Vision Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.804525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.804525Z digest=sha256:350025c210621ad8c39c9358f43e147b62f7cb73e8935bddd32481a1e35be027

Observation 06250222-426a-428f-9eb6-6a309faf1aba · inbound

Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI cites this paper.

Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI Learning to Merge Tokens in Vision Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:53:26.621820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:53:26.621820Z digest=sha256:6081c982b22b42c550f37c84834ce161b6407457a7da3808092588c128ec766e

Observation 5bff8415-6a1c-46a9-8a94-14605872bba8 · inbound

Training-free Token Reduction for Vision Mamba cites this paper.

Training-free Token Reduction for Vision Mamba Learning to Merge Tokens in Vision Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:16:49.380798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:16:49.380798Z digest=sha256:e8b49250c185c02ab9f8767ea6cf519ab5b957e981bb4422ba0df5703e9e1bb1

Observation 78e4a7eb-1f02-45fb-8b56-410e0465d0d5 · inbound

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer cites this paper.

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer Learning to Merge Tokens in Vision Transformers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:36:06.967355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T23:36:06.870759Z digest=sha256:002b641c5f88619228364b22e7761173001bd4ad603262ea7012b31c940d5f60

Observation 396f89f9-983c-400e-95c0-1009c09e924e · inbound

Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding cites this paper.

Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding Learning to Merge Tokens in Vision Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.375051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:31:37.375051Z digest=sha256:d9faf52b6bd496f1a6f06a9a956862c9a807b4f08557edf99e7f85b0c01542d6

Observation 64a6bd45-73c4-433e-8137-78834c6d13f6 · inbound

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors cites this paper.

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors Learning to Merge Tokens in Vision Transformers

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.956741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:31:52.126027Z digest=sha256:836757b68b9fec08b9ab33f21e94268af34c32570ee86b395c13065c11e5efea

Observation 934e1f69-e30d-4e48-bf54-9d239a0a191a · inbound

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference cites this paper.

ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference Learning to Merge Tokens in Vision Transformers

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:29.913691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T18:17:53.013043Z digest=sha256:8fd7c9363c30c81b087249faa1ac0fa69e25bcd61165bbddfe8a095ba1b83203

Observation 43b2b66b-a933-4f03-ab63-9851d71c198b · inbound

ConsisFormer: Compute-Efficient Transformer for Wireless Foundation Models Based on Channel Consistency cites this paper.

ConsisFormer: Compute-Efficient Transformer for Wireless Foundation Models Based on Channel Consistency Learning to Merge Tokens in Vision Transformers

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.501400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T16:16:12.898027Z digest=sha256:98573cd4ce2ca06a1a28e365741038d5d0e1c5575169aa38c3d8b716f6c21bfe

Observation 1943de30-9f1e-4531-82f7-11f2f0aaf79b · inbound

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs cites this paper.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Learning to Merge Tokens in Vision Transformers

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:22.142467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:54ce9fe095f0a45d80f2323d290580a8b3bd10b9da5c335d3af1dbc35231e314

Observation f15b98f5-94d5-4933-a1a3-6035195e3a36 · inbound

REDI: Corpus Aware Patch Ranking for DINOv3 Token Reduction cites this paper.

REDI: Corpus Aware Patch Ranking for DINOv3 Token Reduction Learning to Merge Tokens in Vision Transformers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:55:41.643978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:57:55.575289Z digest=sha256:f0c8b234c907b7036ab711efa8ca5e1475f22728bc4b475c54b0726216414b0c