Pith. sign in

Paper Citation Record · LEDGER

How Do Vision Transformers Work?

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2202.06709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.06709 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:00:16.123817Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

204
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1e2d7390-831f-40a1-a33d-2091a0f01ffd · inbound

D-Cube: Exploiting Hyper-Features of Diffusion Model for Robust Medical Classification cites this paper.

D-Cube: Exploiting Hyper-Features of Diffusion Model for Robust Medical Classification How Do Vision Transformers Work?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T19:00:16.123817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:00:16.123817Z digest=sha256:ab120083cbed95fe3fdfee8d02bf54e0a889a34bf3f36e3de0c43993d05254e6

Observation 7650f84e-c0e3-44f3-b0a4-4ba29f4e6669 · inbound

SatVision-TOA: A Geospatial Foundation Model for Coarse-Resolution All-Sky Remote Sensing Imagery cites this paper.

SatVision-TOA: A Geospatial Foundation Model for Coarse-Resolution All-Sky Remote Sensing Imagery How Do Vision Transformers Work?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:41:01.233642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:41:01.233642Z digest=sha256:d19423c20e9eb0967f8e5324f6d0dc1a48b52983a7b994f757bda9c373b356cf

Observation f48bb99c-1587-4458-8070-ccf9bf6eabf2 · inbound

Enhancing Parameter-Efficient Fine-Tuning of Vision Transformers through Frequency-Based Adaptation cites this paper.

Enhancing Parameter-Efficient Fine-Tuning of Vision Transformers through Frequency-Based Adaptation How Do Vision Transformers Work?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:01.262851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:01.262851Z digest=sha256:5bf97d96b95e11749fbe83ee6b3672c8db902981580de5fdfc42cd660a2625f3

Observation 19b47c2d-2810-4f1b-a698-59836a879e85 · inbound

Adaptive High-Pass Kernel Prediction for Efficient Video Deblurring cites this paper.

Adaptive High-Pass Kernel Prediction for Efficient Video Deblurring How Do Vision Transformers Work?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:21:34.702751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:21:34.702751Z digest=sha256:121363efdf8dc649e45cde90bea45f062a118d7c9ac499c831d09eaa36281113

Observation b2ca3d00-ea22-4df4-ab9f-b0f2c8ec27b2 · inbound

Understanding Transformer-based Vision Models through Inversion cites this paper.

Understanding Transformer-based Vision Models through Inversion How Do Vision Transformers Work?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T19:39:39.182545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:39:39.182545Z digest=sha256:6be5c5d0774e6c9df168a8d2228ebc296e811a74d3ebf53bc4c0cd8f126dfaa8

Observation 72bb7e87-3ac8-41ff-b091-1070b4e882a2 · inbound

Boosting ViT-based MRI Reconstruction from the Perspectives of Frequency Modulation, Spatial Purification, and Scale Diversification cites this paper.

Boosting ViT-based MRI Reconstruction from the Perspectives of Frequency Modulation, Spatial Purification, and Scale Diversification How Do Vision Transformers Work?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:00.556711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:42:00.556711Z digest=sha256:0a363699ed5883070bb2bb8035ea3a4988cb9fa7dac65ea2c2ac04e7100ecc60

Observation 4a6a50fc-4cfb-4b32-bc6b-ad6063d83578 · inbound

SP$^2$T: Sparse Proxy Attention for Dual-stream Point Transformer cites this paper.

SP$^2$T: Sparse Proxy Attention for Dual-stream Point Transformer How Do Vision Transformers Work?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T14:53:36.321211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:53:36.321211Z digest=sha256:d9056ae5dd181e445f022d17cbde6278d0437469d724b8373d2005d03f180ba3

Observation adc93ca0-f0bd-419c-bb9f-691a76683825 · inbound

DH-Mamba: Exploring Dual-domain Hierarchical State Space Models for MRI Reconstruction cites this paper.

DH-Mamba: Exploring Dual-domain Hierarchical State Space Models for MRI Reconstruction How Do Vision Transformers Work?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:05.702588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:05.702588Z digest=sha256:bff215cbf5a817a062ee5c3ce30b6b30fc4e7aed86d4de2352c954b3a6314ecc

Observation 50ead4e3-8886-4ec9-9447-d9d80d9ae90e · inbound

Bridging the Sim2Real Gap: Vision Encoder Pre-Training for Visuomotor Policy Transfer cites this paper.

Bridging the Sim2Real Gap: Vision Encoder Pre-Training for Visuomotor Policy Transfer How Do Vision Transformers Work?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:11.982373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:11.982373Z digest=sha256:c29bd73fdb86aab9809b4515cc76bfa278747873b43601c0bf85aec781868ceb

Observation 5aa6e507-018f-42ba-b478-ca5a1f3a2ebd · inbound

Beyond the Permutation Symmetry of Transformers: The Role of Rotation for Model Fusion cites this paper.

Beyond the Permutation Symmetry of Transformers: The Role of Rotation for Model Fusion How Do Vision Transformers Work?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T19:44:20.891764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:44:20.891764Z digest=sha256:2029a8f762165ce8e07b5ce672a017c9ae8c4654b215d1b1b2fd5fae0f57cad1

Observation bda6bfa4-05b9-49aa-ae3c-9d989db58cf9 · inbound

Cross Paradigm Representation and Alignment Transformer for Image Deraining cites this paper.

Cross Paradigm Representation and Alignment Transformer for Image Deraining How Do Vision Transformers Work?

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:01:54.628626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T17:57:05.010694Z digest=sha256:8ccff2771742cd5386b2657608b653a026c8ea12fecdca0a5795ca24a0eeadc3

Observation dd66d163-2f4f-417e-804b-e63a0360b37e · inbound

Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting cites this paper.

Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting How Do Vision Transformers Work?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:12:12.005496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:12:12.005496Z digest=sha256:bce077dfe7fb2b770062039b893da2bd77c5de6eb089a4ffdb702050716f35ac

Observation 544ab5a6-473d-431d-97d0-72d8e1bee1bd · inbound

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer cites this paper.

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer How Do Vision Transformers Work?

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:01.490975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:07:41.426142Z digest=sha256:a58bd4e91eef551597621162a259529c5c3f9700b294a0e186edbcb5ae092e48

Observation 440e064b-0a4d-49e3-b3f9-30365f4aedd3 · inbound

FreqTrack: Frequency Learning based Vision Transformer for RGB-Event Object Tracking cites this paper.

FreqTrack: Frequency Learning based Vision Transformer for RGB-Event Object Tracking How Do Vision Transformers Work?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:05:22.452546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T12:01:05.411686Z digest=sha256:49f415a0454796b11584450547ca954d117ba91d820825a65c3ce54037f171d3

Observation 4692f353-432d-49d6-9805-1adb7080aac9 · inbound

Scale-Aware Adversarial Analysis: A Diagnostic for Generative AI in Multiscale Complex Systems cites this paper.

Scale-Aware Adversarial Analysis: A Diagnostic for Generative AI in Multiscale Complex Systems How Do Vision Transformers Work?

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T19:56:16.519623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T19:56:13.911743Z digest=sha256:ba8c05b42dc1620badba05543e97d5fa512011d6e8cbc12f994f00af720a8600

Observation 220c00fd-10c2-4456-9652-d5b28448e968 · inbound

PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting cites this paper.

PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting How Do Vision Transformers Work?

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:21:26.219648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T03:27:14.464987Z digest=sha256:88f384378ccc478610eb8751c95d5c7529c40db979da3307c99a85ddf61c2007

Observation 299616c6-00ad-440c-9001-f93df9d438dd · inbound

PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting cites this paper.

PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting How Do Vision Transformers Work?

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:22:59.182144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T21:21:32.256476Z digest=sha256:ddbedbd558c1a64ea162734ce030921ef4c7d62d94ab47b0ed696301d8c36655

Observation 09948d8c-3c39-45c1-bc5c-3fc16827ce0a · inbound

Rethinking Random Transformers as Adaptive Sequence Smoothers for Sleep Staging cites this paper.

Rethinking Random Transformers as Adaptive Sequence Smoothers for Sleep Staging How Do Vision Transformers Work?

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.200349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T04:41:48.085125Z digest=sha256:2a8c180f9fd7e5df43442c601ccbeaf7b3d18fb902f439766ab378f7f1f8864e

Observation 27a182b0-741e-4648-8420-d61fea2a778e · inbound

Elastic Attention Cores for Scalable Vision Transformers cites this paper.

Elastic Attention Cores for Scalable Vision Transformers How Do Vision Transformers Work?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:22.512019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T06:02:40.158866Z digest=sha256:4b07adaafa9dbdb4d018b2ee0017a3e20a1ddfe795686693ea19d5ffe24f381e

Observation a0764c37-b79e-4187-8c10-b8f5152f1a1d · inbound

CineMatte: Background Matting for Virtual Production and Beyond cites this paper.

CineMatte: Background Matting for Virtual Production and Beyond How Do Vision Transformers Work?

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:38:14.674704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T11:34:37.193510Z digest=sha256:dffd68cabcb449cabca179992bba8b83877384b11cae26a160f52f6efe15d91e

Observation 836ebc44-a2f0-4a56-9f51-331eefa78b30 · inbound

FUSE: Frequency-domain Unification and Spectral Energy Alignment for Multi-modal Object Re-Identification cites this paper.

FUSE: Frequency-domain Unification and Spectral Energy Alignment for Multi-modal Object Re-Identification How Do Vision Transformers Work?

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:30.617056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T18:18:27.816492Z digest=sha256:b0fb317adeb4ece4d226c62f45249fbe135ad4b9cbf5180c92a2bdda5477a331

Observation d67ecfaa-d20f-41f2-a3bd-060dbb4a3c22 · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM How Do Vision Transformers Work?

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:29.182535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:63aa761d11b918af6776fc408e42540a75962248cbba766abf88c44a8f9a4733

Observation e5c4ed93-bac7-4d17-8a6b-41ae496d217b · inbound

SqLinear: Balanced Square Partitioning Makes Linear Interaction Sufficient for Large-Scale Traffic Forecasting cites this paper.

SqLinear: Balanced Square Partitioning Makes Linear Interaction Sufficient for Large-Scale Traffic Forecasting How Do Vision Transformers Work?

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:59:38.406344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T14:55:59.023533Z digest=sha256:cc1f3e7f2c8a6b88aa2822702cdb05dfbd6041faa321930a0552a1a28c19d279

Observation 86da79d6-8736-4b29-840d-ebccddcce57f · inbound

SqLinear: Balanced Square Partitioning Makes Linear Interaction Sufficient for Large-Scale Traffic Forecasting cites this paper.

SqLinear: Balanced Square Partitioning Makes Linear Interaction Sufficient for Large-Scale Traffic Forecasting How Do Vision Transformers Work?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T02:09:39.904070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:09:39.904070Z digest=sha256:6b451849173a3ef3bf9a5c462885c61d0bd038918fe00a240ea3e0c7a51c7107

Observation cdbb24d2-c8f0-42a3-b67d-3fc0868c7e87 · inbound

Does Your ViT Still Need U-Net for Segmentation? cites this paper.

Does Your ViT Still Need U-Net for Segmentation? How Do Vision Transformers Work?

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:17:17.696415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-02T19:15:07.611358Z digest=sha256:c8d496307a4b0107a41ace77387288c088aa80f6e4693f451262ac9601daed89

Observation 0b647541-45ed-4121-bd76-6999407e8728 · inbound

Identifying Latent Concepts and Structures for Generalized Category Discovery cites this paper.

Identifying Latent Concepts and Structures for Generalized Category Discovery How Do Vision Transformers Work?

Reference 213

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:47:03.332595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-02T14:42:01.334822Z digest=sha256:c885a24592899cab09c8294970b632fbe4e6d0bbcaa3f816c164d825333b757d

Observation 31affc47-fa6f-4d0a-b944-70e6e0874057 · inbound

SFKD: Spatial--Frequency Joint-Aware Heterogeneous Knowledge Distillation via Multi-Level Wavelet Spectral Interaction cites this paper.

SFKD: Spatial--Frequency Joint-Aware Heterogeneous Knowledge Distillation via Multi-Level Wavelet Spectral Interaction How Do Vision Transformers Work?

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:08:36.800516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T16:02:36.425762Z digest=sha256:596a87bf3c247aa24dad0b838924b5dcf293a0d40e0f03f22717a6a64c0a438c

Observation 8770de0d-8ddd-477e-8fb7-dcb13f6b0beb · inbound

OmniDS: Dual-Stream Context Fusion for Omnidirectional Depth from Fisheye Cameras cites this paper.

OmniDS: Dual-Stream Context Fusion for Omnidirectional Depth from Fisheye Cameras How Do Vision Transformers Work?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T05:17:21.435075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:17:21.435075Z digest=sha256:ed6107908dbc963349009fef98f3f4333fe16c4305b2b9c57a18c5f5e793cc51

Observation a12d4c63-85c2-490e-9b1b-ccbb096033d8 · inbound

GeoSAM-Lite: A Lightweight Foundation Model for Onboard Remote Sensing Segmentation cites this paper.

GeoSAM-Lite: A Lightweight Foundation Model for Onboard Remote Sensing Segmentation How Do Vision Transformers Work?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T00:08:46.276767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:08:46.276767Z digest=sha256:fd6cced0afdcd4139aaf185e23374249d96f163c5e3bbc200f94a30c63f21aaf

Observation d25ae062-4547-449f-9a14-bbad2cd86fac · inbound

FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small Object Detection cites this paper.

FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small Object Detection How Do Vision Transformers Work?

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.482739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T01:13:05.253777Z digest=sha256:0900a17d45992dd398cef9c631dbdc5f1e3a3bbf4313aa7aec50645d99368c81

Observation d75ce1ee-b0e0-4468-97f1-3096d6b11345 · inbound

FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small Object Detection cites this paper.

FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small Object Detection How Do Vision Transformers Work?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T07:49:40.867148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:49:40.867148Z digest=sha256:702273a71cf3f594a9fdb889084728a70932c6228c37ead719a8b92f71064252

Observation 8b5802a2-59e6-42b1-8124-ef089b4dca47 · inbound

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control cites this paper.

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control How Do Vision Transformers Work?

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:45.229506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:45.229506Z digest=sha256:15a77f535390b2c7c1b50b743fc44f235461adca851176a61f5e1da1e4d851d1