Pith. sign in

Paper Citation Record · LEDGER

Visual Transformers: Token-based Image Representation and Processing for Computer Vision

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2006.03677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.03677 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:22:59.772724Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:09:57.476599Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6577d602-f489-4662-b633-ea4472b74b5f · inbound

Defending against Backdoor Attacks via Module Switching cites this paper.

Defending against Backdoor Attacks via Module Switching Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:32:04.763637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T20:28:24.169272Z digest=sha256:616ecfb95c79506baaeb30f6119d5236ba6ac2c7f6c1f8db62753362c86d232d

Observation e34d4855-55ad-4fd9-b450-aebd848e2419 · inbound

Can Vision Transformers with ResNet's Global Features Fairly Authenticate Demographic Faces? cites this paper.

Can Vision Transformers with ResNet's Global Features Fairly Authenticate Demographic Faces? Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:22:59.772724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:22:59.772724Z digest=sha256:6486906911b5a351873f2794652667b411e6b37fc7cf173a8424abc8e17c5b28

Observation b99e02f5-99a0-4e77-a5c0-064d61fb9abf · inbound

Hierarchical Deep Feature Fusion and Ensemble Learning for Enhanced Brain Tumor MRI Classification cites this paper.

Hierarchical Deep Feature Fusion and Ensemble Learning for Enhanced Brain Tumor MRI Classification Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:35.034653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:55:35.034653Z digest=sha256:eec2bad10da87411daaaf2096a470915ed654cc2ebab1e8b0d7f13d146bef978

Observation 4c5aeb60-c08f-4187-af7a-0ab4da22106d · inbound

Hybrid Ensemble Approaches: Optimal Deep Feature Fusion and Hyperparameter-Tuned Classifier Ensembling for Enhanced Brain Tumor Classification cites this paper.

Hybrid Ensemble Approaches: Optimal Deep Feature Fusion and Hyperparameter-Tuned Classifier Ensembling for Enhanced Brain Tumor Classification Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:33.510651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:33.510651Z digest=sha256:e00703ec476a47b5db9b67db0eae62493124ef5292ae51f7452bcaa303a0472b

Observation 384dad4f-8e88-4557-ab0b-b3e25bf0b2e5 · inbound

On the Performance of Concept Probing: The Influence of the Data (Extended Version) cites this paper.

On the Performance of Concept Probing: The Influence of the Data (Extended Version) Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:21.395406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:21.395406Z digest=sha256:3a829d16e7aca4aa2b56835d759e7843c70e4d8795134ade2c8f901dd03cece6

Observation e306f25c-a077-4e99-91eb-7661ce9e69f1 · inbound

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems cites this paper.

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T12:39:41.651354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:39:41.651354Z digest=sha256:227964ebb834e7d9f12d2493952dcaafe33677aeee3fde55b10a70ae1f5d8f3e

Observation 3a97b8a1-aea6-48be-ae4e-aec3f4f4c7dd · inbound

Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models cites this paper.

Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:42:26.151279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:42:26.151279Z digest=sha256:ff89444a7b4f001252e396f16e3da46260b913e3d13204b1da4e34be842ee3f1

Observation babc0f36-4720-4574-bf17-7d6a1ba7dd3b · inbound

On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations cites this paper.

On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:43.996459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:29:43.996459Z digest=sha256:ab3fab33086fec96a46cb29c07246609f031c7052ec02e841199e738c1eb4122

Observation 838fa267-a067-487e-889a-15b41a2cd9a2 · inbound

Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection cites this paper.

Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T10:40:00.899017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:40:00.899017Z digest=sha256:62c3f7855c6cb13f952a151a6dd56955395b4c66466f2f68c5640b221b76e7fe

Observation 3b8d841e-b0ad-4195-8b38-0b4617977183 · inbound

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World cites this paper.

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-03T17:02:40.945788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:02:40.945788Z digest=sha256:6f8a8f8beeba57c641b870edfc388b3f6102daf714a42cb2ad7b8dd8fe0b97ed

Observation a3c569f9-f741-41f5-ad29-97d634cdb446 · inbound

MoDora: Tree-Based Semi-Structured Document Analysis System cites this paper.

MoDora: Tree-Based Semi-Structured Document Analysis System Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:10:15.528788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T19:09:41.482713Z digest=sha256:8fe270ae7c2bad6ecdce95e049c929ee2335ce92224997862831f14165255c7a

Observation 3c2b3914-1d60-4f3e-98bc-aa448f7b5a53 · inbound

Rethinking Intrinsic Dimension Estimation in Neural Representations cites this paper.

Rethinking Intrinsic Dimension Estimation in Neural Representations Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:54:48.758900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T00:50:49.134622Z digest=sha256:d1c870208f476fe3cf2ab5a0db2b93ba69f1f1d19e5f15290b9d6f6ce61cb613

Observation 5dfddba1-9fa1-4cb7-9865-0301af09c7ba · inbound

Modeling Subjective Urban Perception with Human Gaze cites this paper.

Modeling Subjective Urban Perception with Human Gaze Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:47:15.929843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T19:12:54.038217Z digest=sha256:870171b3f0b45f39ffc6ef17442cbdcf133d1525d7694b6838a5ba09e450c622

Observation 3ae6afd6-d479-48c3-9d89-a954915cdb15 · inbound

SAIL: Structure-Aware Interpretable Learning for Anatomy-Aligned Post-hoc Explanations in OCT cites this paper.

SAIL: Structure-Aware Interpretable Learning for Anatomy-Aligned Post-hoc Explanations in OCT Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:41.076077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:16:35.927272Z digest=sha256:c6def371ed48d54f754dade02ecd899a559ce30bc073836c9e613761feeef5e0

Observation fb19430e-446d-4162-a391-142c9a35af9e · inbound

M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification cites this paper.

M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:59:27.369856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:55:57.769186Z digest=sha256:30d6e2a59fc7a882c6575a75823cf9f4a3c64f198366d5f90634456f3bfa7f76

Observation aac02284-2aa9-466d-8eb3-fea5bf884b38 · inbound

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI cites this paper.

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:43:15.188039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:40:19.027987Z digest=sha256:63bdd0291de76f729548e284bfda160a3c707f8628090e8a47bf978f51ceb013

Observation 608cde3a-c82e-4b76-bedb-069ef682970b · inbound

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations cites this paper.

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.768756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:04:59.177437Z digest=sha256:147cfdefac6ec9c0d58e21987ea7dd96a8f39fe4381ecf47a714b7d6a087def2

Observation 098a2994-ae5d-49be-9b8e-8abbe2e85b91 · inbound

Forged Calamity: Benchmark for Cross-Domain Synthetic Disaster Detection in the Age of Diffusion cites this paper.

Forged Calamity: Benchmark for Cross-Domain Synthetic Disaster Detection in the Age of Diffusion Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.327251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:32:27.296146Z digest=sha256:392fabc6ab7e40575c6836a40d52d4d901fb2eea6bc9859362cd87af770431d1

Observation e82670d2-c508-4a32-b824-a848a5d8f388 · inbound

Co-occurring associated retained concepts in Diffusion Unlearning cites this paper.

Co-occurring associated retained concepts in Diffusion Unlearning Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.478300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T00:50:58.689718Z digest=sha256:eff262e515add77e7575b87fc11382ac98e0076949d0b0aa5d6d3a51c0e7d368

Observation 32837a37-b408-4ae6-9442-6e54b233b313 · inbound

One Framework for All: Cross-Modal Membership Inference for Generative Models cites this paper.

One Framework for All: Cross-Modal Membership Inference for Generative Models Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T19:58:05.484186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:58:05.484186Z digest=sha256:f93f6c635bedf852ff252a9db2a754dacbc069c9f2dc80f31ecc21dade2f91b7

Observation cf25f6a4-819f-4763-a036-9e8d143a63d5 · inbound

Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity cites this paper.

Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T00:38:04.305629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:38:04.305629Z digest=sha256:26c43263ba5052a33525fcb58d78d921d98b0e51ac2d912145b159b5b439b398