Pith. sign in

Paper Citation Record · LEDGER

Visual Transformers: Token-based Image Representation and Processing for Computer Vision

As of 23 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2006.03677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.03677 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T19:58:05.484186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:09:57.476599Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6577d602-f489-4662-b633-ea4472b74b5f · inbound

Defending against Backdoor Attacks via Module Switching cites this paper.

Defending against Backdoor Attacks via Module Switching Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:32:04.763637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-22T20:28:24.169272Z digest=sha256:964091082258d0aa758de263f7a5ce5103c7a87b6b6f868e6d45149e8b89a09d

Observation a3c569f9-f741-41f5-ad29-97d634cdb446 · inbound

MoDora: Tree-Based Semi-Structured Document Analysis System cites this paper.

MoDora: Tree-Based Semi-Structured Document Analysis System Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:10:15.528788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T19:09:41.482713Z digest=sha256:28be2237c2d30fef5793259fa221ba6a572fefe224917ac2bbd9b4e647d6ad9b

Observation 3c2b3914-1d60-4f3e-98bc-aa448f7b5a53 · inbound

Rethinking Intrinsic Dimension Estimation in Neural Representations cites this paper.

Rethinking Intrinsic Dimension Estimation in Neural Representations Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:54:48.758900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-10T00:50:49.134622Z digest=sha256:f6dc5ac626e6b18637cf8cfff8ddd14b35b59963fa2dd473ddf01fbbea7effcb

Observation 5dfddba1-9fa1-4cb7-9865-0301af09c7ba · inbound

Modeling Subjective Urban Perception with Human Gaze cites this paper.

Modeling Subjective Urban Perception with Human Gaze Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:47:15.929843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-09T19:12:54.038217Z digest=sha256:9c3cf473a06c7a878318e382c4dcafaed1425d5522e7010ae793e3d15d941deb

Observation 3ae6afd6-d479-48c3-9d89-a954915cdb15 · inbound

SAIL: Structure-Aware Interpretable Learning for Anatomy-Aligned Post-hoc Explanations in OCT cites this paper.

SAIL: Structure-Aware Interpretable Learning for Anatomy-Aligned Post-hoc Explanations in OCT Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:41.076077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T18:16:35.927272Z digest=sha256:749a43b7fb3305559610ecb929e2e008d2b70b231fa3bc9090a964f06faefd47

Observation fb19430e-446d-4162-a391-142c9a35af9e · inbound

M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification cites this paper.

M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:59:27.369856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-14T20:55:57.769186Z digest=sha256:5ecc9fa71d4f4fd635eedb839b346c130b381c16288df14485e43be8e9173f37

Observation aac02284-2aa9-466d-8eb3-fea5bf884b38 · inbound

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI cites this paper.

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:43:15.188039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-20T11:40:19.027987Z digest=sha256:76e0b4ae08f521f7251082f46b713f16de7aefd271fa20e47fa2dedb62e149ab

Observation 608cde3a-c82e-4b76-bedb-069ef682970b · inbound

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations cites this paper.

Deep Attention Reweighting: Post-Hoc Attention-Based Feature Aggregation in CNNs for Disentangling Core and Spurious Features under Spurious Correlations Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.768756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-21T05:04:59.177437Z digest=sha256:fda9603562430964a850bb8f76cfe739c3a7bb8ea395c98c159e6bd0e8256518

Observation 098a2994-ae5d-49be-9b8e-8abbe2e85b91 · inbound

Forged Calamity: Benchmark for Cross-Domain Synthetic Disaster Detection in the Age of Diffusion cites this paper.

Forged Calamity: Benchmark for Cross-Domain Synthetic Disaster Detection in the Age of Diffusion Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.327251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-26T21:32:27.296146Z digest=sha256:3ee05d2a81c2a1047c84cb1588e4ec216088dbe8bab3183147805e37003383b7

Observation e82670d2-c508-4a32-b824-a848a5d8f388 · inbound

Co-occurring associated retained concepts in Diffusion Unlearning cites this paper.

Co-occurring associated retained concepts in Diffusion Unlearning Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.478300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T00:50:58.689718Z digest=sha256:e112fb7eb84639ba59e8fe4f3bfdf83451b6fbfb7564f28fab8c4dfdaf77fcc7

Observation 32837a37-b408-4ae6-9442-6e54b233b313 · inbound

One Framework for All: Cross-Modal Membership Inference for Generative Models cites this paper.

One Framework for All: Cross-Modal Membership Inference for Generative Models Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T19:58:05.484186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:58:05.484186Z digest=sha256:fa276a30e79d120bcbed56b2466b7504a1b561d48ffedee8d9bae191d06914cf