Pith. sign in

Paper Citation Record · LEDGER

How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2106.10270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.10270 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:48:35.044872Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T18:15:59.170440Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c20ce74c-047f-410c-ac72-52b7bcc383ed · inbound

Sigmoid Loss for Language Image Pre-Training cites this paper.

Sigmoid Loss for Language Image Pre-Training How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:05:36.491069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:05:36.460932Z digest=sha256:568f59f230dde30c38b06a7c58c8fa9936441c0b6c26619159c9726c7d01b533

Observation 88d9375c-4005-4576-9c64-077b5c5fa2bc · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:20:20.323505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:f985f1170ffbbd7eb9f309770a56f7a2f69f8a7f4be75e5690906d7eb52c8821

Observation 1a7378d4-cace-4508-b514-cf3741ec5691 · inbound

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey cites this paper.

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:32:36.926712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T11:32:36.738536Z digest=sha256:ce8ab78de3a8e8a2d1c491941e1449a97ebbf56db00f98fd62f857a465393f24

Observation 0a682dcd-0704-4295-8db7-3ac96b14c373 · inbound

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control cites this paper.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:38:24.532658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:ca9a93e6e30ccc1daeab72927617f76f895f31645dd1377d5b63ee3161a70e3b

Observation 42585f21-2d84-489f-8b2a-e98b3c9bf0e0 · inbound

Moment Alignment: Unifying Gradient and Hessian Matching for Domain Generalization cites this paper.

Moment Alignment: Unifying Gradient and Hessian Matching for Domain Generalization How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:35.044872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:48:35.044872Z digest=sha256:3395f80036c9a8f325205d4600f8799284fc8e21a156ed00455bab888f6a2f65

Observation 96bd96aa-b993-4134-9747-392bb262feb0 · inbound

Hidden in plain sight: VLMs overlook their visual representations cites this paper.

Hidden in plain sight: VLMs overlook their visual representations How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.940685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.940685Z digest=sha256:26c4240078fb3bf8ff6d17e96d29e22391f53362a987d5a5c9a0c47d1d7a527b

Observation a7c6cb3a-4163-4dd1-afa1-20b9344c294f · inbound

InceptionMamba: An Efficient Hybrid Network with Large Band Convolution and Bottleneck Mamba cites this paper.

InceptionMamba: An Efficient Hybrid Network with Large Band Convolution and Bottleneck Mamba How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:14.444800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:07:14.444800Z digest=sha256:fa2dddee9917a3ad53da4940b093745c329896d4bce9011e802513bde602d867

Observation 3534d411-7f0c-4e82-8175-a40d15f136fb · inbound

ReStNet: A Reusable & Stitchable Network for Dynamic Adaptation on IoT Devices cites this paper.

ReStNet: A Reusable & Stitchable Network for Dynamic Adaptation on IoT Devices How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:44:09.801949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:44:09.801949Z digest=sha256:f12ccb740eeea756b833f2a68abe09a89bd104c3a36ffa489d57eabbd61b57ab

Observation f3e75e37-b0af-4a03-8caf-d19a2825939c · inbound

DeepTraverse: A Depth-First Search Inspired Network for Algorithmic Visual Understanding cites this paper.

DeepTraverse: A Depth-First Search Inspired Network for Algorithmic Visual Understanding How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:06.578252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:06.578252Z digest=sha256:04d1cc521b0dc513bb3ab18a59cf1d3ecbc99ec0a9f4e55a0d4877fd9b8ea513

Observation 7c5922de-eae6-47b2-9b43-a83d6eb5ce41 · inbound

Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets cites this paper.

Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:07.020403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:08:07.020403Z digest=sha256:59cd5549bb2325bec206bfd6d5dfbc0db3654a1d4f1c6b1d415fc03e659609ed

Observation 6b183a73-cc2c-4838-b1c5-c9976c695e2e · inbound

Underwater Monocular Metric Depth Estimation: Real-World Benchmarks and Synthetic Fine-Tuning with Vision Foundation Models cites this paper.

Underwater Monocular Metric Depth Estimation: Real-World Benchmarks and Synthetic Fine-Tuning with Vision Foundation Models How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:38.631360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:38.631360Z digest=sha256:aed3b481ecdf185809e319c58a8068c54207da57e694a89f6478bb77041e01d2

Observation a2560c90-33d3-4495-af18-358e1fb372a8 · inbound

Elastic ViTs from Pretrained Models without Retraining cites this paper.

Elastic ViTs from Pretrained Models without Retraining How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T09:03:09.342772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:03:09.342772Z digest=sha256:6574472aae289124e1f1e71f138a3ea144bcfcb6696cdaa61a449f90d4b08ca1

Observation 272e2d83-974f-48ea-9120-5bf5c4981d82 · inbound

CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging cites this paper.

CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:01:12.434996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:01:12.434996Z digest=sha256:9d63d51536a51b27dda3e7c8c414c3e771677393ecc605941168a122da746c1c

Observation f6043aef-f7c3-442b-9bf2-838d20253c18 · inbound

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment cites this paper.

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T23:34:16.095023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:34:16.095023Z digest=sha256:a912714c5e83e02822a2b3e15607c8cddc08ff4390cd1c7a98638b5871dcc2b0

Observation f42dac14-2c5d-4b95-924e-1dc28f5ab6fc · inbound

Causal Attribution via Activation Patching cites this paper.

Causal Attribution via Activation Patching How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:02.082743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:09:50.326389Z digest=sha256:77696bc0d94fab098d2e3d84d7ae9bcf5208c6c2cbc2a4f0fdfe2ea1ca2f1454

Observation 376ee791-ad4f-4850-a5d1-1d93161eae35 · inbound

Human-like Object Grouping in Self-supervised Vision Transformers cites this paper.

Human-like Object Grouping in Self-supervised Vision Transformers How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T21:34:48.709465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:34:48.709465Z digest=sha256:a3deb89f1fc3e5eb7b75aef70c07c8d76ff6ee7555be561c1ff126afb23f0f5b

Observation a1b08e9a-7280-4fcb-9bf2-3e635174cd1d · inbound

Decision-Aware Attention Propagation for Vision Transformer Explainability cites this paper.

Decision-Aware Attention Propagation for Vision Transformer Explainability How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:01:01.755658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:23:04.096579Z digest=sha256:724208065ab3c0c6e2d59f0a2e57a6364b73116d7b4c7740b8926c09e1e1a270

Observation b3c06afe-72a1-4860-af0a-567a789be8ec · inbound

Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm cites this paper.

Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:33:32.625105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T02:29:49.803834Z digest=sha256:c891273eae7cb373137e78ae67217215dd86b51e597c5a1e9e7683d71cd102a9

Observation 7494ba04-06cb-448f-aa2a-656cd19c5252 · inbound

ASAP: Attention Sink Anchored Pruning cites this paper.

ASAP: Attention Sink Anchored Pruning How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:14:45.564776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T08:12:13.451406Z digest=sha256:dc0a37ee155bf2360cf0a30cfe7dabf59fd0519712be78d8022dd982f553d9f8

Observation ae73bce2-3402-4ca3-b4fe-61ef2d1c7484 · inbound

Weierstrass Positional Encoding for Vision Transformers cites this paper.

Weierstrass Positional Encoding for Vision Transformers How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:50:23.743705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:48:36.533090Z digest=sha256:f44843cfce889802c37070a3148c10f3f09a47d752f40826168b9487a0835046

Observation 20056288-d1b1-4b65-933b-1b176bd3358e · inbound

Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge cites this paper.

Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:15:59.172041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T02:07:27.600563Z digest=sha256:f843209a025601d136d68774f7a809304b5594bbba5ea3af85370378288dfeb1

Observation ebba460c-3cad-45ec-bfdd-c603b8daea9e · inbound

Screening Is Effective for Visual Recognition cites this paper.

Screening Is Effective for Visual Recognition How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:08:45.376016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:08:45.376016Z digest=sha256:d1edc62f0a00279c3b4cb358f6a2e155aac5f58076f10927036f0d5683caa0e7

Observation cbcb636e-fb69-4db7-89e7-e96e75eb7dac · inbound

Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention cites this paper.

Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T13:33:15.789562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:33:15.789562Z digest=sha256:964764d058b899feddda8b73a7b7dec0a0e7a22e129fc1edf703f2bf26bb2ff8

Observation f1edc681-1762-4852-a8af-99802565b854 · inbound

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI cites this paper.

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 186

Resolution
unresolved
no resolver link, observed 2026-07-31T23:27:36.860612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:27:36.860612Z digest=sha256:7f31cc43c47feb263ac10490690faa1ff05093a3dc0ee39db8db257915f1bc50

Observation 8a185ae9-54d4-4b74-bfab-00f1a50c0bd0 · inbound

Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification cites this paper.

Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T13:20:38.168263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:20:38.168263Z digest=sha256:560621d3038dd64ba8b0738bc03c12d68a0812116457a365ea87e6a7b1127c1a