Pith. sign in

Paper Citation Record · LEDGER

ImageNet-21K Pretraining for the Masses

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2104.10972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.10972 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:57:57.751315Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32d51985-98c8-4073-87a5-01aa653416f0 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment ImageNet-21K Pretraining for the Masses

Reference 208

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.136192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:3b10cfbc478ccd4671abc2ffe9c5a52b03fe0c73f0b9962ee50736bb91deb5db

Observation bdb84ae6-aeaf-49be-8af4-16da472a8658 · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices ImageNet-21K Pretraining for the Masses

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:35:38.079877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:2bad86edec2a26b57ce39e9d5d986844068e6f38b10bd04b87c4bedc26999fad

Observation e9405bb7-a69d-49ef-9676-0926a939234a · inbound

Textile Analysis for Recycling Automation using Transfer Learning and Zero-Shot Foundation Models cites this paper.

Textile Analysis for Recycling Automation using Transfer Learning and Zero-Shot Foundation Models ImageNet-21K Pretraining for the Masses

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:57:57.751315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:57:57.751315Z digest=sha256:b23642998ee639f6ead52503d4f047cbe8b3617a3c0e89894291ade44eca24a0

Observation d46087ef-f70c-4d11-83e1-910194c6e1e2 · inbound

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer cites this paper.

RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer ImageNet-21K Pretraining for the Masses

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:41.086921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:41.086921Z digest=sha256:7cd5d387a2d498768189cae35c5416b255de39c018bda54070c4b283958abb2e

Observation 61e08455-777a-485e-922f-b50865cd6976 · inbound

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer cites this paper.

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer ImageNet-21K Pretraining for the Masses

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:34.691781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:34.691781Z digest=sha256:7d5c71b9a7bbd335220e2c2bf692996b1824f2afe1f3f805a51fc0fe0c39d602

Observation d1176881-a22d-4c8d-b9b1-462021c21c57 · inbound

Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics cites this paper.

Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics ImageNet-21K Pretraining for the Masses

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:54:55.009314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:54:55.009314Z digest=sha256:19a3add9648183420a720d943394e70080a7d79043b63d0e43e95b76e067890c

Observation 5cc46fc6-e0af-4ed7-87e3-17525577ec55 · inbound

Towards Continuous Home Cage Monitoring: An Evaluation of Tracking and Identification Strategies for Laboratory Mice cites this paper.

Towards Continuous Home Cage Monitoring: An Evaluation of Tracking and Identification Strategies for Laboratory Mice ImageNet-21K Pretraining for the Masses

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:37.546443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:37.546443Z digest=sha256:e0c3ffdfa1a986643f049dd2d71e6a5ed4497a91e497abb3bc7280aec9795e5a

Observation 01a8da17-0e19-4411-8b6e-b90d6a2969c2 · inbound

Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset cites this paper.

Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset ImageNet-21K Pretraining for the Masses

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T18:25:20.674966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:25:20.674966Z digest=sha256:c8cf1f1ca79957a6b9021c46d060dbfbd5069895b3b72507cecd7c64f04186e0

Observation 02754b74-4417-4d6c-9eb8-31a0afb277cf · inbound

ViT-ProtoNet for Few-Shot Image Classification: A Multi-Benchmark Evaluation cites this paper.

ViT-ProtoNet for Few-Shot Image Classification: A Multi-Benchmark Evaluation ImageNet-21K Pretraining for the Masses

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:33.053004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:33.053004Z digest=sha256:98615aec2e4fa3287f56a8747709c2d58125c2adefa49109dc2bda1c5a90a7aa

Observation 23983688-5fea-4bee-8774-9be931c53e71 · inbound

GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences cites this paper.

GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences ImageNet-21K Pretraining for the Masses

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:32.195913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:37:32.195913Z digest=sha256:def71bf6eb1f85ebc9c00985d18423ff9d935c0aad03a7d7fc1c29decdb86418

Observation 94e879dd-ac65-4130-b61b-a2a5b9be83cb · inbound

Page image classification for content-specific data processing cites this paper.

Page image classification for content-specific data processing ImageNet-21K Pretraining for the Masses

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:02:07.647676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T06:00:46.843420Z digest=sha256:b401e8ad461d120efc44ae56c9eafe55e82a468cdcc0896215ca6f2d6f03789c

Observation c73d2521-2941-46a6-b06b-16332c09f1b4 · inbound

CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition cites this paper.

CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition ImageNet-21K Pretraining for the Masses

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:13.653762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:13.653762Z digest=sha256:4c99586228f289c75a2ee0e58dafc025ca6ab0700423f58a8158edaf9f65a612

Observation 40f2d71d-1d61-46c5-8a1f-6909817928e5 · inbound

Revisiting Deepfake Detection: Chronological Continual Learning and the Limits of Generalization cites this paper.

Revisiting Deepfake Detection: Chronological Continual Learning and the Limits of Generalization ImageNet-21K Pretraining for the Masses

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T14:12:29.948773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:12:29.948773Z digest=sha256:7559ec875957b80754b81b7556d81f4ada8dd1f907528ed308f1bdd8821e7831

Observation ce1f79a4-4625-4815-a8c5-c41238e0c480 · inbound

H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification cites this paper.

H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification ImageNet-21K Pretraining for the Masses

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T22:31:35.719971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:31:35.719971Z digest=sha256:72711db0ffc2b015d9bc8e5d2948916c68946a140cd4050f08ecd7b802d0f274

Observation 6834d710-eeef-4c62-8529-f6fe3a45c3a8 · inbound

HandyLabel: Towards Post-Processing to Real-Time Annotation Using Skeleton Based Hand Gesture Recognition cites this paper.

HandyLabel: Towards Post-Processing to Real-Time Annotation Using Skeleton Based Hand Gesture Recognition ImageNet-21K Pretraining for the Masses

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:09:04.066271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T05:05:24.411291Z digest=sha256:950c7ce804a9a92d8b46595c52b2eef8f318c7b8896557002e7acc37f25a61a7

Observation 6afc2b4a-3fac-435a-a54a-eb8103f685fa · inbound

MePo: Meta Post-Refinement for Rehearsal-Free General Continual Learning cites this paper.

MePo: Meta Post-Refinement for Rehearsal-Free General Continual Learning ImageNet-21K Pretraining for the Masses

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:20:40.881579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T06:20:18.021294Z digest=sha256:d7c234cf11a37b666819cb95fecb107264ddd06586b3e2bc90fe9bfac815a8e9

Observation 4d064ed4-4093-4dff-957d-89d99d0cf494 · inbound

CanViT: Toward Active-Vision Foundation Models cites this paper.

CanViT: Toward Active-Vision Foundation Models ImageNet-21K Pretraining for the Masses

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:34:06.899210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T10:33:29.023955Z digest=sha256:f6fae0b90a0d112daefc7c1882edddfaa20bb44ebb0afa61c48c7b7f07ab1e2e

Observation b4ae7957-f184-4287-a003-5734e38f2935 · inbound

StableTTA: Improving Vision Model Performance by Training-free Test-Time Adaptation Methods cites this paper.

StableTTA: Improving Vision Model Performance by Training-free Test-Time Adaptation Methods ImageNet-21K Pretraining for the Masses

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:48.411580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:40:24.850441Z digest=sha256:796bb615266e2c1da0c13f1eadfda9b09a878d90320d1d098e551c7ce86f6b7c

Observation a3e5ab55-180b-4b23-a54b-ca17041d427c · inbound

4th Workshop on Maritime Computer Vision (MaCVi): Challenge Overview cites this paper.

4th Workshop on Maritime Computer Vision (MaCVi): Challenge Overview ImageNet-21K Pretraining for the Masses

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:01:00.658333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:18:34.717325Z digest=sha256:6d76cea49577123780c07d8047e05efcd2cc0a40aa3386f20cceed8665993101

Observation d535ac1e-eed2-402a-b86e-40f7675c8ba6 · inbound

Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples cites this paper.

Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples ImageNet-21K Pretraining for the Masses

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:06:10.151812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:40:39.953351Z digest=sha256:e8e193268aa912452cc060e255d3e3ba14ab5372ec65d0c1e04fd0446c0fc8db

Observation fa154e1d-885d-4e6f-8a4d-b8226d9504a2 · inbound

Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising cites this paper.

Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising ImageNet-21K Pretraining for the Masses

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T08:02:31.950023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T07:58:21.350747Z digest=sha256:c435450cc9d7031dbe322d14128e7e4b696c27853cf1cc4d54e75b19e5f57be9

Observation 4f1722eb-e688-4bdb-af02-11db97b3427d · inbound

Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs cites this paper.

Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs ImageNet-21K Pretraining for the Masses

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.789189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:26:02.947081Z digest=sha256:46595e00fbad1ed2a09e78708805e49b498766cb55883db6d1c3109a523a9eee

Observation e944fe83-b74b-4ac2-b6b9-99012df630cd · inbound

Weierstrass Positional Encoding for Vision Transformers cites this paper.

Weierstrass Positional Encoding for Vision Transformers ImageNet-21K Pretraining for the Masses

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:50:23.761775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:48:36.533090Z digest=sha256:c0209024d056227852bd7d30fe338ec760507881499bc550b53ce4217771af9c

Observation ea4a2e2c-6963-4ab4-a97d-1a41fd8fbdb0 · inbound

JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search cites this paper.

JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search ImageNet-21K Pretraining for the Masses

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:48.405563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:12:21.124678Z digest=sha256:72b8a1a72c6fae35a44f108e8b61792c50ee045970e1cb24454c6ae5d65feead

Observation 732d3010-6621-4ce0-8c2b-a428c0fc4245 · inbound

Energy-Structured Low-Rank Adaptation for Continual Learning cites this paper.

Energy-Structured Low-Rank Adaptation for Continual Learning ImageNet-21K Pretraining for the Masses

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:54.816900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T19:38:04.181839Z digest=sha256:f8202537cccd3cd5437e61625e044d98c33850c28587a58e96759a04d0f11adf

Observation 5bc5ad43-1aa2-4bc4-a926-492004439a8b · inbound

LV-OSD: Language-Vision-Complementary Open-Set Object Detection cites this paper.

LV-OSD: Language-Vision-Complementary Open-Set Object Detection ImageNet-21K Pretraining for the Masses

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.939534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T13:44:59.042817Z digest=sha256:b8a72b8a83341819363d325cc4ccfc5de4541a2381138ea47039a19fe8d1fdc9

Observation aa4c9160-869c-476a-aa14-8fce30c641aa · inbound

VISReg: Variance-Invariance-Sketching Regularization for JEPA training cites this paper.

VISReg: Variance-Invariance-Sketching Regularization for JEPA training ImageNet-21K Pretraining for the Masses

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:36:17.080856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T15:17:59.919440Z digest=sha256:d85afcea9336f516faf51de18531f4eb4b8827c1872f48e1dd17f2031e691789

Observation 40ad3c0e-de80-4da6-998e-ed62e8ff5426 · inbound

Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing cites this paper.

Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing ImageNet-21K Pretraining for the Masses

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:54:01.074649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:50:29.393207Z digest=sha256:ee07d42a5ef42f6ae7b647d08671dc7a3ebc7326e96a04ca1ec301f5882540b2

Observation 74204860-2cd5-423d-8409-a90c5ba0ff06 · inbound

ExDet: Open-Domain Open-Vocabulary Detection with Cross-modal Extrapolation and Rectification cites this paper.

ExDet: Open-Domain Open-Vocabulary Detection with Cross-modal Extrapolation and Rectification ImageNet-21K Pretraining for the Masses

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.396703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:53:37.521755Z digest=sha256:a285823b0ca4c476634aa9f12f545c91e8e5b2437401e610e71c35f06777f187

Observation a3dac1e9-fd39-4d00-8d19-fc9988a1b0f6 · inbound

The Edge-on Galaxies in the DESI survey (EGIDE): sample building and photometry cites this paper.

The Edge-on Galaxies in the DESI survey (EGIDE): sample building and photometry ImageNet-21K Pretraining for the Masses

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.695732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:24:51.042931Z digest=sha256:5eca7329a010e37a6fc079404ba040841f972bb69246307352fd3bf2688e2b7f

Observation 03836d3b-efdb-40cd-a471-8243035c081f · inbound

The Edge-on Galaxies in the DESI survey (EGIDE): sample building and photometry cites this paper.

The Edge-on Galaxies in the DESI survey (EGIDE): sample building and photometry ImageNet-21K Pretraining for the Masses

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-02T11:12:41.905093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:12:41.905093Z digest=sha256:57b8d5e34fe40ef1a95a276063a44405853fc2dac001cbc0e2f49a17d67ddd83

Observation 18c8b1b4-0bc3-4036-9009-fc1b5502a2e4 · inbound

Identifying Latent Concepts and Structures for Generalized Category Discovery cites this paper.

Identifying Latent Concepts and Structures for Generalized Category Discovery ImageNet-21K Pretraining for the Masses

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:47:03.336260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-02T14:42:01.334822Z digest=sha256:b53372933df21f47181f3620195d304b878f0941ec329c6af0d75775f1ce7f8f

Observation 3998c07f-7f0a-4624-8570-6e4eb3fcf47f · inbound

Condensing Large-Scale Datasets Directly with Minimal Information Loss cites this paper.

Condensing Large-Scale Datasets Directly with Minimal Information Loss ImageNet-21K Pretraining for the Masses

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:07:02.387711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T13:59:59.488216Z digest=sha256:ecd87bd1c6476975212cd53166c3ce97b5e20fc81d363c5a13dfbf0e8d76efea