Pith. sign in

Paper Citation Record · LEDGER

Multi-Token Enhancing for Vision Representation Learning

As of 13 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2411.15787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15787 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:59:36.446502Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy52
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c0c18bca-5231-4214-b82b-0e51ceb643c7 · outbound

This paper cites Asano, Christian Rupprecht, and Andrea Vedaldi.

Multi-Token Enhancing for Vision Representation Learning Asano, Christian Rupprecht, and Andrea Vedaldi

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.518865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.162993Z digest=sha256:03712ba7baf8fa691ce149ce862690b3207ed61fbd0028a53d5e36e8856f44d3

Observation fb9c483a-0a2e-4aa0-95d1-237f9ae77cf2 · outbound

This paper cites BEit: BERT pre-training of image transformers.

Multi-Token Enhancing for Vision Representation Learning BEit: BERT pre-training of image transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.507105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.167722Z digest=sha256:13910f5ccf425d6914b21d3ccea9c7aaea6ec2f3d024f45b41a1a0f2cb5f552a

Observation 81363a80-e5b7-4c4c-93aa-0416cae652cb · outbound

This paper cites Cascade r-cnn: Delving into high quality object detection.

Multi-Token Enhancing for Vision Representation Learning Cascade r-cnn: Delving into high quality object detection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.493756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.171857Z digest=sha256:1d901cd25dcf31a299c993776603391c6c104359eaaee511fb34126c7ea448f3

Observation 17b47991-8b28-4e0b-9141-b7f273e49974 · outbound

This paper cites Unsupervised learn- ing of visual features by contrasting cluster assignments.

Multi-Token Enhancing for Vision Representation Learning Unsupervised learn- ing of visual features by contrasting cluster assignments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.482258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.175892Z digest=sha256:b1cc824939e4465762ae510064d90681d14847b454f1de9725ad6b9e2c58f0b1

Observation 0eb08d6c-8a62-4bb8-87af-39508d7990a7 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Multi-Token Enhancing for Vision Representation Learning Emerg- ing properties in self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.472053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.180754Z digest=sha256:ca808f226c69109a1e52adbb03a1b7f450ca6713dababb00a589a0a6bb84d29a

Observation e5b8be40-8107-4031-972d-4d6a660ce06f · outbound

This paper cites Mixed autoencoder for self-supervised visual representation learning.

Multi-Token Enhancing for Vision Representation Learning Mixed autoencoder for self-supervised visual representation learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.459500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.185131Z digest=sha256:b005633c432a4953d4ac04c3cf8a93a1a4d260c6351c1610ede9bc0ac4fbf38f

Observation bae4412f-d772-4e0f-80d8-54654691c16e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Multi-Token Enhancing for Vision Representation Learning A simple framework for contrastive learning of visual representations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.446793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.189448Z digest=sha256:377189a97798979ad02b02add53dfba147e729f33d8a8ad6bb165a10aa40cd66

Observation 4987cec4-5022-4483-86b2-723da00bcfe6 · outbound

This paper cites Exploring simple siamese rep- resentation learning.

Multi-Token Enhancing for Vision Representation Learning Exploring simple siamese rep- resentation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.434514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.193352Z digest=sha256:148942e5bf6f1a1846472cf7d74c041cff1db0bb261a56d11ad7889560e28ea0

Observation 7b450849-3c64-41c3-b76d-bf7cd4ad6571 · outbound

This paper cites An empiri- cal study of training self-supervised vision transformers.

Multi-Token Enhancing for Vision Representation Learning An empiri- cal study of training self-supervised vision transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.423402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.197964Z digest=sha256:5d4f55dd2535fff9bf011a7590279703abf9d640ea7194e0c1408c1e7f2b1439

Observation 1c7a8955-0b96-42f2-b548-25581cf233fe · outbound

This paper cites Convit: Improving vision transformers with soft convolutional inductive biases.

Multi-Token Enhancing for Vision Representation Learning Convit: Improving vision transformers with soft convolutional inductive biases

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.410405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.202247Z digest=sha256:500e2e96fbc5b2eefa10fb050fe0813cc48395a134d95071f1be763bcb3dc9cf

Observation d5345f7c-f55e-466c-8da6-f9b683ef5111 · outbound

This paper cites Ensemble methods in machine learn- ing.

Multi-Token Enhancing for Vision Representation Learning Ensemble methods in machine learn- ing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.397794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.205999Z digest=sha256:a5caad4715aede8d3ed685e12c60791d9ba3f0dcb6489fd2fc47632d6758a635

Observation 7c14abbd-a1ad-4f6c-9577-b8b7051979d7 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Multi-Token Enhancing for Vision Representation Learning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.385594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.210402Z digest=sha256:8edaac18e28445a5794976eef340890683350fcad14e6b5c71d12c19629f68b4

Observation 59e7332b-e982-491d-9d09-8a3b37c71d86 · outbound

This paper cites Whitening for self-supervised representation learning.

Multi-Token Enhancing for Vision Representation Learning Whitening for self-supervised representation learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.371780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.215043Z digest=sha256:bde00b5479be9c42ccf5f8256fb88fd479ec708699772f40910b87238a1a3139

Observation f324fea3-b88c-4cbd-95d9-b320c9086d2e · outbound

This paper cites Seed: Self-supervised dis- tillation for visual representation.

Multi-Token Enhancing for Vision Representation Learning Seed: Self-supervised dis- tillation for visual representation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.358369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.218838Z digest=sha256:98cfa330d726b7544a2ee34c5ffb607c6506e7fc0adcbec3647eec812e757551

Observation f8a93134-6acc-4218-9dec-85847ba44a80 · outbound

This paper cites Evolved part masking for self-supervised learning.

Multi-Token Enhancing for Vision Representation Learning Evolved part masking for self-supervised learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.346320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.222715Z digest=sha256:a5b3a2fe01faa7085e0cb7dce5a5c44acbf3e9fa59ff5a982c126de30af70de5

Observation 1df1351e-b688-472e-8008-58ad63218e85 · outbound

This paper cites Ganaie, Minghui Hu, A.K.

Multi-Token Enhancing for Vision Representation Learning Ganaie, Minghui Hu, A.K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.330615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.227572Z digest=sha256:7f55e79921d86d8547a8ec442f5771d17ef7fa62584f0323ce516f8f09614058

Observation cc5b256c-db5d-4777-a2d4-2980d658d30e · outbound

This paper cites Large-scale un- supervised semantic segmentation.

Multi-Token Enhancing for Vision Representation Learning Large-scale un- supervised semantic segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.316715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.232442Z digest=sha256:31fb87ba867f44b8900e92d13cf5e048bd92bd6b240bc5f24be38825dfd2398d

Observation 87d73ed9-4a22-4f77-a574-d3764fe50431 · outbound

This paper cites Richemond, Elena Buchatskaya, Carl Doersch, Bernardo ´Avila Pires, Zhaohan Guo, Moham- mad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, R´emi Munos, and Michal Valko.

Multi-Token Enhancing for Vision Representation Learning Richemond, Elena Buchatskaya, Carl Doersch, Bernardo ´Avila Pires, Zhaohan Guo, Moham- mad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, R´emi Munos, and Michal Valko

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.245865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.236321Z digest=sha256:173bbbc96ac4ff6df4888615f2f467097b6f761e12f26d380f007c1053481e89

Observation 0007e657-5c59-4277-9ac2-129af3be0f57 · outbound

This paper cites Visual Attention Network.

Multi-Token Enhancing for Vision Representation Learning Visual Attention Network

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.240463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.240463Z digest=sha256:dae326caff8b41855e4dc24e04c701d9a0b1b66a2938a26b67f3140d09c40488

Observation 202ee373-4dee-40f8-b622-5c150b8e821b · outbound

This paper cites Hansen and P.

Multi-Token Enhancing for Vision Representation Learning Hansen and P

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.177669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.245057Z digest=sha256:ea645c782fb433b09fb60c324c332b31e02f49fef6f065861b93295785788de6

Observation ec2b9bb3-731d-42c2-8a6a-304c77232e7a · outbound

This paper cites Training independent subnetworks for robust prediction.

Multi-Token Enhancing for Vision Representation Learning Training independent subnetworks for robust prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.165226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.249656Z digest=sha256:55bb7fbfe2714e44458f3025748b23fd6a449bbc8470ff2c4418155d40bc352d

Observation 6f0a9f0a-c6a9-420a-b565-579b2f1db22f · outbound

This paper cites Deep residual learning for image recognition.

Multi-Token Enhancing for Vision Representation Learning Deep residual learning for image recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.253645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.253645Z digest=sha256:1d4063ff105e620c50cafa0234b5f8681859f72b1bc7d2d600348588f85da66a

Observation 7307fe2b-58c4-4e12-be64-f897cce54176 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

Multi-Token Enhancing for Vision Representation Learning Momentum contrast for unsupervised visual rep- resentation learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.144942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.257724Z digest=sha256:71931e7cb58094795b0d313b24a93b37787033e002950ac427ef8adc572c27e9

Observation fec79ee1-7fa5-4f2b-9c1b-320cee2d22b1 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Multi-Token Enhancing for Vision Representation Learning Masked autoencoders are scalable vision learners

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.132864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.261961Z digest=sha256:f81eeb26dca94140a064537e90b8fbddf0112aa2ad14e049a94b4c283946fa8a

Observation ae5b7a43-a9ee-4420-8f01-4c65f282ca8a · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Multi-Token Enhancing for Vision Representation Learning Distilling the Knowledge in a Neural Network

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.265938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.265938Z digest=sha256:e710d0578fc53ca30e56e380c97509097b45c7edeabc4f724d008baa66858552

Observation f91d5ffd-e3db-4dee-ac1b-1705c7950631 · outbound

This paper cites Conv2Former: A Simple Transformer-Style ConvNet for Visual Recognition.

Multi-Token Enhancing for Vision Representation Learning Conv2Former: A Simple Transformer-Style ConvNet for Visual Recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.270510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.270510Z digest=sha256:2ac6482c47a7569c294dae67584220d026ad347a975b06db6df28ac322998582

Observation 2c8310df-939a-47f3-ba8c-54e875f15300 · outbound

This paper cites Averaging Weights Leads to Wider Optima and Better Generalization.

Multi-Token Enhancing for Vision Representation Learning Averaging Weights Leads to Wider Optima and Better Generalization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.274517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.274517Z digest=sha256:e57cac3491b6097c529cb189af3161de85a52a59f9d3425c4777dcc77df2b7a6

Observation 008a7f77-e186-44e7-8e9b-595500a28f5c · outbound

This paper cites Similarity of neural network representa- tions revisited.

Multi-Token Enhancing for Vision Representation Learning Similarity of neural network representa- tions revisited

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.120725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.279665Z digest=sha256:fd4c7389a22c290042e2d2775327262cbb9730080c19e9bdaeaf8edad1103395

Observation 6b864517-9da5-4755-9925-88bcbaf7e41c · outbound

This paper cites Fractalnet: Ultra-deep neural networks without residuals.

Multi-Token Enhancing for Vision Representation Learning Fractalnet: Ultra-deep neural networks without residuals

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.108217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.283495Z digest=sha256:72b1525b080b00aa6127e9f89796793655c6938877df45fcced3ce7550c9fbbf

Observation cf0c168f-5f8b-47c9-b486-193d522a2456 · outbound

This paper cites Why M Heads are Better than One: Training a Diverse Ensemble of Deep Networks.

Multi-Token Enhancing for Vision Representation Learning Why M Heads are Better than One: Training a Diverse Ensemble of Deep Networks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.287012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.287012Z digest=sha256:e66dfc0ca410521c058042d5f19e89759fd4e600643bee1bad63137cfa3d90c7

Observation 49c894ea-0fc9-4cbd-be3b-b64fc9434d9c · outbound

This paper cites Sere: Exploring feature self-relation for self-supervised trans- former.

Multi-Token Enhancing for Vision Representation Learning Sere: Exploring feature self-relation for self-supervised trans- former

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.096601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.291248Z digest=sha256:79e2bd2816b26572456a7087eef2d08c7e62b03d06ebccd08cf2736bc84d7e1f

Observation 06f74ad0-5d71-4f4f-8297-eb3141445aee · outbound

This paper cites Enhancing representa- tions through heterogeneous self-supervised learning.

Multi-Token Enhancing for Vision Representation Learning Enhancing representa- tions through heterogeneous self-supervised learning

Reference 32

Resolution
verified exact
raw_fallback, observed 2026-08-12T13:59:36.630796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.294927Z digest=sha256:fffb7573a6ed89377abb8fd542ea77a5e7bff2bc0d4143b6bc32cd7d84affc6a

Observation 7de05d14-dcb8-4fab-be4f-7cea56036150 · outbound

This paper cites Microsoft coco: Common objects in context.

Multi-Token Enhancing for Vision Representation Learning Microsoft coco: Common objects in context

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.084032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.298601Z digest=sha256:32fc3fa58e313d7896cbb7db66fd492632e97a2eebd7470e0298f34bb2a2ad52

Observation 1e928978-8500-4391-873d-2a7a81ef28c1 · outbound

This paper cites Swin trans- former: Hierarchical vision transformer using shifted win- dows.

Multi-Token Enhancing for Vision Representation Learning Swin trans- former: Hierarchical vision transformer using shifted win- dows

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.071789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.302744Z digest=sha256:a56f6e6f6768c6fbe37f5bd8b40e1d968334670f748d1563531623604455ac22

Observation cd04de15-3d21-4fa2-bd93-99c50f2f8caa · outbound

This paper cites A convnet for the 2020s.

Multi-Token Enhancing for Vision Representation Learning A convnet for the 2020s

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.058450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.306788Z digest=sha256:47abef336d231fa0c50da3127bf3d8441108a1898f28e246fae0f4c8b226daa4

Observation fe97687c-bdf2-4f79-a15a-b052e55841bd · outbound

This paper cites Decoupled weight decay regularization.

Multi-Token Enhancing for Vision Representation Learning Decoupled weight decay regularization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.044230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.310991Z digest=sha256:df084afb932f943450ee4fb4302c6e8f602d50de8ca58ec451fcb9910db2e3f2

Observation 81876898-bc47-4c23-9244-b503f4ac24ce · outbound

This paper cites Representation uncertainty in self-supervised learning as variational inference.

Multi-Token Enhancing for Vision Representation Learning Representation uncertainty in self-supervised learning as variational inference

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.030120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.314763Z digest=sha256:c8d63db4e2a49f014e6272dae70b7a3ea7111344922ecbbc0f3c184aef32e78e

Observation 474571af-d00b-4c85-9f9e-7093000d1393 · outbound

This paper cites Simreg: Regression as a sim- ple yet effective tool for self-supervised knowledge distilla- tion.

Multi-Token Enhancing for Vision Representation Learning Simreg: Regression as a sim- ple yet effective tool for self-supervised knowledge distilla- tion

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.017322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.318518Z digest=sha256:fac539833314c119099da26e7bd1f9fa8bccd5a06b78c69205c978ca9e1b3070

Observation 58b883a7-a95f-485c-965d-8b9157c8bdb5 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Multi-Token Enhancing for Vision Representation Learning Representation Learning with Contrastive Predictive Coding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.323425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.323425Z digest=sha256:e8934121d11259f9b08d535d038b3ab8e31f38a063a295d481bb50995c02e7d2

Observation e92f2de2-92a5-4f0b-90d8-4b6528726863 · outbound

This paper cites Imagenet large scale visual recognition challenge.

Multi-Token Enhancing for Vision Representation Learning Imagenet large scale visual recognition challenge

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.327999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.327999Z digest=sha256:1df7a99f3e0ba9c9ee38dc7445ad3439a7d3039746d06890ab5e93411183c1b2

Observation 30511b61-22e1-4182-bfb2-190f132ed2ca · outbound

This paper cites Learning common rationale to improve self-supervised rep- resentation for fine-grained visual recognition problems.

Multi-Token Enhancing for Vision Representation Learning Learning common rationale to improve self-supervised rep- resentation for fine-grained visual recognition problems

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.995993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.332762Z digest=sha256:8eb67c223b79001ade7f7bb170dd8f3a3f306b401cb7e7cd880fab2aa8371522

Observation 6c11bba4-cc9b-4d59-b780-d51ed92533c1 · outbound

This paper cites an unresolved cited work.

Multi-Token Enhancing for Vision Representation Learning Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:59:36.984006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.337204Z digest=sha256:75cace94c2168811cc0bc161fd75b21bdcece1a1cf0d43f2f73b967d52b9a2b0

Observation 1397c6fb-8ffa-482b-9707-6664ac7e44ae · outbound

This paper cites Multi- mode online knowledge distillation for self-supervised visual representation learning.

Multi-Token Enhancing for Vision Representation Learning Multi- mode online knowledge distillation for self-supervised visual representation learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.971471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.341166Z digest=sha256:6a868d8dcd9371769513ced476de6c151e4ea34a0882bfa6d67f120619331a4a

Observation 39d4be1e-013c-4f9a-ac70-79b84717ee06 · outbound

This paper cites Semantics-consistent feature search for self-supervised visual representation learning.

Multi-Token Enhancing for Vision Representation Learning Semantics-consistent feature search for self-supervised visual representation learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.958797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.345473Z digest=sha256:8eae3f336df04429b409142865cb8a355404f7cef41b2cd89a84bbbaf2cb0f52

Observation 01ba14a2-f08d-40cd-be01-5faafd0dab63 · outbound

This paper cites Dropout: A simple way to prevent neural networks from overfitting.

Multi-Token Enhancing for Vision Representation Learning Dropout: A simple way to prevent neural networks from overfitting

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.946527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.349178Z digest=sha256:232121f85fe19161e9b4ad52c072044751afd59f9714942b75d6c636e8e92a1f

Observation 0ffb2b9b-743a-455d-9305-745c44fb2bd1 · outbound

This paper cites Siamese image modeling for self-supervised vision represen- tation learning.

Multi-Token Enhancing for Vision Representation Learning Siamese image modeling for self-supervised vision represen- tation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.934827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.353103Z digest=sha256:c2fdeab5c8655e60f1a352e0e3a9f7efde9edab6fa306a3b794b7104adcfb7a3

Observation 5160aa41-7918-4047-808c-b87342537b04 · outbound

This paper cites Un- derstanding self-supervised learning dynamics without con- trastive pairs.

Multi-Token Enhancing for Vision Representation Learning Un- derstanding self-supervised learning dynamics without con- trastive pairs

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.921604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.357880Z digest=sha256:12696e9f1141fa5251e76f103f3a919cfb9ccdb3dba9f638814e44c30e5b80bb

Observation d4306c5c-5350-40be-a1a1-e9d2f647d87e · outbound

This paper cites The inaturalist species classification and de- tection dataset.

Multi-Token Enhancing for Vision Representation Learning The inaturalist species classification and de- tection dataset

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.908604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.362071Z digest=sha256:8539ef09e78052d13783431f7872d2fc05cdf869591764323259faeb248c960e

Observation fd3cc78a-db4f-4275-9230-f92fded43f8e · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

Multi-Token Enhancing for Vision Representation Learning Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.365916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.365916Z digest=sha256:5a75e5ee35a51777c1abc4a01040576150ed71d5cfcadb25aaa52e0e54fec5ef

Observation c21843a7-1cc4-4fb6-b6d0-a03cbfb301fd · outbound

This paper cites Dense contrastive learning for self-supervised visual pre-training.

Multi-Token Enhancing for Vision Representation Learning Dense contrastive learning for self-supervised visual pre-training

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.888226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.370044Z digest=sha256:b50c3d0be8dcac64078639b2074d16a8485c8faa246ef6405d7d966c8da8a448

Observation 0b6c0f23-d613-4c04-a131-7ab4fed5a8bf · outbound

This paper cites Masked Feature Prediction for Self-Supervised Visual Pre-Training.

Multi-Token Enhancing for Vision Representation Learning Masked Feature Prediction for Self-Supervised Visual Pre-Training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.374197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.374197Z digest=sha256:0a2da7ebf745b8184586348bd7b11ec25b8a601a8051fe01242e6926be9275c1

Observation 8f63cfc6-8ea1-4a29-8c9c-fbd761f90e77 · outbound

This paper cites Batchensemble: an alternative approach to efficient ensemble and lifelong learning.

Multi-Token Enhancing for Vision Representation Learning Batchensemble: an alternative approach to efficient ensemble and lifelong learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.877196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.379007Z digest=sha256:5517786a7bc1defeada844ba0ade4f387b9969633be7643d2ec75786e044c9fe

Observation fd8cabfb-f044-4e51-ae64-d8cf6949b0e4 · outbound

This paper cites Con- vnext v2: Co-designing and scaling convnets with masked autoencoders.

Multi-Token Enhancing for Vision Representation Learning Con- vnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.865107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.383605Z digest=sha256:5e57c2c689e1ededea95b5643c650abe63e984c1c96427d43ca785d308767332

Observation 3f8ca473-1b5c-482f-ad48-83bae95b1767 · outbound

This paper cites Cvt: Introducing con- volutions to vision transformers.

Multi-Token Enhancing for Vision Representation Learning Cvt: Introducing con- volutions to vision transformers

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.851238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.387576Z digest=sha256:d27f19ec8b98c17e7880f5f81b0e98f8f12f4dc68a84e9d7c71a92797b071f94

Observation d1933b00-fa25-46b6-9444-27b306a7c594 · outbound

This paper cites P2T: Pyramid pooling transformer for scene understanding.

Multi-Token Enhancing for Vision Representation Learning P2T: Pyramid pooling transformer for scene understanding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.837973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.391428Z digest=sha256:791a6b962aa759d4cd406d14c51b4256364e11764c318a942247c8b93b10c069

Observation 85ef4907-c63c-486a-81d0-55e8f7d0a51e · outbound

This paper cites Unified perceptual parsing for scene understand- ing.

Multi-Token Enhancing for Vision Representation Learning Unified perceptual parsing for scene understand- ing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.824794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.395599Z digest=sha256:7826fd41c9e520e0480262310e5be77c8f83365d634f707506fca56b5779f19a

Observation b3a14ab8-cde0-4d9a-b440-18f5bb7d8396 · outbound

This paper cites Detco: Unsu- pervised contrastive learning for object detection.

Multi-Token Enhancing for Vision Representation Learning Detco: Unsu- pervised contrastive learning for object detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.813107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.399528Z digest=sha256:d23e6432944db6288b0dcb0b098e0a41e3901c66a031b835abad3e5b5b102912

Observation e9c44f02-2f1e-4160-ba2c-222945df8fcb · outbound

This paper cites Self-Supervised Learning with Swin Transformers.

Multi-Token Enhancing for Vision Representation Learning Self-Supervised Learning with Swin Transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.404002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.404002Z digest=sha256:361cfee2ca7d890b945b94a2b39f5560823e8e33488a883d99d2a95a6ca57c98

Observation 586afe26-a476-4a99-b125-432adefc9de2 · outbound

This paper cites Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning.

Multi-Token Enhancing for Vision Representation Learning Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.801568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.408074Z digest=sha256:05e5b1b9620000f16ca31330c7d6b21dbecb2e0df23fb90d87ae8f2383b5a1a3

Observation 2262c200-fa96-41d5-9fda-4b9c8c84af89 · outbound

This paper cites Simmim: A simple framework for masked image modeling.

Multi-Token Enhancing for Vision Representation Learning Simmim: A simple framework for masked image modeling

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.412172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.412172Z digest=sha256:dd4b88eaab86316ae1fca05f9799b502b6abcc34d504a9ee06fe9022b00fd57d

Observation 6aab7386-a01d-4f1e-a02d-c4deaeeaa838 · outbound

This paper cites Bag of instances aggregation boosts self-supervised distillation.

Multi-Token Enhancing for Vision Representation Learning Bag of instances aggregation boosts self-supervised distillation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.782035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.416426Z digest=sha256:af720255c0d0c1d52d2bb6f6f67e7a95f00801b6465ed65a2432ab637c1c6a74

Observation d73d5289-4ec1-4238-b808-479c9960f467 · outbound

This paper cites Joint unsuper- vised learning of deep representations and image clusters.

Multi-Token Enhancing for Vision Representation Learning Joint unsuper- vised learning of deep representations and image clusters

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.420129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.420129Z digest=sha256:464de2569f83fd38a25d1ed15b74e1e91a0dfcf63157389b26b53890dfa074a5

Observation b611e978-ca13-4602-a3a5-b61a99a31d0e · outbound

This paper cites Decoupled contrastive learning.

Multi-Token Enhancing for Vision Representation Learning Decoupled contrastive learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.759788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.424067Z digest=sha256:45f5280dc6994566fc0d37becfa6a416ca4a4f56fd559286723ed06e99f88b55

Observation da304176-fd51-44a5-8855-0610b6db12f6 · outbound

This paper cites Online deep clustering for unsupervised representation learning.

Multi-Token Enhancing for Vision Representation Learning Online deep clustering for unsupervised representation learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.745952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.428108Z digest=sha256:6d3e7aba6cb72ffc29fed7840c6deb22309ffce98126ed59c102d9b175ba4af7

Observation d104ba81-a887-4d37-9d54-85f9fd059ddd · outbound

This paper cites Scene parsing through ade20k dataset.

Multi-Token Enhancing for Vision Representation Learning Scene parsing through ade20k dataset

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.732631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.432014Z digest=sha256:7234e42650750db7dfa24f90914774bb6754d1a306486c25b3ec030f43e919ca

Observation 59d1a300-238d-4504-bd43-6c4df54e38fb · outbound

This paper cites ibot: Image bert pre-training with online tokenizer.

Multi-Token Enhancing for Vision Representation Learning ibot: Image bert pre-training with online tokenizer

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.719476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.436927Z digest=sha256:69f754fb46779433a8bbf45a83d3436f74fb8ce31adca5e89f641f410cbc1a20

Observation 3209ecb5-4892-4949-a2d5-7eb2ad445a7b · outbound

This paper cites Mugs: A Multi-Granular Self-Supervised Learning Framework.

Multi-Token Enhancing for Vision Representation Learning Mugs: A Multi-Granular Self-Supervised Learning Framework

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.441292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.441292Z digest=sha256:32797bd7477d4a8162f5e9a9efb2a946e8c9dac83de9a037561c94c206e8dcf0

Observation a1d707bd-7ebc-4ab5-baa6-8d2b9186b074 · outbound

This paper cites Multi-label self- supervised learning with scene images.

Multi-Token Enhancing for Vision Representation Learning Multi-label self- supervised learning with scene images

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.707319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:59:36.446502Z digest=sha256:d019dabd279b9a54b886144bffae676aaeed731e1dc8a3e1d09326d5b7da1640

Pith citing papers

No inbound Pith citation observations are available.