Pith. sign in

Paper Citation Record · LEDGER

SimMIM: A Simple Framework for Masked Image Modeling

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2111.09886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.09886 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:41:10.617330Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

15
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0a5c7945-cd9b-41dc-9b6e-6d4085175ca2 · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models SimMIM: A Simple Framework for Masked Image Modeling

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T10:53:08.409786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:57bdba03f33f73f4309e919ceccfada3109ee5ca09a6c7ce9d1fc5c19a183b7e

Observation acc5e171-64f3-431b-b163-fde77b341ce5 · inbound

Vision Transformers Need Registers cites this paper.

Vision Transformers Need Registers SimMIM: A Simple Framework for Masked Image Modeling

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:41:38.236334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T09:41:37.937046Z digest=sha256:94fbcd1283e604812dfc376aff64fc6fe998e1bd01bdc3fee03478c753dae6cf

Observation bbfc41ed-c0fb-4213-9f46-c6bfd06aad24 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video SimMIM: A Simple Framework for Masked Image Modeling

Reference 290

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T12:40:24.075377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:d2fb10951dfc7c1db78923f09aedcda759630a871639ea29c0a925c70c3fd9d4

Observation 89988654-f3ab-4129-8c52-d4cec6306de0 · inbound

A Self-supervised Learning Method for Raman Spectroscopy based on Masked Autoencoders cites this paper.

A Self-supervised Learning Method for Raman Spectroscopy based on Masked Autoencoders SimMIM: A Simple Framework for Masked Image Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:41:10.617330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:41:10.617330Z digest=sha256:fbf9540de39a2dd688f7fbe46099aa51a0d1fa6c7b5780fe123d25ea5c08fedb

Observation 4c801cb5-4993-4f15-936b-2da4a8541fb8 · inbound

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation cites this paper.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation SimMIM: A Simple Framework for Masked Image Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.394168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.394168Z digest=sha256:c7bc3a0b1dd82dbc41380fcbfb154e003ddcf7e84bd53eceee088bfe4156b973

Observation 14039709-2f8a-4b23-b34f-7ad03ac69994 · inbound

Thoughts on Objectives of Sparse and Hierarchical Masked Image Model cites this paper.

Thoughts on Objectives of Sparse and Hierarchical Masked Image Model SimMIM: A Simple Framework for Masked Image Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:15.049428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:11:15.049428Z digest=sha256:99d1e339b833f38d49d7b15604d7c023f68af7cf765458d233640e9a85ae23c7

Observation 6f119842-b305-4082-a2a1-3b1d854aeec2 · inbound

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm cites this paper.

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm SimMIM: A Simple Framework for Masked Image Modeling

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T01:05:34.862848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:05:34.862848Z digest=sha256:520bc8802213a6f6739c31c42d1a16cae6ae2c4713369577ce35171052a76d9b

Observation c40d9ca7-a5bf-440c-93e2-cfd801e2816b · inbound

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity cites this paper.

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity SimMIM: A Simple Framework for Masked Image Modeling

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:40:03.197599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T11:36:06.967549Z digest=sha256:54630a048d3c49b675e5fdc9236b0f20ffb81438890513f461c00fae7ddeb202

Observation 217eab2f-fac3-4a7d-958c-f66bb3443bf0 · inbound

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training cites this paper.

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training SimMIM: A Simple Framework for Masked Image Modeling

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:26.797647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T05:02:02.367347Z digest=sha256:76166f7ba36825971502d2824dc9a7eb1e891b7649ef62f7c75726acd5e483a3

Observation cc81ad20-096b-40bf-b0ae-cf70c7e93e2c · inbound

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation cites this paper.

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation SimMIM: A Simple Framework for Masked Image Modeling

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:11.403617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-25T19:34:02.046104Z digest=sha256:51ee0818dc14ccc918bf8c7c674422a6898d962656ee8e76781a2dbec9b037dc