Pith. sign in

Paper Citation Record · LEDGER

SimMIM: A Simple Framework for Masked Image Modeling

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2111.09886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.09886 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:41:10.617330Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

15
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0a5c7945-cd9b-41dc-9b6e-6d4085175ca2 · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models SimMIM: A Simple Framework for Masked Image Modeling

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T10:53:08.409786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:9a77f5f1b1586228211d3afbf9dd6132497a4cd15cfcc1e5121701c697ca6f35

Observation acc5e171-64f3-431b-b163-fde77b341ce5 · inbound

Vision Transformers Need Registers cites this paper.

Vision Transformers Need Registers SimMIM: A Simple Framework for Masked Image Modeling

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:41:38.236334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T09:41:37.937046Z digest=sha256:35eab2b1b18f3f81d13f1feee69f0b39c3e5c1725d5d801f8cf9675d2eb7afe4

Observation bbfc41ed-c0fb-4213-9f46-c6bfd06aad24 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video SimMIM: A Simple Framework for Masked Image Modeling

Reference 290

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T12:40:24.075377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:1f50c67645a6d52262db7cb044c02c1c256491502bcc8ee9c71ba75c1838e294

Observation 89988654-f3ab-4129-8c52-d4cec6306de0 · inbound

A Self-supervised Learning Method for Raman Spectroscopy based on Masked Autoencoders cites this paper.

A Self-supervised Learning Method for Raman Spectroscopy based on Masked Autoencoders SimMIM: A Simple Framework for Masked Image Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:41:10.617330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:41:10.617330Z digest=sha256:5a791a722d26d555fb9aa6f561956b0a297429757c0e06129394ae993cd55ff2

Observation 4c801cb5-4993-4f15-936b-2da4a8541fb8 · inbound

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation cites this paper.

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation SimMIM: A Simple Framework for Masked Image Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:06:08.394168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:06:08.394168Z digest=sha256:15a00469cd6cf7ca7d11972a92902883d9858e30ecb27f8a3f36c69382c79adc

Observation 14039709-2f8a-4b23-b34f-7ad03ac69994 · inbound

Thoughts on Objectives of Sparse and Hierarchical Masked Image Model cites this paper.

Thoughts on Objectives of Sparse and Hierarchical Masked Image Model SimMIM: A Simple Framework for Masked Image Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:15.049428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:11:15.049428Z digest=sha256:f1de0fc032c876155e16df4ca259ee81148372a4bf1f10cc29a98638c2de0e17

Observation 6f119842-b305-4082-a2a1-3b1d854aeec2 · inbound

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm cites this paper.

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm SimMIM: A Simple Framework for Masked Image Modeling

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T01:05:34.862848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:05:34.862848Z digest=sha256:520bc8802213a6f6739c31c42d1a16cae6ae2c4713369577ce35171052a76d9b

Observation c40d9ca7-a5bf-440c-93e2-cfd801e2816b · inbound

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity cites this paper.

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity SimMIM: A Simple Framework for Masked Image Modeling

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:40:03.197599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T11:36:06.967549Z digest=sha256:25e61be9027d1a90deb6b23020e47b902800c7c44785331362b96eee347fbaec

Observation 217eab2f-fac3-4a7d-958c-f66bb3443bf0 · inbound

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training cites this paper.

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training SimMIM: A Simple Framework for Masked Image Modeling

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:26.797647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T05:02:02.367347Z digest=sha256:078a578bd494bc14fac009bc8802481a8c6e46b64ea9a8bbb3a9cee7c7aa1c77

Observation cc81ad20-096b-40bf-b0ae-cf70c7e93e2c · inbound

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation cites this paper.

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation SimMIM: A Simple Framework for Masked Image Modeling

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:11.403617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-25T19:34:02.046104Z digest=sha256:b029f469f8804d58743281168d1cd185a407a3f98e32680960a5db4df1865760