Pith. sign in

REVIEW 2 cited by

Stochastic positional embeddings improve masked image modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.00566 v2 pith:MF7DHFCV submitted 2023-07-31 cs.CV cs.AIcs.LG

Stochastic positional embeddings improve masked image modeling

classification cs.CV cs.AIcs.LG
keywords learninglocationmaskedstochasticstopdownstreamembeddingsfeatures
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
abstract

Masked Image Modeling (MIM) is a promising self-supervised learning approach that enables learning from unlabeled images. Despite its recent success, learning good representations through MIM remains challenging because it requires predicting the right semantic content in accurate locations. For example, given an incomplete picture of a dog, we can guess that there is a tail, but we cannot determine its exact location. In this work, we propose to incorporate location uncertainty into MIM by using stochastic positional embeddings (StoP). Specifically, we condition the model on stochastic masked token positions drawn from a Gaussian distribution. StoP reduces overfitting to location features and guides the model toward learning features that are more robust to location uncertainties. Quantitatively, StoP improves downstream MIM performance on a variety of downstream tasks, including $+1.7\%$ on ImageNet linear probing using ViT-B, and $+2.5\%$ for ViT-H using $1\%$ of the data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics

    cs.LG 2025-11 conditional novelty 6.0

    LeJEPA derives an optimal isotropic Gaussian target for embeddings and enforces it via sketched regularization to deliver scalable, heuristics-free self-supervised pretraining with 79% ImageNet linear accuracy on ViT-H/14.

  2. Joint-Embedding Predictive Architecture for Solar PV Panel Fault Classification

    eess.IV 2026-07 accept novelty 5.5

    JEFFNet fuses StoP-JEPA semantic embeddings with EfficientNetV2-S features for thermal IR PV fault classification, beating GEPFNet on F1 for multiclass and binary tasks with 47% fewer parameters.