Pith. sign in

REVIEW 1 cited by

Exploring Long-Sequence Masked Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.07224 v1 pith:6CAUMZU6 submitted 2022-10-13 cs.CV

classification cs.CV
keywords long-sequencepre-trainingsizeacrossdetectionduringimageinput
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Masked Autoencoding (MAE) has emerged as an effective approach for pre-training representations across multiple domains. In contrast to discrete tokens in natural languages, the input for image MAE is continuous and subject to additional specifications. We systematically study each input specification during the pre-training stage, and find sequence length is a key axis that further scales MAE. Our study leads to a long-sequence version of MAE with minimal changes to the original recipe, by just decoupling the mask size from the patch size. For object detection and semantic segmentation, our long-sequence MAE shows consistent gains across all the experimental setups without extra computation cost during the transfer. While long-sequence pre-training is discerned most beneficial for detection and segmentation, we also achieve strong results on ImageNet-1K classification by keeping a standard image size and only increasing the sequence length. We hope our findings can provide new insights and avenues for scaling in computer vision.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Masked Autoencoders for Ultrasound Signals: Robust Representation Learning for Downstream Applications

    cs.LG 2025-08 conditional novelty 6.0 of 10

    MAE pre-training on synthetic ultrasound signals transfers to real measured signals and beats from-scratch and CNN baselines on time-of-flight classification, with the biggest gains in low-label regimes.

Pith tools