Pith. sign in

REVIEW 2 cited by

Reconstruction Target Matters in Masked Image Modeling for Cross-Domain Few-Shot Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.19101 v1 pith:Q7SSYDTS submitted 2024-12-26 cs.CV

classification cs.CV
keywords cdfsldomainreconstructionfeaturesimagelearningtargetfind
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cross-Domain Few-Shot Learning (CDFSL) requires the model to transfer knowledge from the data-abundant source domain to data-scarce target domains for fast adaptation, where the large domain gap makes CDFSL a challenging problem. Masked Autoencoder (MAE) excels in effectively using unlabeled data and learning image's global structures, enhancing model generalization and robustness. However, in the CDFSL task with significant domain shifts, we find MAE even shows lower performance than the baseline supervised models. In this paper, we first delve into this phenomenon for an interpretation. We find that MAE tends to focus on low-level domain information during reconstructing pixels while changing the reconstruction target to token features could mitigate this problem. However, not all features are beneficial, as we then find reconstructing high-level features can hardly improve the model's transferability, indicating a trade-off between filtering domain information and preserving the image's global structure. In all, the reconstruction target matters for the CDFSL task. Based on the above findings and interpretations, we further propose Domain-Agnostic Masked Image Modeling (DAMIM) for the CDFSL task. DAMIM includes an Aggregated Feature Reconstruction module to automatically aggregate features for reconstruction, with balanced learning of domain-agnostic information and images' global structure, and a Lightweight Decoder module to further benefit the encoder's generalizability. Experiments on four CDFSL datasets demonstrate that our method achieves state-of-the-art performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revisiting Continuity of Image Tokens for Cross-domain Few-shot Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Disrupting the spatial continuity of image tokens hurts source-domain accuracy far more than target-domain accuracy, and using such disruption during source training improves cross-domain few-shot classification.

  2. Random Registers for Cross-Domain Few-Shot Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Random registers in a vision transformer improve cross-domain few-shot transfer, and REAP strengthens this by replacing clustered image patches with random noise.

Pith tools