Pith. sign in

REVIEW 3 cited by

PAD: Self-Supervised Pre-Training with Patchwise-Scale Adapter for Infrared Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.08192 v1 pith:RCBSZGPD submitted 2023-12-13 cs.CV

classification cs.CV
keywords pre-trainingimagesinfraredadapterdatasetfeaturesimagefeature
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised learning (SSL) for RGB images has achieved significant success, yet there is still limited research on SSL for infrared images, primarily due to three prominent challenges: 1) the lack of a suitable large-scale infrared pre-training dataset, 2) the distinctiveness of non-iconic infrared images rendering common pre-training tasks like masked image modeling (MIM) less effective, and 3) the scarcity of fine-grained textures making it particularly challenging to learn general image features. To address these issues, we construct a Multi-Scene Infrared Pre-training (MSIP) dataset comprising 178,756 images, and introduce object-sensitive random RoI cropping, an image preprocessing method, to tackle the challenge posed by non-iconic images. To alleviate the impact of weak textures on feature learning, we propose a pre-training paradigm called Pre-training with ADapter (PAD), which uses adapters to learn domain-specific features while freezing parameters pre-trained on ImageNet to retain the general feature extraction capability. This new paradigm is applicable to any transformer-based SSL method. Furthermore, to achieve more flexible coordination between pre-trained and newly-learned features in different layers and patches, a patchwise-scale adapter with dynamically learnable scale factors is introduced. Extensive experiments on three downstream tasks show that PAD, with only 1.23M pre-trainable parameters, outperforms other baseline paradigms including continual full pre-training on MSIP. Our code and dataset are available at https://github.com/casiatao/PAD.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    SpectraDINO extends DINOv2 with lightweight per-modality adapters and staged distillation to handle NIR, SWIR, and LWIR in one backbone, but its SWIR gain is weakened by using the evaluation dataset for pretraining.

  2. UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic Segmentation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A pre-training framework for infrared segmentation that distills hybrid attention patterns from large RGB teachers and reports large mIoU gains for small ViTs.

  3. Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Reweighting visible-infrared pre-training patches by infrared structural reliability improves downstream segmentation, detection, and retrieval by small, mostly consistent margins.

Pith tools