Pith. sign in

REVIEW 1 cited by

InfMAE: A Foundation Model in the Infrared Modality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.00407 v2 pith:V7LIOM3V submitted 2024-02-01 cs.CV

classification cs.CV
keywords infraredfoundationlearningtasksdesigndownstreamimagesinfmae
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, the foundation models have swept the computer vision field and facilitated the development of various tasks within different modalities. However, it remains an open question on how to design an infrared foundation model. In this paper, we propose InfMAE, a foundation model in infrared modality. We release an infrared dataset, called Inf30 to address the problem of lacking large-scale data for self-supervised learning in the infrared vision community. Besides, we design an information-aware masking strategy, which is suitable for infrared images. This masking strategy allows for a greater emphasis on the regions with richer information in infrared images during the self-supervised learning process, which is conducive to learning the generalized representation. In addition, we adopt a multi-scale encoder to enhance the performance of the pre-trained encoders in downstream tasks. Finally, based on the fact that infrared images do not have a lot of details and texture information, we design an infrared decoder module, which further improves the performance of downstream tasks. Extensive experiments show that our proposed method InfMAE outperforms other supervised methods and self-supervised learning methods in three downstream tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic Segmentation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A pre-training framework for infrared segmentation that distills hybrid attention patterns from large RGB teachers and reports large mIoU gains for small ViTs.

Pith tools