AnyMatch synthesizes large-scale geometrically consistent multi-modal image pairs from single-view images, enabling fine-tuned matching networks to achieve substantial gains on benchmarks.
Diffv2ir: visible-to- infrared diffusion model via vision-language understanding
5 Pith papers cite this work. Polarity classification is still indexing.
abstract
The task of translating visible-to-infrared images (V2IR) is inherently challenging due to three main obstacles: 1) achieving semantic-aware translation, 2) managing the diverse wavelength spectrum in infrared imagery, and 3) the scarcity of comprehensive infrared datasets. Current leading methods tend to treat V2IR as a conventional image-to-image synthesis challenge, often overlooking these specific issues. To address this, we introduce DiffV2IR, a novel framework for image translation comprising two key elements: a Progressive Learning Module (PLM) and a Vision-Language Understanding Module (VLUM). PLM features an adaptive diffusion model architecture that leverages multi-stage knowledge learning to infrared transition from full-range to target wavelength. To improve V2IR translation, VLUM incorporates unified Vision-Language Understanding. We also collected a large infrared dataset, IR-500K, which includes 500,000 infrared images compiled by various scenes and objects under various environmental conditions. Through the combination of PLM, VLUM, and the extensive IR-500K dataset, DiffV2IR markedly improves the performance of V2IR. Experiments validate DiffV2IR's excellence in producing high-quality translations, establishing its efficacy and broad applicability. The code, dataset, and DiffV2IR model will be available at https://github.com/LidongWang-26/DiffV2IR.
fields
cs.CV 5years
2026 5representative citing papers
UniTriGen uses unified diffusion in a shared latent space plus lightweight adapters and scene-balanced sampling to produce high-quality aligned VIS-IR-Label triplets from limited paired data, improving few-shot RGB-T semantic segmentation.
T-CLIP introduces a physics-aware thermal captioning dataset (IR-Cap) and a decoupled dual-LoRA adaptation of CLIP that improves cross-modal retrieval on thermal benchmarks by separating scene-level and object-level thermal understanding.
SpectraDINO extends DINOv2 with lightweight per-modality adapters and staged distillation to handle NIR, SWIR, and LWIR in one backbone, but its SWIR gain is weakened by using the evaluation dataset for pretraining.
MonoIR-RS synthesizes 600K infrared remote-sensing images from visible sources, rewrites captions to be IR-aware, and shows that IR-aware fine-tuning improves CLIP retrieval by up to 12.8 points and drives VLM infrared-cue coverage to 100% with near-zero RGB-color leakage.
citing papers explorer
-
AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images
AnyMatch synthesizes large-scale geometrically consistent multi-modal image pairs from single-view images, enabling fine-tuned matching networks to achieve substantial gains on benchmarks.
-
UniTriGen: Unified Triplet Generation of Aligned Visible-Infrared-Label for Few-Shot RGB-T Semantic Segmentation
UniTriGen uses unified diffusion in a shared latent space plus lightweight adapters and scene-balanced sampling to produce high-quality aligned VIS-IR-Label triplets from limited paired data, improving few-shot RGB-T semantic segmentation.
-
T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining
T-CLIP introduces a physics-aware thermal captioning dataset (IR-Cap) and a decoupled dual-LoRA adaptation of CLIP that improves cross-modal retrieval on thermal benchmarks by separating scene-level and object-level thermal understanding.
-
SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands
SpectraDINO extends DINOv2 with lightweight per-modality adapters and staged distillation to handle NIR, SWIR, and LWIR in one backbone, but its SWIR gain is weakened by using the evaluation dataset for pretraining.
-
MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation
MonoIR-RS synthesizes 600K infrared remote-sensing images from visible sources, rewrites captions to be IR-aware, and shows that IR-aware fine-tuning improves CLIP retrieval by up to 12.8 points and drives VLM infrared-cue coverage to 100% with near-zero RGB-color leakage.