Pith. sign in

REVIEW 7 cited by

MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.07556 v1 pith:M43VZBEV submitted 2025-01-13 cs.CV

classification cs.CV
keywords imagematchingacrosscross-modalityimagestasksvariousalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image matching, which aims to identify corresponding pixel locations between images, is crucial in a wide range of scientific disciplines, aiding in image registration, fusion, and analysis. In recent years, deep learning-based image matching algorithms have dramatically outperformed humans in rapidly and accurately finding large amounts of correspondences. However, when dealing with images captured under different imaging modalities that result in significant appearance changes, the performance of these algorithms often deteriorates due to the scarcity of annotated cross-modal training data. This limitation hinders applications in various fields that rely on multiple image modalities to obtain complementary information. To address this challenge, we propose a large-scale pre-training framework that utilizes synthetic cross-modal training signals, incorporating diverse data from various sources, to train models to recognize and match fundamental structures across images. This capability is transferable to real-world, unseen cross-modality image matching tasks. Our key finding is that the matching model trained with our framework achieves remarkable generalizability across more than eight unseen cross-modality registration tasks using the same network weight, substantially outperforming existing methods, whether designed for generalization or tailored for specific tasks. This advancement significantly enhances the applicability of image matching technologies across various scientific disciplines and paves the way for new applications in multi-modality human and artificial intelligence analysis and beyond.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Pairing VLM-generated action proposals with rollouts from a pose-image-conditioned video world model yields high success rates in novel simulated manipulation tasks without end-to-end policy retraining.

  2. RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs

    cs.CV 2026-07 conditional novelty 6.0 of 10

    One frozen DINOv2 token field, shared between a frozen retrieval head and a distilled local-descriptor decoder, performs UAV 6-DoF localization 1.8× faster than separate encoders with re-ranking within ~1 pp.

  3. Precise Action-to-Video Generation Through Visual Action Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.

  4. PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A training-free geometry-only pipeline matches RGB images against rendered depth/normal maps to estimate 6D poses of unseen, textureless, and slightly defective objects from one image.

  5. OpenNavMap: Multi-Session Appearance-Based Topometric Mapping for Scalable Visual Navigation

    cs.RO 2026-01 conditional novelty 5.0 of 10

    OpenNavMap shows that an image graph plus on-demand 3D reconstruction can match structure-based maps for visual localization and navigation.

  6. Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching

    cs.CV 2025-06 reject novelty 5.0 of 10

    A LiDAR-intensity projection, matched to the camera image with an attention-based detector-free network and a repeatability score, achieves state-of-the-art point-pixel registration using only single-frame LiDAR.

  7. Deep Learning Reforms Image Matching: A Survey and Outlook

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A pipeline-aligned survey and benchmark of image matching, finding end-to-end dense matchers dominate pose and homography tasks.

Pith tools