REVIEW 7 cited by
MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Image matching, which aims to identify corresponding pixel locations between images, is crucial in a wide range of scientific disciplines, aiding in image registration, fusion, and analysis. In recent years, deep learning-based image matching algorithms have dramatically outperformed humans in rapidly and accurately finding large amounts of correspondences. However, when dealing with images captured under different imaging modalities that result in significant appearance changes, the performance of these algorithms often deteriorates due to the scarcity of annotated cross-modal training data. This limitation hinders applications in various fields that rely on multiple image modalities to obtain complementary information. To address this challenge, we propose a large-scale pre-training framework that utilizes synthetic cross-modal training signals, incorporating diverse data from various sources, to train models to recognize and match fundamental structures across images. This capability is transferable to real-world, unseen cross-modality image matching tasks. Our key finding is that the matching model trained with our framework achieves remarkable generalizability across more than eight unseen cross-modality registration tasks using the same network weight, substantially outperforming existing methods, whether designed for generalization or tailored for specific tasks. This advancement significantly enhances the applicability of image matching technologies across various scientific disciplines and paves the way for new applications in multi-modality human and artificial intelligence analysis and beyond.
Forward citations
Cited by 7 Pith papers
-
World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
Pairing VLM-generated action proposals with rollouts from a pose-image-conditioned video world model yields high success rates in novel simulated manipulation tasks without end-to-end policy retraining.
-
RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs
One frozen DINOv2 token field, shared between a frozen retrieval head and a distilled local-descriptor decoder, performs UAV 6-DoF localization 1.8× faster than separate encoders with re-ranking within ~1 pp.
-
Precise Action-to-Video Generation Through Visual Action Prompts
Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.
-
PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects
A training-free geometry-only pipeline matches RGB images against rendered depth/normal maps to estimate 6D poses of unseen, textureless, and slightly defective objects from one image.
-
OpenNavMap: Multi-Session Appearance-Based Topometric Mapping for Scalable Visual Navigation
OpenNavMap shows that an image graph plus on-demand 3D reconstruction can match structure-based maps for visual localization and navigation.
-
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
A LiDAR-intensity projection, matched to the camera image with an attention-based detector-free network and a repeatability score, achieves state-of-the-art point-pixel registration using only single-frame LiDAR.
-
Deep Learning Reforms Image Matching: A Survey and Outlook
A pipeline-aligned survey and benchmark of image matching, finding end-to-end dense matchers dominate pose and homography tasks.
Discussion (0). Continue with ORCID to comment.