Pith. sign in

REVIEW 8 cited by

GIM: Learning Generalizable Image Matcher From Internet Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.11095 v1 pith:USBM4ZII submitted 2024-02-16 cs.CV

classification cs.CV
keywords imagematchingdatamethodsperformancevideoszero-shotmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image matching is a fundamental computer vision problem. While learning-based methods achieve state-of-the-art performance on existing benchmarks, they generalize poorly to in-the-wild images. Such methods typically need to train separate models for different scene types and are impractical when the scene type is unknown in advance. One of the underlying problems is the limited scalability of existing data construction pipelines, which limits the diversity of standard image matching datasets. To address this problem, we propose GIM, a self-training framework for learning a single generalizable model based on any image matching architecture using internet videos, an abundant and diverse data source. Given an architecture, GIM first trains it on standard domain-specific datasets and then combines it with complementary matching methods to create dense labels on nearby frames of novel videos. These labels are filtered by robust fitting, and then enhanced by propagating them to distant frames. The final model is trained on propagated data with strong augmentations. We also propose ZEB, the first zero-shot evaluation benchmark for image matching. By mixing data from diverse domains, ZEB can thoroughly assess the cross-domain generalization performance of different methods. Applying GIM consistently improves the zero-shot performance of 3 state-of-the-art image matching architectures; with 50 hours of YouTube videos, the relative zero-shot performance improves by 8.4%-18.1%. GIM also enables generalization to extreme cross-domain data such as Bird Eye View (BEV) images of projected 3D point clouds (Fig. 1(c)). More importantly, our single zero-shot model consistently outperforms domain-specific baselines when evaluated on downstream tasks inherent to their respective domains. The video presentation is available at https://www.youtube.com/watch?v=FU_MJLD8LeY.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

    cs.CV 2026-07 conditional novelty 7.0 of 10

    MV-Forcing composes temporal and view-sequential autoregression in a single diffusion model, using a recurrent 3D reconstruction model as a geometric bridge to generate arbitrarily long, multi-view consistent videos.

  2. Structure-Detail Decoupled Autoregressive Generation for Fast and High-Fidelity Virtual Try-On

    cs.CV 2026-07 conditional novelty 6.5 of 10

    STAR-VTON decouples latent VAR structure synthesis from pixel-space matching-based detail recovery, yielding faster high-fidelity virtual try-on than diffusion baselines.

  3. VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment

    cs.CV 2025-01 conditional novelty 6.0 of 10

    VideoLifter is a fragment-based, SfM-free video-to-3D pipeline using learned stereo priors and hierarchical Gaussian merging to cut training time by over 80% while matching or beating CF-3DGS quality.

  4. 4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SaRO-GS models dynamic scenes with 4D Gaussians plus a scale-aware residual field and adaptive per-Gaussian optimization, achieving state-of-the-art PSNR at real-time frame rates on D-NeRF and Plenoptic Video datasets.

  5. 4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion

    cs.CV 2024-12 conditional novelty 6.0 of 10

    4Real-Video generates consistent 4D video grids, frames across time and viewpoint, with a parallel two-stream diffusion transformer that synchronizes temporal and viewpoint token streams, achieving faster and higher-q...

  6. Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives

    cs.LG 2026-08 conditional novelty 4.0 of 10

    A literature survey and benchmark of cross-view feature matching methods, organized by a new taxonomy and evaluated under partially consistent protocols.

  7. Dynamic View Synthesis as an Inverse Problem

    cs.CV 2025-06 reject novelty 3.0 of 10

    Dynamic view synthesis from a monocular video is achieved by redesigning the noise initialization of a pretrained video diffusion model using a recursive interpolation and a stochastic latent modulation.

  8. Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A survey organizing feature matching research by modality, from SIFT to transformer-based dense matchers and vision-language models.

Pith tools