Pith. sign in

REVIEW 8 cited by

Emergent Correspondence from Image Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03881 v2 pith:JSA5DZ6B submitted 2023-06-06 cs.CV

Emergent Correspondence from Image Diffusion

classification cs.CV
keywords diffusioncorrespondencecorrespondencesdiftfeaturesimageableimages
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Finding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explicit supervision. We propose a simple strategy to extract this implicit knowledge out of diffusion networks as image features, namely DIffusion FeaTures (DIFT), and use them to establish correspondences between real images. Without any additional fine-tuning or supervision on the task-specific data or annotations, DIFT is able to outperform both weakly-supervised methods and competitive off-the-shelf features in identifying semantic, geometric, and temporal correspondences. Particularly for semantic correspondence, DIFT from Stable Diffusion is able to outperform DINO and OpenCLIP by 19 and 14 accuracy points respectively on the challenging SPair-71k benchmark. It even outperforms the state-of-the-art supervised methods on 9 out of 18 categories while remaining on par for the overall performance. Project page: https://diffusionfeatures.github.io

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking

    cs.CV 2026-05 unverdicted novelty 7.0

    Diff-Tracking learns and updates text prompts for diffusion models so that cross-attention maps locate arbitrary targets across video frames without any ground-truth annotations.

  2. MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement

    cs.CV 2025-09 conditional novelty 7.0

    MOSAIC improves multi-subject personalized image generation by supervising attention maps with semantic point correspondences and a disentanglement loss, and introduces the SemAlign-MS dataset for training.

  3. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  4. MatRes: Zero-Shot Test-Time Model Adaptation for Simultaneous Matching and Restoration

    cs.CV 2026-04 unverdicted novelty 6.0

    MatRes jointly optimizes restoration and correspondence estimation at test time by enforcing conditional similarity on a single image pair and adapting lightweight modules without offline training.

  5. Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation

    cs.CV 2026-02 conditional novelty 6.0

    A diffusion-transformer framework with VLM-grounded masked attention and VAE dropout improves identity and prompt fidelity for multi-subject image generation.

  6. BLINK: Multimodal Large Language Models Can See but Not Perceive

    cs.CV 2024-04 accept novelty 6.0

    BLINK benchmark shows multimodal LLMs reach only 45-51 percent accuracy on core visual perception tasks where humans achieve 95 percent, indicating these abilities have not emerged.

  7. A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

    cs.RO 2026-07 conditional novelty 5.0

    Interactive LLM program synthesis plus a persistent skill library from natural-language corrections outperforms zero-shot VLAs and one-shot code policies on complex real-robot manipulation.

  8. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 5.0

    Hallo4D mitigates 3D/4D generation hallucinations via LMM-based detection, multi-model voting correction, and motion-aware optimization without retraining base generators.