REVIEW 8 cited by
Emergent Correspondence from Image Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Emergent Correspondence from Image Diffusion
read the original abstract
Finding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explicit supervision. We propose a simple strategy to extract this implicit knowledge out of diffusion networks as image features, namely DIffusion FeaTures (DIFT), and use them to establish correspondences between real images. Without any additional fine-tuning or supervision on the task-specific data or annotations, DIFT is able to outperform both weakly-supervised methods and competitive off-the-shelf features in identifying semantic, geometric, and temporal correspondences. Particularly for semantic correspondence, DIFT from Stable Diffusion is able to outperform DINO and OpenCLIP by 19 and 14 accuracy points respectively on the challenging SPair-71k benchmark. It even outperforms the state-of-the-art supervised methods on 9 out of 18 categories while remaining on par for the overall performance. Project page: https://diffusionfeatures.github.io
Forward citations
Cited by 8 Pith papers
-
Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking
Diff-Tracking learns and updates text prompts for diffusion models so that cross-attention maps locate arbitrary targets across video frames without any ground-truth annotations.
-
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
MOSAIC improves multi-subject personalized image generation by supervising attention maps with semantic point correspondences and a disentanglement loss, and introduces the SemAlign-MS dataset for training.
-
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.
-
MatRes: Zero-Shot Test-Time Model Adaptation for Simultaneous Matching and Restoration
MatRes jointly optimizes restoration and correspondence estimation at test time by enforcing conditional similarity on a single image pair and adapting lightweight modules without offline training.
-
Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation
A diffusion-transformer framework with VLM-grounded masked attention and VAE dropout improves identity and prompt fidelity for multi-subject image generation.
-
BLINK: Multimodal Large Language Models Can See but Not Perceive
BLINK benchmark shows multimodal LLMs reach only 45-51 percent accuracy on core visual perception tasks where humans achieve 95 percent, indicating these abilities have not emerged.
-
A Few Words Go a Long Way: Language Guided Robot Policy Synthesis
Interactive LLM program synthesis plus a persistent skill library from natural-language corrections outperforms zero-shot VLAs and one-shot code policies on complex real-robot manipulation.
-
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D mitigates 3D/4D generation hallucinations via LMM-based detection, multi-model voting correction, and motion-aware optimization without retraining base generators.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.