REVIEW 3 major objections 6 minor 63 references
Edicho: Consistent Image Editing in the Wild
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Edicho claims that consistent multi-image editing in the wild is achievable training-free by injecting explicit, pre-computed correspondence into self-attention and classifier-free guidance.
desk verdict Edicho's explicit-correspondence-guided attention and CFG is a clean, novel twist on zero-shot consistent editing; the mechanism is plausible and the qualitative work is strong, but the quantitative evidence is thinner than the claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is correspondence-warped self-attention combined with correspondence-guided classifier-free guidance. In the attention module, the query feature of the target image is warped into the source image's coordinate frame using the prestored correspondence map, and attention reads keys and values from the source, so coherent features are transferred across images. In the guidance module, the unconditional branch of the denoising step is modified by randomly injecting a portion of the source noisy latent into the target's unconditional latent, keeping the conditional branch untouched; this design choice is intended to preserve the generative prior. Both manipulations are applied only in a middle range of denoising steps and from a given attention layer onward, preserving the untouched generative prior.
What would settle it
Replace the correspondence map generated by the extractor with a random or corrupted map and run Edicho on the same inputs; if editing-consistency scores (e.g., the paper's EC metric) do not visibly degrade, then the claimed mechanism is not carrying the result, whereas a sharp drop would confirm that the method's success is tied to the accuracy of explicit correspondence.
Extended reading notes
Core claim
The central claim is that injecting explicit pre-estimated correspondence into the diffusion denoising process yields consistent edits across in-the-wild images, outperforming implicit-correspondence methods. Concretely, Edicho computes a correspondence map between source and target images, then in self-attention warps the target query by that map and queries the source keys and values, so that each edited location attends to its matching source location rather than to wherever attention happens to land. In classifier-free guidance, it modifies only the unconditional branch, fusing source latents into the target's unconditional noise estimate, which is reported to preserve the model's generative priors while improving consistency. When plugged into a Stable Diffusion backbone with existing editing backbones, this produces text-aligned edits that are consistent in texture, object counts, and style across image pairs and sets, as measured by CLIP-based scores and user studies.
Load-bearing premise
The method assumes the pre-computed correspondence map is accurate and remains aligned with the diffusion model's latent feature space at the layers and steps where editing is applied, so that warped queries and injected latents copy the right features.
Editorial extensions
If this is right
- The method is model-agnostic: it can be layered onto any Stable-Diffusion-based editing backbone without retraining, so future editing models inherit the consistency mechanism.
- Explicit correspondence prediction is more reliable than implicit attention-derived correspondence, so consistent editing should be built on precomputed matching rather than on attention similarity.
- Consistent edits can serve as training data for concept customization and as inputs for 3D reconstruction, extending the usefulness of the editing outputs.
- Because the method is zero-shot, it can be applied to arbitrary image sets without per-set optimization, making batch photo editing practical.
- The method's consistency quality is directly tied to the accuracy of the correspondence extractor, so improving extractors should directly improve editing consistency.
Reading between the lines
- If the mechanism is as robust as claimed, a natural extension is to apply it to video frames or multi-view images of the same scene, where correspondence is easier to compute and temporal consistency is a similar problem.
- The method effectively decouples what to edit from where to edit consistently, suggesting that any edit operator—not just text-to-image backbones—could be made consistent by the same two injection points.
- A possible failure mode not fully explored in the paper is occlusion: because the warp assumes one-to-one correspondence, images with occlusions may copy features from the occluder; a visibility-aware warp would be a testable improvement.
- The fixed denoising-step and layer range is likely task-dependent; an adaptive schedule chosen based on where the correspondence extractor is most reliable could improve results further.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Edicho, a training-free framework for consistent image editing across multiple images. The method first computes explicit semantic correspondence between input images using a pre-trained extractor (DIFT). It then injects this correspondence into a pre-trained diffusion model in two ways: (i) Corr-Attention, which warps the target's query features to the source according to the correspondence and borrows source keys and values in self-attention (Eqs. 3–4), and (ii) Corr-CFG, which fuses source latents into the unconditional branch of classifier-free guidance (Eqs. 7–9) to enhance consistency. Edicho is designed to be plug-and-play with existing editing models such as ControlNet and BrushNet. The paper reports qualitative results, a user study, and quantitative TA/EC scores for local and global editing, plus applications to customization and 3D reconstruction.
Significance. Edicho addresses a relevant problem, and the core idea—using explicit correspondence to guide attention and CFG—is well motivated and distinct from prior implicit-correspondence approaches. The method is training-free, model-agnostic, and appears to work across diverse in-the-wild inputs. The qualitative results in Figs. 4–5 are often convincing, and the user study (App. E) provides some evidence of preference. However, the quantitative evaluation is currently too weak to establish the claimed improvements: the EC metric is partially circular with the editing mechanism, Table 1 lacks error bars and sample sizes, and the user study lacks statistical detail. For these reasons, the empirical support for the central claim requires strengthening, though the underlying approach is plausible and likely to be useful.
major comments (3)
- [§4.2, Table 1] The editing consistency (EC) metric measures CLIP feature similarity between edited images. Edicho's design directly promotes such similarity by copying source features into the target: Eq. (4) makes the target's attention output at each selected step attend only to source keys and values, and Eq. (8) injects source latents into the target's unconditional branch. Both operations push the target's internal representation, and therefore the decoded image, toward the source in any feature space, including CLIP. Consequently, the reported EC margins (e.g., 0.9355 vs. 0.9258 for global editing) cannot be interpreted as independent evidence of better perceptual consistency. To support the claim of outperforming implicit-correspondence baselines, please add control experiments (e.g., ablating the query warping or latent injection and reporting EC), compute a consistency metric that is not trivially increased by feature copying, and report error bars and significance tests.
- [§4.1–4.2, App. E] The quantitative evaluation is statistically underspecified. Table 1 reports single point estimates with no indication of the number of test image pairs, no variance, and no significance testing, yet the text in §4.2 calls the improvements 'significant.' The user study in App. E reports only overall vote percentages (over 60% for the proposed method) from 30 participants answering up to 20 questions, but does not state the number of underlying image sets, the randomization procedure, or the variability across participants. Please provide the full evaluation protocol, per-participant breakdown, and appropriate statistical tests so the central quantitative claims can be verified.
- [§3.3–3.4, Conclusion] The method's correctness rests on the accuracy of the precomputed correspondence C_ij from DIFT. If the correspondence is misaligned due to large pose changes, occlusions, or lighting differences, the warped queries in Eq. (3) and the latent injection in Eq. (8) transport mismatched content, which the authors partially concede in the Conclusion ('sometimes the generated textures would be inconsistent due to the correspondence misalignment'). The manuscript does not quantify the frequency or severity of such failures, nor does it provide a sensitivity analysis to correspondence quality. Please add a failure analysis or a sensitivity study (e.g., varying the correspondence extractor, or artificially corrupting C_ij) to delineate when the method is reliable.
minor comments (6)
- [Eqs. (7)–(8)] The notation for the fusing function T is inconsistent: Eq. (7) calls T with three arguments (two noise predictions and C_ij), while Eq. (8) defines T with four arguments (z_i, z_j, C_ij, γ). Please align the notation and explicitly define the injection function Inj.
- [Section 3.4 title] The section title contains a typo: 'Classifer-free Guidance' should be 'Classifier-free Guidance'.
- [Reference [45]] The reference entry for Adobe Firefly contains a typo: 'Adobe reseachers' should be 'Adobe researchers'.
- [Fig. 2 and App. D] The implicit correspondence is extracted at different layer/step pairs for different cases (e.g., (1,10), (2,15), (4,25)), so the qualitative comparison in Fig. 2 is not apples-to-apples. Please use a fixed set of layers/steps or report the best-performing configuration for the implicit method.
- [Eq. (3)] The phrase 'Where the Warpfunction' should be 'where the Warp function' for correct capitalization and spacing.
- [App. B] The hyperparameters λ=0.8 and γ=0.9 are set without any sensitivity analysis; a small ablation over these values would help readers understand the method's robustness.
Circularity Check
No significant circularity: Edicho is an inference-time feature-injection method validated against external baselines and human studies; the EC-similarity caveat is a metric-validity concern, not a derivation-level reduction.
full rationale
The paper's contribution is an inference-time algorithm rather than a derived theorem: explicit correspondence is precomputed by external extractors (DIFT, Dust3R) via Eq. (2), and the two proposed mechanisms, Corr-Attention (Eqs. (3)-(4)) and Corr-CFG (Eqs. (7)-(8)), are explicit design operations with fixed hyperparameters lambda and gamma, not fitted parameters. No quantity is fitted to a subset of data and then reported as a prediction, and no uniqueness theorem is imported from the authors' prior work. The evaluation uses external benchmarks: CLIP-based TA and EC following Custom Diffusion, plus a human preference study (Appendix E). The only plausible worry is that the EC metric, defined as CLIP feature similarity of edited images, is partly aligned with Edicho's mechanism because Eq. (4) borrows source keys/values and Eq. (8) injects source latents, mechanically raising similarity. However, this is a metric-validity caveat rather than a circular derivation: the EC values are measured, not computed from the method's equations, and the same metric is applied uniformly to all baselines. The user study provides an independent (if underspecified) human signal. Self-citations appear in the references but are not load-bearing for the central claim, which rests on externally sourced correspondence extractors and editing backbones. Therefore no step reduces to its own inputs by construction.
Assumptions & free parameters
free parameters (4)
- lambda (fusion weight in Corr-CFG) =
0.8
- gamma (injection ratio in Corr-CFG) =
0.9
- denoising step range for correspondence guidance =
steps 4 to 40 of 50
- attention layer threshold =
from the 8th attention layer onward
assumptions (4)
- domain assumption Intermediate features of pretrained diffusion models are spatially aligned with the generated image space, so warping query tokens by image correspondences transfers meaningful content.
- domain assumption The chosen explicit correspondence extractor (DIFT, and Dust3R for 3D cases) yields sufficiently accurate, dense correspondences for in-the-wild image pairs.
- domain assumption Copying source attention values and injecting source latents into the unconditional CFG branch does not destroy the pretrained generative prior.
- domain assumption Zero-shot and plug-and-play integration into ControlNet and BrushNet works without adapting those models.
Cite this review
Pith. "Pith review of Edicho: Consistent Image Editing in the Wild." pith.science (2026). https://pith.science/paper/OLFOJAWE
@misc{pith2026241221079,
author = {Pith},
title = {Pith review of: Edicho: Consistent Image Editing in the Wild},
year = {2026},
howpublished = {\url{https://pith.science/paper/OLFOJAWE}},
note = {Machine review of arXiv:2412.21079}
}
read the original abstract
As a verified need, consistent editing across in-the-wild images remains a technical challenge arising from various unmanageable factors, like object poses, lighting conditions, and photography environments. Edicho steps in with a training-free solution based on diffusion models, featuring a fundamental design principle of using explicit image correspondence to direct editing. Specifically, the key components include an attention manipulation module and a carefully refined classifier-free guidance (CFG) denoising strategy, both of which take into account the pre-estimated correspondence. Such an inference-time algorithm enjoys a plug-and-play nature and is compatible to most diffusion-based editing methods, such as ControlNet and BrushNet. Extensive results demonstrate the efficacy of Edicho in consistent cross-image editing under diverse settings. We will release the code to facilitate future studies.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Cross-image attention for zero-shot appearance transfer
Yuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch- Elor, and Daniel Cohen-Or. Cross-image attention for zero-shot appearance transfer. In ACM SIGGRAPH 2024 Conference Papers, 2024. 1, 2, 4, 5, 6, 7
work page 2024
-
[2]
Omri Avrahami, Ohad Fried, and Dani Lischinski. Blended latent diffusion. ACM Trans. Graph., 2023. 2
work page 2023
-
[3]
Real-time 3d-aware portrait editing from a single image
Qingyan Bai, Zifan Shi, Yinghao Xu, Hao Ouyang, Qiuyu Wang, Ceyuan Yang, Xuan Wang, Gordon Wetzstein, Yujun Shen, and Qifeng Chen. Real-time 3d-aware portrait editing from a single image. In Eur. Conf. Comput. Vis., 2024
work page 2024
-
[4]
Stable video diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023
arXiv 2023
-
[5]
Align your latents: High-resolution video synthesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dock- horn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 2
work page 2023
-
[6]
In- structPix2Pix: Learning to follow image editing instructions
Tim Brooks, Aleksander Holynski, and Alexei A Efros. In- structPix2Pix: Learning to follow image editing instructions. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 2
work page 2023
-
[7]
MasaCtrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, and Yinqiang Zheng. MasaCtrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. In Int. Conf. Comput. Vis., 2023. 2, 4, 5, 6, 7, 1
work page 2023
-
[8]
Instruction-based image manipulation by watching how things move
Mingdeng Cao, Xuaner Zhang, Yinqiang Zheng, and Zhihao Xia. Instruction-based image manipulation by watching how things move. arXiv preprint arXiv:2412.12087, 2024. 2
arXiv 2024
Show all 63 references
-
[9]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Int. Conf. Comput. Vis., 2021. 3
2021
-
[10]
Videocrafter1: Open diffusion models for high-quality video generation
Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, et al. Videocrafter1: Open diffusion models for high-quality video generation. arXiv preprint arXiv:2310.19512, 2023. 2
-
[11]
Zero-shot image editing with reference imitation
Xi Chen, Yutong Feng, Mengting Chen, Yiyang Wang, Shilong Zhang, Yu Liu, Yujun Shen, and Hengshuang Zhao. Zero-shot image editing with reference imitation. arXiv preprint arXiv:2406.07547, 2024. 2
2024 arXiv
-
[12]
Anydoor: Zero-shot object-level image customization
Xi Chen, Lianghua Huang, Yu Liu, Yujun Shen, Deli Zhao, and Hengshuang Zhao. Anydoor: Zero-shot object-level image customization. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 1, 2, 5, 6, 7
2024
-
[13]
Learning naturally aggregated appearance for efficient 3d editing
Ka Leong Cheng, Qiuyu Wang, Zifan Shi, Kecheng Zheng, Yinghao Xu, Hao Ouyang, Qifeng Chen, and Yujun Shen. Learning naturally aggregated appearance for efficient 3d editing. arXiv preprint arXiv:2312.06657, 2023. 2
2023 arXiv
-
[14]
Diffusion models beat GANs on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat GANs on image synthesis. In Adv. Neural Inform. Process. Syst., 2021. 2, 3, 5
2021
-
[15]
Tapir: Tracking any point with per-frame initialization and temporal refinement
Carl Doersch, Yi Yang, Mel Vecerik, Dilara Gokay, Ankush Gupta, Yusuf Aytar, Joao Carreira, and Andrew Zisserman. Tapir: Tracking any point with per-frame initialization and temporal refinement. In Int. Conf. Comput. Vis., 2023. 3
2023
-
[16]
Structure and content-guided video synthesis with diffusion models
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. In Int. Conf. Comput. Vis., 2023. 2
2023
-
[17]
Proxedit: Improving tuning-free real image editing with proximal guidance
Ligong Han, Song Wen, Qi Chen, Zhixing Zhang, Kunpeng Song, Mengwei Ren, Ruijiang Gao, Anastasis Stathopoulos, Xiaoxiao He, Yuxiao Chen, et al. Proxedit: Improving tuning-free real image editing with proximal guidance. In IEEE Winter Conf. Appl. Comput. Vis., 2024. 2
2024
-
[18]
Style aligned image generation via shared atten- tion
Amir Hertz, Andrey V oynov, Shlomi Fruchter, and Daniel Cohen-Or. Style aligned image generation via shared atten- tion. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 2, 4, 5, 6, 7
2024
-
[19]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Adv. Neural Inform. Process. Syst., 2020. 2
2020
-
[20]
Video diffusion models
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. Adv. Neural Inform. Process. Syst., 2022. 2
2022
-
[21]
Text2room: Extracting textured 3d meshes from 2d text-to-image models
Lukas H ¨ollein, Ang Cao, Andrew Owens, Justin Johnson, and Matthias Nießner. Text2room: Extracting textured 3d meshes from 2d text-to-image models. In Int. Conf. Comput. Vis., 2023. 2
2023
-
[22]
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In Int. Conf. Learn. Represent., 2022. 8
2022
-
[23]
Composer: Creative and controllable image synthesis with composable conditions
Lianghua Huang, Di Chen, Yu Liu, Yujun Shen, Deli Zhao, and Jingren Zhou. Composer: Creative and controllable image synthesis with composable conditions. arXiv preprint arXiv:2302.09778, 2023. 2
2023 arXiv
-
[24]
Cotr: Correspondence transformer for matching across images
Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, and Kwang Moo Yi. Cotr: Correspondence transformer for matching across images. In Int. Conf. Comput. Vis., 2021. 3
2021
-
[25]
Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion
Xuan Ju, Xian Liu, Xintao Wang, Yuxuan Bian, Ying Shan, and Qiang Xu. Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion. InEur. Conf. Comput. Vis., 2024. 2, 3, 5, 1
2024
-
[26]
Pnp inversion: Boosting diffusion-based editing with 3 lines of code
Xuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu, and Qiang Xu. Pnp inversion: Boosting diffusion-based editing with 3 lines of code. In Int. Conf. Learn. Represent., 2024. 2
2024
-
[27]
Co- tracker: It is better to track together
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co- tracker: It is better to track together. In Eur. Conf. Comput. Vis., 2024. 3
2024
-
[28]
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. In Adv. Neural Inform. Process. Syst., 2022. 2
2022
-
[29]
Imagic: Text-based real image editing with diffusion models
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models. In IEEE Conf. Comput. Vis. Pattern Recog., 2023
2023
-
[30]
Text2video-zero: Text- to-image diffusion models are zero-shot video generators
Levon Khachatryan, Andranik Movsisyan, Vahram Tade- vosyan, Roberto Henschel, Zhangyang Wang, Shant 9 Navasardyan, and Humphrey Shi. Text2video-zero: Text- to-image diffusion models are zero-shot video generators. In Int. Conf. Comput. Vis., 2023
2023
-
[31]
Multi-concept customization of text-to-image diffusion
Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 1, 2, 5, 7
2023
-
[32]
Ground- ing image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and J´erˆome Revaud. Ground- ing image matching in 3d with mast3r. InEur. Conf. Comput. Vis., 2024. 3
2024
-
[33]
SDEdit: Guided image synthesis and editing with stochastic differential equa- tions
Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. SDEdit: Guided image synthesis and editing with stochastic differential equa- tions. In Int. Conf. Learn. Represent., 2022. 2
2022
-
[34]
Negative-prompt inversion: Fast image inversion for editing with text-guided diffusion models
Daiki Miyake, Akihiro Iohara, Yu Saito, and Toshiyuki Tanaka. Negative-prompt inversion: Fast image inversion for editing with text-guided diffusion models. arXiv preprint arXiv:2305.16807, 2023. 2
2023 arXiv
-
[35]
Null-text inversion for editing real images using guided diffusion models
Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Null-text inversion for editing real images using guided diffusion models. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 2
2023
-
[36]
Null-text inversion for editing real images using guided diffusion models
Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Null-text inversion for editing real images using guided diffusion models. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 2, 5
2023
-
[37]
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In Assoc. Adv. Artif. Intell., 2024. 2
2024
-
[38]
Edit one for all: Interactive batch image editing
Thao Nguyen, Utkarsh Ojha, Yuheng Li, Haotian Liu, and Yong Jae Lee. Edit one for all: Interactive batch image editing. In IEEE Conf. Comput. Vis. Pattern Recog. , 2024. 2
2024
-
[39]
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In Int. Conf. Mach. Learn., 2021. 2
2021
-
[40]
CoDeF: Content deformation fields for tempo- rally consistent video processing
Hao Ouyang, Qiuyu Wang, Yuxi Xiao, Qingyan Bai, Juntao Zhang, Kecheng Zheng, Xiaowei Zhou, Qifeng Chen, and Yujun Shen. CoDeF: Content deformation fields for tempo- rally consistent video processing. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 3
2023
-
[41]
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. In Int. Conf. Learn. Represent., 2023. 2
2023
-
[42]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In Int. Conf. Mach. Learn., 2021. 5
2021
-
[43]
Segment anything meets point tracking
Frano Raji ˇc, Lei Ke, Yu-Wing Tai, Chi-Keung Tang, Martin Danelljan, and Fisher Yu. Segment anything meets point tracking. arXiv:2307.01197, 2023. 3
2023 arXiv
-
[44]
Optical flow estimation using a spatial pyramid network
Anurag Ranjan and Michael J Black. Optical flow estimation using a spatial pyramid network. InIEEE Conf. Comput. Vis. Pattern Recog., 2017. 3
2017
-
[45]
Adobe firefly: Free generative ai for cre- atives
Adobe reseachers. Adobe firefly: Free generative ai for cre- atives. https://firefly.adobe.com/generate/ inpaint, 2023. 5, 6, 7
2023
-
[46]
The md5 message-digest algorithm
Ronald Rivest. The md5 message-digest algorithm. Techni- cal report, 1992. 4
1992
-
[47]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 2, 3, 5, 1
2022
-
[48]
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In IEEE Conf. Comput. Vis. Pattern Recog. ,
-
[49]
De- noising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. De- noising diffusion implicit models. In Int. Conf. Learn. Represent., 2021. 2, 5, 1
2021
-
[50]
Emergent correspondence from image diffusion
Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan. Emergent correspondence from image diffusion. In Adv. Neural Inform. Process. Syst.,
-
[51]
Training-free consistent text-to-image generation
Yoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten, Lior Wolf, Gal Chechik, and Yuval Atzmon. Training-free consistent text-to-image generation. ACM Trans. Graph. ,
-
[52]
Plug-and-play diffusion features for text-driven image-to-image translation
Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. Plug-and-play diffusion features for text-driven image-to-image translation. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 2
2023
-
[53]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdv. Neural Inform. Process. Syst., 2017. 4
2017
-
[54]
Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion
Vikram V oleti, Chun-Han Yao, Mark Boss, Adam Letts, David Pankratz, Dmitry Tochilkin, Christian Laforte, Robin Rombach, and Varun Jampani. Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion. arXiv preprint arXiv:2403.12008, 2024. 2
2024 arXiv
-
[55]
Dust3r: Geometric 3d vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vision made easy. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 1, 3, 4, 8
2024
-
[56]
Rodin: A generative model for sculpting 3d digital avatars using diffusion
Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, et al. Rodin: A generative model for sculpting 3d digital avatars using diffusion. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 2
2023
-
[57]
Spatialtracker: Tracking any 2d pixels in 3d space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou. Spatialtracker: Tracking any 2d pixels in 3d space. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 3
2024
-
[58]
Paint by example: Exemplar-based image editing with diffusion models
Binxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang, Xuejin Chen, Xiaoyan Sun, Dong Chen, and Fang Wen. Paint by example: Exemplar-based image editing with diffusion models. In IEEE Conf. Comput. Vis. Pattern Recog. , 2023. 2, 5, 6, 7
2023
-
[59]
A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence
Junyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Pola- nia Cabrera, Varun Jampani, Deqing Sun, and Ming-Hsuan 10 Yang. A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence. In Adv. Neural Inform. Process. Syst., 2024. 3, 4
2024
-
[60]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Int. Conf. Comput. Vis., 2023. 2, 3, 5, 1
2023
-
[61]
Cross-domain correspondence learning for exemplar- based image translation
Pan Zhang, Bo Zhang, Dong Chen, Lu Yuan, and Fang Wen. Cross-domain correspondence learning for exemplar- based image translation. In IEEE Conf. Comput. Vis. Pattern Recog., 2020. 1, 3
2020
-
[62]
Uni-controlnet: All-in-one control to text-to-image diffusion models
Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao, Shaozhe Hao, Lu Yuan, and Kwan-Yee K Wong. Uni-controlnet: All-in-one control to text-to-image diffusion models. In Adv. Neural Inform. Process. Syst., 2024. 2
2024
-
[63]
Implicit
Xingran Zhou, Bo Zhang, Ting Zhang, Pan Zhang, Jianmin Bao, Dong Chen, Zhongfei Zhang, and Fang Wen. Cocosnet v2: Full-resolution correspondence learning for image trans- lation. In IEEE Conf. Comput. Vis. Pattern Recog., 2021. 1, 3 11 Edicho: Consistent Image Editing in the W...
2021
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.