REVIEW 3 major objections 4 minor 106 references
Distillation of Diffusion Features for Semantic Correspondence
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A distilled 87M-parameter model beats two-teacher baselines in semantic correspondence at 18 times the speed
desk verdict A practical distillation paper with a strong efficiency story and a 3D fine-tuning claim that needs a control to be credible. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a multi-teacher distillation objective over dense similarity maps, combined with a low-rank adapter (LoRA) on the student's query and value projections. Teachers are DINOv2 (layer 11) and SDXL Turbo (layer 1, averaged over timesteps 51, 101, 151, and 201); their features are concatenated to form the target similarity map, and the student is trained to match that map with cross-entropy after a temperature-scaled softmax. The 3D augmentation stage uses CO3D's depth maps and camera intrinsics to compute a mutual-visibility mask (threshold $\epsilon = 0.01$), projects visible pixels between views, smooths the correspondence targets with a Gaussian kernel, and fine-tunes with the same dense cross-entropy loss. This design transfers the complementary world knowledge of both teachers into a single model and replaces human correspondence labels with geometric pseudo-labels from multi-view data.
What would settle it
If the CO3D-derived pseudo-labels are the source of the +3D gain, then corrupting the camera parameters or depth maps during fine-tuning, for example by adding random rotations larger than the $\epsilon = 0.01$ threshold, should erase the 0.35-to-0.69 PCK improvement reported in Table 5; if the gain survives, the geometric supervision is not doing the claimed work.
Extended reading notes
Core claim
The paper's central claim is that the complementary feature knowledge of a diffusion model and a self-supervised Vision Transformer can be transferred into one smaller, faster model without losing accuracy, and that the student can then be improved further using unlabeled 3D data. The student is DINOv2 B/14 with LoRA adapters; its training signal is the dense similarity distribution between image pairs produced by the concatenated features of DINOv2 and SDXL Turbo, matched to the student's own similarity map by a softmax cross-entropy loss. A second fine-tuning stage projects CO3D depth maps and camera parameters to build a mutual-visibility mask, creates Gaussian-blurred correspondence targets, and tunes the student with the same dense objective. The reported result is superior PCK on SPair-71k, PF-WILLOW, and CUB-200, with 87M parameters and roughly 18 times higher throughput than the strongest two-teacher baseline, while supervised fine-tuning of the same model is also competitive at much higher speed.
Load-bearing premise
The whole 3D improvement rests on CO3D's depth and camera parameters being accurate enough that the mutual-visibility mask with threshold $\epsilon = 0.01$ produces correct dense correspondences; if the geometry is wrong, the unlabeled gain is just noise.
Editorial extensions
If this is right
- On SPair-71k, the distilled model posts 65.1 PCK bbox@0.1 in the unsupervised setting, beating the 64.0 of the combined DINOv2+SD1.5 teacher while using 87M instead of 1.1B parameters and 28.6 instead of 0.4 images per second.
- With the pose-align weakly supervised step, the same model reaches 70.6 PCK bbox@0.1 on SPair-71k, outperforming previous pose-align methods.
- The unsupervised 3D fine-tuning from CO3D stacks with timestep ensembling, window soft-argmax, and pose-align, lifting PCK bbox@0.1 from 62.37 to 69.09 on SPair-71k without human keypoint labels.
- Supervised fine-tuning of the same distilled model reaches 80.19 PCK bbox@0.1 on SPair-71k, comparable to prior supervised systems at much higher throughput.
- The method's low input resolution (434x434) and 87M parameters make semantic video correspondence practical at nearly 30 frames per second on an A100.
Reading between the lines
- If dense similarity-distribution distillation is the operative mechanism, the same recipe should transfer to other dense prediction tasks beyond correspondence, such as monocular depth or segmentation, by swapping the task head and teacher features.
- The success of a rank-8 LoRA bottleneck suggests the two teachers' useful knowledge is low-dimensional inside the student, so further compression through quantization or a smaller backbone is a natural untested extension.
- Because the +3D gain relies on CO3D, it may not transfer to object categories absent from CO3D's 50 classes; testing on SPair-71k categories not covered by CO3D would separate geometric learning from category memorization.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a multi-teacher knowledge-distillation framework for semantic correspondence. A DINOv2-B/14 student with LoRA adapters is trained to match the pair-similarity distributions produced by two large teachers, DINOv2 and SDXL Turbo, and is then fine-tuned on CO3D using depth-derived pseudo-correspondences. The authors report state-of-the-art PCK numbers on SPair-71k, PF-WILLOW, and CUB-200, along with large gains in throughput and a reduction in parameter count relative to two-model baselines. The main contributions claimed are the distillation recipe, the 3D-data fine-tuning protocol, and the resulting efficiency-accuracy trade-off.
Significance. If the results hold, this is a practically valuable contribution: it demonstrates that the complementary strengths of DINOv2 and a diffusion model can be compressed into a single smaller model without sacrificing accuracy, and it offers a path toward real-time semantic correspondence. The paper is also notable for its unusually extensive ablation coverage, including LoRA rank, point sampling, image sampling, softmax temperature, 3D threshold, teacher choices, and timestep ensembles, and for shipping code and weights. The central claims are, however, weakened by an inconsistency in the distillation objective as written, an underspecified 3D fine-tuning loss, and a missing control that isolates the contribution of geometric pseudo-labels from the effect of simply training on more CO3D data.
major comments (3)
- [§3.2, Eq. (4)] Equation (4) is written as Ldist = CE(στ(F1·F′1^T), στ(F2·F′2^T)) = CE(S,S′), but Eq. (3) defines S = F1·F2^T, and the prose states that the objective is to align the teacher and student similarity distributions for an image pair. As written, the loss compares within-image teacher–student similarities rather than the cross-image teacher similarity with the corresponding student similarity. The equality to CE(S,S′) is therefore not justified, and the arguments of CE are ambiguous because CE(P,T) was defined as −E_P[log T], which requires a clear choice of which argument is the teacher distribution and which is the student distribution. Please rewrite Eq. (4) to match the described intent, for example as Ldist = CE(στ(F1·F2^T), στ(F′1·F′2^T)), and define the CE arguments consistently.
- [§3.2, Eq. (7)] The 3D fine-tuning objective is not well-defined. Eq. (7) reads Lfine = CE(στ(F1·F2^T), gk(G)), but G is never formally defined; the text says that following [53] a k×k Gaussian kernel is applied to 'the correspondence points' resulting in a (N×H×W) sized correspondence map, which is not the same as specifying G as a matrix of target correspondences. In addition, F1 and F2 in Eq. (7) reuse notation that earlier referred to combined teacher features, whereas the student features were denoted F′1 and F′2; it is unclear whose features enter the 3D fine-tuning loss. Finally, the second argument gk(G) is a blurred correspondence map and is not stated to be normalized into a probability distribution, so the cross-entropy is not well defined as written. Please define G, state which feature extractor produces F1 and F2 in Eq. (7), and specify the normalization of the target.
- [§4.2, Table 5] The claim that 3D data augmentation provides a further performance boost is confounded. Each '+3D (ours)' row differs from its baseline in two ways: it adds a new dataset (CO3D) with additional training iterations, and it uses depth/camera-derived pseudo-correspondences as the training signal. There is no control that trains the same student on CO3D image pairs for the same number of iterations using, for example, retrieval pairs or random pairs but without the mutual-visibility projection of Eq. (6). Without such a control, the observed increments of 0.35 and 0.69 bbox PCK points (and the correspondingly small img-PCK increments) cannot be attributed to the geometric pseudo-labels. Furthermore, no error bars or significance tests are reported anywhere in the paper, so the small margins in Tables 1, 4, and 5 are not established as stable. A no-geometry CO3D control is needed to support the abstract's claim that 3D augmentation is responsible for the improved performance.
minor comments (4)
- [§2] The phrase 'Generative Generative Adversarial Networks' in the second paragraph of the related work section contains a duplicated word; 'Generative Adversarial Networks' is intended.
- [§4, Implementation Details] The implementation details specify the distillation training setup (40 epochs, 12,000 COCO samples, retrieval with k=10) but not the corresponding details for the 3D fine-tuning stage: the number of CO3D frames/videos used, the number of epochs, the learning rate, and the batch size are omitted. Please report these for reproducibility.
- [Table 5] The first column of Table 5 is labeled 'Dataset', but the rows mix training-set identifiers with method names (e.g., 'SPair-71k Full sampling' and 'COCO Retrieval pairs'). The table would be clearer if the training data and the method were listed in separate columns.
- [§4.1, Table 1] The throughput numbers are reported as images per second on a single A100, but it is not stated whether this includes the pose-align test-time procedure or the multi-timestep teacher forward passes for the baselines; a precise measurement protocol would make the 18x speed-up claim easier to verify.
Circularity Check
No significant circularity: the distilled student is trained on teacher similarity pseudo-labels and geometry-derived correspondences, then evaluated on external keypoint benchmarks.
full rationale
The paper's derivation chain is self-contained with respect to the benchmarks it claims to predict. The student is trained by aligning its pair-similarity distribution with the teacher similarity distribution (Eq. 4) built from DINOv2 and SDXL Turbo features, and the 3D fine-tuning objective (Eq. 7) uses correspondences derived from CO3D depth maps, camera parameters, and the mutual-visibility mask (Eq. 6). None of these training targets are the SPair-71k, PF-WILLOW, or CUB-200 keypoint labels used for evaluation, so the reported PCK numbers are not equivalent to the training objective by construction. No fitted parameter is renamed as a prediction, and no load-bearing premise is justified by a self-citation: the complementarity premise is attributed to the external 'A Tale of Two Features' work, while the self-citations in the paper (e.g., a representation-learning survey and a flow-matching paper) are incidental related-work references rather than support for the central result. The skeptic's concern that the '+3D' gain is confounded by the addition of CO3D data and extra training iterations without a no-geometry control is a legitimate experimental-validity issue, but it is not circularity: the ablation still compares models against an external benchmark and does not define the claimed improvement in terms of its own inputs. The paper therefore receives a score of 0 for circularity.
Assumptions & free parameters
free parameters (6)
- Softmax temperature tau =
0.01
- 3D masking threshold epsilon =
0.01
- LoRA rank r =
8
- Gaussian kernel size k for 3D targets =
7
- Number of retrieved training points k =
10
- Diffusion teacher timestep set =
51, 101, 151, 201
assumptions (4)
- domain assumption Teacher similarity maps computed with Eq. 3 and the intended Eq. 4 are a valid learning target: mimicking the cosine-similarity distribution between image pairs transfers dense correspondence ability.
- domain assumption CO3D depth maps and camera parameters are accurate enough to produce correct dense correspondences after the z-difference mask of Eq. 6.
- domain assumption SDXL Turbo features are informative when extracted at layer 1 for the chosen timesteps and concatenated with DINOv2 layer 11 features.
- standard math Standard linear algebra operations, softmax, and cross-entropy are valid for the formulated objectives.
Cite this review
Pith. "Pith review of Distillation of Diffusion Features for Semantic Correspondence." pith.science (2026). https://pith.science/paper/NN4BC436
@misc{pith2026241203512,
author = {Pith},
title = {Pith review of: Distillation of Diffusion Features for Semantic Correspondence},
year = {2026},
howpublished = {\url{https://pith.science/paper/NN4BC436}},
note = {Machine review of arXiv:2412.03512}
}
read the original abstract
Semantic correspondence, the task of determining relationships between different parts of images, underpins various applications including 3D reconstruction, image-to-image translation, object tracking, and visual place recognition. Recent studies have begun to explore representations learned in large generative image models for semantic correspondence, demonstrating promising results. Building on this progress, current state-of-the-art methods rely on combining multiple large models, resulting in high computational demands and reduced efficiency. In this work, we address this challenge by proposing a more computationally efficient approach. We propose a novel knowledge distillation technique to overcome the problem of reduced efficiency. We show how to use two large vision foundation models and distill the capabilities of these complementary models into one smaller model that maintains high accuracy at reduced computational cost. Furthermore, we demonstrate that by incorporating 3D data, we are able to further improve performance, without the need for human-annotated correspondences. Overall, our empirical results demonstrate that our distilled model with 3D data augmentation achieves performance superior to current state-of-the-art methods while significantly reducing computational load and enhancing practicality for real-world applications, such as semantic video correspondence. Our code and weights are publicly available on our project page.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[53]
Simsc: A simple framework for seman- tic correspondence with temperature learning
Xinghui Li, Kai Han, Xingchen Wan, and Victor Adrian Prisacariu. Simsc: A simple framework for seman- tic correspondence with temperature learning. CoRR, abs/2305.02385, 2023. 5
arXiv 2023
-
[1]
Deep vit features as dense visual descriptors
Shir Amir, Yossi Gandelsman, Shai Bagon, and Tali Dekel. Deep vit features as dense visual descriptors. CoRR, abs/2112.05814, 2021. 1, 2
arXiv 2021
-
[2]
Segdiff: Image segmentation with diffusion proba- bilistic models, 2022
Tomer Amit, Tal Shaharbany, Eliya Nachmani, and Lior Wolf. Segdiff: Image segmentation with diffusion proba- bilistic models, 2022. 2
2022
-
[3]
Parameter efficient fine-tuning of self- supervised vits without catastrophic forgetting, 2024
Reza Akbarian Bafghi, Nidhin Harilal, Claire Monteleoni, and Maziar Raissi. Parameter efficient fine-tuning of self- supervised vits without catastrophic forgetting, 2024. 4
2024
-
[4]
Label-efficient semantic segmentation with diffusion models
Dmitry Baranchuk, Andrey V oynov, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. Label-efficient semantic segmentation with diffusion models. In The Tenth International Conference on Learning Representa- tions, ICLR 2022, Virtual Event, April 25-29, 2022 . Open- Review.net, 2022. 2
2022
-
[5]
Surf: Speeded up robust features
Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. Surf: Speeded up robust features. In Ale ˇs Leonardis, Horst Bischof, and Axel Pinz, editors, Computer Vision – ECCV 2006, pages 404–417, Berlin, Heidelberg, 2006. Springer Berlin Heidelberg. 1, 2
2006
-
[6]
Courville, and Pascal Vincent
Yoshua Bengio, Aaron C. Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE Trans. Pattern Anal. Mach. Intell., 35(8):1798–1828,
-
[7]
Dan Biderman, Jose Javier Gonzalez Ortiz, Jacob Portes, Mansheej Paul, Philip Greengard, Connor Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, Cody Blakeney, and John P. Cunningham. Lora learns less and forgets less. CoRR, abs/2405.09673, 2024. 4
arXiv 2024
Show all 106 references
-
[8]
Subpixel heatmap regression for facial landmark local- ization
Adrian Bulat, Enrique Sanchez, and Georgios Tzimiropou- los. Subpixel heatmap regression for facial landmark local- ization. In 32nd British Machine Vision Conference 2021, BMVC 2021, Online, November 22-25, 2021 , page 422. BMV A Press, 2021. 5
2021
-
[9]
Diffu- siondet: Diffusion model for object detection
Shoufa Chen, Peize Sun, Yibing Song, and Ping Luo. Diffu- siondet: Diffusion model for object detection. InIEEE/CVF International Conference on Computer Vision, ICCV 2023, pages 19773–19786. IEEE, 2023. 2
2023
-
[10]
Hinton, and David J
Ting Chen, Lala Li, Saurabh Saxena, Geoffrey E. Hinton, and David J. Fleet. A generalist framework for panoptic segmentation of images and videos. In IEEE/CVF Interna- tional Conference on Computer Vision, ICCV 2023 , pages 909–919. IEEE, 2023. 2
2023
-
[11]
Cats: Cost aggregation transformers for visual correspondence
Seokju Cho, Sunghwan Hong, Sangryul Jeon, Yunsung Lee, Kwanghoon Sohn, and Seungryong Kim. Cats: Cost aggregation transformers for visual correspondence. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, ed- itors, Advances...
2021
-
[12]
Cats++: Boosting cost aggregation with convolutions and transformers
Seokju Cho, Sunghwan Hong, and Seungryong Kim. Cats++: Boosting cost aggregation with convolutions and transformers. CoRR, abs/2202.06817, 2022. 1, 2
2022 arXiv
-
[13]
Custom-edit: Text-guided image editing with customized diffusion models
Jooyoung Choi, Yunjey Choi, Yunji Kim, Junho Kim, and Sungroh Yoon. Custom-edit: Text-guided image editing with customized diffusion models. CoRR, abs/2305.15779,
-
[14]
Choy, JunYoung Gwak, Silvio Savarese, and Manmohan Krishna Chandraker
Christopher B. Choy, JunYoung Gwak, Silvio Savarese, and Manmohan Krishna Chandraker. Universal correspondence network. In Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett, editors, Ad- vances in Neural Information Processing Systems 29, p...
2016
-
[15]
Diffedit: Diffusion-based semantic image editing with mask guidance
Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. Diffedit: Diffusion-based semantic image editing with mask guidance. In The Eleventh International Conference on Learning Representations, ICLR 2023, Ki- gali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. 2
2023
-
[16]
Surgical-dino: Adapter learning of foundation mod- els for depth estimation in endoscopic surgery
Beilei Cui, Mobarakol Islam, Long Bai, and Hongliang Ren. Surgical-dino: Adapter learning of foundation mod- els for depth estimation in endoscopic surgery. CoRR, abs/2401.06013, 2024. 4
2024 arXiv
-
[17]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat gans on image synthesis. In Marc’Aurelio Ran- zato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, editors, Advances in Neu- ral Information Processing Systems 34 , pages 8780–8794,
-
[18]
An im- age is worth 16x16 words: Transformers for image recog- nition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An im- age is worth 16x16 words: Transformers for image recog- nitio...
2021
-
[19]
Diffusiondepth: Diffusion denoising approach for monocular depth estima- tion
Yiqun Duan, Zheng Zhu, and Xianda Guo. Diffusiondepth: Diffusion denoising approach for monocular depth estima- tion. CoRR, abs/2303.05021, 2023. 2
2023 arXiv
-
[20]
Diffusion models and representation learning: A survey
Michael Fuest, Pingchuan Ma, Ming Gui, Johannes Schus- terbauer, Vincent Tao Hu, and Bjorn Ommer. Diffusion models and representation learning: A survey. arXiv preprint arXiv:2407.00783, 2024. 2
2024 arXiv
-
[21]
Sanchit Gandhi, Patrick von Platen, and Alexander M. Rush. Distil-whisper: Robust knowledge distillation via large-scale pseudo labelling. CoRR, abs/2311.00430, 2023. 3
2023 arXiv
-
[22]
Aiatrack: Attention in attention 9 for transformer visual tracking
Shenyuan Gao, Chunluan Zhou, Chao Ma, Xinggang Wang, and Junsong Yuan. Aiatrack: Attention in attention 9 for transformer visual tracking. In Shai Avidan, Gabriel J. Brostow, Moustapha Ciss´e, Giovanni Maria Farinella, and Tal Hassner, editors, Computer Vision - ECCV 2022 , vo...
2022
-
[23]
Do semantic parts emerge in convolutional neural net- works? Int
Abel Gonzalez-Garcia, Davide Modolo, and Vittorio Fer- rari. Do semantic parts emerge in convolutional neural net- works? Int. J. Comput. Vis., 126(5):476–494, 2018. 2
2018
-
[24]
Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial networks. CoRR, abs/1406.2661, 2014. 2
2014 arXiv
-
[25]
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Do- ersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neur...
2020
-
[26]
Susskind
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Lingjie Liu, and Josh M. Susskind. BOOT: data-free distillation of denoising diffusion models with bootstrapping. CoRR, abs/2306.05544, 2023. 3
2023 arXiv
-
[27]
Knowledge distillation of large language models
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. Knowledge distillation of large language models. CoRR, abs/2306.08543, 2023. 3
2023 arXiv
-
[28]
Depthfm: Fast monocular depth estimation with flow matching
Ming Gui, Johannes Schusterbauer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Ste- fan Andreas Baumann, Vincent Tao Hu, and Bj ¨orn Om- mer. Depthfm: Fast monocular depth estimation with flow matching. arXiv preprint arXiv:2403.13788, 2024. 2
2024 arXiv
-
[29]
ASIC: aligning sparse in-the-wild image collec- tions
Kamal Gupta, Varun Jampani, Carlos Esteves, Abhinav Shrivastava, Ameesh Makadia, Noah Snavely, and Ab- hishek Kar. ASIC: aligning sparse in-the-wild image collec- tions. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, pages 4111–4122. IEEE, 2023. 1, 2
2023
-
[30]
Proposal flow
Bumsub Ham, Minsu Cho, Cordelia Schmid, and Jean Ponce. Proposal flow. In 2016 IEEE Conference on Com- puter Vision and Pattern Recognition, CVPR 2016, Las Ve- gas, NV , USA, June 27-30, 2016, pages 3475–3484. IEEE Computer Society, 2016. 5
2016
-
[31]
Rezende, Bumsub Ham, Kwan-Yee K
Kai Han, Rafael S. Rezende, Bumsub Ham, Kwan-Yee K. Wong, Minsu Cho, Cordelia Schmid, and Jean Ponce. Sc- net: Learning semantic correspondence. In IEEE Interna- tional Conference on Computer Vision, ICCV 2017 , pages 1849–1858. IEEE Computer Society, 2017. 1
2017
-
[32]
Unsupervised semantic correspondence using stable diffusion
Eric Hedlin, Gopal Sharma, Shweta Mahajan, Hos- sam Isack, Abhishek Kar, Andrea Tagliasacchi, and Kwang Moo Yi. Unsupervised semantic correspondence using stable diffusion. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Adv...
2023
-
[33]
Prompt-to-prompt im- age editing with cross-attention control
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt im- age editing with cross-attention control. InThe Eleventh In- ternational Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net,
2023
-
[34]
Hinton, Oriol Vinyals, and Jeffrey Dean
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. Distilling the knowledge in a neural network. CoRR, abs/1503.02531, 2015. 2, 3
2015 arXiv
-
[35]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Repre- sentations, ICLR 2022, Virtual Event, April 25-29, 2022 ...
2022
-
[36]
Tao Hu, David W Zhang, Pascal Mettes, Meng Tang, Deli Zhao, and Cees G.M. Snoek. Latent space editing in transformer-based flow matching. In AAAI, 2024. 2
2024
-
[37]
Twigg, Po-Chen Wu, Junsong Yuan, Cem Keskin, and Robert Wang
Lin Huang, Tomas Hodan, Lingni Ma, Linguang Zhang, Luan Tran, Christopher D. Twigg, Po-Chen Wu, Junsong Yuan, Cem Keskin, and Robert Wang. Neural corre- spondence field for object pose estimation. In Shai Avi- dan, Gabriel J. Brostow, Moustapha Ciss´e, Giovanni Maria Farinella...
2022
-
[38]
Allan Jabri, Andrew Owens, and Alexei A. Efros. Space- time correspondence as a contrastive random walk. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria- Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33, 2020. 2
2020
-
[39]
Difnet: Semantic segmentation by diffu- sion networks
Peng Jiang, Fanglin Gu, Yunhai Wang, Changhe Tu, and Baoquan Chen. Difnet: Semantic segmentation by diffu- sion networks. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol `o Cesa-Bianchi, and Roman Garnett, editors, Advances in Neural Information Proce...
2018
-
[40]
COTR: correspondence transformer for matching across images
Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, and Kwang Moo Yi. COTR: correspondence transformer for matching across images. In 2021 IEEE/CVF Interna- tional Conference on Computer Vision, ICCV 2021 , pages 6187–6197. IEEE, 2021. 1, 2
2021
-
[41]
Imagic: Text-based real image editing with diffusion mod- els
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Hui- wen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion mod- els. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June...
2023
-
[42]
Scherer, K
Nikhil Varma Keetha, Avneesh Mishra, Jay Karhade, Kr- ishna Murthy Jatavallabhula, Sebastian A. Scherer, K. Mad- hava Krishna, and Sourav Garg. Anyloc: Towards univer- sal visual place recognition. IEEE Robotics Autom. Lett. , 9(2):1286–1293, 2024. 1
2024
-
[43]
Recurrent transformer net- works for semantic correspondence
Seungryong Kim, Stephen Lin, Sangryul Jeon, Dongbo Min, and Kwanghoon Sohn. Recurrent transformer net- works for semantic correspondence. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol`o Cesa-Bianchi, and Roman Garnett, editors, Ad- vances in Neural ...
2018
-
[44]
FCSS: fully convolutional self- similarity for dense semantic correspondence
Seungryong Kim, Dongbo Min, Bumsub Ham, Stephen Lin, and Kwanghoon Sohn. FCSS: fully convolutional self- similarity for dense semantic correspondence. IEEE Trans. Pattern Anal. Mach. Intell., 41(3):581–595, 2019. 2 10
2019
-
[45]
Transfor- matcher: Match-to-match attention for semantic correspon- dence
Seungwook Kim, Juhong Min, and Minsu Cho. Transfor- matcher: Match-to-match attention for semantic correspon- dence. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 8687–8697. IEEE, 2022. 1, 2
2022
-
[46]
Yoon Kim and Alexander M. Rush. Sequence-level knowl- edge distillation. In Jian Su, Xavier Carreras, and Kevin Duh, editors, Proceedings of the 2016 Conference on Em- pirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016 , pages 13...
2016
-
[47]
To the point: Correspondence-driven monocular 3d category reconstruc- tion
Filippos Kokkinos and Iasonas Kokkinos. To the point: Correspondence-driven monocular 3d category reconstruc- tion. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, ed- itors, Advances in Neural Information Processing Syst...
2021
-
[48]
Sfnet: Learning object-aware semantic correspon- dence
Junghyup Lee, Dohyung Kim, Jean Ponce, and Bumsub Ham. Sfnet: Learning object-aware semantic correspon- dence. In IEEE Conference on Computer Vision and Pat- tern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 2278–2287. Computer Vision Founda- tion / IEE...
2019
-
[49]
Jae Yong Lee, Joseph DeGol, Victor Fragoso, and Sudipta N. Sinha. Patchmatch-based neighborhood con- sensus for semantic correspondence. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 , pages 13153–13163. Computer Vision Fou...
2021
-
[50]
Li, Mihir Prabhudesai, Shivam Duggal, El- lis Brown, and Deepak Pathak
Alexander C. Li, Mihir Prabhudesai, Shivam Duggal, El- lis Brown, and Deepak Pathak. Your diffusion model is secretly a zero-shot classifier. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, pages 2206–
2023
-
[51]
Costain, Henry Howard- Jenkins, and Victor Prisacariu
Shuda Li, Kai Han, Theo W. Costain, Henry Howard- Jenkins, and Victor Prisacariu. Correspondence net- works with adaptive neighbourhood consensus. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 10193...
2020
-
[52]
Probabilistic model distillation for se- mantic correspondence
Xin Li, Deng-Ping Fan, Fan Yang, Ao Luo, Hong Cheng, and Zicheng Liu. Probabilistic model distillation for se- mantic correspondence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7505–7514, June 2021. 3
2021
-
[54]
Sd4match: Learning to prompt stable diffusion model for semantic matching
Xinghui Li, Jingyi Lu, Kai Han, and Victor Prisacariu. Sd4match: Learning to prompt stable diffusion model for semantic matching. CoRR, abs/2310.17569, 2023. 2, 5
2023 arXiv
-
[55]
Data distillation for text classifi- cation
Yongqi Li and Wenjie Li. Data distillation for text classifi- cation. arXiv preprint arXiv:2104.08448, 2021. 2
2021 arXiv
-
[56]
Cycle-consistency based hierarchical dense semantic cor- respondence
Chuang Lin, Hongxun Yao, Wei Yu, and Xiaoshuai Sun. Cycle-consistency based hierarchical dense semantic cor- respondence. In 2018 IEEE International Conference on Image Processing, ICIP 2018, Athens, Greece, October 7- 10, 2018, pages 818–822. IEEE, 2018. 2
2018
-
[57]
Belongie, Lubomir D
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C. Lawrence Zit- nick. Microsoft COCO: common objects in context. CoRR, abs/1405.0312, 2014. 5
2014 arXiv
-
[58]
Scale Invariant Feature Transform, vol- ume 7
Tony Lindeberg. Scale Invariant Feature Transform, vol- ume 7. 05 2012. 1, 2
2012
-
[59]
Sift flow: Dense correspondence across scenes and its applications
Ce Liu, Jenny Yuen, and Antonio Torralba. Sift flow: Dense correspondence across scenes and its applications. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(5):978–994, 2011. 2
2011
-
[60]
Structured knowledge dis- tillation for semantic segmentation
Yifan Liu, Ke Chen, Chris Liu, Zengchang Qin, Zhenbo Luo, and Jingdong Wang. Structured knowledge dis- tillation for semantic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2604–2613, 2019. 2
2019
-
[61]
Do con- vnets learn correspondence? In Zoubin Ghahramani, Max Welling, Corinna Cortes, Neil D
Jonathan Long, Ning Zhang, and Trevor Darrell. Do con- vnets learn correspondence? In Zoubin Ghahramani, Max Welling, Corinna Cortes, Neil D. Lawrence, and Kilian Q. Weinberger, editors, Advances in Neural Information Pro- cessing Systems 27, pages 1601–1609, 2014. 2
2014
-
[62]
Diffusion hyperfeatures: Searching through time and space for semantic correspon- dence
Grace Luo, Lisa Dunlap, Dong Huk Park, Aleksander Holynski, and Trevor Darrell. Diffusion hyperfeatures: Searching through time and space for semantic correspon- dence. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors,Ad- vanc...
2023
-
[63]
Im- proving semantic correspondence with viewpoint-guided spherical maps
Octave Mariotti, Oisin Mac Aodha, and Hakan Bilen. Im- proving semantic correspondence with viewpoint-guided spherical maps. CoRR, abs/2312.13216, 2023. 1, 3, 6, 8
2023 arXiv
-
[64]
Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans
Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik P. Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans. On distillation of guided diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 , pages 1...
2023
-
[65]
Con- ditional teacher-student learning
Zhong Meng, Jinyu Li, Yong Zhao, and Yifan Gong. Con- ditional teacher-student learning. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2019, Brighton, United Kingdom, May 12-17, 2019, pages 6445–6449. IEEE, 2019. 3
2019
-
[66]
Spair-71k: A large-scale benchmark for semantic corre- spondence
Juhong Min, Jongmin Lee, Jean Ponce, and Minsu Cho. Spair-71k: A large-scale benchmark for semantic corre- spondence. CoRR, abs/1908.10543, 2019. 5
1908 arXiv
-
[67]
Learning to compose hypercolumns for visual correspon- dence
Juhong Min, Jongmin Lee, Jean Ponce, and Minsu Cho. Learning to compose hypercolumns for visual correspon- dence. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision - ECCV 2020, volume 12360 ofLecture Notes in Computer Science, pages...
2020
-
[68]
Coordgan: Self-supervised dense correspondences emerge from gans
Jiteng Mu, Shalini De Mello, Zhiding Yu, Nuno Vasconce- los, Xiaolong Wang, Jan Kautz, and Sifei Liu. Coordgan: Self-supervised dense correspondences emerge from gans. 11 In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 1...
2022
-
[69]
Diffusion models beat gans on image classification
Soumik Mukhopadhyay, Matthew Gwilliam, Vatsal Agar- wal, Namitha Padmanabhan, Archana Swaminathan, Srinidhi Hegde, Tianyi Zhou, and Abhinav Shrivastava. Diffusion models beat gans on image classification. CoRR, abs/2307.08702, 2023. 2
2023 arXiv
-
[70]
Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Herv ´e J ´egou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Rus- sell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michae...
-
[71]
FEED: feature-level en- semble for knowledge distillation
Seonguk Park and Nojun Kwak. FEED: feature-level en- semble for knowledge distillation. CoRR, abs/1909.10754,
1909 arXiv
-
[72]
Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotn´y. Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In 2021 IEEE/CVF In- ternational Conference on Computer Vision, ICCV ...
2021
-
[73]
Con- volutional neural network architecture for geometric match- ing
Ignacio Rocco, Relja Arandjelovic, and Josef Sivic. Con- volutional neural network architecture for geometric match- ing. IEEE Trans. Pattern Anal. Mach. Intell., 41(11):2553– 2567, 2019. 1
2019
-
[74]
Effi- cient neighbourhood consensus networks via submanifold sparse convolutions
Ignacio Rocco, Relja Arandjelovic, and Josef Sivic. Effi- cient neighbourhood consensus networks via submanifold sparse convolutions. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision - ECCV 2020 , volume 12354 of Lecture Notes in C...
2020
-
[75]
Neighbourhood consensus networks
Ignacio Rocco, Mircea Cimpoi, Relja Arandjelovic, Aki- hiko Torii, Tom´as Pajdla, and Josef Sivic. Neighbourhood consensus networks. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol `o Cesa-Bianchi, and Roman Garnett, editors, Advances in Neural Inform...
2018
-
[76]
High-resolution im- age synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution im- age synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 , pages 1067...
2022
-
[77]
Fit- nets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. Fit- nets: Hints for thin deep nets. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May ...
2015
-
[78]
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In The Tenth In- ternational Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net,
2022
-
[79]
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter.CoRR, abs/1910.01108,
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter.CoRR, abs/1910.01108,
1910 arXiv
-
[80]
Balasubramanian
Bharat Bhusan Sau and Vineeth N. Balasubramanian. Deep model compression: Distilling knowledge from noisy teachers. CoRR, abs/1610.09650, 2016. 3
2016 arXiv
-
[81]
Saurabh Saxena, Abhishek Kar, Mohammad Norouzi, and David J. Fleet. Monocular depth estimation using diffusion models. CoRR, abs/2302.14816, 2023. 2
2023 arXiv
-
[82]
Baumann, Vincent Tao Hu, and Bj ¨orn Ommer
Johannes Schusterbauer, Ming Gui, Pingchuan Ma, Nick Stracke, Stefan A. Baumann, Vincent Tao Hu, and Bj ¨orn Ommer. Boosting latent diffusion with flow matching. In ECCV, 2024. 2
2024
-
[83]
Shuwei Shao, Zhongcai Pei, Weihai Chen, Dingchi Sun, Peter C. Y . Chen, and Zhengguo Li. Monodiffusion: Self-supervised monocular depth estimation using diffusion model. CoRR, abs/2311.07198, 2023. 2, 3
2023 arXiv
-
[84]
Yujun Shi, Chuhui Xue, Jun Hao Liew, Jiachun Pan, Han- shu Yan, Wenqing Zhang, Vincent Y . F. Tan, and Song Bai. Dragdiffusion: Harnessing diffusion models for interactive point-based image editing, 2023. 2
2023
-
[85]
Dpodv2: Dense correspondence-based 6 dof pose estima- tion
Ivan Shugurov, Sergey Zakharov, and Slobodan Ilic. Dpodv2: Dense correspondence-based 6 dof pose estima- tion. IEEE Trans. Pattern Anal. Mach. Intell., 44(11):7417– 7435, 2022. 1
2022
-
[86]
Consistency models
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, International Con- ference on Machine Learning, ICML 2023 , volume 202 of Procee...
2023
-
[87]
Visual correspondence-based explanations improve AI ro- bustness and human-ai team accuracy
Mohammad Reza Taesiri, Giang Nguyen, and Anh Nguyen. Visual correspondence-based explanations improve AI ro- bustness and human-ai team accuracy. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing S...
2022
-
[88]
Semantic dif- fusion network for semantic segmentation
Haoru Tan, Sitong Wu, and Jimin Pi. Semantic dif- fusion network for semantic segmentation. CoRR, abs/2302.02057, 2023. 2
2023 arXiv
-
[89]
Diffss: Diffu- sion model for few-shot semantic segmentation
Weimin Tan, Siyuan Chen, and Bo Yan. Diffss: Diffu- sion model for few-shot semantic segmentation. CoRR, abs/2307.00773, 2023. 2
2023 arXiv
-
[90]
Emergent correspondence from image diffusion
Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan. Emergent correspondence from image diffusion. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information Processing Sys- t...
2023
-
[91]
Splicing vit features for semantic appearance trans- fer
Narek Tumanyan, Omer Bar-Tal, Shai Bagon, and Tali Dekel. Splicing vit features for semantic appearance trans- fer. In IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 10738–10747. IEEE, 2022. 1 12
2022
-
[92]
Plug-and-play diffusion features for text-driven image-to-image translation
Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. Plug-and-play diffusion features for text-driven image-to-image translation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 , pages 1921–
2023
-
[93]
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. Caltech-ucsd birds-200-2011 (cub-200-2011). Technical Report CNS-TR-2011-001, California Institute of Technology, 2011. 5
2011
-
[94]
Learning feature descriptors using cam- era pose supervision
Qianqian Wang, Xiaowei Zhou, Bharath Hariharan, and Noah Snavely. Learning feature descriptors using cam- era pose supervision. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision - ECCV 2020 , volume 12346 of Lecture Notes in Compute...
2020
-
[95]
Xiaolong Wang, Allan Jabri, and Alexei A. Efros. Learn- ing correspondence from the cycle-consistency of time. In IEEE Conference on Computer Vision and Pattern Recogni- tion, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 2566–2576. Computer Vision Foundation / IEEE,
2019
-
[96]
Julia Wolleb, Robin Sandk ¨uhler, Florentin Bieder, Philippe Valmaggia, and Philippe C. Cattin. Diffusion mod- els for implicit image segmentation ensembles. In En- der Konukoglu, Bjoern H. Menze, Archana Venkatara- man, Christian F. Baumgartner, Qi Dou, and Shadi Albar- qouni...
2022
-
[97]
Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using dif- fusion models
Weijia Wu, Yuzhong Zhao, Mike Zheng Shou, Hong Zhou, and Chunhua Shen. Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using dif- fusion models. In IEEE/CVF International Conference on Computer Vision, ICCV 2023 , pages 1206–1217. IEEE,
2023
-
[98]
Open-vocabulary panoptic segmentation with text-to-image diffusion models
Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xi- aolong Wang, and Shalini De Mello. Open-vocabulary panoptic segmentation with text-to-image diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, ...
2023
-
[99]
A gift from knowledge distillation: Fast optimization, net- work minimization and transfer learning
Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim. A gift from knowledge distillation: Fast optimization, net- work minimization and transfer learning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017 , pages 7...
2017
-
[100]
Paying more at- tention to attention: Improving the performance of convo- lutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis. Paying more at- tention to attention: Improving the performance of convo- lutional neural networks via attention transfer. In 5th In- ternational Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Confere...
2017
-
[101]
Zeiler and Rob Fergus
Matthew D. Zeiler and Rob Fergus. Visualizing and under- standing convolutional networks. In David J. Fleet, Tom ´as Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, Com- puter Vision - ECCV 2014 , volume 8689 of Lecture Notes in Computer Science, pages 818–833. Springer,...
2014
-
[102]
A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence
Junyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Pola- nia Cabrera, Varun Jampani, Deqing Sun, and Ming-Hsuan Yang. A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Morit...
2023
-
[103]
Telling left from right: Identifying geometry-aware seman- tic correspondence
Junyi Zhang, Charles Herrmann, Junhwa Hur, Eric Chen, Varun Jampani, Deqing Sun, and Ming-Hsuan Yang. Telling left from right: Identifying geometry-aware seman- tic correspondence. CoRR, abs/2311.17034, 2023. 1, 3, 5, 6, 7, 8
2023 arXiv
-
[104]
Unleashing text-to-image diffusion models for visual perception
Wenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu, Jie Zhou, and Jiwen Lu. Unleashing text-to-image diffusion models for visual perception. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, pages 5706–
2023
-
[105]
Learning deep features for discrim- inative localization
Bolei Zhou, Aditya Khosla, `Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discrim- inative localization. In 2016 IEEE Conference on Com- puter Vision and Pattern Recognition, CVPR 2016, Las Ve- gas, NV , USA, June 27-30, 2016, pages 2921–2929. I...
2016
-
[106]
a photo of a [category]
Yitao Zhu, Zhenrong Shen, Zihao Zhao, Sheng Wang, Xin Wang, Xiangyu Zhao, Dinggang Shen, and Qian Wang. Melo: Low-rank adaptation is better than fine-tuning for medical image diagnosis. CoRR, abs/2311.08236, 2023. 4 13 Appendix A. Softmax Temperature Ablation In Fig. 7, we abl...
2023 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.