Pith. sign in

REVIEW 3 major objections 4 minor 106 references

Distillation of Diffusion Features for Semantic Correspondence

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A distilled 87M-parameter model beats two-teacher baselines in semantic correspondence at 18 times the speed

desk verdict A practical distillation paper with a strong efficiency story and a 3D fine-tuning claim that needs a control to be credible. read the letter →

arxiv 2412.03512 v1 pith:NN4BC436 submitted 2024-12-04 cs.CV

classification cs.CV
keywords semanticcorrespondenceknowledgedistillationdiffusionmodelsDINOv2low-rankadaptation3Ddataaugmentationmulti-teacherdensefeaturematching
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Semantic correspondence, matching the same semantic part across two images, underpins tasks such as 3D reconstruction, image-to-image translation, and object tracking. Current best systems run two large foundation models, a Vision Transformer and a diffusion model, and combine their features, making them slow and memory-hungry. This paper claims that the complementary knowledge of those two teachers can be distilled into a single 87M-parameter DINOv2 student with low-rank adapters, and that adding unsupervised 3D multi-view data from CO3D pushes accuracy past the two-model state of the art. On SPair-71k the distilled model reaches 65.1 PCK bbox@0.1 unsupervised versus 64.0 for the combined teacher, and 70.6 with a weakly supervised pose-align step, while processing 28.6 images per second instead of 0.4. If correct, this makes high-accuracy semantic correspondence practical for near-real-time video and resource-constrained deployment.

What carries the argument

The load-bearing object is a multi-teacher distillation objective over dense similarity maps, combined with a low-rank adapter (LoRA) on the student's query and value projections. Teachers are DINOv2 (layer 11) and SDXL Turbo (layer 1, averaged over timesteps 51, 101, 151, and 201); their features are concatenated to form the target similarity map, and the student is trained to match that map with cross-entropy after a temperature-scaled softmax. The 3D augmentation stage uses CO3D's depth maps and camera intrinsics to compute a mutual-visibility mask (threshold $\epsilon = 0.01$), projects visible pixels between views, smooths the correspondence targets with a Gaussian kernel, and fine-tunes with the same dense cross-entropy loss. This design transfers the complementary world knowledge of both teachers into a single model and replaces human correspondence labels with geometric pseudo-labels from multi-view data.

What would settle it

If the CO3D-derived pseudo-labels are the source of the +3D gain, then corrupting the camera parameters or depth maps during fine-tuning, for example by adding random rotations larger than the $\epsilon = 0.01$ threshold, should erase the 0.35-to-0.69 PCK improvement reported in Table 5; if the gain survives, the geometric supervision is not doing the claimed work.

Watch

Extended reading notes

Core claim

The paper's central claim is that the complementary feature knowledge of a diffusion model and a self-supervised Vision Transformer can be transferred into one smaller, faster model without losing accuracy, and that the student can then be improved further using unlabeled 3D data. The student is DINOv2 B/14 with LoRA adapters; its training signal is the dense similarity distribution between image pairs produced by the concatenated features of DINOv2 and SDXL Turbo, matched to the student's own similarity map by a softmax cross-entropy loss. A second fine-tuning stage projects CO3D depth maps and camera parameters to build a mutual-visibility mask, creates Gaussian-blurred correspondence targets, and tunes the student with the same dense objective. The reported result is superior PCK on SPair-71k, PF-WILLOW, and CUB-200, with 87M parameters and roughly 18 times higher throughput than the strongest two-teacher baseline, while supervised fine-tuning of the same model is also competitive at much higher speed.

Load-bearing premise

The whole 3D improvement rests on CO3D's depth and camera parameters being accurate enough that the mutual-visibility mask with threshold $\epsilon = 0.01$ produces correct dense correspondences; if the geometry is wrong, the unlabeled gain is just noise.

Editorial extensions

If this is right

  • On SPair-71k, the distilled model posts 65.1 PCK bbox@0.1 in the unsupervised setting, beating the 64.0 of the combined DINOv2+SD1.5 teacher while using 87M instead of 1.1B parameters and 28.6 instead of 0.4 images per second.
  • With the pose-align weakly supervised step, the same model reaches 70.6 PCK bbox@0.1 on SPair-71k, outperforming previous pose-align methods.
  • The unsupervised 3D fine-tuning from CO3D stacks with timestep ensembling, window soft-argmax, and pose-align, lifting PCK bbox@0.1 from 62.37 to 69.09 on SPair-71k without human keypoint labels.
  • Supervised fine-tuning of the same distilled model reaches 80.19 PCK bbox@0.1 on SPair-71k, comparable to prior supervised systems at much higher throughput.
  • The method's low input resolution (434x434) and 87M parameters make semantic video correspondence practical at nearly 30 frames per second on an A100.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If dense similarity-distribution distillation is the operative mechanism, the same recipe should transfer to other dense prediction tasks beyond correspondence, such as monocular depth or segmentation, by swapping the task head and teacher features.
  • The success of a rank-8 LoRA bottleneck suggests the two teachers' useful knowledge is low-dimensional inside the student, so further compression through quantization or a smaller backbone is a natural untested extension.
  • Because the +3D gain relies on CO3D, it may not transfer to object categories absent from CO3D's 50 classes; testing on SPair-71k categories not covered by CO3D would separate geometric learning from category memorization.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a multi-teacher knowledge-distillation framework for semantic correspondence. A DINOv2-B/14 student with LoRA adapters is trained to match the pair-similarity distributions produced by two large teachers, DINOv2 and SDXL Turbo, and is then fine-tuned on CO3D using depth-derived pseudo-correspondences. The authors report state-of-the-art PCK numbers on SPair-71k, PF-WILLOW, and CUB-200, along with large gains in throughput and a reduction in parameter count relative to two-model baselines. The main contributions claimed are the distillation recipe, the 3D-data fine-tuning protocol, and the resulting efficiency-accuracy trade-off.

Significance. If the results hold, this is a practically valuable contribution: it demonstrates that the complementary strengths of DINOv2 and a diffusion model can be compressed into a single smaller model without sacrificing accuracy, and it offers a path toward real-time semantic correspondence. The paper is also notable for its unusually extensive ablation coverage, including LoRA rank, point sampling, image sampling, softmax temperature, 3D threshold, teacher choices, and timestep ensembles, and for shipping code and weights. The central claims are, however, weakened by an inconsistency in the distillation objective as written, an underspecified 3D fine-tuning loss, and a missing control that isolates the contribution of geometric pseudo-labels from the effect of simply training on more CO3D data.

major comments (3)
  1. [§3.2, Eq. (4)] Equation (4) is written as Ldist = CE(στ(F1·F′1^T), στ(F2·F′2^T)) = CE(S,S′), but Eq. (3) defines S = F1·F2^T, and the prose states that the objective is to align the teacher and student similarity distributions for an image pair. As written, the loss compares within-image teacher–student similarities rather than the cross-image teacher similarity with the corresponding student similarity. The equality to CE(S,S′) is therefore not justified, and the arguments of CE are ambiguous because CE(P,T) was defined as −E_P[log T], which requires a clear choice of which argument is the teacher distribution and which is the student distribution. Please rewrite Eq. (4) to match the described intent, for example as Ldist = CE(στ(F1·F2^T), στ(F′1·F′2^T)), and define the CE arguments consistently.
  2. [§3.2, Eq. (7)] The 3D fine-tuning objective is not well-defined. Eq. (7) reads Lfine = CE(στ(F1·F2^T), gk(G)), but G is never formally defined; the text says that following [53] a k×k Gaussian kernel is applied to 'the correspondence points' resulting in a (N×H×W) sized correspondence map, which is not the same as specifying G as a matrix of target correspondences. In addition, F1 and F2 in Eq. (7) reuse notation that earlier referred to combined teacher features, whereas the student features were denoted F′1 and F′2; it is unclear whose features enter the 3D fine-tuning loss. Finally, the second argument gk(G) is a blurred correspondence map and is not stated to be normalized into a probability distribution, so the cross-entropy is not well defined as written. Please define G, state which feature extractor produces F1 and F2 in Eq. (7), and specify the normalization of the target.
  3. [§4.2, Table 5] The claim that 3D data augmentation provides a further performance boost is confounded. Each '+3D (ours)' row differs from its baseline in two ways: it adds a new dataset (CO3D) with additional training iterations, and it uses depth/camera-derived pseudo-correspondences as the training signal. There is no control that trains the same student on CO3D image pairs for the same number of iterations using, for example, retrieval pairs or random pairs but without the mutual-visibility projection of Eq. (6). Without such a control, the observed increments of 0.35 and 0.69 bbox PCK points (and the correspondingly small img-PCK increments) cannot be attributed to the geometric pseudo-labels. Furthermore, no error bars or significance tests are reported anywhere in the paper, so the small margins in Tables 1, 4, and 5 are not established as stable. A no-geometry CO3D control is needed to support the abstract's claim that 3D augmentation is responsible for the improved performance.
minor comments (4)
  1. [§2] The phrase 'Generative Generative Adversarial Networks' in the second paragraph of the related work section contains a duplicated word; 'Generative Adversarial Networks' is intended.
  2. [§4, Implementation Details] The implementation details specify the distillation training setup (40 epochs, 12,000 COCO samples, retrieval with k=10) but not the corresponding details for the 3D fine-tuning stage: the number of CO3D frames/videos used, the number of epochs, the learning rate, and the batch size are omitted. Please report these for reproducibility.
  3. [Table 5] The first column of Table 5 is labeled 'Dataset', but the rows mix training-set identifiers with method names (e.g., 'SPair-71k Full sampling' and 'COCO Retrieval pairs'). The table would be clearer if the training data and the method were listed in separate columns.
  4. [§4.1, Table 1] The throughput numbers are reported as images per second on a single A100, but it is not stated whether this includes the pose-align test-time procedure or the multi-timestep teacher forward passes for the baselines; a precise measurement protocol would make the 18x speed-up claim easier to verify.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the distilled student is trained on teacher similarity pseudo-labels and geometry-derived correspondences, then evaluated on external keypoint benchmarks.

full rationale

The paper's derivation chain is self-contained with respect to the benchmarks it claims to predict. The student is trained by aligning its pair-similarity distribution with the teacher similarity distribution (Eq. 4) built from DINOv2 and SDXL Turbo features, and the 3D fine-tuning objective (Eq. 7) uses correspondences derived from CO3D depth maps, camera parameters, and the mutual-visibility mask (Eq. 6). None of these training targets are the SPair-71k, PF-WILLOW, or CUB-200 keypoint labels used for evaluation, so the reported PCK numbers are not equivalent to the training objective by construction. No fitted parameter is renamed as a prediction, and no load-bearing premise is justified by a self-citation: the complementarity premise is attributed to the external 'A Tale of Two Features' work, while the self-citations in the paper (e.g., a representation-learning survey and a flow-matching paper) are incidental related-work references rather than support for the central result. The skeptic's concern that the '+3D' gain is confounded by the addition of CO3D data and extra training iterations without a no-geometry control is a legitimate experimental-validity issue, but it is not circularity: the ablation still compares models against an external benchmark and does not define the claimed improvement in terms of its own inputs. The paper therefore receives a score of 0 for circularity.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The method rests on several domain assumptions: teacher similarity maps are a valid distillation target, CO3D depth and pose data produce reliable pseudo-labels, and the fixed layer and timestep choices for SDXL Turbo are informative. No new physical entities are introduced; the distilled model is a fine-tuned DINOv2 with LoRA adapters.

free parameters (6)
  • Softmax temperature tau = 0.01
    Chosen by ablation evaluated on SPair-71k (Appendix Fig. 1, main Fig. 7).
  • 3D masking threshold epsilon = 0.01
    Threshold for mutual visibility in Eq. 6; tuned via ablation on SPair-71k (Appendix Fig. 2, main Fig. 8).
  • LoRA rank r = 8
    Set by ablation on SPair-71k (Fig. 5), balancing trainable parameters at 290K and PCK.
  • Gaussian kernel size k for 3D targets = 7
    Used in Eq. 7 to blur 3D-derived correspondence targets, following [53] and not independently validated.
  • Number of retrieved training points k = 10
    Top-10 DINOv2 retrieval neighbors per source image; no ablation for this choice is reported.
  • Diffusion teacher timestep set = 51, 101, 151, 201
    Fixed set of timesteps over which SDXL Turbo teacher features are averaged; only on/off comparison is reported in Tab. 5, not a per-timestep ablation.
assumptions (4)
  • domain assumption Teacher similarity maps computed with Eq. 3 and the intended Eq. 4 are a valid learning target: mimicking the cosine-similarity distribution between image pairs transfers dense correspondence ability.
    The method's training signal is the teacher's predicted similarity map; no independent verification is provided that this target is sufficient to learn correspondence.
  • domain assumption CO3D depth maps and camera parameters are accurate enough to produce correct dense correspondences after the z-difference mask of Eq. 6.
    The 3D fine-tuning loss (Eq. 7) uses these pseudo-labels; errors in depth or poses directly corrupt targets.
  • domain assumption SDXL Turbo features are informative when extracted at layer 1 for the chosen timesteps and concatenated with DINOv2 layer 11 features.
    The layer and timestep choices are empirical (Appendix Tabs. 1-3) and may not generalize beyond the benchmark categories.
  • standard math Standard linear algebra operations, softmax, and cross-entropy are valid for the formulated objectives.
    Equations 1-7 use conventional matrix multiplication, softmax, and cross-entropy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Distillation of Diffusion Features for Semantic Correspondence." pith.science (2026). https://pith.science/paper/NN4BC436

@misc{pith2026241203512,
  author       = {Pith},
  title        = {Pith review of: Distillation of Diffusion Features for Semantic Correspondence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NN4BC436}},
  note         = {Machine review of arXiv:2412.03512}
}
read the original abstract

Semantic correspondence, the task of determining relationships between different parts of images, underpins various applications including 3D reconstruction, image-to-image translation, object tracking, and visual place recognition. Recent studies have begun to explore representations learned in large generative image models for semantic correspondence, demonstrating promising results. Building on this progress, current state-of-the-art methods rely on combining multiple large models, resulting in high computational demands and reduced efficiency. In this work, we address this challenge by proposing a more computationally efficient approach. We propose a novel knowledge distillation technique to overcome the problem of reduced efficiency. We show how to use two large vision foundation models and distill the capabilities of these complementary models into one smaller model that maintains high accuracy at reduced computational cost. Furthermore, we demonstrate that by incorporating 3D data, we are able to further improve performance, without the need for human-annotated correspondences. Overall, our empirical results demonstrate that our distilled model with 3D data augmentation achieves performance superior to current state-of-the-art methods while significantly reducing computational load and enhancing practicality for real-world applications, such as semantic video correspondence. Our code and weights are publicly available on our project page.

Figures

Figures reproduced from arXiv: 2412.03512 by the authors.

Figure 1
Figure 1. Our method achieves better performance and throughput with less parameters on the SPair-71k dataset. The circle size represents number of parameters. For more details, see Tab. 1. 101] and Vision Transformers [11, 12, 40, 43, 45] have rev￾olutionized the field, providing powerful methods to extract semantically rich features from images in various ways. Since finding semantic correspondences is challenging due to th… view at source ↗
Figure 2
Figure 2. Illustration of our multi-teacher distillation framework (a) and 3D data augmentation method (b). We distill two com￾plementary models, DINOv2 and SDXL Turbo, into one single and more efficient model. Using unsupervised 3D data augmentation we further refine our distilled model to achieve new state-of-the-art in both throughput and performance. into a single student model that provides fast inference with accurate p… view at source ↗
Figure 3
Figure 3. Examples image pairs from SPair-71k with predicted correspondences of different methods. Green indicates correct, while red indicates incorrect according to PCKbbox@0.1. (840 × 840) was used as input resolution for DINOv2. 4.1. Comparison with State-of-the-Art Tab. 1 and [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Video semantic correspondence sample, showing accurate correspondences at a high frame rate. We use source points on the first frame to calculate the corresponding points on all other frames at almost 30 FPS on an NVIDIA A100 80GB [PITH_FULL_IMAGE:figures/full_fig_p00…
Figure 5
Figure 5. Figure 5: Ablation of the rank parameter of LoRA, evaluated on SPair-71k. Trained for 20 epochs on COCO with retrieval sampling. The #Params correspond to the ranks: 4, 8, 16 and 32, respectively. Image DistillDIFT DINOv2 [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Zero-shot foreground-background differentiation us￾ing k-means. Our distilled model produces segmentation masks with less noisy edges. Dataset Method PCKimg@0.1 PCKbbox@0.1 SPair-71k Manually selected pairs 71.67 62.37 Random category pairs 70.11 60.76 Random pairs 68.…
Figure 7
Figure 7. Figure 7: Ablation of the softmax temperature parameter τ , evaluated on SPair-71k. Trained for 10 epochs on COCO with retrieval sampling. B. 3D Threshold Ablation [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: The effect of the 3D data threshold parameter ϵ. Method PCKimg@0.1 PCKbbox@0.1 DINOv2 + SD 71.77 63.29 DINOv2 + SD + CLIP 68.87 60.17 [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: Examples of the improved foreground/background segmentation masks with our model. 17 [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]
Figure 1
Figure 1. Figure 1: Ablation of the softmax temperature parameter τ , evaluated on SPair-71k. Trained for 10 epochs on COCO with retrieval sampling. 0.2 3D Threshold Ablation [PITH_FULL_IMAGE:figures/full_fig_p018_1.png]
Figure 2
Figure 2. Figure 2: The effect of the 3D data threshold parameter ϵ. Model SPair-71K PF-WILLOW CUB-200 S T L SD1.5 66.11/56.24 86.58/73.60 90.58/79.18 7682 201 5 SD2.1 65.29/57.87 87.18/74.83 88.63/78.23 7682 261 8 SDXL Base 64.02/55.52 88.37/76.49 92.39/84.20 7682 101 1 SDXL Base 65.64/5…
Figure 3
Figure 3. Figure 3: Examples of the improved foreground/background segmentation masks with our model. 4 [PITH_FULL_IMAGE:figures/full_fig_p021_3.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

106 extracted references · 53 canonical work pages

  1. [53]

    Simsc: A simple framework for seman- tic correspondence with temperature learning

    Xinghui Li, Kai Han, Xingchen Wan, and Victor Adrian Prisacariu. Simsc: A simple framework for seman- tic correspondence with temperature learning. CoRR, abs/2305.02385, 2023. 5

  2. [1]

    Deep vit features as dense visual descriptors

    Shir Amir, Yossi Gandelsman, Shai Bagon, and Tali Dekel. Deep vit features as dense visual descriptors. CoRR, abs/2112.05814, 2021. 1, 2

  3. [2]

    Segdiff: Image segmentation with diffusion proba- bilistic models, 2022

    Tomer Amit, Tal Shaharbany, Eliya Nachmani, and Lior Wolf. Segdiff: Image segmentation with diffusion proba- bilistic models, 2022. 2

  4. [3]

    Parameter efficient fine-tuning of self- supervised vits without catastrophic forgetting, 2024

    Reza Akbarian Bafghi, Nidhin Harilal, Claire Monteleoni, and Maziar Raissi. Parameter efficient fine-tuning of self- supervised vits without catastrophic forgetting, 2024. 4

  5. [4]

    Label-efficient semantic segmentation with diffusion models

    Dmitry Baranchuk, Andrey V oynov, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. Label-efficient semantic segmentation with diffusion models. In The Tenth International Conference on Learning Representa- tions, ICLR 2022, Virtual Event, April 25-29, 2022 . Open- Review.net, 2022. 2

  6. [5]

    Surf: Speeded up robust features

    Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. Surf: Speeded up robust features. In Ale ˇs Leonardis, Horst Bischof, and Axel Pinz, editors, Computer Vision – ECCV 2006, pages 404–417, Berlin, Heidelberg, 2006. Springer Berlin Heidelberg. 1, 2

  7. [6]

    Courville, and Pascal Vincent

    Yoshua Bengio, Aaron C. Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE Trans. Pattern Anal. Mach. Intell., 35(8):1798–1828,

  8. [7]

    Cunningham

    Dan Biderman, Jose Javier Gonzalez Ortiz, Jacob Portes, Mansheej Paul, Philip Greengard, Connor Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, Cody Blakeney, and John P. Cunningham. Lora learns less and forgets less. CoRR, abs/2405.09673, 2024. 4

Show all 106 references
  1. [8]

    Subpixel heatmap regression for facial landmark local- ization

    Adrian Bulat, Enrique Sanchez, and Georgios Tzimiropou- los. Subpixel heatmap regression for facial landmark local- ization. In 32nd British Machine Vision Conference 2021, BMVC 2021, Online, November 22-25, 2021 , page 422. BMV A Press, 2021. 5

  2. [9]

    Diffu- siondet: Diffusion model for object detection

    Shoufa Chen, Peize Sun, Yibing Song, and Ping Luo. Diffu- siondet: Diffusion model for object detection. InIEEE/CVF International Conference on Computer Vision, ICCV 2023, pages 19773–19786. IEEE, 2023. 2

  3. [10]

    Hinton, and David J

    Ting Chen, Lala Li, Saurabh Saxena, Geoffrey E. Hinton, and David J. Fleet. A generalist framework for panoptic segmentation of images and videos. In IEEE/CVF Interna- tional Conference on Computer Vision, ICCV 2023 , pages 909–919. IEEE, 2023. 2

  4. [11]

    Cats: Cost aggregation transformers for visual correspondence

    Seokju Cho, Sunghwan Hong, Sangryul Jeon, Yunsung Lee, Kwanghoon Sohn, and Seungryong Kim. Cats: Cost aggregation transformers for visual correspondence. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, ed- itors, Advances...

  5. [12]

    Cats++: Boosting cost aggregation with convolutions and transformers

    Seokju Cho, Sunghwan Hong, and Seungryong Kim. Cats++: Boosting cost aggregation with convolutions and transformers. CoRR, abs/2202.06817, 2022. 1, 2

  6. [13]

    Custom-edit: Text-guided image editing with customized diffusion models

    Jooyoung Choi, Yunjey Choi, Yunji Kim, Junho Kim, and Sungroh Yoon. Custom-edit: Text-guided image editing with customized diffusion models. CoRR, abs/2305.15779,

  7. [14]

    Choy, JunYoung Gwak, Silvio Savarese, and Manmohan Krishna Chandraker

    Christopher B. Choy, JunYoung Gwak, Silvio Savarese, and Manmohan Krishna Chandraker. Universal correspondence network. In Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett, editors, Ad- vances in Neural Information Processing Systems 29, p...

  8. [15]

    Diffedit: Diffusion-based semantic image editing with mask guidance

    Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. Diffedit: Diffusion-based semantic image editing with mask guidance. In The Eleventh International Conference on Learning Representations, ICLR 2023, Ki- gali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. 2

  9. [16]

    Surgical-dino: Adapter learning of foundation mod- els for depth estimation in endoscopic surgery

    Beilei Cui, Mobarakol Islam, Long Bai, and Hongliang Ren. Surgical-dino: Adapter learning of foundation mod- els for depth estimation in endoscopic surgery. CoRR, abs/2401.06013, 2024. 4

  10. [17]

    Diffusion models beat gans on image synthesis

    Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat gans on image synthesis. In Marc’Aurelio Ran- zato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, editors, Advances in Neu- ral Information Processing Systems 34 , pages 8780–8794,

  11. [18]

    An im- age is worth 16x16 words: Transformers for image recog- nition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An im- age is worth 16x16 words: Transformers for image recog- nitio...

  12. [19]

    Diffusiondepth: Diffusion denoising approach for monocular depth estima- tion

    Yiqun Duan, Zheng Zhu, and Xianda Guo. Diffusiondepth: Diffusion denoising approach for monocular depth estima- tion. CoRR, abs/2303.05021, 2023. 2

  13. [20]

    Diffusion models and representation learning: A survey

    Michael Fuest, Pingchuan Ma, Ming Gui, Johannes Schus- terbauer, Vincent Tao Hu, and Bjorn Ommer. Diffusion models and representation learning: A survey. arXiv preprint arXiv:2407.00783, 2024. 2

  14. [21]

    Sanchit Gandhi, Patrick von Platen, and Alexander M. Rush. Distil-whisper: Robust knowledge distillation via large-scale pseudo labelling. CoRR, abs/2311.00430, 2023. 3

  15. [22]

    Aiatrack: Attention in attention 9 for transformer visual tracking

    Shenyuan Gao, Chunluan Zhou, Chao Ma, Xinggang Wang, and Junsong Yuan. Aiatrack: Attention in attention 9 for transformer visual tracking. In Shai Avidan, Gabriel J. Brostow, Moustapha Ciss´e, Giovanni Maria Farinella, and Tal Hassner, editors, Computer Vision - ECCV 2022 , vo...

  16. [23]

    Do semantic parts emerge in convolutional neural net- works? Int

    Abel Gonzalez-Garcia, Davide Modolo, and Vittorio Fer- rari. Do semantic parts emerge in convolutional neural net- works? Int. J. Comput. Vis., 126(5):476–494, 2018. 2

  17. [24]

    Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C

    Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial networks. CoRR, abs/1406.2661, 2014. 2

  18. [25]

    Bootstrap your own latent-a new approach to self-supervised learning

    Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Do- ersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neur...

  19. [26]

    Susskind

    Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Lingjie Liu, and Josh M. Susskind. BOOT: data-free distillation of denoising diffusion models with bootstrapping. CoRR, abs/2306.05544, 2023. 3

  20. [27]

    Knowledge distillation of large language models

    Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. Knowledge distillation of large language models. CoRR, abs/2306.08543, 2023. 3

  21. [28]

    Depthfm: Fast monocular depth estimation with flow matching

    Ming Gui, Johannes Schusterbauer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Ste- fan Andreas Baumann, Vincent Tao Hu, and Bj ¨orn Om- mer. Depthfm: Fast monocular depth estimation with flow matching. arXiv preprint arXiv:2403.13788, 2024. 2

  22. [29]

    ASIC: aligning sparse in-the-wild image collec- tions

    Kamal Gupta, Varun Jampani, Carlos Esteves, Abhinav Shrivastava, Ameesh Makadia, Noah Snavely, and Ab- hishek Kar. ASIC: aligning sparse in-the-wild image collec- tions. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, pages 4111–4122. IEEE, 2023. 1, 2

  23. [30]

    Proposal flow

    Bumsub Ham, Minsu Cho, Cordelia Schmid, and Jean Ponce. Proposal flow. In 2016 IEEE Conference on Com- puter Vision and Pattern Recognition, CVPR 2016, Las Ve- gas, NV , USA, June 27-30, 2016, pages 3475–3484. IEEE Computer Society, 2016. 5

  24. [31]

    Rezende, Bumsub Ham, Kwan-Yee K

    Kai Han, Rafael S. Rezende, Bumsub Ham, Kwan-Yee K. Wong, Minsu Cho, Cordelia Schmid, and Jean Ponce. Sc- net: Learning semantic correspondence. In IEEE Interna- tional Conference on Computer Vision, ICCV 2017 , pages 1849–1858. IEEE Computer Society, 2017. 1

  25. [32]

    Unsupervised semantic correspondence using stable diffusion

    Eric Hedlin, Gopal Sharma, Shweta Mahajan, Hos- sam Isack, Abhishek Kar, Andrea Tagliasacchi, and Kwang Moo Yi. Unsupervised semantic correspondence using stable diffusion. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Adv...

  26. [33]

    Prompt-to-prompt im- age editing with cross-attention control

    Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt im- age editing with cross-attention control. InThe Eleventh In- ternational Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net,

  27. [34]

    Hinton, Oriol Vinyals, and Jeffrey Dean

    Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. Distilling the knowledge in a neural network. CoRR, abs/1503.02531, 2015. 2, 3

  28. [35]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Repre- sentations, ICLR 2022, Virtual Event, April 25-29, 2022 ...

  29. [36]

    Tao Hu, David W Zhang, Pascal Mettes, Meng Tang, Deli Zhao, and Cees G.M. Snoek. Latent space editing in transformer-based flow matching. In AAAI, 2024. 2

  30. [37]

    Twigg, Po-Chen Wu, Junsong Yuan, Cem Keskin, and Robert Wang

    Lin Huang, Tomas Hodan, Lingni Ma, Linguang Zhang, Luan Tran, Christopher D. Twigg, Po-Chen Wu, Junsong Yuan, Cem Keskin, and Robert Wang. Neural corre- spondence field for object pose estimation. In Shai Avi- dan, Gabriel J. Brostow, Moustapha Ciss´e, Giovanni Maria Farinella...

  31. [38]

    Allan Jabri, Andrew Owens, and Alexei A. Efros. Space- time correspondence as a contrastive random walk. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria- Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33, 2020. 2

  32. [39]

    Difnet: Semantic segmentation by diffu- sion networks

    Peng Jiang, Fanglin Gu, Yunhai Wang, Changhe Tu, and Baoquan Chen. Difnet: Semantic segmentation by diffu- sion networks. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol `o Cesa-Bianchi, and Roman Garnett, editors, Advances in Neural Information Proce...

  33. [40]

    COTR: correspondence transformer for matching across images

    Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, and Kwang Moo Yi. COTR: correspondence transformer for matching across images. In 2021 IEEE/CVF Interna- tional Conference on Computer Vision, ICCV 2021 , pages 6187–6197. IEEE, 2021. 1, 2

  34. [41]

    Imagic: Text-based real image editing with diffusion mod- els

    Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Hui- wen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion mod- els. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June...

  35. [42]

    Scherer, K

    Nikhil Varma Keetha, Avneesh Mishra, Jay Karhade, Kr- ishna Murthy Jatavallabhula, Sebastian A. Scherer, K. Mad- hava Krishna, and Sourav Garg. Anyloc: Towards univer- sal visual place recognition. IEEE Robotics Autom. Lett. , 9(2):1286–1293, 2024. 1

  36. [43]

    Recurrent transformer net- works for semantic correspondence

    Seungryong Kim, Stephen Lin, Sangryul Jeon, Dongbo Min, and Kwanghoon Sohn. Recurrent transformer net- works for semantic correspondence. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol`o Cesa-Bianchi, and Roman Garnett, editors, Ad- vances in Neural ...

  37. [44]

    FCSS: fully convolutional self- similarity for dense semantic correspondence

    Seungryong Kim, Dongbo Min, Bumsub Ham, Stephen Lin, and Kwanghoon Sohn. FCSS: fully convolutional self- similarity for dense semantic correspondence. IEEE Trans. Pattern Anal. Mach. Intell., 41(3):581–595, 2019. 2 10

  38. [45]

    Transfor- matcher: Match-to-match attention for semantic correspon- dence

    Seungwook Kim, Juhong Min, and Minsu Cho. Transfor- matcher: Match-to-match attention for semantic correspon- dence. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 8687–8697. IEEE, 2022. 1, 2

  39. [46]

    Yoon Kim and Alexander M. Rush. Sequence-level knowl- edge distillation. In Jian Su, Xavier Carreras, and Kevin Duh, editors, Proceedings of the 2016 Conference on Em- pirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016 , pages 13...

  40. [47]

    To the point: Correspondence-driven monocular 3d category reconstruc- tion

    Filippos Kokkinos and Iasonas Kokkinos. To the point: Correspondence-driven monocular 3d category reconstruc- tion. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, ed- itors, Advances in Neural Information Processing Syst...

  41. [48]

    Sfnet: Learning object-aware semantic correspon- dence

    Junghyup Lee, Dohyung Kim, Jean Ponce, and Bumsub Ham. Sfnet: Learning object-aware semantic correspon- dence. In IEEE Conference on Computer Vision and Pat- tern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 2278–2287. Computer Vision Founda- tion / IEE...

  42. [49]

    Jae Yong Lee, Joseph DeGol, Victor Fragoso, and Sudipta N. Sinha. Patchmatch-based neighborhood con- sensus for semantic correspondence. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 , pages 13153–13163. Computer Vision Fou...

  43. [50]

    Li, Mihir Prabhudesai, Shivam Duggal, El- lis Brown, and Deepak Pathak

    Alexander C. Li, Mihir Prabhudesai, Shivam Duggal, El- lis Brown, and Deepak Pathak. Your diffusion model is secretly a zero-shot classifier. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, pages 2206–

  44. [51]

    Costain, Henry Howard- Jenkins, and Victor Prisacariu

    Shuda Li, Kai Han, Theo W. Costain, Henry Howard- Jenkins, and Victor Prisacariu. Correspondence net- works with adaptive neighbourhood consensus. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 10193...

  45. [52]

    Probabilistic model distillation for se- mantic correspondence

    Xin Li, Deng-Ping Fan, Fan Yang, Ao Luo, Hong Cheng, and Zicheng Liu. Probabilistic model distillation for se- mantic correspondence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7505–7514, June 2021. 3

  46. [54]

    Sd4match: Learning to prompt stable diffusion model for semantic matching

    Xinghui Li, Jingyi Lu, Kai Han, and Victor Prisacariu. Sd4match: Learning to prompt stable diffusion model for semantic matching. CoRR, abs/2310.17569, 2023. 2, 5

  47. [55]

    Data distillation for text classifi- cation

    Yongqi Li and Wenjie Li. Data distillation for text classifi- cation. arXiv preprint arXiv:2104.08448, 2021. 2

  48. [56]

    Cycle-consistency based hierarchical dense semantic cor- respondence

    Chuang Lin, Hongxun Yao, Wei Yu, and Xiaoshuai Sun. Cycle-consistency based hierarchical dense semantic cor- respondence. In 2018 IEEE International Conference on Image Processing, ICIP 2018, Athens, Greece, October 7- 10, 2018, pages 818–822. IEEE, 2018. 2

  49. [57]

    Belongie, Lubomir D

    Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C. Lawrence Zit- nick. Microsoft COCO: common objects in context. CoRR, abs/1405.0312, 2014. 5

  50. [58]

    Scale Invariant Feature Transform, vol- ume 7

    Tony Lindeberg. Scale Invariant Feature Transform, vol- ume 7. 05 2012. 1, 2

  51. [59]

    Sift flow: Dense correspondence across scenes and its applications

    Ce Liu, Jenny Yuen, and Antonio Torralba. Sift flow: Dense correspondence across scenes and its applications. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(5):978–994, 2011. 2

  52. [60]

    Structured knowledge dis- tillation for semantic segmentation

    Yifan Liu, Ke Chen, Chris Liu, Zengchang Qin, Zhenbo Luo, and Jingdong Wang. Structured knowledge dis- tillation for semantic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2604–2613, 2019. 2

  53. [61]

    Do con- vnets learn correspondence? In Zoubin Ghahramani, Max Welling, Corinna Cortes, Neil D

    Jonathan Long, Ning Zhang, and Trevor Darrell. Do con- vnets learn correspondence? In Zoubin Ghahramani, Max Welling, Corinna Cortes, Neil D. Lawrence, and Kilian Q. Weinberger, editors, Advances in Neural Information Pro- cessing Systems 27, pages 1601–1609, 2014. 2

  54. [62]

    Diffusion hyperfeatures: Searching through time and space for semantic correspon- dence

    Grace Luo, Lisa Dunlap, Dong Huk Park, Aleksander Holynski, and Trevor Darrell. Diffusion hyperfeatures: Searching through time and space for semantic correspon- dence. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors,Ad- vanc...

  55. [63]

    Im- proving semantic correspondence with viewpoint-guided spherical maps

    Octave Mariotti, Oisin Mac Aodha, and Hakan Bilen. Im- proving semantic correspondence with viewpoint-guided spherical maps. CoRR, abs/2312.13216, 2023. 1, 3, 6, 8

  56. [64]

    Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans

    Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik P. Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans. On distillation of guided diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 , pages 1...

  57. [65]

    Con- ditional teacher-student learning

    Zhong Meng, Jinyu Li, Yong Zhao, and Yifan Gong. Con- ditional teacher-student learning. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2019, Brighton, United Kingdom, May 12-17, 2019, pages 6445–6449. IEEE, 2019. 3

  58. [66]

    Spair-71k: A large-scale benchmark for semantic corre- spondence

    Juhong Min, Jongmin Lee, Jean Ponce, and Minsu Cho. Spair-71k: A large-scale benchmark for semantic corre- spondence. CoRR, abs/1908.10543, 2019. 5

  59. [67]

    Learning to compose hypercolumns for visual correspon- dence

    Juhong Min, Jongmin Lee, Jean Ponce, and Minsu Cho. Learning to compose hypercolumns for visual correspon- dence. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision - ECCV 2020, volume 12360 ofLecture Notes in Computer Science, pages...

  60. [68]

    Coordgan: Self-supervised dense correspondences emerge from gans

    Jiteng Mu, Shalini De Mello, Zhiding Yu, Nuno Vasconce- los, Xiaolong Wang, Jan Kautz, and Sifei Liu. Coordgan: Self-supervised dense correspondences emerge from gans. 11 In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 1...

  61. [69]

    Diffusion models beat gans on image classification

    Soumik Mukhopadhyay, Matthew Gwilliam, Vatsal Agar- wal, Namitha Padmanabhan, Archana Swaminathan, Srinidhi Hegde, Tianyi Zhou, and Abhinav Shrivastava. Diffusion models beat gans on image classification. CoRR, abs/2307.08702, 2023. 2

  62. [70]

    Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Herv ´e J ´egou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski

    Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Rus- sell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michae...

  63. [71]

    FEED: feature-level en- semble for knowledge distillation

    Seonguk Park and Nojun Kwak. FEED: feature-level en- semble for knowledge distillation. CoRR, abs/1909.10754,

  64. [72]

    Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction

    Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotn´y. Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In 2021 IEEE/CVF In- ternational Conference on Computer Vision, ICCV ...

  65. [73]

    Con- volutional neural network architecture for geometric match- ing

    Ignacio Rocco, Relja Arandjelovic, and Josef Sivic. Con- volutional neural network architecture for geometric match- ing. IEEE Trans. Pattern Anal. Mach. Intell., 41(11):2553– 2567, 2019. 1

  66. [74]

    Effi- cient neighbourhood consensus networks via submanifold sparse convolutions

    Ignacio Rocco, Relja Arandjelovic, and Josef Sivic. Effi- cient neighbourhood consensus networks via submanifold sparse convolutions. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision - ECCV 2020 , volume 12354 of Lecture Notes in C...

  67. [75]

    Neighbourhood consensus networks

    Ignacio Rocco, Mircea Cimpoi, Relja Arandjelovic, Aki- hiko Torii, Tom´as Pajdla, and Josef Sivic. Neighbourhood consensus networks. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol `o Cesa-Bianchi, and Roman Garnett, editors, Advances in Neural Inform...

  68. [76]

    High-resolution im- age synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution im- age synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 , pages 1067...

  69. [77]

    Fit- nets: Hints for thin deep nets

    Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. Fit- nets: Hints for thin deep nets. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May ...

  70. [78]

    Progressive distillation for fast sampling of diffusion models

    Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In The Tenth In- ternational Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net,

  71. [79]

    Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter.CoRR, abs/1910.01108,

    Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter.CoRR, abs/1910.01108,

  72. [80]

    Balasubramanian

    Bharat Bhusan Sau and Vineeth N. Balasubramanian. Deep model compression: Distilling knowledge from noisy teachers. CoRR, abs/1610.09650, 2016. 3

  73. [81]

    Saurabh Saxena, Abhishek Kar, Mohammad Norouzi, and David J. Fleet. Monocular depth estimation using diffusion models. CoRR, abs/2302.14816, 2023. 2

  74. [82]

    Baumann, Vincent Tao Hu, and Bj ¨orn Ommer

    Johannes Schusterbauer, Ming Gui, Pingchuan Ma, Nick Stracke, Stefan A. Baumann, Vincent Tao Hu, and Bj ¨orn Ommer. Boosting latent diffusion with flow matching. In ECCV, 2024. 2

  75. [83]

    Shuwei Shao, Zhongcai Pei, Weihai Chen, Dingchi Sun, Peter C. Y . Chen, and Zhengguo Li. Monodiffusion: Self-supervised monocular depth estimation using diffusion model. CoRR, abs/2311.07198, 2023. 2, 3

  76. [84]

    Yujun Shi, Chuhui Xue, Jun Hao Liew, Jiachun Pan, Han- shu Yan, Wenqing Zhang, Vincent Y . F. Tan, and Song Bai. Dragdiffusion: Harnessing diffusion models for interactive point-based image editing, 2023. 2

  77. [85]

    Dpodv2: Dense correspondence-based 6 dof pose estima- tion

    Ivan Shugurov, Sergey Zakharov, and Slobodan Ilic. Dpodv2: Dense correspondence-based 6 dof pose estima- tion. IEEE Trans. Pattern Anal. Mach. Intell., 44(11):7417– 7435, 2022. 1

  78. [86]

    Consistency models

    Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, International Con- ference on Machine Learning, ICML 2023 , volume 202 of Procee...

  79. [87]

    Visual correspondence-based explanations improve AI ro- bustness and human-ai team accuracy

    Mohammad Reza Taesiri, Giang Nguyen, and Anh Nguyen. Visual correspondence-based explanations improve AI ro- bustness and human-ai team accuracy. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing S...

  80. [88]

    Semantic dif- fusion network for semantic segmentation

    Haoru Tan, Sitong Wu, and Jimin Pi. Semantic dif- fusion network for semantic segmentation. CoRR, abs/2302.02057, 2023. 2

  81. [89]

    Diffss: Diffu- sion model for few-shot semantic segmentation

    Weimin Tan, Siyuan Chen, and Bo Yan. Diffss: Diffu- sion model for few-shot semantic segmentation. CoRR, abs/2307.00773, 2023. 2

  82. [90]

    Emergent correspondence from image diffusion

    Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan. Emergent correspondence from image diffusion. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information Processing Sys- t...

  83. [91]

    Splicing vit features for semantic appearance trans- fer

    Narek Tumanyan, Omer Bar-Tal, Shai Bagon, and Tali Dekel. Splicing vit features for semantic appearance trans- fer. In IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 10738–10747. IEEE, 2022. 1 12

  84. [92]

    Plug-and-play diffusion features for text-driven image-to-image translation

    Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. Plug-and-play diffusion features for text-driven image-to-image translation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 , pages 1921–

  85. [93]

    C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. Caltech-ucsd birds-200-2011 (cub-200-2011). Technical Report CNS-TR-2011-001, California Institute of Technology, 2011. 5

  86. [94]

    Learning feature descriptors using cam- era pose supervision

    Qianqian Wang, Xiaowei Zhou, Bharath Hariharan, and Noah Snavely. Learning feature descriptors using cam- era pose supervision. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision - ECCV 2020 , volume 12346 of Lecture Notes in Compute...

  87. [95]

    Xiaolong Wang, Allan Jabri, and Alexei A. Efros. Learn- ing correspondence from the cycle-consistency of time. In IEEE Conference on Computer Vision and Pattern Recogni- tion, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 2566–2576. Computer Vision Foundation / IEEE,

  88. [96]

    Julia Wolleb, Robin Sandk ¨uhler, Florentin Bieder, Philippe Valmaggia, and Philippe C. Cattin. Diffusion mod- els for implicit image segmentation ensembles. In En- der Konukoglu, Bjoern H. Menze, Archana Venkatara- man, Christian F. Baumgartner, Qi Dou, and Shadi Albar- qouni...

  89. [97]

    Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using dif- fusion models

    Weijia Wu, Yuzhong Zhao, Mike Zheng Shou, Hong Zhou, and Chunhua Shen. Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using dif- fusion models. In IEEE/CVF International Conference on Computer Vision, ICCV 2023 , pages 1206–1217. IEEE,

  90. [98]

    Open-vocabulary panoptic segmentation with text-to-image diffusion models

    Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xi- aolong Wang, and Shalini De Mello. Open-vocabulary panoptic segmentation with text-to-image diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, ...

  91. [99]

    A gift from knowledge distillation: Fast optimization, net- work minimization and transfer learning

    Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim. A gift from knowledge distillation: Fast optimization, net- work minimization and transfer learning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017 , pages 7...

  92. [100]

    Paying more at- tention to attention: Improving the performance of convo- lutional neural networks via attention transfer

    Sergey Zagoruyko and Nikos Komodakis. Paying more at- tention to attention: Improving the performance of convo- lutional neural networks via attention transfer. In 5th In- ternational Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Confere...

  93. [101]

    Zeiler and Rob Fergus

    Matthew D. Zeiler and Rob Fergus. Visualizing and under- standing convolutional networks. In David J. Fleet, Tom ´as Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, Com- puter Vision - ECCV 2014 , volume 8689 of Lecture Notes in Computer Science, pages 818–833. Springer,...

  94. [102]

    A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence

    Junyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Pola- nia Cabrera, Varun Jampani, Deqing Sun, and Ming-Hsuan Yang. A tale of two features: Stable diffusion complements DINO for zero-shot semantic correspondence. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Morit...

  95. [103]

    Telling left from right: Identifying geometry-aware seman- tic correspondence

    Junyi Zhang, Charles Herrmann, Junhwa Hur, Eric Chen, Varun Jampani, Deqing Sun, and Ming-Hsuan Yang. Telling left from right: Identifying geometry-aware seman- tic correspondence. CoRR, abs/2311.17034, 2023. 1, 3, 5, 6, 7, 8

  96. [104]

    Unleashing text-to-image diffusion models for visual perception

    Wenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu, Jie Zhou, and Jiwen Lu. Unleashing text-to-image diffusion models for visual perception. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, pages 5706–

  97. [105]

    Learning deep features for discrim- inative localization

    Bolei Zhou, Aditya Khosla, `Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discrim- inative localization. In 2016 IEEE Conference on Com- puter Vision and Pattern Recognition, CVPR 2016, Las Ve- gas, NV , USA, June 27-30, 2016, pages 2921–2929. I...

  98. [106]

    a photo of a [category]

    Yitao Zhu, Zhenrong Shen, Zihao Zhao, Sheng Wang, Xin Wang, Xiangyu Zhao, Dinggang Shen, and Qian Wang. Melo: Low-rank adaptation is better than fine-tuning for medical image diagnosis. CoRR, abs/2311.08236, 2023. 4 13 Appendix A. Softmax Temperature Ablation In Fig. 7, we abl...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.