Pith. sign in

REVIEW 4 major objections 2 minor 57 references

Towards Robust Semantic Correspondence: A Benchmark and Insights

T0 review · 4 major / 2 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper's abstract claims a 14-scenario benchmark with robustness insights; the full text is an unrelated gamma-ray burst study.

desk verdict The submission is two different papers: the abstract promises a computer-vision benchmark, the full text is a GRB astrophysics paper, so the claimed benchmark cannot be reviewed. read the letter →

arxiv 2508.00272 v1 pith:A3SHJ5HD submitted 2025-08-01 cs.CV

classification cs.CV
keywords semanticcorrespondencerobustnessbenchmarkadverseconditionslarge-scalevisionmodelsDINOStableDiffusiondataaugmentationcomputer
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The submission announces a benchmark for evaluating semantic correspondence under adverse conditions: 14 scenarios spanning geometric distortion, blurring, digital artifacts, and occlusion, and it reports three insights about robustness. A sympathetic reading takes these as the paper's claims: all methods degrade under adverse conditions; large-scale models improve overall robustness but fine-tuning erodes relative robustness; DINO beats Stable Diffusion on relative robustness while their fusion wins on absolute robustness, and general data augmentation fails. None of this can be checked in the submitted full text, because the body is an unrelated astrophysics study of the X-ray emission of GRB 220711B. The reader is left with an abstract that promises results and a manuscript that does not contain them.

What carries the argument

The central object in the abstract is the benchmark dataset itself: 14 adverse-condition scenarios organized into four families—geometric distortion, image blurring, digital artifacts, and environmental occlusion—used to measure robustness of semantic correspondence methods. The argumentative machinery is a comparative evaluation protocol pitting large-scale vision models (DINO and Stable Diffusion) and their fusion against each other, plus an augmentation study. In the submitted full text, none of this machinery appears; the only machinery present belongs to the magnetar spin-down and precession model proposed for GRB 220711B.

What would settle it

Open the submitted full text and search for the terms 'semantic correspondence', 'DINO', 'Stable Diffusion', or 'benchmark'; none appear, and the body instead discusses GRB 220711B's X-ray light curves and a magnetar precession model. That direct observation suffices to show that the abstract's central claims are not established by this submission.

Watch

Extended reading notes

Core claim

The abstract's discovery is that semantic correspondence methods, despite strong performance on clean images, degrade noticeably under common imaging problems, and that model choice and fine-tuning reshape that degradation in a specific way: DINO's features are relatively more robust than Stable Diffusion's, and combining them improves absolute accuracy. The paper also claims that standard data augmentation does not fix robustness, implying task-specific designs are needed. That is the core claim the author forwards. The attached full text, however, contains no semantic-correspondence experiments, no benchmark construction, and no DINO or Stable Diffusion comparisons; it presents an analysis of GRB 220711B's X-ray light curves and proposes a precessing magnetar central-engine model.

Load-bearing premise

The load-bearing premise is that the attached manuscript body actually is the benchmark study described in the title and abstract; without that, the abstract's claims have no supporting methods, data, or results.

Editorial extensions

If this is right

  • If the benchmark is sound, robustness should be reported alongside accuracy in future semantic-correspondence work, not as an afterthought.
  • The claimed fine-tuning robustness drop would urge caution in adapting large vision models for correspondence tasks, since better absolute accuracy can come with more fragile behavior under adverse conditions.
  • The reported advantage of fusing DINO and Stable Diffusion would point toward feature ensembling as a practical robustness lever over any single model.
  • The ineffectiveness of general data augmentation, if true, would redirect effort toward augmentation and training schemes designed specifically around geometric distortion, blur, digital artifacts, and occlusion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A reader who acts on this abstract alone would be making decisions about semantic-correspondence method selection on evidence this submission does not actually provide.
  • The mismatch suggests either a file-attachment error at submission or a metadata mix-up; a corrected resubmission pairing the abstract with its benchmark study is needed before the claims can be evaluated.
  • The abstract's phrasing leaves open a testable research program: measuring whether synthetic adverse-condition training or test-time adaptation closes the gap between absolute and relative robustness.
  • The reported 'relative robustness' is not defined in the abstract; an editor's inference is that it presumably normalizes each method's adverse-condition performance by its clean performance, but the submission does not state this.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 2 minor

Summary. The submission as supplied consists of an abstract titled "Towards Robust Semantic Correspondence: A Benchmark and Insights" (arXiv:2508.00272, cs.CV) followed by a full-text preprint that is a different paper: an MNRAS-style astrophysics article titled "Signature of a magnetar central engine with precession motion in the X-ray emission of GRB 220711B" (arXiv:2508.00284v1, astro-ph.HE). The body text contains no dataset definition, no evaluation protocol, no experimental tables, no error bars, and no analysis of semantic correspondence in adverse conditions. Consequently, every substantive claim in the abstract—the 14-scenario benchmark, the robustness insights for large-scale vision models, and the evaluation of augmentation strategies—is unsupported by the content of this submission.

Significance. If the claimed benchmark and robustness study existed and were properly validated, the work could be a valuable resource for the semantic correspondence community, particularly for evaluating large-scale visual features under degradations such as geometric distortion, blur, digital artifacts, and occlusion. However, the submitted manuscript provides no evidence for any of these contributions: there is no dataset, no method description, no results, and no code. The only content available for review is an unrelated astrophysics preprint, so no scientific claim from the abstract can currently be assessed or credited.

major comments (4)
  1. [Full text (entire manuscript)] The full text is not the paper described by the title and abstract. The body is an astrophysics preprint about the X-ray emission of GRB 220711B, a magnetar central engine, and quasi-periodic oscillations, and its header explicitly identifies it as arXiv:2508.00284v1 [astro-ph.HE]. The target manuscript is arXiv:2508.00272 (cs.CV). No part of the supplied full text discusses semantic correspondence, the 14 challenging scenarios, DINO, Stable Diffusion, or data augmentation, so the central claim in the abstract has zero supporting content in this submission.
  2. [Full text, dataset definition] The claimed benchmark dataset is never defined. The abstract refers to a dataset comprising 14 challenging scenarios in geometric distortion, image blurring, digital artifacts, and environmental occlusion, but nowhere in the supplied text is there a description of how these images were collected or generated, what annotations or ground-truth correspondences are provided, what splits are used, or how the dataset can be accessed. Without this material, the core contribution of a benchmark cannot be evaluated.
  3. [Full text, evaluation protocol and results] No evaluation protocol or experimental results appear in the manuscript. There are no tables reporting performance drops, no error bars, no baseline comparisons, no definitions of the metrics, and no descriptions of the augmentation strategies tested. The only figures and fitted models concern the GRB light curves, QPO searches, and magnetar spin-down, which are irrelevant to the abstract's claims. The abstract's specific statements that fine-tuning large-scale models reduces relative robustness, that DINO outperforms Stable Diffusion in relative robustness, and that their fusion gives better absolute robustness are therefore entirely unverifiable from this submission.
  4. [Full text, Data Availability statement] The Data Availability section of the supplied full text states that 'There are no new data associated with this article.' This directly contradicts the abstract's claim that the paper establishes a new benchmark dataset. Even if this statement belongs to the astrophysics paper rather than to the intended semantic-correspondence paper, its presence in the submitted manuscript is a document-level inconsistency that must be resolved by supplying the correct full text.
minor comments (2)
  1. [Abstract] The abstract has inconsistent capitalization in 'Moreover, We evaluate' where 'We' should be lowercase, and it mixes 'we' and 'We' across the paragraph.
  2. [References] The reference list contains only astrophysics literature; there are no citations to semantic correspondence benchmarks, DINO, Stable Diffusion, or augmentation methods, which would be required for the paper described by the title and abstract.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation is visible; the attached full text is an unrelated astrophysics paper, so the benchmark abstract's claims are unverifiable rather than circular.

full rationale

The submitted full text is an unrelated MNRAS astrophysics preprint, 'Signature of a magnetar central engine with precession motion in the X-ray emission of GRB 220711B' (arXiv:2508.00284v1 [astro-ph.HE]), while the title and abstract describe a computer-vision benchmark for semantic correspondence under adverse conditions (arXiv:2508.00272 cs.CV). Within the astrophysical text, the magnetar spin-down and precession parameters are inferred from the same observed X-ray light curve that they are used to interpret, which is ordinary model fitting rather than an equivalence-by-construction prediction, and no equation-level reduction can be exhibited from the provided excerpts. The computer-vision abstract's claims about 14 scenarios, large-scale vision models (DINO, Stable Diffusion), and robustness insights have no supporting methods, evaluation protocol, or results anywhere in the attached manuscript. That absence is a document-level integrity or packaging failure, not a circularity pattern of the kind this pass is designed to detect (self-definition, fitted input renamed as prediction, load-bearing self-citation, or ansatz smuggled via citation). Therefore the circularity score is 0, and the empty steps list reflects that no circular step could be quoted with a specific reduction. This score should not be read as endorsing the benchmark claims, which are unverifiable from the submitted text.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central claim, taken from the abstract, depends on the existence of a 14-scenario dataset and evaluation results. None of these appear in the supplied text, so the ledger is empty except for the failed document-correspondence assumption. For the unrelated GRB text, no benchmark parameters are definable.

assumptions (1)
  • ad hoc to paper The submitted full text corresponds to the semantic-correspondence benchmark described in the abstract.
    The entire evaluation depends on this document-level premise. The supplied full text is an unrelated MNRAS paper on GRB 220711B, so the premise is violated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Robust Semantic Correspondence: A Benchmark and Insights." pith.science (2026). https://pith.science/paper/A3SHJ5HD

@misc{pith2026250800272,
  author       = {Pith},
  title        = {Pith review of: Towards Robust Semantic Correspondence: A Benchmark and Insights},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A3SHJ5HD}},
  note         = {Machine review of arXiv:2508.00272}
}
read the original abstract

Semantic correspondence aims to identify semantically meaningful relationships between different images and is a fundamental challenge in computer vision. It forms the foundation for numerous tasks such as 3D reconstruction, object tracking, and image editing. With the progress of large-scale vision models, semantic correspondence has achieved remarkable performance in controlled and high-quality conditions. However, the robustness of semantic correspondence in challenging scenarios is much less investigated. In this work, we establish a novel benchmark for evaluating semantic correspondence in adverse conditions. The benchmark dataset comprises 14 distinct challenging scenarios that reflect commonly encountered imaging issues, including geometric distortion, image blurring, digital artifacts, and environmental occlusion. Through extensive evaluations, we provide several key insights into the robustness of semantic correspondence approaches: (1) All existing methods suffer from noticeable performance drops under adverse conditions; (2) Using large-scale vision models can enhance overall robustness, but fine-tuning on these models leads to a decline in relative robustness; (3) The DINO model outperforms the Stable Diffusion in relative robustness, and their fusion achieves better absolute robustness; Moreover, We evaluate common robustness enhancement strategies for semantic correspondence and find that general data augmentations are ineffective, highlighting the need for task-specific designs. These results are consistent across both our dataset and real-world benchmarks.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

57 extracted references · 45 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Berthelot, D.; Carlini, N.; Goodfellow, I.; Papernot, N.; Oliver, A.; and Raffel, C. A. 2019. Mixmatch: A holistic approach to semi-supervised learning. Advances in neural information processing systems, 32

  4. [4]

    Caron, M.; Touvron, H.; Misra, I.; J\'egou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021. Emerging Properties in Self-Supervised Vision Transformers. In Proceedings of the International Conference on Computer Vision (ICCV)

  5. [5]

    Chen, W.-T.; Fang, H.-Y.; Hsieh, C.-L.; Tsai, C.-C.; Chen, I.; Ding, J.-J.; Kuo, S.-Y.; et al. 2021. All snow removed: Single image desnowing algorithm using hierarchical dual-tree complex wavelet representation and contradict channel loss. In Proceedings of the IEEE/CVF international conference on computer vision, 4196--4205

  6. [6]

    Chen, X.; Pan, J.; Dong, J.; and Tang, J. 2025. Towards unified deep image deraining: A survey and a new benchmark. IEEE Transactions on Pattern Analysis and Machine Intelligence

  7. [7]

    Cho, S.; Hong, S.; Jeon, S.; Lee, Y.; Sohn, K.; and Kim, S. 2021. Cats: Cost aggregation transformers for visual correspondence. Advances in Neural Information Processing Systems, 34: 9011--9023

  8. [8]

    Cho, S.; Hong, S.; and Kim, S. 2022. CATs++: Boosting Cost Aggregation with Convolutions and Transformers. arXiv preprint arXiv:2202.06817

Show all 57 references
  1. [9]

    D.; Zoph, B.; Mane, D.; Vasudevan, V.; and Le, Q

    Cubuk, E. D.; Zoph, B.; Mane, D.; Vasudevan, V.; and Le, Q. V. 2019. Autoaugment: Learning augmentation strategies from data. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 113--123

  2. [10]

    D.; Zoph, B.; Shlens, J.; and Le, Q

    Cubuk, E. D.; Zoph, B.; Shlens, J.; and Le, Q. V. 2020. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 702--703

  3. [11]

    Dalal, N.; and Triggs, B. 2005. Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05), volume 1, 886--893. Ieee

  4. [12]

    F.; and Dundar, A

    Dalva, Y.; Pehlivan, H.; Alt ndi s , S. F.; and Dundar, A. 2023. Benchmarking the robustness of instance segmentation models. IEEE Transactions on Neural Networks and Learning Systems

  5. [13]

    Dong, Y.; Kang, C.; Zhang, J.; Zhu, Z.; Wang, Y.; Yang, X.; Su, H.; Wei, X.; and Zhu, J. 2023. Benchmarking robustness of 3d object detection to common corruptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1022--1032

  6. [14]

    Du, X.; Sun, Y.; Zhu, J.; and Li, Y. 2023. Dream the impossible: Outlier imagination with diffusion models. Advances in Neural Information Processing Systems, 36: 60878--60901

  7. [15]

    A.; Van Gool, L.; Williams, C

    Everingham, M.; Eslami, S. A.; Van Gool, L.; Williams, C. K.; Winn, J.; and Zisserman, A. 2015. The pascal visual object classes challenge: A retrospective. International journal of computer vision, 111(1): 98--136

  8. [16]

    Fan, Y.; Kukleva, A.; Dai, D.; and Schiele, B. 2023. Revisiting consistency regularization for semi-supervised learning. International Journal of Computer Vision, 131(3): 626--643

  9. [17]

    T.; and Ommer, B

    Fundel, F.; Schusterbauer, J.; Hu, V. T.; and Ommer, B. 2025. Distillation of Diffusion Features for Semantic Correspondence. WACV

  10. [18]

    Gao, S.; Zhou, C.; Ma, C.; Wang, X.; and Yuan, J. 2022. Aiatrack: Attention in attention for transformer visual tracking. In European conference on computer vision, 146--164. Springer

  11. [19]

    Gupta, K.; Jampani, V.; Esteves, C.; Shrivastava, A.; Makadia, A.; Snavely, N.; and Kar, A. 2023. Asic: Aligning sparse in-the-wild image collections. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 4134--4145

  12. [20]

    Ham, B.; Cho, M.; Schmid, C.; and Ponce, J. 2016. Proposal flow. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3475--3484

  13. [21]

    Ham, B.; Cho, M.; Schmid, C.; and Ponce, J. 2017. Proposal flow: Semantic correspondences from object proposals. IEEE transactions on pattern analysis and machine intelligence, 40(7): 1711--1725

  14. [22]

    S.; Ham, B.; Wong, K.-Y

    Han, K.; Rezende, R. S.; Ham, B.; Wong, K.-Y. K.; Cho, M.; Schmid, C.; and Ponce, J. 2017. Scnet: Learning semantic correspondence. In Proceedings of the IEEE international conference on computer vision, 1831--1840

  15. [23]

    Hassani, A.; Walton, S.; Li, J.; Li, S.; and Shi, H. 2023. Neighborhood attention transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 6185--6194

  16. [24]

    He, K.; Chen, X.; Xie, S.; Li, Y.; Doll \'a r, P.; and Girshick, R. 2022. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 16000--16009

  17. [25]

    Hendrycks, D.; and Dietterich, T. 2019 a . Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. Proceedings of the International Conference on Learning Representations

  18. [26]

    Hendrycks, D.; and Dietterich, T. 2019 b . Benchmarking neural network robustness to common corruptions and perturbations. arXiv preprint arXiv:1903.12261

  19. [27]

    D.; Zoph, B.; Gilmer, J.; and Lakshminarayanan, B

    Hendrycks, D.; Mu, N.; Cubuk, E. D.; Zoph, B.; Gilmer, J.; and Lakshminarayanan, B. 2020. AugMix : A Simple Data Processing Method to Improve Robustness and Uncertainty. Proceedings of the International Conference on Learning Representations (ICLR)

  20. [28]

    Jiang, W.; Trulls, E.; Hosang, J.; Tagliasacchi, A.; and Yi, K. M. 2021. Cotr: Correspondence transformer for matching across images. In Proceedings of the IEEE/CVF international conference on computer vision, 6207--6217

  21. [29]

    Kim, S.; Min, D.; Ham, B.; Jeon, S.; Lin, S.; and Sohn, K. 2017. Fcss: Fully convolutional self-similarity for dense semantic correspondence. In Proceedings of the IEEE conference on computer vision and pattern recognition, 6560--6569

  22. [30]

    Kim, S.; Min, J.; and Cho, M. 2022. Transformatcher: Match-to-match attention for semantic correspondence. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8697--8707

  23. [31]

    Kornblith, S.; Norouzi, M.; Lee, H.; and Hinton, G. 2019. Similarity of neural network representations revisited. In International conference on machine learning, 3519--3529. PMlR

  24. [32]

    Lee, J.; Kim, D.; Ponce, J.; and Ham, B. 2019. Sfnet: Learning object-aware semantic correspondence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2278--2287

  25. [33]

    Li, B.; Ren, W.; Fu, D.; Tao, D.; Feng, D.; Zeng, W.; and Wang, Z. 2018. Benchmarking single-image dehazing and beyond. IEEE Transactions on Image Processing, 28(1): 492--505

  26. [34]

    Li, S.; Wang, Z.; Juefei-Xu, F.; Guo, Q.; Li, X.; and Ma, L. 2023 a . Common corruption robustness of point cloud detectors: Benchmark and enhancement. IEEE Transactions on Multimedia

  27. [35]

    Li, X.; Lu, J.; Han, K.; and Prisacariu, V. 2023 b . SD4Match: Learning to Prompt Stable Diffusion Model for Semantic Matching. arXiv:2310.17569

  28. [36]

    Liu, Y.; Shen, Z.; Lin, Z.; Peng, S.; Bao, H.; and Zhou, X. 2019. Gift: Learning transformation-invariant dense visual descriptors via group cnns. Advances in Neural Information Processing Systems, 32

  29. [37]

    Liu, Y.; Zhu, L.; Yamada, M.; and Yang, Y. 2020. Semantic Correspondence as an Optimal Transport Problem. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4463--4472

  30. [38]

    Lowe, D. G. 2004. Distinctive image features from scale-invariant keypoints. International journal of computer vision, 60: 91--110

  31. [39]

    H.; Holynski, A.; and Darrell, T

    Luo, G.; Dunlap, L.; Park, D. H.; Holynski, A.; and Darrell, T. 2023. Diffusion hyperfeatures: Searching through time and space for semantic correspondence. Advances in Neural Information Processing Systems, 36: 47500--47510

  32. [40]

    Min, J.; and Cho, M. 2021. Convolutional Hough Matching Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2940--2950

  33. [41]

    Min, J.; Lee, J.; Ponce, J.; and Cho, M. 2019. Spair-71k: A large-scale benchmark for semantic correspondence. arXiv preprint arXiv:1908.10543

  34. [42]

    Ofri-Amar, D.; Geyer, M.; Kasten, Y.; and Dekel, T. 2023. Neural congealing: Aligning images to a joint semantic atlas. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 19403--19412

  35. [43]

    Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H. V.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; Howes, R.; Huang, P.-Y.; Xu, H.; Sharma, V.; Li, S.-W.; Galuba, W.; Rabbat, M.; Assran, M.; Ballas, N.; Synnaeve, G.; Misra, I.; Jegou, H.; Maira...

  36. [44]

    Rocco, I.; Arandjelovic, R.; and Sivic, J. 2017. Convolutional neural network architecture for geometric matching. In Proceedings of the IEEE conference on computer vision and pattern recognition, 6148--6157

  37. [45]

    Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 10684--10695

  38. [46]

    Schneider, S.; Rusak, E.; Eck, L.; Bringmann, O.; Brendel, W.; and Bethge, M. 2020. Improving robustness against common corruptions by covariate shift adaptation. Advances in neural information processing systems, 33: 11539--11551

  39. [47]

    L.; and Frahm, J.-M

    Schonberger, J. L.; and Frahm, J.-M. 2016. Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4104--4113

  40. [48]

    H.; Lee, J.; Jung, D.; Han, B.; and Cho, M

    Seo, P. H.; Lee, J.; Jung, D.; Han, B.; and Cho, M. 2018. Attentive semantic alignment with offset-aware correlation kernels. In Proceedings of the European Conference on Computer Vision (ECCV), 349--364

  41. [49]

    Shu, M.; Nie, W.; Huang, D.-A.; Yu, Z.; Goldstein, T.; Anandkumar, A.; and Xiao, C. 2022. Test-time prompt tuning for zero-shot generalization in vision-language models. Advances in Neural Information Processing Systems, 35: 14274--14289

  42. [50]

    A.; Cubuk, E

    Sohn, K.; Berthelot, D.; Carlini, N.; Zhang, Z.; Zhang, H.; Raffel, C. A.; Cubuk, E. D.; Kurakin, A.; and Li, C.-L. 2020. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. Advances in neural information processing systems, 33: 596--608

  43. [51]

    P.; and Hariharan, B

    Tang, L.; Jia, M.; Wang, Q.; Phoo, C. P.; and Hariharan, B. 2023. Emergent Correspondence from Image Diffusion. In Thirty-seventh Conference on Neural Information Processing Systems

  44. [52]

    Wang, Q.; Chang, Y.-Y.; Cai, R.; Li, Z.; Hariharan, B.; Holynski, A.; and Snavely, N. 2023. Tracking everything everywhere all at once. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 19795--19806

  45. [53]

    D.; and Fergus, R

    Zeiler, M. D.; and Fergus, R. 2014. Visualizing and understanding convolutional networks. In Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part I 13, 818--833. Springer

  46. [54]

    P.; Jampani, V.; Sun, D.; and Yang, M.-H

    Zhang, J.; Herrmann, C.; Hur, J.; Cabrera, L. P.; Jampani, V.; Sun, D.; and Yang, M.-H. 2023. A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondence . arXiv preprint arxiv:2305.15347

  47. [55]

    Zhang, J.; Herrmann, C.; Hur, J.; Chen, E.; Jampani, V.; Sun, D.; and Yang, M.-H. 2024. Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  48. [56]

    Zhao, D.; Song, Z.; Ji, Z.; Zhao, G.; Ge, W.; and Yu, Y. 2021. Multi-scale matching networks for semantic correspondence. In Proceedings of the IEEE/CVF international conference on computer vision, 3354--3364

  49. [57]

    Zhou, Y.; Barnes, C.; Shechtman, E.; and Amirghodsi, S. 2021. Transfill: Reference-guided image inpainting by merging multiple color and spatial transformations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2266--2276

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.