REVIEW 4 major objections 6 minor 54 references
Exploring Few-Shot Defect Segmentation in General Industrial Scenarios with Metric Learning and Vision Foundation Models
T0 review · 4 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This paper claims that on a new 12-product benchmark, meta-learning-based few-shot defect segmentation generally fails in cross-domain settings, while SAM2's video track mode reaches 47.9 average mIoU and a proposed feature-matching…
desk verdict A useful FDS benchmark and a plausible but under-verified SAM2 video-mode result; the benchmark deserves referee time, the headline needs more controls. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two mechanisms carry the argument. First, the benchmark itself: a newly contributed rubber-ring dataset with nine defect classes is combined with reorganized public datasets covering textures, single-component objects, and multi-component objects, and split into cross-domain and in-domain folds. Second, the method-side machinery: for feature matching, a ViT-S/8 DINO teacher trained on ImageNet distills high-resolution features into a shallow CNN student; query features are matched against patch-averaged foreground prototypes and dense background vectors by cosine similarity, and the raw mask is fused with FastSAM's zero-shot masks through morphological connected-component analysis. For the SAM2 route, the support image and mask are encoded by SAM2's memory encoder, and the resulting memory feature is combined with the query image features through memory attention before the mask decoder—repurposing a mechanism designed for video tracking to link a static support-query image pair.
What would settle it
Run SAM2's video track mode on a held-out industrial product not in the benchmark, or on support-query pairs with large changes in viewpoint, scale, lighting, or product instance; if average accuracy drops toward the meta-learning baselines rather than staying in the 40-48 mIoU range, the claimed transfer does not generalize. A second direct check is to reverse the frame order or insert synthetic motion between support and query and measure how much the memory-attention result changes.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the hard part of few-shot defect segmentation in general industrial settings is the domain gap between training and test products, and that pre-trained vision foundation models can cross that gap where meta-learning cannot. On the assembled benchmark, the four meta-learning methods (PFENet, SSP, HSNet, PATNet) saturate at low accuracy under the cross-domain protocol—PATNet, the best, reaches 19.2 mIoU—and in-domain training does not reliably fix object-based products. Existing VFM-based one-shot methods PerSAM and Matcher also perform poorly on industrial images. The clear exception is SAM2 applied through its video track mode: using the support image and mask as the memory frame and the query as the current frame, SAM2-s at 1024x1024 reaches 47.9 average mIoU across the benchmark's 54 defect categories. The paper also establishes that high-resolution features are the key to representing small industrial defects, and that distilling a DINO ViT-S/8 teacher into a shallow CNN student yields a feature extractor that is faster and better than its teacher, giving 36.5 mIoU at 277.8 FPS without refinement and 39.3 mIoU with FastSAM refinement.
Load-bearing premise
The paper's headline result assumes that SAM2's memory attention, which was trained on videos in which objects move continuously, will still work when the 'previous frame' and 'current frame' are two unrelated still images of an industrial product, and the paper only tests this transfer on its own benchmark.
Editorial extensions
If this is right
- With one labeled defect image per category, SAM2's video track mode can segment defects across a wide range of industrial products without any model training.
- A factory can deploy the feature-matching method on a production line because 277.8 FPS makes it fast enough for real-time inspection, and its 36.5 mIoU beats every trained meta-learning method in the study.
- Meta-learning FSS methods would need in-domain training data to be useful, which defeats their purpose where defects are scarce; the paper's data-usage experiments make this limitation explicit.
- High-resolution features, not just large pretrained models, drive FDS performance, so future industrial feature extractors should be designed with resolution in mind.
- The SAM2 result suggests that video segmentation models with memory mechanisms are a natural fit for few-shot defect segmentation, even when the input images are unrelated stills.
Reading between the lines
- Treating few-shot segmentation as video tracking may be the general recipe here: any future video segmentation model with a memory mechanism could inherit SAM2's FDS ability, a consequence the paper only gestures at.
- The benchmark's 1-way 1-shot protocol and deliberately consistent category granularity are choices; under coarser or finer defect taxonomies, or with 5-shot support sets, the ranking of methods could shift.
- The distilled high-resolution feature extractor could be tested outside defects, for example in medical or satellite small-object segmentation, where resolution is also the limiting factor.
- If SAM2's memory attention truly works across large viewpoint and scale changes, it would also enable few-shot segmentation from only a single reference image captured under very different imaging conditions, which is an industrial scenario the paper does not directly test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses few-shot defect segmentation (FDS) in general industrial scenarios. The authors contribute a new real-world rubber-ring defect dataset, reorganize several existing industrial datasets into a 12-product benchmark spanning textures, single-component objects, and multi-component objects, and evaluate metric-learning-based approaches from two paradigms: classical meta-learning FSS methods (PFENet, SSP, HSNet, PATNet) and vision foundation model (VFM) based methods (PerSAM, Matcher, SAM2, and a proposed feature-matching method). The paper reports that meta-learning methods generally fail under the proposed cross-domain setting, that SAM2's video track mode yields the best segmentation accuracy (47.9 mIoU with SAM2-s at 1024 resolution), and that the proposed feature-matching method with knowledge distillation and FastSAM fusion attains 39.3 mIoU at 90.9 FPS while remaining more efficient than SAM2. The central claims are that existing FDS benchmarks are too texture-centric, that VFMs, especially SAM2, are promising for FDS, and that high-resolution features are important for small industrial defects.
Significance. If the conclusions hold, the paper makes a useful empirical contribution: the new rubber-ring dataset broadens the scope of FDS beyond textures, the systematic comparison of meta-learning and VFM methods on a multi-product benchmark is practically relevant, and the proposed efficient feature-matching method offers a plausible speed-accuracy trade-off for production-line deployment. The authors also deserve credit for using official implementations of baselines, promising release of code and data, and explicitly questioning whether meta-learning assumptions transfer to industrial domains. However, the main comparative conclusions currently rest on single-run evaluations without statistical support, and the headline SAM2 video-track result lacks the ablations needed to show that the memory mechanism, rather than the image encoder and mask decoder, is responsible for the observed performance.
major comments (4)
- [Table 4; Sec. 6.2] Every quantitative result in Table 4 is a single mIoU/FB-IoU value with no variance estimates, number of episodes, or random seeds. Because the meta-learning methods are trained from scratch (Sec. 6.1.1) and several benchmark categories are very small (e.g., glue: 11 images; rough: 15 images in Table 1), the observed ordering — for example PATNet at 19.2 vs. SSP at 18.8 average mIoU in the cross-domain setting, or SAM2-s at 47.9 vs. SAM2-t at 45.1 — could easily be within run-to-run noise. Please report mean and standard deviation over at least three seeds or episode draws, or otherwise justify that the comparative conclusions are statistically meaningful.
- [Sec. 5.2, Eqs. (15)-(16); Table 4] The claim that SAM2's video track mode is 'particularly effective' for FDS is underdetermined. The experiments use a fixed protocol in which the support image is always treated as the first frame and the query as the second frame, under matched imaging conditions, and no ablation isolates the contribution of memory attention. The reported 47.9 mIoU could in principle be achieved by the image encoder and mask decoder alone, with the memory mechanism adding little. Please add (i) an image-mode SAM2 baseline with the support mask provided as a prompt, (ii) a memory-ablated variant of the video track, (iii) reversed frame order, and (iv) support/query pairs under visible domain shift (e.g., different illumination, resolution, or a different product) to support the generalization claim.
- [Sec. 4; Table 2; Table 1] The evaluation protocol for the 1-way 1-shot setting is underspecified. The paper does not state how many query images are paired with each support image, whether the support/query split is randomized or fixed for each defect category, or how categories with very few images (e.g., 11-16 images in Table 1) are partitioned without leakage. The fold division in Table 2 also requires clarification: the 'Quantity' column mixes category counts and image counts, and it is unclear which products are used for meta-training versus testing in each fold. This detail is necessary for reproducibility and for interpreting the 54-category average mIoU in Table 4.
- [Sec. 5.1.3, Eqs. (11) and (13); Sec. 6.1.2] The proposed fusion method depends on several hyperparameters — τ1=0.2, τ2=0.9, dilation kernel size 21, FastSAM confidence and IoU thresholds, foreground patch size 3, and student width c — and the paper does not describe how these were selected. If they were tuned on the same benchmark used for the final evaluation, the reported 39.3 mIoU / 68.2 FB-IoU may overstate performance on new industrial products. Please state the selection procedure or provide a sensitivity analysis demonstrating that the conclusions are stable across reasonable hyperparameter ranges.
minor comments (6)
- [Sec. 5.1.2, after Eq. (8)] The text says 'We concatenate F b∗ q and F b∗ q', but the second term should be F f∗ q; this typo makes the sentence nonsensical.
- [Sec. 5.1.3, Eq. (10)] The notation 'I ∈ R^{H×W} refers to the identity matrix' is incorrect in this context; the condition should use an all-ones array or a partition of the image domain, not the identity matrix.
- [Tables 2 and 4; Sec. 3.2] The abbreviation 'Kol' appears in Tables 2 and 4 without being defined in the captions; Sec. 3.2 introduces the dataset as KolektorSDD. Please define the abbreviation where it first appears in the tables.
- [Sec. 6.3] The FPS measurement excludes the processing time of the support image. For SAM2's video track mode, support-image processing includes memory encoding (Eq. (15)), so excluding it may understate the true cost of the method. Please state explicitly for each method which components are included in the FPS measurement.
- [Table 1] The 'Category(quantity)' row is difficult to parse because multiple categories are listed without clear separators; a per-row breakdown or a structured table would improve readability and make the dataset composition unambiguous.
- [Sec. 7.1.1, Table 5] The table reports FPS for feature extractors, but it is unclear whether these speeds use the same hardware and precision settings as the main experiments in Sec. 6.3; please clarify to ensure fair comparison with the proposed method's FPS.
Circularity Check
No significant circularity: the proposed feature-matching method and the SAM2 video-mode evaluation are empirically measured on a held-out benchmark, with no load-bearing step reducing to fitted inputs or self-citations.
full rationale
The paper's central derivation chain is self-contained. The proposed feature-matching method distills a student feature extractor from DINO/ImageNet and then matches support and query features using cosine similarity, patch-averaged foreground prototypes, and a FastSAM-based fusion; the equations in Sec. 5.1 define a segmentation procedure rather than encoding the benchmark outcomes. The SAM2 video-mode result is obtained by treating the support image as a previous frame and the query image as a current frame (Eqs. 15-16) using the off-the-shelf SAM2 model, so the reported mIoU is an empirical measurement of zero-shot transfer, not a quantity fitted to the benchmark. No load-bearing premise is justified by a self-citation: citations to DINO, DINOv2, SAM2, and FastSAM refer to external, independently trained models, and the paper does not invoke a uniqueness theorem or prior work by the same authors to force its choices. Hyperparameters such as tau1=0.2 and tau2=0.9 are fixed settings of the proposed fusion algorithm and are not presented as predictions; even if they were tuned on the same benchmark, that would be a generalization or overfitting concern, not circularity. No step reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (6)
- foreground patch size for PatchAvg =
3 with stride 1
- FastSAM mask selection threshold tau1 =
0.2
- dilation coverage threshold tau2 =
0.9
- dilation kernel size for FastSAM masks =
21
- FastSAM iou and confidence thresholds =
iou=0.5, confidence=0.1
- student model base channel width c =
256 for FM-l, 128 for FM-s
assumptions (5)
- domain assumption Support and query images for a test task come from the same product and same defect category
- domain assumption Selected existing datasets have defect category definitions consistent enough for FDS evaluation
- ad hoc to paper Treating an independent support image as the previous frame of a video makes SAM2's memory attention applicable to FDS
- domain assumption Knowledge distillation on ImageNet yields features that transfer to industrial defect segmentation
- ad hoc to paper The two-fold cross-domain split is representative of industrial domain shift
Cite this review
Pith. "Pith review of Exploring Few-Shot Defect Segmentation in General Industrial Scenarios with Metric Learning and Vision Foundation Models." pith.science (2026). https://pith.science/paper/BRLIO34T
@misc{pith2026250201216,
author = {Pith},
title = {Pith review of: Exploring Few-Shot Defect Segmentation in General Industrial Scenarios with Metric Learning and Vision Foundation Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/BRLIO34T}},
note = {Machine review of arXiv:2502.01216}
}
read the original abstract
Industrial defect segmentation is critical for manufacturing quality control. Due to the scarcity of training defect samples, few-shot semantic segmentation (FSS) holds significant value in this field. However, existing studies mostly apply FSS to tackle defects on simple textures, without considering more diverse scenarios. This paper aims to address this gap by exploring FSS in broader industrial products with various defect types. To this end, we contribute a new real-world dataset and reorganize some existing datasets to build a more comprehensive few-shot defect segmentation (FDS) benchmark. On this benchmark, we thoroughly investigate metric learning-based FSS methods, including those based on meta-learning and those based on Vision Foundation Models (VFMs). We observe that existing meta-learning-based methods are generally not well-suited for this task, while VFMs hold great potential. We further systematically study the applicability of various VFMs in this task, involving two paradigms: feature matching and the use of Segment Anything (SAM) models. We propose a novel efficient FDS method based on feature matching. Meanwhile, we find that SAM2 is particularly effective for addressing FDS through its video track mode. The contributed dataset and code will be available at: https://github.com/liutongkun/GFDS.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Y. Bao, K. Song, J. Liu, Y. Wang, Y. Yan, H. Yu, X. Li, Triplet- graph reasoning network for few-shot metal generic surface de- fect segmentation, IEEE Transactions on Instrumentation and Measurement 70 (2021) 1–11
work page 2021
- [2]
-
[3]
W. Ren, Y. Tang, Q. Sun, C. Zhao, Q.-L. Han, Visual semantic segmentation based on few/zero-shot learning: An overview, IEEE/CAA Journal of Automatica Sinica (2023)
work page 2023
-
[4]
Z. Tian, H. Zhao, M. Shu, Z. Yang, R. Li, J. Jia, Prior guided feature enrichment network for few-shot segmentation, IEEE transactions on pattern analysis and machine intelligence 44 (2) (2020) 1050–1065
work page 2020
-
[5]
J. Min, D. Kang, M. Cho, Hypercorrelation squeeze for few-shot segmentation, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 6941–6952
work page 2021
-
[6]
S. Lei, X. Zhang, J. He, F. Chen, B. Du, C.-T. Lu, Cross-domain few-shot semantic segmentation, in: European Conference on Computer Vision, Springer, 2022, pp. 73–90
work page 2022
- [7]
-
[8]
Y. Liu, M. Zhu, H. Li, H. Chen, X. Wang, C. Shen, Matcher: Segment anything with one shot using all-purpose feature matching, arXiv preprint arXiv:2305.13310 (2023)
arXiv 2023
Show all 54 references
-
[9]
M. Kaya, H. S ¸. Bilge, Deep metric learning: A survey, Symme- try 11 (9) (2019) 1066
2019
-
[10]
Hospedales, A
T. Hospedales, A. Antoniou, P. Micaelli, A. Storkey, Meta- learning in neural networks: A survey, IEEE transactions on pattern analysis and machine intelligence 44 (9) (2021) 5149– 5169
2021
-
[11]
G.-S. Xie, H. Xiong, J. Liu, Y. Yao, L. Shao, Few-shot semantic segmentation with cyclic memory network, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 7293–7302
2021
-
[12]
Shaban, S
A. Shaban, S. Bansal, Z. Liu, I. Essa, B. Boots, One-shot learn- ing for semantic segmentation, arXiv preprint arXiv:1709.03410 (2017)
2017 arXiv
-
[13]
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ra- manan, P. Doll´ ar, C. L. Zitnick, Microsoft coco: Common ob- jects in context, in: Computer Vision–ECCV 2014: 13th Eu- ropean Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer, ...
2014
-
[14]
H. Yao, W. Luo, W. Yu, X. Zhang, Z. Qiang, D. Luo, H. Shi, Few-shot unseen defect segmentation for polycrystalline sili- con panels with an interpretable dual subspace attention vari- ational learning framework, Advanced Engineering Informatics 62 (2024) 102613
2024
-
[15]
R. Yu, B. Guo, K. Yang, Selective prototype network for few- shot metal surface defect segmentation, IEEE Transactions on Instrumentation and Measurement 71 (2022) 1–10
2022
-
[16]
J. Zhu, Y. Qi, J. Wu, Medical sam 2: Segment medical im- ages as video via segment anything model 2, arXiv preprint arXiv:2408.00874 (2024)
2024 arXiv
-
[17]
L. Zhao, X. Chen, E. Z. Chen, Y. Liu, T. Chen, S. Sun, Retrieval-augmented few-shot medical image segmentation with foundation models, arXiv preprint arXiv:2408.08813 (2024)
2024 arXiv
-
[18]
Y. Bai, Q. Yu, B. Yun, D. Jin, Y. Xia, Y. Wang, Fs-medsam2: Exploring the potential of sam2 for few-shot medical image seg- mentation without fine-tuning, arXiv preprint arXiv:2409.04298 (2024)
2024 arXiv
-
[19]
Bergmann, M
P. Bergmann, M. Fauser, D. Sattlegger, C. Steger, Mvtec ad–a comprehensive real-world dataset for unsupervised anomaly de- tection, in: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition, 2019, pp. 9592–9600
2019
-
[20]
Y. Zou, J. Jeong, L. Pemula, D. Zhang, O. Dabeer, Spot-the- difference self-supervised pre-training for anomaly detection and segmentation, in: European Conference on Computer Vision, Springer, 2022, pp. 392–408
2022
-
[21]
Jeong, Y
J. Jeong, Y. Zou, T. Kim, D. Zhang, A. Ravichandran, O. Dabeer, Winclip: Zero-/few-shot anomaly classification and segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 19606– 19616
2023
-
[22]
Y. Cao, X. Xu, C. Sun, X. Huang, W. Shen, Towards generic anomaly detection and understanding: Large-scale visual-linguistic model (gpt-4v) takes the lead, arXiv preprint arXiv:2311.02782 (2023)
2023 arXiv
-
[23]
Kirillov, E
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al., Segment anything, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 4015– 4026
2023
-
[24]
X. Zhao, W. Ding, Y. An, Y. Du, T. Yu, M. Li, M. Tang, J. Wang, Fast segment anything, arXiv preprint arXiv:2306.12156 (2023)
2023 arXiv
-
[25]
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R¨ adle, C. Rolland, L. Gustafson, et al., Sam 2: Segment anything in images and videos, arXiv preprint arXiv:2408.00714 (2024)
2024 arXiv
-
[26]
J. Liu, G. Xie, J. Wang, S. Li, C. Wang, F. Zheng, Y. Jin, Deep industrial image anomaly detection: A survey, Machine Intelligence Research 21 (1) (2024) 104–135
2024
-
[27]
Bergmann, X
P. Bergmann, X. Jin, D. Sattlegger, C. Steger, The mvtec 3d-ad dataset for unsupervised 3d anomaly detection and localization, arXiv preprint arXiv:2112.09045 (2021)
2021 arXiv
-
[28]
H. Yao, Y. Cao, W. Luo, W. Zhang, W. Yu, W. Shen, Prior normality prompt transformer for multiclass industrial image anomaly detection, IEEE Transactions on Industrial Infor- matics 20 (10) (2024) 11866–11876. doi:10.1109/TII.2024. 12 3413322
2024 doi
-
[29]
Z. You, L. Cui, Y. Shen, K. Yang, X. Lu, Y. Zheng, X. Le, A unified model for multi-class anomaly detection, Advances in Neural Information Processing Systems 35 (2022) 4571–4584
2022
-
[30]
K. Roth, L. Pemula, J. Zepeda, B. Sch¨ olkopf, T. Brox, P. Gehler, Towards total recall in industrial anomaly detection, in: Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, 2022, pp. 14318–14328
2022
-
[31]
S. Li, J. Cao, P. Ye, Y. Ding, C. Tu, T. Chen, Clipsam: Clip and sam collaboration for zero-shot anomaly segmentation, arXiv preprint arXiv:2401.12665 (2024)
2024 arXiv
-
[32]
Nakamura, T
A. Nakamura, T. Harada, Revisiting fine-tuning for few-shot learning, arXiv preprint arXiv:1910.00216 (2019)
2019 arXiv
-
[33]
N. Dong, E. P. Xing, Few-shot semantic segmentation with pro- totype learning., in: BMVC, Vol. 3, 2018, p. 4
2018
-
[34]
Q. Fan, W. Pei, Y.-W. Tai, C.-K. Tang, Self-support few-shot semantic segmentation, in: European Conference on Computer Vision, Springer, 2022, pp. 701–719
2022
-
[35]
G. Li, V. Jampani, L. Sevilla-Lara, D. Sun, J. Kim, J. Kim, Adaptive prototype learning and allocation for few-shot seg- mentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 8334–8343
2021
-
[36]
Y. Guo, N. C. Codella, L. Karlinsky, J. V. Codella, J. R. Smith, K. Saenko, T. Rosing, R. Feris, A broader study of cross-domain few-shot learning, in: Computer Vision–ECCV 2020: 16th Eu- ropean Conference, Glasgow, UK, August 23–28, 2020, Proceed- ings, Part XXVII 16, Springe...
2020
-
[37]
X. Li, T. Wei, Y. P. Chen, Y.-W. Tai, C.-K. Tang, Fss-1000: A 1000-class dataset for few-shot segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2869–2878
2020
-
[38]
Candemir, S
S. Candemir, S. Jaeger, K. Palaniappan, J. P. Musco, R. K. Singh, Z. Xue, A. Karargyris, S. Antani, G. Thoma, C. J. Mc- Donald, Lung segmentation in chest radiographs using anatom- ical atlases with nonrigid registration, IEEE transactions on medical imaging 33 (2) (2013) 577–590
2013
-
[39]
Jaeger, A
S. Jaeger, A. Karargyris, S. Candemir, L. Folio, J. Siegelman, F. Callaghan, Z. Xue, K. Palaniappan, R. K. Singh, S. Antani, et al., Automatic tuberculosis screening using chest radiographs, IEEE transactions on medical imaging 33 (2) (2013) 233–245
2013
-
[40]
Demir, K
I. Demir, K. Koperski, D. Lindenbaum, G. Pang, J. Huang, S. Basu, F. Hughes, D. Tuia, R. Raskar, Deepglobe 2018: A challenge to parse the earth through satellite images, in: Pro- ceedings of the IEEE conference on computer vision and pattern recognition workshops, 2018, pp. 172–181
2018
-
[41]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Ima- genet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee, 2009, pp. 248–255
2009
-
[42]
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778
2016
-
[43]
Alexey, An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv: 2010.11929 (2020)
D. Alexey, An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv: 2010.11929 (2020)
2020 arXiv
-
[44]
Caron, H
M. Caron, H. Touvron, I. Misra, H. J´ egou, J. Mairal, P. Bo- janowski, A. Joulin, Emerging properties in self-supervised vi- sion transformers, in: Proceedings of the IEEE/CVF interna- tional conference on computer vision, 2021, pp. 9650–9660
2021
-
[45]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al., Dinov2: Learning robust visual features without super- vision, arXiv preprint arXiv:2304.07193 (2023)
2023 arXiv
-
[46]
Zhang, D
C. Zhang, D. Han, Y. Qiao, J. U. Kim, S.-H. Bae, S. Lee, C. S. Hong, Faster segment anything: Towards lightweight sam for mobile applications, arXiv preprint arXiv:2306.14289 (2023)
2023 arXiv
-
[47]
Xiong, B
Y. Xiong, B. Varadarajan, L. Wu, X. Xiang, F. Xiao, C. Zhu, X. Dai, D. Wang, F. Sun, F. Iandola, et al., Efficientsam: Lever- aged masked image pretraining for efficient segment anything, in: Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, ...
2024
-
[48]
Songa, B
Y. Songa, B. Pua, P. Wanga, H. Jiang, D. Donga, Y. Shen, Sam- lightening: A lightweight segment anything model with dilated flash attention to achieve 30 times acceleration, arXiv preprint arXiv:2403.09195 (2024)
2024 arXiv
-
[49]
Tabernik, S
D. Tabernik, S. ˇSela, J. Skvarˇ c, D. Skoˇ caj, Segmentation-based deep-learning approach for surface-defect detection, Journal of Intelligent Manufacturing 31 (3) (2020) 759–776
2020
-
[50]
Wieler, T
M. Wieler, T. Hahn, Weakly supervised learning for industrial optical inspection, in: DAGM symposium in, Vol. 6, 2007, p. 11
2007
-
[51]
Zhang, R
J. Zhang, R. Ding, M. Ban, T. Guo, Fdsnet: An accurate real- time surface defect segmentation network, in: ICASSP 2022- 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2022, pp. 3803–3807
2022
-
[52]
Jocher, A
G. Jocher, A. Chaurasia, J. Qiu, Yolo by ultralytics (2023)
2023
-
[53]
Batzner, L
K. Batzner, L. Heckler, R. K¨ onig, Efficientad: Accurate visual anomaly detection at millisecond-level latencies, in: Proceed- ings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV), 2024, pp. 128–138
2024
-
[54]
Cohen, Y
N. Cohen, Y. Hoshen, Sub-image anomaly detection with deep pyramid correspondences, arXiv preprint arXiv:2005.02357 (2020). 13
2020 arXiv
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.