Pith. sign in

REVIEW 4 major objections 5 minor 56 references

DefFiller: Mask-Conditioned Diffusion for Salient Steel Surface Defect Generation

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A diffusion model fine-tuned on 900 steel defect pairs generates mask-matched defects and improves saliency detection in low-data regimes.

desk verdict A clean GLIGEN adaptation for defect generation, but its central data-expansion result is undermined by a training/test leakage that needs a rerun before the claims can be trusted. read the letter →

arxiv 2412.15570 v1 pith:BCW6C2W4 submitted 2024-12-20 cs.CV

classification cs.CV
keywords mask-conditioneddefectgenerationsteelsurfacesaliency-baseddetectiondiffusionmodeldataaugmentationlayout-to-imageSD-Saliency-900FIDevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a diffusion model pre-trained on natural images can be adapted, with a small fine-tuned mask encoder, to generate realistic steel-surface defect images that obey a given binary mask. If true, this gives defect-detection researchers a data augmentation method that needs only mask conditions, not pixel-level manual annotation, and that improves saliency-based detectors in low-data regimes. The paper reports that on SD-Saliency-900, generated images achieve an average FID of 54.40 versus 106.22 for AdaBLDM, and adding 900 generated pairs raises S-measure from 0.822 to 0.834 for CSEPNet, 0.793 to 0.845 for TSERNet, and 0.787 to 0.833 for MINet. The method is framed as the first mask-conditioned defect generation built on a layout-to-image diffusion prior.

What carries the argument

The load-bearing mechanism is a layout-conditioned latent diffusion model: a pre-trained GLIGEN is fine-tuned with a trainable mask encoder that converts a binary mask into 64 layout tokens, and a gated self-attention layer that injects these tokens into the U-Net; the mask is also downsampled and concatenated to the noisy latent at the U-Net input. Training minimizes the standard noise-prediction mean squared error while freezing the autoencoder and text encoder, and inference uses classifier-free guidance with the guidance scale set to 3. This design lets the model keep the natural-image diffusion prior while adding pixel-level control.

What would settle it

A direct test would be to fine-tune DefFiller on a different steel defect dataset with equally scarce data and check whether generated images both match masks and improve a detector trained on augmented data; if FID fails to drop below the GAN baseline or detection S-measure does not improve, the transfer claim is refuted. A more controlled test is to train DefFiller from random initialization on the same 900 pairs: if the resulting FID is comparable to the fine-tuned version, the natural-image prior is not load-bearing, whereas if FID collapses, the transfer assumption is supported.

Watch

Extended reading notes

Core claim

The paper's central claim is that DefFiller, a fine-tuned layout-to-image diffusion model with an added mask encoder, can synthesize steel defect images whose defective regions align with user-supplied masks, and that these synthetic mask-image pairs are faithful enough to substitute for real training data and to improve saliency-based defect detectors when the real training set is small. The evidence is FID comparisons (54.40 average versus 106.22 for AdaBLDM on original masks, and 87.68 versus 297.21 on newly generated masks) and detector experiments showing that replacing the real training set with DefFiller images keeps performance close to the original, while expanding a 1:5 training split with 900 DefFiller pairs improves all three tested detectors.

Load-bearing premise

The load-bearing premise is that the pre-trained diffusion prior, learned mostly on natural images, can be transferred to steel surface defect textures by fine-tuning on only 900 mask-image pairs for 30,000 iterations.

Editorial extensions

If this is right

  • Mask-conditioned synthetic pairs can substitute for real training data with minimal performance loss; CSEPNet's S-measure stays at 0.860 against 0.887 on real data.
  • Expanding a scarce training set with DefFiller pairs improves saliency detection, with the average S-measure across the three tested detectors rising by roughly 4 percent.
  • The method requires only masks rather than pixel-level defect annotations, so annotators can draw coarse shapes instead of labeling every defective pixel.
  • DefFiller handles multiple defect classes in one model and generates 300 images per category per run, in contrast to AdaBLDM which generates one defect type per training session.
  • The evaluation framework ties generation quality, measured by FID, to downstream detection utility, giving a template for judging other mask-conditioned generative models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the transfer assumption holds beyond SD-Saliency-900, the same recipe could be applied to other industrial surfaces, such as fabrics or semiconductors, using only mask-image pairs and thereby reducing annotation cost.
  • The reported S-measure gains might partly reflect added training quantity rather than the fidelity of generated textures; a controlled experiment that adds real mask-image pairs or random image crops could disentangle quantity from quality.
  • Because the new masks are produced by a separate DDPM, the pipeline's upper bound depends on mask realism, so improving the mask producer could yield further detection gains.
  • FID uses features trained on natural images, which may not capture steel texture realism; texture-specific perceptual metrics could produce a different ranking of generation methods.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes DefFiller, a mask-conditioned diffusion method for generating steel surface defect images paired with masks. The method fine-tunes a pre-trained GLIGEN model with an added mask encoder and a downsampling network, using a training objective that combines noise prediction with classifier-free guidance. The authors evaluate generation quality with FID and assess downstream utility by training three saliency-based defect detectors (CSEPNet, TSERNet, MINet) on datasets augmented with generated mask-image pairs. They report that DefFiller achieves lower FID than AdaBLDM and improves detection performance after data expansion, e.g., TSERNet S-measure from 0.793 to 0.845 in a 1:5 split experiment.

Significance. If the results hold, DefFiller would be a practical data-augmentation tool for low-data industrial inspection, contributing a mask-conditioned generation method that leverages a natural-image diffusion prior for a specialized defect domain. The paper includes useful components: an ablation of the fine-tuning strategy, a comparison with AdaBLDM, and an evaluation framework that combines FID with downstream detection performance. The code is made available, which supports reproducibility. However, the central data-expansion claim is currently undermined by a train/test leakage problem, and the FID numbers are tuned on the same evaluation data used to report them. These issues need to be resolved before the claimed improvements can be trusted.

major comments (4)
  1. [§4.4.3 (with §4.1.2, §4.4.1)] The central data-expansion claim is undermined by a train/test leakage. Section 4.4.3 splits SD-Saliency-900 into a 1:5 train/test split for detector evaluation, but Section 4.1.2 states that DefFiller is fine-tuned on mask-image pairs from the full SD-Saliency-900 dataset, and Section 4.4.1 states that the DDPM mask producer is trained on the ground truth from the dataset (all 900 masks). Both generative models therefore see the test masks and the image content of the test pairs before the synthetic pairs are added to the 150-sample training set. The improvements in Table 8 (e.g., TSERNet Sα from 0.793 to 0.845, CSEPNet from 0.822 to 0.834) may then reflect information about the test set encoded in the synthetic pairs rather than DefFiller's augmentation quality. Please rerun the data-expansion experiment with DefFiller and the mask DDPM trained only on the 1:5 training split; the same ambiguity affects Section 4.3, where the generator is trained on the full dataset before the 9:1 substitution split.
  2. [§3.2.1, Table 2] The FID comparison in Tables 3 and 7 is not a fair out-of-sample estimate. Section 3.2.1 states that the guidance scale ω_cfg is iteratively adjusted for each defect category to achieve lower FID, and Table 2 reports FID values across the candidate scales; the final ω_cfg=3 is selected on the same data used to report the headline FID scores. This is selection on the evaluation metric, which inflates the apparent generation quality. Please use a held-out split for guidance-scale selection (or report all candidate FID values and the selection rule) before claiming that DefFiller achieves lower FID than AdaBLDM.
  3. [Abstract and Section 5] The abstract claims that DefFiller 'eliminates the need for pixel-level annotations,' but the method is trained on mask-image pairs from SD-Saliency-900 (Section 4.1.2) and the mask producer is trained on the ground-truth masks (Section 4.4.1). The method therefore requires a seed set of pixel-level annotations for training both components; it can produce new paired masks at inference time, but it does not eliminate the need for pixel-level annotations in the development pipeline. Please rephrase the contribution (e.g., 'generates paired masks without additional manual annotation after training') or justify the stronger claim.
  4. [§4.4.3, Table 8] The reported detection gains are based on single training runs for each network, with no error bars or significance tests. The gains are modest in some cases (CSEPNet Sα from 0.822 to 0.834; MINet from 0.787 to 0.833), and without repeated seeds it is unclear whether the differences are within training noise. Please report means and standard deviations over at least three seeds, and include standard classical augmentation baselines (e.g., random flips/crops/color jitter, Copy-Paste) at the same added-sample count to contextualize the benefit.
minor comments (5)
  1. [Section 3.2] The phrase 'a evaluation framework' should be 'an evaluation framework'.
  2. [Section 4.1.2 and Fig. 1] The word 'downsmpling' appears to be a typo for 'downsampling', and Table 1 header 'Adapation' should be 'Adaptation'.
  3. [Section 4.4.3] The claim of 'approximately 4%' average S-measure improvement should be stated more precisely: the table shows absolute improvements of 0.012, 0.052, and 0.046, and the percentage depends on the chosen baseline; please clarify the calculation.
  4. [Table 3] The DFMGAN rows are empty because it does not accept a mask condition; please add a footnote in the table itself for readability.
  5. [Figure 7] The text refers to 'blue bars', but the figure appears to be grayscale in the preprint; please ensure that color labels are legible in the final version.

Circularity Check

2 steps flagged · score 4.0 of 10

FID quality numbers are tuned on the evaluation metric; detection gains may inherit test-set leakage from whole-dataset generator training.

  1. fitted input called prediction [Section 3.2.1 (Generation quality), Eq. (6) and Table 2]
    "To enhance the quality of the generated defect images, we iteratively adjust the guidance scale ωcf gfor each defect category to achieve lower FID scores, as detailed in Table 2."

    The FID in Eq. (6) is computed between generated images and the real SD-Saliency-900 images. The guidance scale is then selected by scanning values (1, 3, 5, 7) and retaining the setting that gives the lowest FID on the same evaluation data. Consequently, the FID values reported in Tables 3 and 7 are not independent measurements of generation quality; they are the fitted optimum of a hyperparameter search over the evaluation statistic itself. The comparison against AdaBLDM therefore partly compares a tuned configuration against an untuned one, so the headline quality advantage (e.g., 54.40 average FID vs 106.22 for AdaBLDM) is inflated by construction.

  2. other [Sections 4.1.2, 4.4.1, and 4.4.3]
    "fine-tunes the parameters in the gated self-attention layers, the mask encoder and the downsmpling network using the Adam optimizer [53] with mask-image pairs from the SD-Saliency-900 dataset. ... We train a DDPM from scratch for 200 epochs using the ground truth from the SD-Saliency-900 dataset. ... we first split the SD-Saliency-900 dataset into a training set and a testing set with a 1:5 ratio."

    The implementation text does not restrict generator fine-tuning or mask-DDPM training to the 1:5 training split introduced in Section 4.4.3. As written, both the image generator and the mask producer are trained on the whole SD-Saliency-900 dataset, which includes the images and masks later used as the testing set. When synthetic mask-image pairs generated by these models are added to the 150-pair training set, the detector can indirectly learn statistics of the test split, so the Table 8 improvements (e.g., TSERNet Sα 0.793→0.845, CSEPNet Sα 0.822→0.834) cannot be attributed cleanly to DefFiller's augmentation quality. This is an evaluation-leakage circularity rather than a definitional one, but it affects the paper's central data-expansion claim.

full rationale

Two concrete issues prevent a clean non-circular reading. First, the generation-quality evidence is partially self-fulfilling: the guidance scale is explicitly tuned to minimize FID on the same real-image distribution used to report FID, so the reported FID advantage is a fitted statistic rather than an independent prediction. Second, the data-expansion evaluation in Section 4.4.3 splits SD-Saliency-900 1:5 into train and test, but the generator fine-tuning (4.1.2) and the mask DDPM training (4.4.1) are described as using the whole SD-Saliency-900 dataset, with no stated exclusion of the test split; synthetic pairs added to the small training set can therefore leak test-set appearance into training. The detector-improvement experiment is still meaningfully independent in form—actual CSEPNet/TSERNet/MINet models are retrained on augmented data—so the central claim is not definitionally forced. There is no load-bearing self-citation: reference [17] is prior related work by the same first author but is not used to justify DefFiller's design or evaluation. Overall score 4 reflects partial circularity in the quality metric and a serious protocol leak for the augmentation claim, without reducing the entire derivation to its inputs.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

No new physical or conceptual entities are introduced; the mask encoder, gated self-attention layers, and downsampling network are architectural components inherited from GLIGEN. The ledger is dominated by domain assumptions about transfer, mask quality, and metric reliability.

free parameters (1)
  • guidance scale omega_cfg = 3 overall; per-category candidates 1, 3, 5, 7 tested
    Selected per defect category to minimize FID in Table 2, then fixed to 3 for final experiments. This tunes the reported quality metric, so the final FID values are not independent measurements.
assumptions (4)
  • domain assumption Pre-trained GLIGEN/Stable Diffusion prior transfers to steel surface defect images with only 900 mask-image pairs and 30,000 fine-tuning iterations.
    Invoked in Section 3.1 and supported only by the training-strategy ablation in Section 4.2.1; all generation quality and detection results depend on this transfer.
  • domain assumption Ground-truth masks in SD-Saliency-900 are accurate and sufficient as training conditions and as test conditions for controlled generation.
    Used as mask-image pairs in Section 4.1.1 and as conditions in the data substitution experiments of Section 4.3.
  • domain assumption FID computed on 300 generated images per category, without error bars, is a reliable metric for model selection and comparison.
    The guidance scale is tuned against per-category FID in Section 4.2.2, and FID is the primary quality evidence in Tables 3 and 7.
  • ad hoc to paper Classifier-free guidance with a text prompt improves or at least does not harm mask-conditioned defect generation.
    Equation (5) includes a text prompt c, but defects are defined primarily by the mask; the marginal benefit of the text condition is not ablated in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DefFiller: Mask-Conditioned Diffusion for Salient Steel Surface Defect Generation." pith.science (2026). https://pith.science/paper/BCW6C2W4

@misc{pith2026241215570,
  author       = {Pith},
  title        = {Pith review of: DefFiller: Mask-Conditioned Diffusion for Salient Steel Surface Defect Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BCW6C2W4}},
  note         = {Machine review of arXiv:2412.15570}
}
read the original abstract

Current saliency-based defect detection methods show promise in industrial settings, but the unpredictability of defects in steel production environments complicates dataset creation, hampering model performance. Existing data augmentation approaches using generative models often require pixel-level annotations, which are time-consuming and resource-intensive. To address this, we introduce DefFiller, a mask-conditioned defect generation method that leverages a layout-to-image diffusion model. DefFiller generates defect samples paired with mask conditions, eliminating the need for pixel-level annotations and enabling direct use in model training. We also develop an evaluation framework to assess the quality of generated samples and their impact on detection performance. Experimental results on the SD-Saliency-900 dataset demonstrate that DefFiller produces high-quality defect images that accurately match the provided mask conditions, significantly enhancing the performance of saliency-based defect detection models trained on the augmented dataset.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 38 canonical work pages

  1. [1]

    The Visual Computer, 1–15 (2024)

    Wu, C., He, T.: Efficient minor defects detection on steel surface via res-attention and position encoding. The Visual Computer, 1–15 (2024)

  2. [2]

    Virtual Reality & Intelligent Hardware 5(1), 57–67 (2023)

    Zhang, M., Tian, X.: Transformer architecture based on mutual attention for image-anomaly detection. Virtual Reality & Intelligent Hardware 5(1), 57–67 (2023)

  3. [3]

    Virtual Reality & Intelligent Hardware 5(6), 471–489 (2023)

    Tian, X., Wu, Z., Cao, J., Chen, S., Dong, X.: Ilidviz: An incremental learning- based visual analysis system for network anomaly detection. Virtual Reality & Intelligent Hardware 5(6), 471–489 (2023)

  4. [4]

    The Visual Computer, 1–15 (2024)

    Sun, W., Zhang, J., Liu, Y.: Adversarial-based refinement dual-branch network for semi-supervised salient object detection of strip steel surface defects. The Visual Computer, 1–15 (2024)

  5. [5]

    Measurement 199, 111429 (2022)

    Ding, T., Li, G., Liu, Z., Wang, Y.: Cross-scale edge purification network for salient object detection of steel defect images. Measurement 199, 111429 (2022)

  6. [6]

    IEEE Transactions on Instrumentation and Measurement 71, 1–12 (2022)

    Han, C., Li, G., Liu, Z.: Two-stage edge reuse network for salient object detec- tion of strip steel surface defects. IEEE Transactions on Instrumentation and Measurement 71, 1–12 (2022)

  7. [7]

    IEEE Transactions on Industrial Informatics (2024)

    Shen, K., Zhou, X., Liu, Z.: Minet: Multiscale interactive network for real-time salient object detection of strip steel surface defects. IEEE Transactions on Industrial Informatics (2024)

  8. [8]

    IEEE Transactions on Instrumentation and Measurement 72, 1–12 (2023) 16

    Wan, B., Zhou, X., Zheng, B., Yin, H., Zhu, Z., Wang, H., Sun, Y., Zhang, J., Yan, C.: Lfrnet: Localizing, focus, and refinement network for salient object detection of surface defects. IEEE Transactions on Instrumentation and Measurement 72, 1–12 (2023) 16

Show all 56 references
  1. [9]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Wei, J., Shen, F., Lv, C., Zhang, Z., Zhang, F., Yang, H.: Diversified and multi- class controllable industrial defect synthesis for data augmentation and transfer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4444–4452 (2023)

  2. [10]

    In: AAAI (2023)

    Duan, Y., Hong, Y., Niu, L., Zhang, L.: Few-shot defect image generation via defect-aware feature manipulation. In: AAAI (2023)

  3. [11]

    arXiv preprint arXiv:2402.19330 (2024)

    Li, H., Zhang, Z., Chen, H., Wu, L., Li, B., Liu, D., Wang, M.: A novel approach to industrial defect generation through blended latent diffusion model with online adaptation. arXiv preprint arXiv:2402.19330 (2024)

  4. [12]

    Advances in neural information processing systems 33, 6840–6851 (2020)

    Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. Advances in neural information processing systems 33, 6840–6851 (2020)

  5. [13]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695 (2022)

  6. [14]

    arXiv preprint arXiv:2211.01324 (2022)

    Balaji, Y., Nah, S., Huang, X., Vahdat, A., Song, J., Kreis, K., Aittala, M., Aila, T., Laine, S., Catanzaro, B., et al.: ediffi: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324 (2022)

  7. [15]

    The Visual Computer, 1–15 (2024)

    Wu, H., Li, B., Tian, L., Dong, C.: Ddfa: a displacement and diffusion-based feature augmentation method for imbalanced image recognition. The Visual Computer, 1–15 (2024)

  8. [16]

    IEEE Transactions on Industrial Informatics (2024)

    Yang, X., Ye, T., Yuan, X., Zhu, W., Mei, X., Zhou, F.: A novel data augmentation method based on denoising diffusion probabilistic model for fault diagnosis under imbalanced data. IEEE Transactions on Industrial Informatics (2024)

  9. [17]

    IEEE Transactions on Automation Science and Engineering, 1–13 (2024)

    Tai, Y., Yang, K., Peng, T., Huang, Z., Zhang, Z.: Defect image sample generation with diffusion prior for steel surface defect recognition. IEEE Transactions on Automation Science and Engineering, 1–13 (2024)

  10. [18]

    In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp

    Chen, M., Laina, I., Vedaldi, A.: Training-free layout control with cross-attention guidance. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 5343–5353 (2024)

  11. [19]

    In: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, pp

    Xie, J., Li, Y., Huang, Y., Liu, H., Zhang, W., Zheng, Y., Shou, M.Z.: Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion. In: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7452–7461 (2023)

  12. [20]

    CVPR (2023) 17

    Li, Y., Liu, H., Wu, Q., Mu, F., Yang, J., Gao, J., Li, C., Lee, Y.J.: Gligen: Open-set grounded text-to-image generation. CVPR (2023) 17

  13. [21]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp

    Wang, X., Darrell, T., Rambhatla, S.S., Girdhar, R., Misra, I.: Instancediffusion: Instance-level control for image generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6232–6242 (2024)

  14. [22]

    The Visual Computer 40(9), 6033–6045 (2024)

    Endo, Y.: Masked-attention diffusion guidance for spatially controlling text-to- image generation. The Visual Computer 40(9), 6033–6045 (2024)

  15. [23]

    Advances in neural information processing systems 27 (2014)

    Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial nets. Advances in neural information processing systems 27 (2014)

  16. [24]

    In: Proceedings of the IEEE International Conference on Computer Vision, pp

    Zhu, J.-Y., Park, T., Isola, P., Efros, A.A.: Unpaired image-to-image transla- tion using cycle-consistent adversarial networks. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 2223–2232 (2017)

  17. [25]

    In: ACM SIGGRAPH 2022 Conference Proceedings, pp

    Sauer, A., Schwarz, K., Geiger, A.: Stylegan-xl: Scaling stylegan to large diverse datasets. In: ACM SIGGRAPH 2022 Conference Proceedings, pp. 1–10 (2022)

  18. [26]

    Computer Animation and Virtual Worlds 35(3), 2248 (2024)

    Zhao, W., Zhu, J., Huang, J., Li, P., Sheng, B.: Gan-based multi-decomposition photo cartoonization. Computer Animation and Virtual Worlds 35(3), 2248 (2024)

  19. [27]

    IEEE Transactions on Visualization and Computer Graphics (2024)

    Hu, X., Yang, C., Fang, F., Huang, J., Li, P., ShengB, B., Lee, T.-Y.: Msemb- gan: Multi-stitch embroidery synthesis via region-aware texture generation. IEEE Transactions on Visualization and Computer Graphics (2024)

  20. [28]

    In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp

    Zhang, G., Cui, K., Hung, T.-Y., Lu, S.: Defect-gan: High-fidelity defect synthe- sis for automated defect inspection. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 2524–2534 (2021)

  21. [29]

    IEEE Transactions on Instrumentation and Measurement (2023)

    Zhao, C., Xue, W., Fu, W., Li, Z., Fang, X.: Defect sample image generation method based on gans in diamond tool defect detection. IEEE Transactions on Instrumentation and Measurement (2023)

  22. [30]

    IEEE Transactions on Automation Science and Engineering (2023)

    Li, W., Gu, C., Chen, J., Ma, C., Zhang, X., Chen, B., Wan, S.: Dls-gan: gen- erative adversarial nets for defect location sensitive data augmentation. IEEE Transactions on Automation Science and Engineering (2023)

  23. [31]

    Measurement Science and Technology 35(4), 045408 (2024)

    Ran, G., Yao, X., Wang, K., Ye, J., Ou, S.: Sketch-guided spatial adaptive nor- malization and high-level feature constraints based gan image synthesis for steel strip defect detection data augmentation. Measurement Science and Technology 35(4), 045408 (2024)

  24. [32]

    arXiv preprint arXiv:2204.06125 1(2), 3 (2022) 18

    Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., Chen, M.: Hierarchical text- conditional image generation with clip latents. arXiv preprint arXiv:2204.06125 1(2), 3 (2022) 18

  25. [33]

    arXiv preprint arXiv:2402.05712 (2024)

    Ma, Z., Zhu, X., Qi, G., Qian, C., Zhang, Z., Lei, Z.: Diffspeaker: Speech-driven 3d facial animation with diffusion transformer. arXiv preprint arXiv:2402.05712 (2024)

  26. [34]

    In: European Conference on Computer Vision, pp

    Ma, Z., Wei, Y., Zhang, Y., Zhu, X., Lei, Z., Zhang, L.: Scaledreamer: Scal- able text-to-3d synthesis with asynchronous score distillation. In: European Conference on Computer Vision, pp. 1–19 (2025). Springer

  27. [35]

    arXiv preprint arXiv:2406.08177 (2024)

    Wu, R., Sun, L., Ma, Z., Zhang, L.: One-step effective diffusion network for real- world image super-resolution. arXiv preprint arXiv:2406.08177 (2024)

  28. [36]

    Measurement Science and Technology 35(10), 106111 (2024)

    Xiao, Z., Li, C., Liu, T., Liu, W., Mo, S., Houjoh, H.: Parameter sharing fault data generation method based on diffusion model under imbalance data. Measurement Science and Technology 35(10), 106111 (2024)

  29. [37]

    Computer Animation and Virtual Worlds 35(3), 2252 (2024)

    Zhang, M., Yang, J., Xian, Y., Li, W., Gu, J., Meng, W., Zhang, J., Zhang, X.: Ag-sdm: Aquascape generation based on stable diffusion model with low-rank adaptation. Computer Animation and Virtual Worlds 35(3), 2252 (2024)

  30. [38]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

    Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3836–3847 (2023)

  31. [39]

    ACM transac- tions on graphics (TOG) 42(4), 1–11 (2023)

    Avrahami, O., Fried, O., Lischinski, D.: Blended latent diffusion. ACM transac- tions on graphics (TOG) 42(4), 1–11 (2023)

  32. [40]

    In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp

    Achanta, R., Hemami, S., Estrada, F., Susstrunk, S.: Frequency-tuned salient region detection. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1597–1604 (2009). IEEE

  33. [41]

    The Visual Computer 39(10), 4887–4899 (2023)

    Ma, C., He, T., Gao, J.: Skin scar segmentation based on saliency detection. The Visual Computer 39(10), 4887–4899 (2023)

  34. [42]

    IEEE Transactions on Instrumentation and Measurement 71, 1–14 (2021)

    Zhou, X., Fang, H., Liu, Z., Zheng, B., Sun, Y., Zhang, J., Yan, C.: Dense attention-guided cascaded network for salient object detection of strip steel sur- face defects. IEEE Transactions on Instrumentation and Measurement 71, 1–14 (2021)

  35. [43]

    IEEE Access 9, 149465–149476 (2021)

    Zhou, X., Fang, H., Fei, X., Shi, R., Zhang, J.: Edge-aware multi-level interactive network for salient object detection of strip steel surface defects. IEEE Access 9, 149465–149476 (2021)

  36. [44]

    IEEE Transactions on Industrial Informatics16(12), 7448–7458 (2019)

    Dong, H., Song, K., He, Y., Xu, J., Yan, Y., Meng, Q.: Pga-net: Pyramid fea- ture fusion and global context attention network for automated surface defect detection. IEEE Transactions on Industrial Informatics16(12), 7448–7458 (2019)

  37. [45]

    Engineering Applications of Artificial Intelligence 123, 106474 (2023)

    Wan, B., Zhou, X., Sun, Y., Zhu, Z., Yin, H., Hu, J., Zhang, J., Yan, C.: Sminet: Semantics-aware multi-level feature interaction network for surface defect 19 detection. Engineering Applications of Artificial Intelligence 123, 106474 (2023)

  38. [46]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022)

    Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., Xie, S.: A convnet for the 2020s. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022)

  39. [47]

    In: International Conference on Machine Learning, pp

    Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748–8763 (2021). PMLR

  40. [48]

    arXiv preprint arXiv:2207.12598 (2022)

    Ho, J., Salimans, T.: Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022)

  41. [49]

    Advances in neural information processing systems 30 (2017)

    Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S.: Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30 (2017)

  42. [50]

    In: Proceedings of the IEEE International Conference on Computer Vision, pp

    Fan, D.-P., Cheng, M.-M., Liu, Y., Li, T., Borji, A.: Structure-measure: A new way to evaluate foreground maps. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 4548–4557 (2017)

  43. [51]

    arXiv preprint arXiv:1805.10421 (2018)

    Fan, D.-P., Gong, C., Cao, Y., Ren, B., Cheng, M.-M., Borji, A.: Enhanced- alignment measure for binary foreground map evaluation. arXiv preprint arXiv:1805.10421 (2018)

  44. [52]

    IEEE Transactions on Instrumentation and Measurement 69(12), 9709–9719 (2020)

    Song, G., Song, K., Yan, Y.: Edrnet: Encoder–decoder residual network for salient object detection of strip steel surface defects. IEEE Transactions on Instrumentation and Measurement 69(12), 9709–9719 (2020)

  45. [53]

    arXiv preprint arXiv:1412.6980 (2014)

    Kingma, D.P.: Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)

  46. [54]

    https://huggingface.co/gligen/gligen-generation-sem

    Li, Y., Liu, H., Wu, Q., Mu, F., Yang, J., Gao, J., Li, C., Lee, Y.J.: GLIGEN-sem. https://huggingface.co/gligen/gligen-generation-sem

  47. [55]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017)

    Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A.: Scene parsing through ade20k dataset. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017)

  48. [56]

    https://huggingface.co/CompVis/stable-diffusion-v1-4 20

    Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: Stable-Diffusion- v1-4. https://huggingface.co/CompVis/stable-diffusion-v1-4 20

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.