{"id":"57caebd1-14f4-4ae1-a40f-ad94d933dd90","arxiv_id":"2504.12689","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"HSS-IAD is a new benchmark of 8,580 metallic-part images, curated from five existing datasets, on which current unsupervised anomaly detection methods perform substantially worse than on MVTec-AD and related benchmarks.","lead":"This paper introduces HSS-IAD, a new industrial anomaly detection dataset with 8,580 images of metal-like parts, built by reclassifying and re-annotating images from five existing datasets. It shows that current multi-class anomaly detection methods score much lower on HSS-IAD than on established benchmarks, suggesting the dataset captures harder, more realistic conditions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Without a source-domain confound check, the reported AUROC drop cannot be attributed to same-sort realism; HSS-IAD's factory interpretation is not yet established.","rationale":"The reader's weakest_assumption identifies the same underlying issue: HSS-IAD's realism claim depends on the assumption that five public source datasets can be homogenized into a same-sort, same-factory-like benchmark. My concern is more specific: the benchmark's headline evidence, the large performance drop in Table II, is confounded by source-domain shift. The paper provides the dataset at a public URL and evaluates six methods under two protocols, which is useful, but it does not provide curation code, exact split specifications, or any analysis of the domain gap among the five source datasets. The paper also reports no variance across runs, and the manual reclassification criteria for the Casting categories are qualitative, so the dataset's composition is not fully reproducible as described. None of this invalidates HSS-IAD as a hard multi-class anomaly detection benchmark; the measured drop is real given the released images. However, the central claim that it represents a realistic same-factory setting requires that source-specific imaging differences are not the main driver of difficulty. The proposed source-classification test directly separates these explanations: if source identity is trivially predictable, the dataset is primarily a mixture of domains rather than a coherent same-sort product family. This supports keeping the reader's CONDITIONAL verdict: the paper should be accepted only if the authors either provide evidence that domain shift is not the dominant factor or reframe the contribution as a challenging heterogeneous multi-source benchmark rather than a same-factory dataset.","tokens_in":10156,"tokens_out":3806,"duration_ms":47244,"concrete_test":"Run a source-domain discriminability check: train a small classifier (e.g., ResNet-18) on a held-out split of HSS-IAD images to predict the original source dataset (KolektorSDD2, MTD, Casting_C1/C2/C3, KolektorSDD, STEEL). If held-out source classification AUROC or accuracy is near-perfect (e.g., AUROC > 0.95), then source identity is almost perfectly recoverable from the images, meaning the five acquisition pipelines are strong domain cues. That would show the 67.0% mean AUROC likely reflects cross-dataset domain shift rather than the intended same-sort structural/appearance variation, and the same-factory realism claim would need to be weakened or re-argued.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central contribution is that HSS-IAD is harder and more realistic because its categories are the same sort of industrial product with heterogeneous structure/appearance and subtle, material-like defects. The evidence is the large multi-class AUROC drop (94.0% on MVTec-AD to 67.0% on HSS-IAD in Table II). However, HSS-IAD is assembled from five independently acquired source datasets (KolektorSDD2, MTD, Casting, KolektorSDD, STEEL) with different cameras, lighting, resolutions, and object scales, then manually filtered and reclassified as described in Section III.A. Nothing in the paper rules out that a substantial part of the measured performance gap is caused by cross-source domain shift: a unified model can exploit source-specific low-level cues (illumination, color, texture, resolution) or fail because the five sources are not a coherent product family. The claim that the benchmark bridges the gap to real factory conditions specifically requires that same-sort structural/appearance variation and defect subtlety drive the difficulty, not the heterogeneity of the original acquisition pipelines. The paper itself states that the original Casting six-class setup 'lacked meaningful distinctions' and that manual reclassification and filtering were needed, so it is important to verify that the resulting categories are stable and that the source composition is not the dominant factor behind the benchmark results.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces HSS-IAD, a benchmark for multi-class unsupervised industrial anomaly detection (MUAD) containing 8,580 images of metallic-like parts organized into seven categories, assembled from five existing datasets (KolektorSDD2, Casting, MTD, KolektorSDD, and STEEL). The construction includes reclassification of Casting images into three finer categories, removal of blurry/duplicate/low-quality images, manual label filtering and refinement, and generation of pixel-level masks for STEEL; foreground masks are also released. The authors benchmark six methods (DeSTSeg, SimpleNet, DRÆM, UniAD, RD4AD, Dinomaly) under multi-class and class-separated settings, reporting image-level AUROC, pixel-level AP, and AUPRO. The central empirical claim is that HSS-IAD is much harder than existing benchmarks: average I-AUROC drops from 94.0% on MVTec-AD to 67.0% on HSS-IAD and P-AUPRO from 83.2% to 43.8% (Table II). The paper interprets this gap as evidence that the dataset captures same-sort structural/appearance variation and subtle, material-like defects found in real factories.","tokens_in":10473,"tokens_out":9396,"duration_ms":92029,"significance":"If the central claim is established, HSS-IAD fills a real need: most MUAD benchmarks mix semantically unrelated categories, while industrial settings require detecting anomalies across variants of the same product family. The paper's concrete strengths are a publicly released dataset, pixel-precise anomaly annotations, a reproducible benchmark protocol using official implementations, and results under two settings with detailed per-category tables. The reported performance gap is large and plausible. The main weakness is that the dataset is assembled from five independently collected source datasets, and the paper does not yet quantify whether the difficulty comes from same-sort realism or from cross-source acquisition shift; the 'same sort' claim is asserted rather than demonstrated with data. With additional diagnostic experiments and curation documentation, this could be a valuable benchmark.","major_comments":[{"comment":"The headline result — average I-AUROC dropping from 94.0% on MVTec-AD to 67.0% on HSS-IAD and P-AUPRO from 83.2% to 43.8% — is presented as evidence that HSS-IAD is harder because of same-sort structural/appearance variation and subtle defects. However, HSS-IAD is built from five independently acquired datasets (KolektorSDD2, Casting, MTD, KolektorSDD, STEEL) with different cameras, lighting, resolutions, and backgrounds, and no control for cross-source domain shift is reported. A unified model could suffer or benefit from source-specific low-level cues, so the measured gap cannot yet be attributed to the claimed same-sort realism. Please add source-diagnostic experiments, for example per-source performance within HSS-IAD, source-controlled train/test splits, cross-source transfer tests, or low-level image statistics, and state explicitly which fraction of the performance drop persists when acquisition heterogeneity is controlled.","section":"Section III.A, Table II"},{"comment":"The central motivation depends on the notions of 'same sort' products and 'the same factory', but neither is operationalized. The five sources are not shown to originate from a common production environment: steel sheets, magnetic tiles, electrical commutators, and castings can plausibly come from different plants and manufacturing processes. Please define 'same sort' with explicit inclusion criteria (e.g., shared material, product family, or manufacturing process), provide evidence that the selected categories can co-occur in one factory, or revise the claim to describe the dataset as 'metallic-like industrial parts from heterogeneous sources'.","section":"Sections I and III.A"},{"comment":"The dataset construction relies on several qualitative curation steps — 'carefully evaluated' image quality, 'zero image check', 'manual label filtering', and reclassification of the Casting images — but no quantitative documentation is provided. For a dataset paper these steps are the method, and manual selection can bias both measured difficulty and the perceived similarity between defects and backgrounds. Please report the number of images removed or reclassified at each stage, the number and type of label corrections, annotator counts and agreement, and release the filtering/split lists or scripts so that the curation is reproducible.","section":"Section III.A"},{"comment":"Table I lists 'Similarity (between defect and background)' as a dataset attribute with values Low/High, and Section III.B states that the similarity in HSS-IAD is 'notably high', but no quantitative definition or measurement is given. This is a load-bearing attribute for the paper's claim that defects 'closely resemble the base materials'. Please define and report a reproducible similarity measure (e.g., low-level pixel statistics, feature-space distances, or human ratings) for HSS-IAD and for the comparison datasets.","section":"Table I and Section III.B"}],"minor_comments":[{"comment":"The text states that Fig. 4(c) shows 'the range of defect ratios for each class', but the caption describes Fig. 4(c) as the aspect ratio of the minimum bounding rectangles of defects; please align the text with the figure caption.","section":"Section III.B and Fig. 4"},{"comment":"In the caption of Fig. 6 the last panel is labeled '(e)', while the text refers to 'Fig. 6(c)' for the anomaly-localization visualizations; please make the panel labels consistent.","section":"Fig. 6"},{"comment":"Section II contains the stray characters 'small.xs' at the end of the BTAD sentence; please remove them.","section":"Section II"},{"comment":"Table II leaves the DRÆM entry for Real-IAD blank without explanation; please state whether the method was omitted for resource reasons or could not be evaluated, since this affects the comparability of the Mean±Std row.","section":"Table II"},{"comment":"No information is given about the number of random seeds or runs used to produce the results; the reported 'Mean±Std' is across methods rather than across runs, which can be misread as run variability. Please state the experimental repetition protocol.","section":"Tables II–IV"},{"comment":"The abstract and Section III mention foreground images for synthetic anomaly generation, but no experiment uses synthetic anomalies; please clarify explicitly that these masks are released for future use.","section":"Abstract and Section III"}],"recommendation":"major_revision","confidential_remarks":"The dataset is potentially useful and the benchmark numbers are reported in detail. The key risk is that the 'same-factory' realism claim is not yet supported because the data come from five independent public datasets; I recommend requesting source-domain diagnostics and quantitative curation documentation as conditions for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: HSS-IAD is a genuinely useful dataset contribution, and the same-sort framing fills a real gap in the MUAD benchmark landscape. But the paper overinterprets its main number: the 27-point AUROC drop relative to MVTec-AD is not evidence that same-sort heterogeneity and subtle defects are the cause, because HSS-IAD also brings in five different source domains with different cameras, lighting, and resolutions. That confound is not addressed.\n\nWhat's new: the dataset itself. Reclassifying casting into three distinct categories, adding refined pixel masks to STEEL, weeding out blurry/duplicate images, and providing DIS foreground masks for synthetic generation is real curation work. Seven metallic-like categories, 8,580 images, with defect-to-background similarity explicitly considered: this is a solid benchmark resource. The evaluation is reasonably thorough, covering six methods across embedding-, augmentation-, and reconstruction-based families, with two protocols and per-category numbers.\n\nWhere it wobbles: the source-domain issue is not a minor quibble. If you train a unified model on a mix of KolektorSDD2, MTD, Casting, and STEEL, the model can exploit low-level cues like illumination or resolution across sources, or fail because the sources aren't a coherent product family. The paper claims the benchmark bridges the gap to real factory conditions, but nothing rules out that a good chunk of the drop is just domain shift. They don't report per-source AUROC, don't run a domain-adaptation or source-label baseline, and don't quantify acquisition differences. The manual reclassification and 'excessive background noise' filtering are also subjective; the paper should document the criteria and show inter-annotator agreement or at least examples of excluded images. There are no variance estimates across training runs, which is standard for benchmark papers. The GitHub link is good, but the paper doesn't give exact split specs or license info.\n\nProportionately, these are fixable issues. The dataset remains valuable even if the difficulty is partly domain shift; the same-sort hypothesis is plausible but unproven. The paper deserves peer review, but with a conditional recommendation: require a source-domain analysis (leave-one-source-out, domain adaptation baseline, or per-source metrics) and proper reproducibility files.\n\nSerious thinker: yes. The curation logic is coherent and the authors are honest about manual steps. This is a resource paper, not a theoretical contribution, but it is a useful one.\n\nRecommendation: send to review, and ask for the confound analysis and reproducibility details before acceptance.","headline":"Useful same-sort IAD benchmark, but the headline AUROC drop is confounded by source-domain shift; the dataset is worth reviewing, the interpretation needs tightening.","tokens_in":10933,"tokens_out":2648,"would_cite":true,"duration_ms":28777,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By recombining five existing defect-image collections into 8,580 same-sort metal-part images, HSS-IAD cuts the average multiclass anomaly-detection AUROC from 94.0% on MVTec-AD to 67.0%, and pixel-level AUPRO from 83.2% to 43.8%.","keywords":["HSS-IAD","industrial anomaly detection","unsupervised anomaly detection","multi-class anomaly detection","same-sort heterogeneity","surface defect detection","anomaly localization","metallic parts"],"falsifier":"Run the six methods on a matched control dataset built by randomly mixing the same five source collections without same-sort reclassification; if the AUROC drop relative to MVTec-AD is the same as on HSS-IAD, the challenge is explained by cross-dataset domain shift rather than same-sort factory realism. Collecting images from a single actual production line and comparing transfer from HSS-IAD-trained versus MVTec-AD-trained models would settle it directly.","tokens_in":9925,"feed_emoji":"🔩","tokens_out":10427,"duration_ms":96450,"temperature":0.7,"pith_summary":"The paper argues that current multiclass unsupervised anomaly detection (MUAD) benchmarks overstate real-world readiness because they mix product classes no single factory would make and use defects that are too obvious. To test the gap, it introduces HSS-IAD, a dataset of 8,580 images of metallic-like parts that are \"same sort\"—electrical commutators, magnetic tiles, steel sheet, engine castings—but vary in structure and appearance, with subtle defects that resemble machining marks, oil stains, and oxide scale. The paper reports that six reproduced state-of-the-art methods average a 94.0% image-level AUROC on MVTec-AD but only 67.0% on HSS-IAD, with pixel-level AUPRO falling from 83.2% to 43.8%, showing the dataset is markedly harder and better able to separate method strengths. It also provides foreground segmentations for synthetic anomaly generation, so the benchmark can be used for augmentation-based training as well as evaluation.","feed_headline":"Same-sort part benchmark drops anomaly detection to 67% AUROC","feed_subtitle":"Eight thousand metal-part images with subtle, material-like defects expose where multi-class detectors fail.","key_machinery":"The central object is the HSS-IAD dataset and its construction pipeline. Images are selected and reclassified from five existing defect datasets, deduplicated, and quality-filtered; castings are split into three categories defined by surface complexity and interference; steel-defect bounding boxes are converted into pixel masks; labeled regions that are too large, too small, or misclassified are manually corrected; and normal samples with confusable attached elements are deliberately kept. This curation is what creates the paper's intended test: the same-sort relationship supplies a coherent \"normal\" distribution, while structural variation, process features, and defect-material similarity force models to separate true anomalies from benign manufacturing variation. The authors also use a dichotomous image segmentation method to provide foreground masks, enabling synthetic anomaly generation for augmentation-based training.","core_discovery":"On the paper's own terms, the core discovery is that \"same-sort heterogeneity\" is a distinct failure mode for MUAD. HSS-IAD's seven categories are all metallic-like industrial parts, but within the same sort the appearance changes—castings have machined and unmachined surfaces, threaded holes, oil stains, oxide scale—and the defects are subtle enough to be mistaken for these normal features. When six unsupervised methods are evaluated under the multi-class protocol, the average image-level AUROC is 67.0% and the average pixel-level AUPRO is 43.8%, versus 94.0% and 83.2% on MVTec-AD; similar drops appear relative to VisA, Real-IAD, BTAD, and MPDD. The paper interprets this as evidence that the dataset captures real factory conditions that existing benchmarks miss, and that a more discriminative benchmark is needed to drive algorithmic progress.","pith_inferences":["If the same-sort heterogeneity is the active ingredient, then fine-tuning a method on HSS-IAD should transfer to a single factory's inspection line better than training on category-diverse datasets; the paper does not test this transfer, but it is a direct consequence of its motivation.","The drop in performance could partly come from cross-dataset domain shift (different cameras, lighting, resolutions) rather than same-sort difficulty, because the source images come from five independent collections; a control experiment mixing arbitrary categories from the same sources would separate these effects.","The provided foreground masks make it possible to generate synthetic anomalies of controlled subtlety; one could measure AUROC as a function of defect contrast or size to map exactly where current methods start failing, an experiment the paper does not run.","The class-separated versus multi-class gap suggests that unified MUAD models struggle mainly with learning one shared normal distribution across heterogeneous same-sort parts; class-conditional or prompt-guided normalization might close part of the gap."],"forward_implications":["A model that scores well on existing MUAD benchmarks cannot be assumed ready for a real factory; the same methods lose roughly 27 points in image-level AUROC and 39 points in pixel-level AUPRO on HSS-IAD.","The high defect-to-background similarity and 3.0% anomalous-pixel ratio make HSS-IAD a stress test for small, subtle defect localization, so methods that rely on high-level features at the expense of spatial detail will be exposed.","Because foreground masks are provided, data-augmentation methods can be trained with synthetic anomalies on HSS-IAD, making the dataset useful both as an evaluation benchmark and as a training resource.","Method ranking on HSS-IAD changes relative to existing benchmarks, with feature- and reconstruction-based approaches outperforming latent-noise approaches, suggesting the benchmark can reveal which design choices matter for realistic industrial defects."],"supporting_citations":[{"why":"Supplies the KolektorSDD2 electrical-commutator images that become one HSS-IAD category.","marker":"[12]"},{"why":"Source of the casting images, reselected and split into three categories by surface complexity.","marker":"[1]"},{"why":"Source of magnetic-tile images for the MTD category.","marker":"[13]"},{"why":"Source of KolektorSDD images, including normal samples with confusing attached elements.","marker":"[11]"},{"why":"Source of steel-defect images; its bounding-box annotations are converted into pixel-level masks.","marker":"[14]"},{"why":"MVTec-AD is the main comparison benchmark, setting the baseline performance that HSS-IAD drops well below.","marker":"[4]"},{"why":"Real-IAD supplies the large-scale multi-view comparison case in the benchmark table.","marker":"[5]"},{"why":"UniAD defines the multi-class unified-model setting and serves as a reproduced baseline.","marker":"[2]"},{"why":"Dinomaly is the strongest reproduced method on HSS-IAD, used to show the benchmark's difficulty and spread.","marker":"[20]"},{"why":"The dichotomous image segmentation method used to produce foreground images for synthetic anomaly generation.","marker":"[10]"}],"fun_headline_variants":["Same-sort metal parts fool anomaly detectors, AUROC drops to 67%","HSS-IAD: same-sort parts trip MUAD to 67% AUROC","Subtle same-sort defects push MUAD AUROC down to 67%","HSS-IAD: same-sort heterogeneity cuts MUAD AUROC to 67%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset can stand in for a single factory's same-sort production line even though its images are drawn from five independent public datasets with different imaging conditions and were reorganized by manual reclassification and filtering.","fun_headline_variants_meta":{"raw":{"variants":["Same-sort metal parts fool anomaly detectors, AUROC drops to 67%","HSS-IAD: same-sort parts trip MUAD to 67% AUROC","Subtle same-sort defects push MUAD AUROC down to 67%","HSS-IAD: same-sort heterogeneity cuts MUAD AUROC to 67%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001151,"raw_usage":{"total_tokens":4762,"prompt_tokens":925,"completion_tokens":3837,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":3746}},"tokens_in":541,"tokens_out":3837,"duration_ms":26857,"temperature":1.0,"reasoning_tokens":3746,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:24:26.604610+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the six methods on a matched control dataset built by randomly mixing the same five source collections without same-sort reclassification; if the AUROC drop relative to MVTec-AD is the same as on HSS-IAD, the challenge is explained by cross-dataset domain shift rather than same-sort factory realism. Collecting images from a single actual production line and comparing transfer from HSS-IAD-trained versus MVTec-AD-trained models would settle it directly.","supporting_citations":[{"cited_title":"Mixed supervision for surface-defect detection: From weakly to fully supervised learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the KolektorSDD2 electrical-commutator images that become one HSS-IAD category."},{"cited_title":"A casting surface dataset and benchmark for subtle and confusable defect detection in complex contexts,","cited_arxiv_id":null,"evidence_quote":"Source of the casting images, reselected and split into three categories by surface complexity."},{"cited_title":"Surface defect saliency of magnetic tile,","cited_arxiv_id":null,"evidence_quote":"Source of magnetic-tile images for the MTD category."},{"cited_title":"Segmentation-based deep-learning approach for surface-defect detec- tion,","cited_arxiv_id":null,"evidence_quote":"Source of KolektorSDD images, including normal samples with confusing attached elements."},{"cited_title":"Severstal: Steel defect detection,","cited_arxiv_id":null,"evidence_quote":"Source of steel-defect images; its bounding-box annotations are converted into pixel-level masks."},{"cited_title":"Mvtec ad–a comprehensive real-world dataset for unsupervised anomaly detection,","cited_arxiv_id":null,"evidence_quote":"MVTec-AD is the main comparison benchmark, setting the baseline performance that HSS-IAD drops well below."},{"cited_title":"Real-iad: A real-world multi-view dataset for benchmarking versatile industrial anomaly detection,","cited_arxiv_id":null,"evidence_quote":"Real-IAD supplies the large-scale multi-view comparison case in the benchmark table."},{"cited_title":"A unified model for multi-class anomaly detection,","cited_arxiv_id":null,"evidence_quote":"UniAD defines the multi-class unified-model setting and serves as a reproduced baseline."},{"cited_title":"Highly accurate dichotomous image segmentation,","cited_arxiv_id":null,"evidence_quote":"The dichotomous image segmentation method used to produce foreground images for synthetic anomaly generation."}],"review_version":1}