{"id":"6db2f65a-ebb3-40ac-8a22-071168f9e456","arxiv_id":"2505.10996","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"M2AD, a large-scale benchmark with 120 view-illumination configurations per object, shows that state-of-the-art visual anomaly detection methods drop markedly when viewpoint and lighting vary together.","lead":"This paper introduces M2AD, a new image dataset of 119,880 photos that pair 12 camera angles with 10 lighting setups to test how well automated defect detectors cope when both viewpoint and lighting change at once. It shows that leading anomaly detection models lose significant accuracy on these harder images, and proposes two evaluation tracks for future work.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The M2AD-Invariant protocol's supervised detectability filter (Mask R-CNN consensus at IoU≥0.3, p≥0.5) is an unvalidated selection step; Table 4 and the 81.3% headline number inherit any bias it introduces, so the central 'profound challenge' claim is not yet secure.","rationale":"The reader's weakest assumption correctly identifies the supervised detectability filter as the linchpin of the M2AD-Invariant benchmark. I considered two alternatives: (i) the 'synchronized' wording is inaccurate because the setup uses a single camera plus motorized turntable, but temporal synchronization is irrelevant for static specimens and does not affect the benchmark's validity; (ii) M2AD-Synergy's O-AUROC averaging is not a true fusion baseline, but that affects only the synergy protocol, not the invariant result. The invariant number is the strongest evidence for the paper's central claim, and it is entirely downstream of the filter. The internal ablations (Fig. 4) and resolution comparisons (Tables 2-3) are internally consistent and provide partial support that configuration count degrades performance, but they do not validate the filter or rule out selection bias. Therefore the concrete test should be a sensitivity analysis of IoU/confidence thresholds plus a human-detectability comparison. No ad hominem is involved; the concern is purely about methodological transparency and the interpretability of the headline quantitative result.","tokens_in":15490,"tokens_out":5892,"duration_ms":64594,"concrete_test":"Re-run M2AD-Invariant for Dinomaly and RD++ on (a) the full unfiltered abnormal set and (b) subsets generated with IoU thresholds 0.1/0.5 and confidence thresholds 0.3/0.7, using identical training splits. If I-AUROC/AUPRO change by more than about 2-3 points across reasonable thresholds, then the reported 81.3% is filter-dependent and the paper should present sensitivity bounds or a validated filter. As a second check, have human annotators independently rate detectability on a random 200-image sample and compare against Mask R-CNN consensus; this directly tests whether the filter tracks human-visible anomalies.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 defines 'detectable' abnormal images as those where three supervised Mask R-CNN models agree with manual annotation at IoU >= 0.3 and confidence >= 0.5; Section 3.2 reports this retains only about 75% of abnormal images. The paper offers no sensitivity analysis or external validation of these thresholds. Because M2AD-Invariant is the basis for the strongest evidence (Dinomaly 99.6 on MVTec AD dropping to 81.3 I-AUROC, Table 4), any systematic bias in the filter directly changes the headline. The direction of bias is not obvious: if the filter selects high-contrast anomalies that supervised models find easy, the retained subset may be easier than the full distribution, making the current numbers conservative; if it selects anomalies with texture cues that supervised models exploit but unsupervised methods miss, the challenge is overstated. Without knowing which, the central quantitative claim is not fully interpretable. Additionally, the manual-annotation procedure for images where a defect is visually absent is not specified, so the ground truth used to train the filter itself is not independently documented. This is a load-bearing assumption, not a side detail.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces M2AD, a large-scale visual anomaly detection benchmark with 119,880 high-resolution images of 999 specimens across 10 categories, captured under 12 synchronized viewpoints and 10 illumination conditions (120 configurations per specimen). The authors define two evaluation protocols: M2AD-Synergy, which aggregates predictions across all configurations to test multi-view/multi-illumination fusion, and M2AD-Invariant, which evaluates single-image robustness on a subset of anomalies deemed 'detectable' by a supervised Mask R-CNN consensus filter. They benchmark five recent unsupervised methods (CDO, RD++, MSFlow, Dinomaly, INP-Former) and report substantial performance drops compared with published scores on MVTec AD, e.g., Dinomaly drops from 99.6% to 81.3% I-AUROC on M2AD-Invariant. Additional experiments examine input resolution, the number of configurations, and the number of illumination conditions.","tokens_in":15702,"tokens_out":4537,"duration_ms":47584,"significance":"If the dataset construction and protocols are sound, M2AD is a valuable contribution: it is the first large-scale VAD benchmark to jointly control viewpoint and illumination, it provides high-resolution imagery with sub-millimeter defects, and its two complementary protocols target different failure modes of current methods. The consistent performance drops across five independent methods and two protocols, the ablations with means and standard deviations, and the fact that the authors' own methods (CDO, INP-Former) perform relatively poorly all support the claim that existing VAD methods are not robust to view-illumination interplay. The planned public release of data, test suite, and imaging prototype design further strengthens reproducibility. However, the headline quantitative claims rest on two under-specified components: the supervised detectability filter that defines the M2AD-Invariant subset, and the train/test split used for both protocols. These issues must be resolved before the benchmark's central 'profound challenge' claim is fully secure.","major_comments":[{"comment":"The M2AD-Invariant protocol retains only abnormal images where three supervised Mask R-CNN models agree with manual annotations at IoU >= 0.3 and confidence >= 0.5, yet no sensitivity analysis or external validation of these thresholds is provided. Because Table 4 and the headline 81.3% I-AUROC for Dinomaly are computed exclusively on this retained subset, a miscalibrated filter directly changes the central claim. The direction of the bias is unknown: if the filter preferentially retains high-contrast anomalies that supervised models find easy, the reported difficulty is understated; if it retains texture-based cues that supervised models exploit but unsupervised methods do not, the difficulty is overstated. I request a sensitivity analysis over threshold values, and ideally a control experiment on the full unfiltered set or a comparison with human detectability ratings.","section":"Section 3.1, 'Detectability Assessment' and Table 4"},{"comment":"The paper never specifies the exact train/test split for either protocol. The statement in Section 4.3 that 'all imaging configurations are utilized for training' is ambiguous: if all normal images from all 120 configurations are used in training and the same normal images are then used to compute I-AUROC, the evaluation violates the standard unsupervised VAD protocol of held-out normal test images and would artificially inflate scores. The benchmark setup in Section 4.1 also lacks a clear statement of how normal and abnormal specimens or images are partitioned. This is a load-bearing specification issue, as the reported performance gaps are only interpretable if the split is standard and reproducible.","section":"Section 4.1, 'Benchmark Setups' and Section 4.3, 'M2AD-Invariant Results'"},{"comment":"The manual annotation procedure for images in which a defect is visually absent or undetectable is not documented. The figure refers to 'Manual Labeling' and 'Human Checking', but the paper does not state how annotators decided that an anomaly is not visible under a given view-illumination configuration, nor how inter-annotator disagreements were resolved. Since these manual labels are the ground truth against which the supervised detectability filter is trained, the reliability of the entire M2AD-Invariant subset depends on this undocumented annotation step.","section":"Section 3.1, 'Detectability Assessment' and Fig. 2(c)"}],"minor_comments":[{"comment":"The text contains inconsistent spacing in the abbreviation 'V AD' instead of 'VAD' in several places (e.g., the abstract and Section 1); this should be harmonized.","section":"Throughout"},{"comment":"The caption and legend for Fig. 3(a) are difficult to parse; the relationship between 'All/Detectable Abnormal' counts and the text's statement that about 75% of abnormal images are detectable should be made explicit with per-category numbers.","section":"Fig. 3(a)"},{"comment":"The ablation in Fig. 4 randomly selects configurations and illumination conditions three times, but the sampling procedure (whether subsets are per specimen, whether they are stratified by view/illumination, and what random seeds are used) is not described, which limits reproducibility.","section":"Section 4.2, Fig. 4"},{"comment":"The sentence stating that 'all imaging configurations are utilized for training' belongs in the protocol definition in Section 4.1, not in the results section; currently it is easy to miss when reading the benchmark setup.","section":"Section 4.3, 'Quantitative Results'"},{"comment":"The row for PAD lists '20' under the main category column and '30' under total number of categories, which is confusing; the column semantics for 'Main' and 'Sub.' should be clarified, and the counts should be double-checked against the cited source.","section":"Table 1"},{"comment":"Reference [8] (Dinomaly) is an arXiv preprint without a year or venue; since the paper is from CVPR 2025, the citation should be updated to the published version if available.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a timely and potentially influential benchmark paper, and the core dataset appears genuinely useful to the VAD community. The revision hinges on two fixable but central issues: a sensitivity analysis for the detectability filter, and an unambiguous specification of the train/test split. I would also encourage the authors to provide the MVTec AD baseline numbers obtained with the same codebase and resolutions, rather than citing published scores, to make the performance-drop comparison controlled. If these points are addressed, the paper could be a strong acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The punchline: this is a genuinely useful benchmark resource, and the core qualitative claim holds up, but the headline invariant numbers rest on a filtering step that needs validation before I'd trust it as a standard.\n\nWhat's new: M2AD is the first VAD benchmark I know of with synchronized multi-view and multi-illumination capture at this scale: 12 views × 10 illuminations = 120 configurations, 999 specimens, ~120k high-res images. That fills a real gap. Existing datasets vary one axis or the other; none combine them. The capture rig is cheap and described well enough to reproduce. The evaluation is also honest in one important way: the authors' own methods are not cherry-picked winners — CDO and INP-Former sit near the bottom. That cuts against self-selection.\n\nWhat it does well: two complementary protocols make sense; Synergy tests fusion over configurations, Invariant tests single-image robustness. The ablations on configuration count and illumination count include mean±std over three random draws. Using official implementations for baselines is good practice.\n\nSoft spots: The main one is Section 3.1's detectability filter. M2AD-Invariant keeps only abnormal images where three Mask R-CNN models agree with manual annotations at IoU≥0.3 and confidence≥0.5, about 75% of the abnormal images. Those thresholds look arbitrary, and the paper gives no sensitivity analysis and no external validation that Mask R-CNN consensus tracks what an unsupervised method can detect. The direction of any bias is unclear: it could make the retained subset easier (supervised models find high-contrast flaws) or harder (they exploit texture cues unsupervised methods miss). So the exact I-AUROC numbers in Table 4 — including the 81.3% for Dinomaly — should be read as provisional. That said, the stress-test note overstates things a bit. The 'profound challenge' claim does not rest only on Table 4. The Synergy results, which skip the filter, show the same story across five methods and two resolutions. Even if the invariant numbers shift after re-filtering, the qualitative conclusion that current methods degrade badly under view-illumination interplay is likely to survive.\n\nOther issues are minor: the main benchmark tables have no error bars across seeds; the manual annotation procedure for images where a defect is visually absent is not specified; and the dataset is not yet released, so reproducibility is impossible today. None of these sink the benchmark concept. The paper would benefit from either relaxing or better justifying the filter, publishing the per-class retained ratios, and doing a sensitivity sweep over the thresholds.\n\nWho this is for: anyone building or evaluating industrial VAD methods. It deserves a serious referee — not a desk reject — with the expectation that the filter gets addressed in revision.\n\nRecommendation: send it to review; ask for the sensitivity analysis and the release. I'd bring it to reading group and would cite it once the data is out.","headline":"A useful new benchmark that fills a real gap in VAD evaluation, with a load-bearing detectability filter that needs sensitivity analysis before the invariant-protocol numbers become standard.","tokens_in":16326,"tokens_out":3008,"would_cite":true,"duration_ms":30121,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark shows anomaly detectors fail when viewpoint and lighting vary together.","keywords":["visual anomaly detection","benchmark dataset","multi-view imaging","multi-illumination","view-illumination interplay","industrial inspection","unsupervised anomaly detection","robustness evaluation"],"falsifier":"Take the abnormal images that the detectability filter excluded and ask independent human annotators to mark the anomalies without seeing the supervised-model predictions; if humans can reliably label a large fraction of the excluded images, then the IoU and confidence thresholds are not measuring human-detectable anomalies, and the M2AD-Invariant numbers would need to be recomputed on a differently filtered set.","tokens_in":15243,"feed_emoji":"📊","tokens_out":7850,"duration_ms":76261,"temperature":0.7,"pith_summary":"Visual anomaly detection — finding defects in images of manufactured objects — looks nearly solved on standard datasets, where the top method reports 99.6% AUROC. The paper argues this is misleading because current benchmarks keep viewpoint and illumination fixed or vary only one at a time, while real inspections face both changing together. To test that, it introduces M2AD, a benchmark of 119,880 high-resolution images of 999 specimens, each captured under 120 combinations of 12 viewpoints and 10 illumination conditions. On the two proposed protocols, M2AD-Synergy and M2AD-Invariant, state-of-the-art unsupervised methods degrade sharply; the same top method falls to 81.3% image-level AUROC on M2AD-Invariant. The paper concludes that benchmark saturation on standard industrial data overstates real-world readiness.","feed_headline":"M2AD benchmark exposes blind spot in anomaly detection","feed_subtitle":"Top detector drops from 99.6% AUROC to 81.3% when viewpoint and lighting vary together","key_machinery":"The load-bearing object is the M2AD dataset itself, built by a capture rig that combines a motorized turntable providing 12 angular views in 30-degree steps with ten programmable illumination configurations, giving 120 calibrated images per specimen at 3,648 by 5,472 resolution. The paper adds a detectability filter for the M2AD-Invariant protocol: only abnormal images that three supervised detection models independently flag at IoU of at least 0.3 with confidence of at least 0.5 are retained, so the invariant benchmark measures robustness on anomalies that are arguably visible. The two evaluation protocols then define what counts as success, with M2AD-Synergy aggregating predictions across configurations to test fusion and M2AD-Invariant testing single-image robustness. The synchronized factorial design is what carries the argument, because without it the joint effect of view and illumination could not be separated from ordinary dataset difficulty.","core_discovery":"The central discovery the paper tries to establish is that the view-illumination interplay, not any single imaging factor, is what breaks current visual anomaly detectors. On M2AD-Synergy, where a model can use all 120 configurations of a specimen, the best evaluated method, Dinomaly, reaches 90.0% object-level AUROC and 83.0% image-level AUROC; on M2AD-Invariant, which keeps single images but includes realistic view-illumination variation, the best method reaches only 81.3% image-level AUROC and 83.3% pixel-level AUPRO. These numbers contrast with the same method's 99.6% AUROC on MVTec AD. The paper also finds that simply averaging scores over more configurations does not help and can hurt, and that raising input resolution recovers several points, implying the remaining gap is partly about fine defect detail as well as about fusion.","pith_inferences":["The 12 by 10 factorial design also allows isolating how much of the failure is due to view changes versus illumination changes; the paper reports ablations on configuration count but not a full variance decomposition, so that decomposition is a natural next analysis.","If the detectability filter's IoU and confidence thresholds are miscalibrated, the M2AD-Invariant numbers would shift; a human re-annotation study of the excluded images would settle whether the filter is fair.","The synchronized multi-light capture means M2AD could double as a photometric-stereo benchmark, letting anomaly detectors consume estimated surface normals rather than raw pixel images; that is an extension the paper does not itself explore."],"forward_implications":["State-of-the-art methods that saturate standard benchmarks cannot be assumed deployment-ready; M2AD-style evaluation is needed to expose the gap.","Current score-averaging fusion strategies over multiple views and illuminations are insufficient, making feature-level or physics-informed fusion a concrete target.","High-resolution inputs recover meaningful performance, around 3 to 6 percentage points in object-level AUROC, so resolution cannot be treated as a free parameter in industrial inspection.","The dual sub-categories in each of the ten classes make M2AD a usable testbed for cross-category generalization and zero-shot or few-shot anomaly detection.","Methods that explicitly model illumination and geometry, such as photometric-stereo or multi-view-stereo inspired approaches, become testable on real synchronized data."],"supporting_citations":[{"why":"Supplies the standard benchmark whose near-saturated AUROC is the baseline that M2AD's difficulty is contrasted against.","marker":"[1]"},{"why":"Supplies VisA, another standard single-view dataset whose evaluation practice M2AD-Invariant follows.","marker":"[2]"},{"why":"Supplies Real-IAD, a multi-view dataset that M2AD compares against to show no previous dataset combines view and illumination.","marker":"[3]"},{"why":"Supplies PAD, a multi-view pose-agnostic benchmark representing the view-only line of prior work.","marker":"[4]"},{"why":"Supplies MVTec AD 2, a multi-illumination benchmark representing the illumination-only line of prior work.","marker":"[5]"},{"why":"Supplies Eyecandies, a synthetic multi-illumination dataset that M2AD contrasts with real captures.","marker":"[6]"},{"why":"Supplies Dinomaly, the top-performing method whose 99.6% MVTec AD result is the direct comparison point for M2AD degradation.","marker":"[8]"},{"why":"Supplies CDO, one of the evaluated knowledge-distillation methods whose scores on both M2AD protocols are tabulated.","marker":"[24]"},{"why":"Supplies RD++, a second evaluated method and one of the current state-of-the-art checks on M2AD-Synergy and M2AD-Invariant.","marker":"[37]"}],"fun_headline_variants":["M2AD benchmark: view-illumination interplay stumps anomaly detectors","Anomaly detection drops to 81% AUROC when view and light vary together","New M2AD dataset: 120 view-light combos crush state-of-the-art VAD","M2AD exposes anomaly detectors' blind spot: view-illumination interplay","Best detector 81.3% on M2AD vs 99.6% on standard benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The M2AD-Invariant protocol assumes that anomalies agreed on by three supervised detection models at IoU of at least 0.3 and confidence of at least 0.5 are exactly the anomalies an unsupervised method should be expected to detect; if that filter is wrong, all M2AD-Invariant difficulty numbers inherit the error.","fun_headline_variants_meta":{"raw":{"variants":["M2AD benchmark: view-illumination interplay stumps anomaly detectors","Anomaly detection drops to 81% AUROC when view and light vary together","New M2AD dataset: 120 view-light combos crush state-of-the-art VAD","M2AD exposes anomaly detectors' blind spot: view-illumination interplay","Best detector 81.3% on M2AD vs 99.6% on standard benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000265,"raw_usage":{"total_tokens":1613,"prompt_tokens":956,"completion_tokens":657,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":544}},"tokens_in":572,"tokens_out":657,"duration_ms":6753,"temperature":1.0,"reasoning_tokens":544,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:58:55.755540+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the abnormal images that the detectability filter excluded and ask independent human annotators to mark the anomalies without seeing the supervised-model predictions; if humans can reliably label a large fraction of the excluded images, then the IoU and confidence thresholds are not measuring human-detectable anomalies, and the M2AD-Invariant numbers would need to be recomputed on a differently filtered set.","supporting_citations":[{"cited_title":"The MVTec anomaly detection dataset: A comprehensive real-world dataset for unsupervised anomaly detection","cited_arxiv_id":null,"evidence_quote":"Supplies the standard benchmark whose near-saturated AUROC is the baseline that M2AD's difficulty is contrasted against."},{"cited_title":"Spot-the-difference self-supervised pre-training for anomaly detection and segmentation","cited_arxiv_id":null,"evidence_quote":"Supplies VisA, another standard single-view dataset whose evaluation practice M2AD-Invariant follows."},{"cited_title":"Real-iad: A real-world multi-view dataset for benchmarking versatile industrial anomaly detection","cited_arxiv_id":null,"evidence_quote":"Supplies Real-IAD, a multi-view dataset that M2AD compares against to show no previous dataset combines view and illumination."},{"cited_title":"Pad: A dataset and benchmark for pose-agnostic anomaly detection","cited_arxiv_id":null,"evidence_quote":"Supplies PAD, a multi-view pose-agnostic benchmark representing the view-only line of prior work."},{"cited_title":"The eyecandies dataset for unsupervised multimodal anomaly detection and localization","cited_arxiv_id":null,"evidence_quote":"Supplies Eyecandies, a synthetic multi-illumination dataset that M2AD contrasts with real captures."},{"cited_title":"Dinomaly: The less is more philosophy in multi-class unsupervised anomaly detection","cited_arxiv_id":null,"evidence_quote":"Supplies Dinomaly, the top-performing method whose 99.6% MVTec AD result is the direct comparison point for M2AD degradation."},{"cited_title":"Collaborative discrepancy optimization for reliable image anomaly localization","cited_arxiv_id":null,"evidence_quote":"Supplies CDO, one of the evaluated knowledge-distillation methods whose scores on both M2AD protocols are tabulated."}],"review_version":1}