{"id":"258ac206-ee34-4fae-a0f5-73df22cfd6b6","arxiv_id":"2607.12352","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"Filtering out low-quality images with an IQA metric and threshold beats standard denoising for traffic-sign (93.8%) and object recognition (84.9%) accuracy.","lead":"The paper proposes discarding poor-quality images via an image-quality score and a threshold, rather than denoising them, before training recognition models. If the reported gains hold, it is a simple data-prep alternative for traffic-sign and object recognition under real-world noise.","discovery_kind":"incremental","skeptic_critique":{"model":"grok-4.5","headline":"Abstract-only review cannot verify the load-bearing empirical claim; threshold selection and metric identity remain unspecified, so circularity and generalizability cannot be assessed.","rationale":"The Reader correctly flags that an abstract-only review leaves the empirical claim unverifiable and that the unspecified optimum threshold is the weakest assumption. No additional internal contradiction can be demonstrated without the full text; manufacturing one would violate the good-faith rule. The appropriate posture is therefore to leave the verdict UNVERDICTED with low confidence, exactly as the Reader did. The concrete test above is the minimal check that would settle whether the circularity concern actually lands once methods become available. If that check later shows a fixed, transferable threshold and held-out gains, the verdict can be revised upward; until then it stays UNVERDICTED.","tokens_in":2114,"tokens_out":507,"duration_ms":4385,"concrete_test":"Obtain the full paper (or author code/data). Identify the exact IQA metric and the procedure used to set the threshold. Re-run the traffic-sign and object-recognition pipelines with that threshold fixed a priori (or chosen only on a held-out noise split) and report accuracy on a completely unseen noise condition. If accuracy drops below the denoising baselines or the retained set becomes too small for stable training, the headline claim does not hold under non-circular evaluation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that discarding images below an 'optimum threshold' on an IQA metric yields higher recognition accuracy than SOTA denoising (93.8% traffic-sign, 84.9% object recognition) while retaining enough data for DL training. Because only the abstract is available, three conditions required for that claim cannot be checked: (1) which IQA metric is used and whether it is fixed or chosen post-hoc; (2) how the threshold is selected—held-out validation, fixed a priori, or tuned on the same accuracy numbers reported; (3) whether the retained subset is large enough and representative under diverse noise, or whether the gains simply reflect easier remaining examples. Without methods, baselines, error bars, or code, the reported percentages cannot be treated as evidence that filtering generalizes better than denoising. The reader's circularity concern is therefore still material and is the single most load-bearing soft spot.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript (available only as an abstract) proposes replacing conventional image denoising with a discard-based data-preparation step: images are scored by an (unnamed) image quality assessment (IQA) metric, those below an “optimum threshold” are removed, and the retained set is required to remain large enough to train a deep recognition model. On real and simulated traffic-sign and object-recognition data the authors report average recognition accuracies of 93.8% and 84.9%, respectively, claimed to exceed state-of-the-art denoising baselines, with intended application to autonomous driving.","tokens_in":2282,"tokens_out":922,"duration_ms":20250,"significance":"If the filtering pipeline were shown to be metric-specified, threshold-validated on held-out data, and superior to named denoising baselines under controlled discard rates, the result would be a simple, practically useful alternative to denoising for recognition pipelines that can tolerate reduced sample size. The abstract’s emphasis on retaining enough data for DL training and on real-life AV use cases is therefore potentially consequential. Because the full methods, baselines, and validation protocol are not supplied, that significance cannot yet be credited as demonstrated.","major_comments":[{"comment":"The central control is an “optimum threshold” on an IQA score. The abstract neither names the IQA metric nor states how the threshold is chosen (a priori, held-out validation, or tuned to the same recognition accuracy later reported). Without that procedure the headline gains (93.8%, 84.9%) cannot be distinguished from post-hoc cutoff fitting on the evaluation metric, which is the load-bearing circularity risk for the claim that filtering outperforms denoising.","section":"Abstract"},{"comment":"The claim of “performance supremacy … compared with the state-of-the-art approaches” is unsupported in the provided text: no denoising baselines are named, no train/val/test protocol or split is given, and no error bars, significance tests, or discard-rate / retained-N figures are reported. The two point accuracies therefore cannot be treated as evidence that filtering generalizes better than denoising across environmental and camera noise.","section":"Abstract"},{"comment":"The abstract asserts that “a sufficient number of images remain to develop the deep learning (DL) model” but supplies no retained-set sizes, class-balance checks, or ablation of accuracy versus discard rate. Without those quantities it is impossible to verify that the retained subset is both large enough and representative rather than an easier residual of high-quality examples.","section":"Abstract"},{"comment":"Only the abstract was available for review; methods, equations, tables, figures, and code are absent. Under these conditions the empirical central claim cannot be verified or falsified, so a definitive accept/reject decision on the full manuscript is not possible from the material provided.","section":"Abstract (scope of review)"}],"minor_comments":[{"comment":"The phrase “for the first time” is a strong novelty claim that should be backed by a brief related-work contrast once the full text is available; as written it is unsubstantiated.","section":"Abstract"},{"comment":"“Real and simulated traffic and object recognition data” should name the datasets (or describe the simulation) so that the reported percentages can be contextualized against published numbers.","section":"Abstract"},{"comment":"The abstract conflates “filtering noise” (denoising) with “filtering out poor-quality images” (sample rejection); a clearer terminological distinction would reduce reader confusion.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"Full text was not available; this is an abstract-only review. The circularity concern around the optimum threshold is material and cannot be cleared without methods. If the editor can obtain the full PDF, a re-review is warranted; if the submission is effectively abstract-only, desk rejection for insufficient methodological disclosure is appropriate. Fit for a serious cs.CV venue is doubtful until metric identity, threshold protocol, named baselines, and retained-N ablations are shown."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is an abstract-only cs.CV data-prep paper. The one thing to know is that the load-bearing claim is empirical and currently unverifiable: discard images below an “optimum” IQA threshold (while keeping enough for DL training) and you beat SOTA denoising on traffic-sign and object recognition (93.8% and 84.9% average accuracy). Filtering bad samples is not a new principle—standard hygiene—but the concrete package (IQA + threshold as a substitute for denoising on these domains, with those numbers) is what they are selling, and the “for the first time” framing is overstated relative to common practice.\n\nWhat they do well, on the face of it, is state a clear practical alternative and report point results on real and simulated traffic/object data with an eye toward AV-style use. If the full paper later names a fixed metric, a held-out or transferable threshold rule, baselines, and enough retained data under diverse noise, that would be a legitimate engineering contribution worth citing for data-prep pipelines.\n\nThe soft spots are real and proportional to what we can see. The free parameter is the optimum quality threshold. The abstract does not say which IQA metric, how the threshold is chosen, whether it is tuned on the same accuracy numbers, or how much data remains under different noise regimes. That is the circularity risk the stress-test flags, and it lands: without that protocol the reported gains could partly be “easier remaining examples.” No error bars, ablations, or named baselines either. Soundness and generalizability are therefore open, not refuted.\n\nWho it is for: people building recognition under environmental/camera noise who care about prep recipes more than theory. A serious referee should see the full methods if the journal wants applied CV work; I would not desk-reject solely on the abstract, but I would not cite or bring it to reading group until the threshold story and metric are fixed and checkable. Treat the percentages as claims, not evidence, until then.","headline":"Abstract-only empirical recipe for IQA-based discard vs denoising; useful idea, but threshold selection and metric identity are unchecked so the accuracy claims cannot be treated as settled evidence.","tokens_in":2930,"tokens_out":518,"would_cite":false,"duration_ms":4738,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Filtering out low-quality images beats denoising for recognition accuracy on traffic and object data.","keywords":["image quality assessment","data preparation","image filtering","deep learning","traffic sign recognition","object recognition","noise handling","autonomous vehicles"],"falsifier":"Re-run the identical recognition pipelines on the same traffic-sign and object datasets after selecting the quality threshold only on a held-out validation split, then compare accuracy against the same denoisers; if the filtered approach no longer outperforms, the claim fails.","tokens_in":2960,"feed_emoji":"📷","tokens_out":499,"duration_ms":3814,"temperature":0.7,"pith_summary":"This paper argues that the usual first step of cleaning noisy images for deep learning is the wrong move. Instead of applying median, Gaussian, bilateral, or CNN denoisers that only handle some kinds of environmental and camera noise and force resizing, the authors filter out poor-quality images entirely. Quality is scored with an image-quality metric; an optimum threshold drops the bad ones while still leaving enough samples to train a model. On real and simulated traffic-sign and object-recognition data the approach yields higher average recognition accuracy than state-of-the-art denoisers (93.8 % and 84.9 % respectively). If the claim holds, practitioners preparing vision datasets for autonomous vehicles and similar systems can simply discard rather than repair, gaining both accuracy and simplicity.","feed_headline":"Filter bad images, don't denoise them, for higher recognition accuracy","feed_subtitle":"Discarding low-quality samples beats state-of-the-art denoisers on traffic-sign and object data","key_machinery":"An image-quality-assessment (IQA) metric paired with a single optimum threshold that discards low-scoring images yet keeps a sufficient training set for the downstream deep-learning model.","core_discovery":"Filtering out poor-quality images with an image-quality-assessment metric and an optimum threshold, while retaining enough images for deep-learning training, produces higher recognition accuracy than state-of-the-art denoising methods across diverse environmental and camera noise.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Filter poor images out, skip denoising, for higher recognition","Discard low-quality images via IQA threshold to beat denoisers","Filter bad samples with quality metric, outperform SOTA denoising","IQA threshold drops poor images, lifts traffic and object accuracy","Keep enough good images after filtering to top denoising methods"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That a single (or per-dataset) optimum threshold on the quality metric can be chosen so that the remaining images both suffice for training and deliver fair, generalizable gains over denoising without being tuned to the same evaluation that reports the accuracy numbers.","fun_headline_variants_meta":{"raw":{"variants":["Filter poor images out, skip denoising, for higher recognition","Discard low-quality images via IQA threshold to beat denoisers","Filter bad samples with quality metric, outperform SOTA denoising","IQA threshold drops poor images, lifts traffic and object accuracy","Keep enough good images after filtering to top denoising methods"]},"model":"grok-4.5","effort":"low","cost_usd":0.004056,"raw_usage":{"total_tokens":1274,"prompt_tokens":801,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":40560000,"prompt_tokens_details":{"text_tokens":801,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":399,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":801,"tokens_out":74,"duration_ms":4039,"temperature":1.0,"reasoning_tokens":399,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T06:46:38.540888+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the identical recognition pipelines on the same traffic-sign and object datasets after selecting the quality threshold only on a held-out validation split, then compare accuracy against the same denoisers; if the filtered approach no longer outperforms, the claim fails.","supporting_citations":[],"review_version":1}