{"id":"e87a5a2d-33d7-4ad5-801f-2a06071ca00f","arxiv_id":"1908.08433","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A co-occurrence texture metric with block-level spatial structure is reported to match human perceptual rankings of facial sketches better than SSIM, FSIM, and other standard metrics on the authors' new human-judgment dataset.","lead":"Facial sketch similarity judgments from humans often disagree with standard image metrics. This paper introduces Scoot, a metric that combines block-level spatial layout with co-occurrence texture statistics, plus a new human-judgment dataset, and reports that Scoot matches human rankings better than existing metrics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Hyperparameters tuned on the same human judgments used for the final comparison make Scoot's reported advantage a fit, not a prediction; even in Table 2, Gabor beats Scoot on RCUFSF.","rationale":"The reader's weakest assumption already flags that hyperparameters were selected on the same evaluation criteria, and that is the load-bearing issue I would press. I do not need to question whether human judgments are valid in general: even granting the dataset, the final comparison is circular because the configuration was chosen to do well on it. A second, independent observation strengthens the concern: Table 2 shows Gabor at 80.9% versus Scoot at 78.8% on RCUFSF, so the claimed superiority is not consistent even within the reported numbers. A held-out or pre-registered evaluation with fixed hyperparameters would settle whether Scoot genuinely generalizes; without it, the central claim remains unsupported. I therefore keep the reader's reject verdict, though the metric and dataset could become useful if the held-out test succeeds.","tokens_in":15778,"tokens_out":7865,"duration_ms":82097,"concrete_test":"Hold out RCUFSF (or a randomly chosen half of each dataset) before any tuning. Select k, Nl, and the feature combination using only RCUFS (or the tuning half), freeze the configuration, and recompute Table 2 on the held-out data, reporting paired bootstrap confidence intervals for the difference in human-judgment agreement. If Scoot no longer beats the best alternative (for example, Gabor's 80.9% on RCUFSF) or the MM1/MM2/MM3 gains shrink materially, the current superiority claim is an artifact of tuning on the evaluation set.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Scoot exceeds prior metrics is not independently tested. Section 4.1 fixes k=4, Nl=6, and the CE feature set; Section 6 reports that these choices were made by applying the three meta-measures and the human judgments to each single feature, each pair, and the triple (H, C, E). These are the same RCUFS/RCUFSF judgments and meta-measures on which Table 2 then shows Scoot as the best row. The reported advantage is therefore partly a selection artifact, not an out-of-sample prediction. The table itself undermines the claim: on RCUFSF, the Gabor feature configuration reaches 80.9% human-judgment agreement while Scoot/CE reaches 78.8%, so Scoot does not exceed the best alternative on the larger judgment set. Because no held-out evaluation with frozen hyperparameters is provided, the evidence for the headline claim is a fit to the test criteria.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Scoot, a perceptual similarity metric for facial sketches that combines gray-level co-occurrence matrix statistics (Contrast and Energy) computed on a k x k block grid, with features averaged over four orientations after quantizing images to N_l gray levels. The authors introduce three meta-measures: stability to slight resizing, rotation sensitivity, and content capture capability. They also collect a two-alternative forced-choice (2AFC) human judgment dataset (RCUFS/RCUFSF) with about 152k judgments. Experiments in Table 2 compare Scoot with classical IQA metrics and with texture/edge features, reporting that Scoot achieves the best performance on the meta-measures and high agreement with human judgments.","tokens_in":15961,"tokens_out":3788,"duration_ms":35642,"significance":"If the central claim were established, the paper would provide a simple and potentially useful perceptual metric for face sketch evaluation, and the human judgment dataset plus the meta-measure suite would be valuable community resources. The manuscript deserves credit for collecting a large-scale human judgment dataset, evaluating a broad set of texture and edge features, and including sensitivity analyses for grid size and quantization. The main obstacle is that the reported advantage of Scoot is partly a selection artifact: its hyperparameters and feature combination were chosen using the same evaluation criteria on which it is subsequently judged. The claim is also contradicted by the paper's own Table 2 for one of the two human-judgment benchmarks. The contribution is therefore promising but not yet supported by an independent test.","major_comments":[{"comment":"The hyperparameters of Scoot are selected using the same benchmarks that later serve as evidence. Section 6 states that the three meta-measures and the human judgments were applied to test every single feature, every feature pair, and the triple, and Fig. 7 shows k=4 and N_l=6 chosen after inspecting MM1-MM3 and Jud curves. Consequently, the favorable results for Scoot/CE in Table 2 are a fit to the evaluation criteria, not an out-of-sample prediction. The paper needs a validation protocol in which k, N_l, and the feature set are frozen before evaluating on held-out datasets, or a nested cross-validation.","section":"Sec. 6, Table 2, Fig. 7"},{"comment":"The claim that Scoot exceeds prior work is not supported by Table 2 on the larger human-judgment benchmark. On RCUFSF, the Gabor feature configuration reaches 80.9% human-judgment agreement while Scoot/CE reaches 78.8%, and HE reaches 80.3%. The abstract and conclusion state unqualified superiority. This discrepancy must be addressed, either by restricting the claims to the specific meta-measures or by reporting a formal comparison that justifies the ordering across all criteria.","section":"Table 2"},{"comment":"The MM3 content-capture meta-measure assumes as ground truth that a complete SOTA synthetic sketch is always perceptually better than a 'light stroke' version produced by thresholding at gray level 170. This assumption is not validated with human judgments and may not hold for all reference sketches or synthesis algorithms. Since MM3 is one of the three pillars of the evaluation, the assumption needs justification, or the incomplete versions should be constructed and validated using human judgments.","section":"Sec. 4.2, Fig. 5"},{"comment":"The caption claims that all differences are statistically significant at the alpha<0.05 level, but no statistical test, sample size, variance estimate, or multiple-comparison correction is described anywhere in the manuscript. Without this information the reader cannot verify the significance claim, and the test would in any case be invalidated by the parameter selection on the same data. The authors should report a concrete test procedure and confidence intervals.","section":"Table 2 caption"}],"minor_comments":[{"comment":"The phrase 'quick assess' should be 'quickly assess', and there is a spacing error in 'metric,called'.","section":"Abstract"},{"comment":"The norm notation in Eq. (6) uses triple vertical bars with a subscript 2; this should be typeset consistently as a standard Euclidean norm.","section":"Sec. 3.3, Eq. (6)"},{"comment":"The caption refers to MM4, but only MM1-MM3 are defined in Sec. 4.2; clarify what MM4 denotes, likely the human-judgment measure.","section":"Fig. 7"},{"comment":"The sentence 'To increase an inherently noisy process' should be revised to something like 'To reduce the effect of noise in human judgments'.","section":"Sec. 5.3"},{"comment":"Step 4 refers to 'CE features' before the choice of C and E is explained; the algorithm should state that p=2 and that the selected statistics are Contrast and Energy, with the selection procedure described in Sec. 6.","section":"Algorithm 1"}],"recommendation":"major_revision","confidential_remarks":"The circularity concern is substantial and would prevent acceptance in the current form. I recommend asking the authors to provide a validation study with hyperparameters fixed a priori, to report per-dataset results including the Gabor comparison, and to validate the MM3 construction with human judgments. The dataset and meta-measure framework are potentially citable resources, so a major revision is warranted rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Name],\n\nQuick take: the paper is worth a look for the dataset and the meta-measure framework, but the headline claim—that Scoot beats existing metrics on human judgments—is not independently supported. The metric's hyperparameters and feature combination were selected on the same human judgments and meta-measures that later serve as evidence, making the reported advantage a fit, not a prediction.\n\nThe genuinely new contributions: Scoot is a simple, reproducible combination of block-level spatial structure and GLCM co-occurrence features; the three meta-measures (resizing stability, rotation sensitivity, content capture) are a reasonable way to probe what a metric responds to; and the RCUFS/RCUFSF dataset with ~152k human judgments is a real community resource. The paper is clearly written and the algorithm description is complete enough to reimplement.\n\nThe weak spot is the evaluation. In Section 6 the authors state that k=4, Nl=6, and the CE feature pair were chosen by applying the meta-measures and human judgments to all feature combinations. The final table then uses those same meta-measures and human judgments to show Scoot as the best. That is explicitly self-referential. Even within Table 2, the Gabor variant reaches 80.9% agreement on the larger RCUFSF set while Scoot/CE reaches 78.8%, so the chosen feature is not consistently best. The 'statistically significant at alpha<0.05' claim is also given without any test details.\n\nThe meta-measures themselves carry assumptions. MM3 presumes a complete synthetic sketch is always perceptually better than a thresholded 'light-stroke' version. That may be reasonable for face sketches, but it is a modeling choice, not a ground truth. The dataset curation filters for easy-to-rank pairs, which may inflate agreement numbers, although the filtering is at least transparent.\n\nWho should read this? Anyone building or evaluating perceptual metrics for sketches or stylized images. The dataset and meta-measure idea can be reused. The specific superiority claim for Scoot should be taken as a hypothesis until an external, held-out evaluation with frozen hyperparameters is done.\n\nMy recommendation: send it to referees. A serious review could push the authors to add a proper out-of-sample test, and the dataset itself justifies a look.\n\nBest.","headline":"Nice dataset and meta-measure idea, but Scoot's reported advantage is a fit to the same human judgments used to pick its hyperparameters—and even then Gabor beats it on the larger set.","tokens_in":16562,"tokens_out":3607,"would_cite":false,"duration_ms":31135,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims a co-occurrence-texture metric, Scoot, tracks human judgments of sketch similarity better than standard image-quality metrics.","keywords":["perceptual metric","face sketch synthesis","co-occurrence texture","spatial structure","image quality assessment","human perception","2AFC judgments","meta-measure"],"falsifier":"A held-out 2AFC sketch-similarity dataset built without the paper's selection rules, where the complete-versus-light-stroke ordering is not assumed in advance, would settle the claim: if a standard metric such as SSIM or FSIM, or a deep-feature metric, agreed with human choices at least as often as Scoot on such a benchmark, the claimed perceptual superiority would fail.","tokens_in":15544,"feed_emoji":"✏️","tokens_out":6804,"duration_ms":64163,"temperature":0.7,"pith_summary":"Facial-sketch comparison is a human perceptual task, yet the metrics commonly used to evaluate synthesized sketches such as SSIM and FSIM were built for photographic distortions and often rank sketches against human intuition. This paper tries to establish that a metric combining block-level spatial structure with co-occurrence texture statistics, named Scoot, tracks human perceptual choices far more closely, reaching about 76% agreement with human judgments on one benchmark and about 79% on another, versus roughly 25–59% for standard metrics. To support that claim, the paper proposes three meta-measures for resizing stability, rotation sensitivity, and content capture, and releases a 152k-judgment human-ranked sketch database. If the claim is right, automatic evaluation of face sketch synthesis can become substantially more perceptually aligned without a learned model.","feed_headline":"Scoot matches human sketch judgments 76% of the time","feed_subtitle":"Block-level spatial structure plus co-occurrence texture beats SSIM, FSIM, and VIF on face sketch similarity.","key_machinery":"The load-bearing device is the gray-level co-occurrence matrix $\\mathbf M$ computed on a quantized sketch $I'$, where $M(i,j)|_d$ counts how often gray value $i$ appears at a displacement $d$ from gray value $j$. The paper's contribution is to apply this classical texture statistic at block level: the image is split into a $k\\times k$ grid, the matrix is normalized within each block, and the Contrast and Energy statistics are concatenated across blocks and averaged over four displacement orientations to form a feature vector; similarity is then $E_s = 1/(1+\\|\\vec\\Psi(X'_s)-\\vec\\Psi(Y'_s)\\|_2)$. This machinery carries the argument because it makes the metric sensitive to stroke direction and regional content while staying insensitive to small resizing and rotation, exactly where pixel-level metrics fail.","core_discovery":"The paper's central discovery is that the perceptual gap in sketch evaluation comes from a mismatch of scale: pixel-level image-quality metrics miss the block-level spatial arrangement and stroke-direction statistics that human viewers use. Scoot measures similarity by quantizing each sketch to six gray levels, computing gray-level co-occurrence matrices inside a 4-by-4 grid of blocks, extracting Contrast and Energy statistics per block, averaging over four orientations, and comparing the resulting feature vectors with a Euclidean-distance-to-similarity mapping $E_s = 1/(1+\\|\\vec\\Psi(X'_s)-\\vec\\Psi(Y'_s)\\|_2)$. On the paper's three meta-measures and on its human two-alternative forced-choice databases, this construction outperforms IFC, SSIM, FSIM, VIF, and GMSD by wide margins. The paper reads that evidence as showing that spatial structure and co-occurrence texture are generally applicable perceptual features in face sketch synthesis.","pith_inferences":["As an editorial extension, the same construction could be tested on non-face line drawings and stylized illustrations; the paper's argument only requires texture and spatial structure to be the relevant perceptual cues, which is plausible but unverified there.","Because the grid size, quantization level, and feature pair were selected on the same benchmarks used for evaluation, an independent evaluation on a separate dataset is needed before treating the reported margin as the metric's true advantage.","The meta-measures themselves could serve as a reusable protocol: any future sketch-similarity metric, learned or hand-crafted, could be reported against the same three properties, making cross-paper comparisons direct.","A natural extension is to let a learned deep network consume the block-level co-occurrence features as input; the paper shows hand-crafted features already carry a large share of the perceptual signal, so a hybrid might combine their stability with learned flexibility."],"forward_implications":["Face sketch synthesis papers could evaluate quality with a metric that, on the paper's human-judgment data, agrees with people roughly 26 percentage points more often than the best standard metric tested.","Slight resizing and rotation of reference sketches, common when artist sketches do not align with photos, would no longer flip the ranking of synthesis algorithms the way they do for SSIM, VIF, and GMSD.","The block-level co-occurrence construction offers a simple, non-learned alternative to learned perceptual metrics for tasks where texture and stroke direction matter.","Complete sketches that preserve hair, eyes, and other facial texture are ranked above incomplete light-stroke versions, matching the expectation that content capture is part of perceptual quality.","The released human-judgment databases give future sketch-synthesis methods a way to test their evaluation choices against actual human perception."],"supporting_citations":[{"why":"Supplies the SSIM baseline whose pixel-level comparisons fail on sketches and which Scoot outperforms.","marker":"[66]"},{"why":"Supplies the FSIM baseline, a widely used structural metric that also underperforms on the human-judgment test.","marker":"[76]"},{"why":"Supplies the VIF baseline, an information-theoretic metric that ranks light-stroke sketches too high.","marker":"[39]"},{"why":"Supplies the IFC baseline, an information-fidelity metric included in the comparison.","marker":"[40]"},{"why":"Supplies the GMSD baseline, a recent gradient-based image-quality metric included in the comparison.","marker":"[72]"},{"why":"Provides the co-occurrence matrix as the texture extractor at the core of Scoot.","marker":"[25]"},{"why":"Motivates gray-level quantization to reduce sensitivity to minor intensity variations.","marker":"[6]"},{"why":"Motivates the 2AFC human-perception database and the use of perceptual similarity judgments.","marker":"[80]"},{"why":"Supplies the meta-measure methodology used to evaluate metrics.","marker":"[37]"},{"why":"Inspires the block-level spatial-structure treatment for perceptual measures.","marker":"[9]"}],"fun_headline_variants":["Scoot metric beats SSIM, FSIM on face sketch perception","Scoot aligns with human sketch judgment 76%","Structure and texture statistics beat pixel metrics for sketches","Scoot uses co-occurrence texture and structure for sketch similarity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on treating the three meta-measures and the newly collected human rankings as valid, unbiased yardsticks for perceptual similarity — in particular, on assuming a complete synthetic sketch is always perceptually better than a light-stroke thresholded version — and on assuming the settings tuned on those same benchmarks will keep their advantage on other sketches.","fun_headline_variants_meta":{"raw":{"variants":["Scoot metric beats SSIM, FSIM on face sketch perception","Scoot aligns with human sketch judgment 76%","Structure and texture statistics beat pixel metrics for sketches","Scoot uses co-occurrence texture and structure for sketch similarity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000937,"raw_usage":{"total_tokens":4001,"prompt_tokens":932,"completion_tokens":3069,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":2999}},"tokens_in":548,"tokens_out":3069,"duration_ms":20345,"temperature":1.0,"reasoning_tokens":2999,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:55:27.822473+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A held-out 2AFC sketch-similarity dataset built without the paper's selection rules, where the complete-versus-light-stroke ordering is not assumed in advance, would settle the claim: if a standard metric such as SSIM or FSIM, or a deep-feature metric, agreed with human choices at least as often as Scoot on such a benchmark, the claimed perceptual superiority would fail.","supporting_citations":[{"cited_title":"Image quality assessment: from error visibility to structural similarity","cited_arxiv_id":null,"evidence_quote":"Supplies the SSIM baseline whose pixel-level comparisons fail on sketches and which Scoot outperforms."},{"cited_title":"FSIM: A feature similarity index for image quality assess- ment","cited_arxiv_id":null,"evidence_quote":"Supplies the FSIM baseline, a widely used structural metric that also underperforms on the human-judgment test."},{"cited_title":"Image information and visual quality","cited_arxiv_id":null,"evidence_quote":"Supplies the VIF baseline, an information-theoretic metric that ranks light-stroke sketches too high."},{"cited_title":"An information ﬁdelity criterion for image quality assess- ment using natural scene statistics","cited_arxiv_id":null,"evidence_quote":"Supplies the IFC baseline, an information-fidelity metric included in the comparison."},{"cited_title":"Gradient magnitude similarity deviation: A highly efﬁcient perceptual image quality index","cited_arxiv_id":null,"evidence_quote":"Supplies the GMSD baseline, a recent gradient-based image-quality metric included in the comparison."},{"cited_title":"Textu- ral features for image classiﬁcation","cited_arxiv_id":null,"evidence_quote":"Provides the co-occurrence matrix as the texture extractor at the core of Scoot."},{"cited_title":"The unreasonable effectiveness of deep features as a perceptual metric","cited_arxiv_id":null,"evidence_quote":"Motivates the 2AFC human-perception database and the use of perceptual similarity judgments."},{"cited_title":"Measures and meta- measures for the supervised evaluation of image segmenta- tion","cited_arxiv_id":null,"evidence_quote":"Supplies the meta-measure methodology used to evaluate metrics."}],"review_version":1}