{"id":"7af845a7-c398-4ba8-9be5-7d1d1f451e2e","arxiv_id":"2607.22783","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"AIC2026 is a large-scale benchmark of 9,618 fine-grained distorted images spanning 17 conventional and learned codec configurations across 20 CVVDP-based perceptual levels.","lead":"A new open dataset, AIC2026, provides 9,618 compressed images from 70 sources and 17 codec configurations for fine-grained image quality assessment. It is designed to test whether objective quality metrics can tell apart very small differences in high-fidelity compressed images.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.2-JND uniform spacing is an extrapolation of an AIC-3–fitted CVVDP power law; without subjective validation on AIC2026 items the central 'perceptually uniform levels' claim is unverified.","rationale":"The reader's weakest assumption identified the same load-bearing concern: the perceptual uniformity of the 20 distortion levels is inherited from a CVVDP-to-JND mapping fitted to AIC-3 and extrapolated beyond its calibration range. My analysis confirms this is the single most critical point on which the strongest claim rests. The mapping is global, two-parameter, and fitted to a small, codec-limited dataset; learning-based codec artifacts in AIC2026 are precisely the kind not represented in the calibration data. The paper explicitly discloses the extrapolation and the contrast-constancy limitation, but disclosure does not remove the gap: without subjective validation on AIC2026 items, 'perceptually uniform' remains a claim about CVVDP spacing, not about human perception. The dataset construction, source selection, codec coverage, and objective analysis are otherwise solid and reproducible; the release includes bitstreams, decoded images, and recipes. The central contribution—a large-scale, diverse, fine-grained compressed-image dataset—remains credible if the JND levels are reinterpreted as CVVDP-uniform rather than perceptually uniform. The planned future subjective study is the natural way to settle this, which is why the verdict should remain conditional rather than moving to acceptance or rejection. No other concern (e.g., IMD-based source selection, codec subset assignment, or inter-metric correlation analysis) appears to threaten the central claim as directly.","tokens_in":23721,"tokens_out":3863,"duration_ms":39920,"concrete_test":"Select a stratified sample of about 10 sources (covering the range of IMD scores and content categories) and 5 codecs spanning conventional and learning-based types (e.g., JPEG, JPEG XL, AVIF, JPEG AI HOP, Cool-Chic Wasserstein). For each source–codec pair, run an AIC-3-style subjective experiment (PTC/BTC) on all 20 distorted images plus the source, reconstruct perceptual JND scale values per pair, and compare them to the target 0.2-JND steps and to Eq. 2 predictions. If the mean absolute deviation between target and measured subjective JND exceeds ~0.1 JND, or if deviations are systematically larger for learning-based codecs or for levels >2.5 JND, the 'perceptually uniform' claim is not supported and the mapping/dataset levels must be recalibrated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of AIC2026 is that it provides 20 perceptually uniform distortion levels spanning 0.2–4.0 JND per source–codec pair. This depends entirely on Eq. 2, a power-law mapping from CVVDP to JND whose parameters (3.1889, 1.0129) were fitted to subjective data from the JPEG AIC-3 dataset. AIC-3 contains only 5 sources, 6 codecs, and about 295 distorted images, with subjective JND values that themselves were reconstructed via the AIC-3 PTC/BTC methodology. The mapping is global: it assumes the same CVVDP-to-JND relationship for all image content and all codecs, including learning-based codecs (Cool-Chic, FTIC, JPEG AI) that produce artifacts unlike those in the calibration set (e.g., spatial displacements, texture resampling, hallucinated details).\n\nThe paper itself acknowledges two critical facts: (1) AIC-3 subjective data span only about 0–2.5 JND, so all levels above 2.5 JND (half the target range) are extrapolated; and (2) Hammou et al. showed that CVVDP does not fully capture contrast constancy for supra-threshold distortions, so the metric is least reliable exactly in the extrapolated range. Because the 20 levels are selected by inverting Eq. 2, any bias in the mapping—for certain codecs, content types, or distortion magnitudes—directly violates the 'perceptually uniform spacing' part of the strongest claim. The dataset would still be a large-scale, diverse collection of compressed images with CVVDP-defined levels, but it would not be a perceptually calibrated fine-grained benchmark as claimed. The inter-metric disagreement analysis in Sec. V also uses CVVDP-estimated JND spacing; if the spacing is perceptually nonuniform, the granularity conclusions (Fig. 10) are relative to the metric's scale, not to human perception.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AIC2026, a large-scale dataset for fine-grained assessment of compressed image quality. It contains 70 source images selected from 2,787 candidates via DINOv2-based semantic clustering, inter-metric disagreement, and manual inspection. Each source is encoded with five base codecs and two codecs from an extended set, yielding 490 source–codec pairs in 17 coding configurations. For each pair, 20 distorted images are selected to approximate uniform 0.2-JND spacing over 0.2–4.0 JND, with JND values obtained from ColorVideoVDP scores through a power-law mapping fitted to the JPEG AIC-3 subjective dataset. The dataset comprises 9,618 distorted images and is publicly released. The paper further reports an extensive objective evaluation using 36 full-reference IQA metrics, documenting substantial inter-metric disagreement for fine-grained quality differences, particularly for learning-based codecs.","tokens_in":24184,"tokens_out":4187,"duration_ms":43550,"significance":"If the perceptual-spacing claim were validated, AIC2026 would be a uniquely large and diverse fine-grained benchmark covering both conventional and learning-based codecs, with substantial potential impact on IQA benchmarking and codec development. The manuscript has notable strengths: the dataset is publicly released with encoding recipes and an interactive visualization platform; the source-selection procedure is documented in detail; and the pairwise metric-correlation table is a useful reference. The paper is also honest about several limitations, including the extrapolated mapping and incomplete coverage for some codecs. However, the headline claim of \"20 perceptually uniform distortion levels\" rests on an unvalidated, extrapolated CVVDP-to-JND mapping; the current evidence supports a CVVDP-calibrated dataset rather than a perceptually calibrated one. With appropriate reframing or validation, this can still be a very valuable resource.","major_comments":[{"comment":"The central claim that the 20 levels are \"approximately perceptually uniform\" and span 0.2–4.0 JND rests entirely on the power-law mapping CVVDP_JND = 3.1889(10−CVVDP)^1.0129, fitted to AIC-3 subjective data. The paper itself notes that AIC-3 data cover only about 0–2.5 JND and that CVVDP does not fully capture supra-threshold contrast constancy [46]; levels above 2.5 JND are therefore extrapolated, and no AIC2026-specific subjective validation is provided. Some codecs (learning-based, WebP, palette PNG) do not even reach 0.2 JND. To support the headline claim, the authors should either add a validation experiment on a subset of AIC2026 using the AIC-3 PTC/BTC protocol, or consistently describe the levels as \"CVVDP-calibrated\" rather than \"perceptually uniform JND levels\" in the abstract, Section IV, and conclusion.","section":"Section IV-A, Eq. (2)"},{"comment":"The inter-metric-disagreement analysis in Fig. 10 is partly self-referential. The 20 distortion levels are selected to be uniformly spaced in CVVDP-mapped JND units via Eq. (2); for any metric that correlates strongly with CVVDP, the observed decrease in disagreement with coarser spacing is partly imposed by the selection procedure rather than an independent property of human perception or of the metric set. Please report the analysis with CVVDP excluded from the metric set, or at least provide per-codec / per-metric-group breakdowns, and state this caveat near Eq. (3). Without this, the conclusion that \"inter-metric disagreement increases as the spacing between distortion levels decreases\" is overstated.","section":"Section V, Eq. (3), Fig. 10"},{"comment":"The dataset coverage is uneven: the five base codecs cover all 70 sources, but each extended codec covers only 11–13 sources. Thus the 490 source–codec pairs are dominated by the base set, and cross-codec comparisons involving extended codecs are possible only on small subsets. The paper should include a precise table of achieved JND-range coverage per codec and per source–codec pair, including how many pairs reach the nominal 0.2 JND lower bound and 4.0 JND upper bound. The current Fig. 5 shows only aggregate medians, which can hide pairs where the target range is not achieved. This information is needed to calibrate all claims about \"20 perceptually uniform distortion levels spanning 0.2–4.0 JND\".","section":"Section IV-A and Table V"}],"minor_comments":[{"comment":"The relationship between the three processing categories and the final dimensions could be clearer. For example, how many of the 70 sources are exactly 840×944 versus larger? A short table listing final resolution ranges and counts would help.","section":"Section III-E / Table IV"},{"comment":"The codec name \"A VIF\" appears with a space in several places (e.g., Table I, Section IV). This should be corrected to \"AVIF\" consistently. Also, consider standardizing the codec acronyms used in Table V and the text.","section":"Throughout"},{"comment":"The notation uses m for both the number of subsets in Eq. (3) and the quality-score vectors m_s^i introduced in Eq. (1). Use a different symbol (e.g., M or R) for the number of subsets to avoid ambiguity.","section":"Section V, Eq. (3)"},{"comment":"The x-axis mixing raw CVVDP scores and JND-mapped values is visually confusing, especially because the mapping in Eq. (2) is nonlinear. Consider separate panels or clear dual-axis labeling.","section":"Fig. 1"},{"comment":"The full 36×36 correlation table is useful but very dense. Consider making it available in the supplementary material and keeping in the main text a compact version with only representative metrics, to improve readability.","section":"Table VI"}],"recommendation":"major_revision","confidential_remarks":"The dataset itself is a valuable contribution and the release is commendable. The main issue is a mismatch between the manuscript's headline claim of perceptually uniform JND spacing and the evidence provided: the spacing is defined by an extrapolated, codec-agnostic CVVDP mapping. This is fixable either by adding a targeted subjective validation or by reframing the claim as CVVDP-calibrated. I would not reject, but the revision must address this before the paper can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious, large-scale dataset that will be useful to the IQA and codec communities, and the paper is unusually candid about its main limitation. The headline claim—20 perceptually uniform distortion levels per source–codec pair spanning 0.2–4.0 JND—currently rests on Equation 2, a global power-law mapping from CVVDP to JND fitted to AIC-3 subjective data. AIC-3 covers only about 0–2.5 JND, so half the target range is extrapolated, and the calibration set did not include the artifacts typical of learned codecs (spatial displacements, texture resampling, hallucinated details). The paper says this out loud, and even cites Hammou et al. on CVVDP's supra-threshold weaknesses. That is honest and it tells readers exactly what to distrust.\n\nWhat is genuinely new: 70 sources, 9,618 distorted images, 17 codec configurations including four learned codecs, and 20 levels per source–codec pair. That is a real step up from AIC-3's 5 sources, 6 codecs, and 295 images. The source-selection pipeline—DINOv2 clustering plus inter-metric disagreement scoring—is sensible for building a stress test, and the release includes full encoding recipes, bitstreams, decoded images, fixed crops for subjective testing, and an interactive viewer. That is reproducible, formal, and exactly what an AIC-4 benchmark needs.\n\nThe soft spots are proportionate to the claims. The 'perceptually uniform' language should be read as 'uniform according to the CVVDP–JND mapping.' The Fig. 10 result that inter-metric disagreement grows as spacing shrinks may partly reflect the CVVDP scale itself, not a psychophysical fact. Also, the IMD-based source selection deliberately over-represents difficult, high-disagreement content; fine for benchmarking, but the dataset is not a random sample of natural images. These are disclosed, but they temper the central claim.\n\nOverall: a solid, important dataset, with a load-bearing caveat that is openly admitted. The authors promise a large-scale subjective study to validate the spacing; until then, downstream users should treat the levels as CVVDP-calibrated rather than psychophysically confirmed. I would send this to a serious referee, expect revisions around claim language, and cite it once it is published.","headline":"AIC2026 is the largest fine-grained compressed-image benchmark so far and it is built carefully and transparently; the main caveat is that its 'perceptually uniform' JND levels are defined by a CVVDP mapping fitted to earlier AIC-3 subjective data, extrapolated beyond the fitted range, so they are CVVDP-uniform until subjective validation lands.","tokens_in":24733,"tokens_out":3696,"would_cite":true,"duration_ms":37536,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces AIC2026 — 9,618 compressed images, 17 codec configurations, 20 fine-grained distortion levels from about 0.2 to 4.0 JND — and reports that objective quality metrics disagree on the subtlest differences, especially for l","keywords":["image quality assessment","image compression","fine-grained quality","just-noticeable difference","learning-based codecs","dataset","inter-metric disagreement","ColorVideoVDP"],"falsifier":"Run a fine-grained subjective JND study, using the same methodology that produced the calibration data, on a random subset of AIC2026 source-codec pairs, and compare the reconstructed perceptual scale values with the dataset's assigned CVVDP-JND levels. If adjacent levels do not come out roughly 0.2 JND apart, or if the subjective ordering disagrees with the assigned ordering for a substantial fraction of pairs, the perceptual-uniformity claim fails.","tokens_in":23615,"feed_emoji":"🖼️","tokens_out":5078,"duration_ms":51205,"temperature":0.7,"pith_summary":"The paper is trying to establish that fine-grained perceptual quality of compressed images can be measured and benchmarked at a scale and granularity not previously available. It presents AIC2026, a public dataset of 70 diverse sources encoded by eight conventional and four learning-based codecs, with 20 distortion levels per source-codec pair spanning roughly 0.2 to 4.0 just-noticeable differences (JND) — 9,618 distorted images in all. The levels are spaced using a power-law mapping from ColorVideoVDP scores to JND units, fitted to subjective data from the AIC-3 dataset. Using 36 objective IQA methods, the paper reports substantial disagreement among metrics when quality differences are small, particularly for artifacts produced by learned codecs. If the perceptual calibration holds, the dataset offers the community a principled testbed for deciding which metrics can actually resolve subtle quality differences.","feed_headline":"9,618 images test 17 codecs at 20 perceptual quality levels","feed_subtitle":"Quality metrics disagree most precisely where compressed images look alike, especially for AI codecs.","key_machinery":"The load-bearing machinery is the dataset-construction pipeline: (1) source selection via semantic clustering of deep visual features plus an inter-metric disagreement score (IMD), a rank-correlation-based measure of how differently objective metrics rank the distorted versions of a candidate image; (2) dense parameter sweeps across 17 coding configurations from eight conventional and four learning-based codecs; and (3) the CVVDP-to-JND mapping, a power law CVVDP_JND = 3.1889 (10 - CVVDP)^1.0129, fitted to AIC-3 subjective data and used to select 20 roughly equally spaced distortion levels per source-codec pair. That mapping is what converts raw metric scores into the claim of perceptually u","core_discovery":"To the authors' knowledge, AIC2026 is the first large-scale and diverse dataset for fine-grained compressed-image assessment covering both conventional and learning-based codecs. It provides 9,618 distorted images from 490 source-codec pairs, each with 20 decoded versions selected to approximate uniform 0.2-JND spacing between 0.2 and 4.0 JND as estimated by ColorVideoVDP. The central empirical finding is that state-of-the-art objective IQA metrics — 24 conventional and 12 learning-based — disagree strongly when ranking these fine-grained differences, and that disagreement grows as the spacing between distortion levels shrinks; the largest discrepancies involve artifacts from learning-based","pith_inferences":["A testable extension: if the 0.2-JND spacing is validated by subjective testing, the dataset could be used to recalibrate existing IQA metrics or train learned metrics specifically on fine-grained quality differences — a step the paper lists as future work but does not itself take.","The source-selection criterion of maximizing inter-metric disagreement could be reused to build fine-grained datasets for other distortion families, such as video, HDR, or screen content, where subtle artifacts also matter.","The observed growth of disagreement at fine spacing suggests that pairwise comparisons among adjacent levels, rather than correlation with mean opinion scores alone, should become a standard evaluation protocol for IQA metrics.","Because the 2.5–4.0 JND range is extrapolated, the high-distortion tail of the dataset is the most likely place for the perceptual-uniformity assumption to break; a targeted subjective check on those levels would be the decisive test."],"forward_implications":["AIC2026 supports fine-grained distortion-rate analysis across conventional and learning-based codecs at a granularity unavailable in earlier datasets.","Objective IQA metrics that agree on coarse distortions may fail to resolve 0.2-JND differences, so fine-grained datasets become necessary for benchmarking metrics.","The public release of bitstreams, decoded images, encoding recipes, and metric scores enables a reproducible benchmark for future subjective and objective studies.","The finding that inter-metric disagreement increases as level spacing decreases indicates that fine-grained evaluation is intrinsically harder and should be treated separately from coarse MOS evaluation.","The dataset is positioned as a testbed for JPEG AIC-4 objective evaluation and for future large-scale subjective studies following the AIC-3 methodology."],"fun_headline_variants":["New dataset: 9,618 images reveal IQA blind spots at fine quality gaps","Fine-grained codec test: 9,618 images, 17 codecs, metrics clash on AI artifacts","Dataset spotlights where image-quality metrics fail on AI codecs","9,618-image benchmark exposes metric disagreement at sub-JND levels","AIC2026: 70 sources, 20 distortion steps, IQA metrics split on AI codecs"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The 20 distortion levels are called 'perceptually uniform' because of a power-law formula that turns ColorVideoVDP scores into JND units, and that formula was fitted to subjective data covering only about 0–2.5 JND; if the mapping is wrong for learning-based codec artifacts or for the extrapolated 2.5–4.0 JND range, the fine-grained levels are CVVDP-uniform rather than truly perceptually uniform.","fun_headline_variants_meta":{"raw":{"variants":["New dataset: 9,618 images reveal IQA blind spots at fine quality gaps","Fine-grained codec test: 9,618 images, 17 codecs, metrics clash on AI artifacts","Dataset spotlights where image-quality metrics fail on AI codecs","9,618-image benchmark exposes metric disagreement at sub-JND levels","AIC2026: 70 sources, 20 distortion steps, IQA metrics split on AI codecs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000815,"raw_usage":{"total_tokens":3432,"prompt_tokens":793,"completion_tokens":2639,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":2527}},"tokens_in":537,"tokens_out":2639,"duration_ms":16027,"temperature":1.0,"reasoning_tokens":2527,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:30:25.690895+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a fine-grained subjective JND study, using the same methodology that produced the calibration data, on a random subset of AIC2026 source-codec pairs, and compare the reconstructed perceptual scale values with the dataset's assigned CVVDP-JND levels. If adjacent levels do not come out roughly 0.2 JND apart, or if the subjective ordering disagrees with the assigned ordering for a substantial fraction of pairs, the perceptual-uniformity claim fails.","supporting_citations":[],"review_version":1}