Pith. sign in

REVIEW 3 major objections 5 minor 79 references

The paper introduces AIC2026 — 9,618 compressed images, 17 codec configurations, 20 fine-grained distortion levels from about 0.2 to 4.0 JND — and reports that objective quality metrics disagree on the subtlest differences, especially for l

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 05:30 UTC pith:MIKUGXQA

load-bearing objection AIC2026 is the largest fine-grained compressed-image benchmark so far and it is built carefully and transparently; the main caveat is that its 'perceptually uniform' JND levels are defined by a CVVDP mapping fitted to earlier AIC-3 subjective data, extrapolated beyond the fitted range, so they are CVVDP-uniform until subjective validation lands. the 3 major comments →

arxiv 2607.22783 v1 pith:MIKUGXQA submitted 2026-07-24 eess.IV cs.CV

JPEG AIC2026: A large-scale dataset for fine-grained assessment of image coding

classification eess.IV cs.CV
keywords image quality assessmentimage compressionfine-grained qualityjust-noticeable differencelearning-based codecsdatasetinter-metric disagreementColorVideoVDP
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper is trying to establish that fine-grained perceptual quality of compressed images can be measured and benchmarked at a scale and granularity not previously available. It presents AIC2026, a public dataset of 70 diverse sources encoded by eight conventional and four learning-based codecs, with 20 distortion levels per source-codec pair spanning roughly 0.2 to 4.0 just-noticeable differences (JND) — 9,618 distorted images in all. The levels are spaced using a power-law mapping from ColorVideoVDP scores to JND units, fitted to subjective data from the AIC-3 dataset. Using 36 objective IQA methods, the paper reports substantial disagreement among metrics when quality differences are small, particularly for artifacts produced by learned codecs. If the perceptual calibration holds, the dataset offers the community a principled testbed for deciding which metrics can actually resolve subtle quality differences.

Core claim

To the authors' knowledge, AIC2026 is the first large-scale and diverse dataset for fine-grained compressed-image assessment covering both conventional and learning-based codecs. It provides 9,618 distorted images from 490 source-codec pairs, each with 20 decoded versions selected to approximate uniform 0.2-JND spacing between 0.2 and 4.0 JND as estimated by ColorVideoVDP. The central empirical finding is that state-of-the-art objective IQA metrics — 24 conventional and 12 learning-based — disagree strongly when ranking these fine-grained differences, and that disagreement grows as the spacing between distortion levels shrinks; the largest discrepancies involve artifacts from learning-based

What carries the argument

The load-bearing machinery is the dataset-construction pipeline: (1) source selection via semantic clustering of deep visual features plus an inter-metric disagreement score (IMD), a rank-correlation-based measure of how differently objective metrics rank the distorted versions of a candidate image; (2) dense parameter sweeps across 17 coding configurations from eight conventional and four learning-based codecs; and (3) the CVVDP-to-JND mapping, a power law CVVDP_JND = 3.1889 (10 - CVVDP)^1.0129, fitted to AIC-3 subjective data and used to select 20 roughly equally spaced distortion levels per source-codec pair. That mapping is what converts raw metric scores into the claim of perceptually u

Load-bearing premise

The 20 distortion levels are called 'perceptually uniform' because of a power-law formula that turns ColorVideoVDP scores into JND units, and that formula was fitted to subjective data covering only about 0–2.5 JND; if the mapping is wrong for learning-based codec artifacts or for the extrapolated 2.5–4.0 JND range, the fine-grained levels are CVVDP-uniform rather than truly perceptually uniform.

What would settle it

Run a fine-grained subjective JND study, using the same methodology that produced the calibration data, on a random subset of AIC2026 source-codec pairs, and compare the reconstructed perceptual scale values with the dataset's assigned CVVDP-JND levels. If adjacent levels do not come out roughly 0.2 JND apart, or if the subjective ordering disagrees with the assigned ordering for a substantial fraction of pairs, the perceptual-uniformity claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • AIC2026 supports fine-grained distortion-rate analysis across conventional and learning-based codecs at a granularity unavailable in earlier datasets.
  • Objective IQA metrics that agree on coarse distortions may fail to resolve 0.2-JND differences, so fine-grained datasets become necessary for benchmarking metrics.
  • The public release of bitstreams, decoded images, encoding recipes, and metric scores enables a reproducible benchmark for future subjective and objective studies.
  • The finding that inter-metric disagreement increases as level spacing decreases indicates that fine-grained evaluation is intrinsically harder and should be treated separately from coarse MOS evaluation.
  • The dataset is positioned as a testbed for JPEG AIC-4 objective evaluation and for future large-scale subjective studies following the AIC-3 methodology.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A testable extension: if the 0.2-JND spacing is validated by subjective testing, the dataset could be used to recalibrate existing IQA metrics or train learned metrics specifically on fine-grained quality differences — a step the paper lists as future work but does not itself take.
  • The source-selection criterion of maximizing inter-metric disagreement could be reused to build fine-grained datasets for other distortion families, such as video, HDR, or screen content, where subtle artifacts also matter.
  • The observed growth of disagreement at fine spacing suggests that pairwise comparisons among adjacent levels, rather than correlation with mean opinion scores alone, should become a standard evaluation protocol for IQA metrics.
  • Because the 2.5–4.0 JND range is extrapolated, the high-distortion tail of the dataset is the most likely place for the perceptual-uniformity assumption to break; a targeted subjective check on those levels would be the decisive test.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces AIC2026, a large-scale dataset for fine-grained assessment of compressed image quality. It contains 70 source images selected from 2,787 candidates via DINOv2-based semantic clustering, inter-metric disagreement, and manual inspection. Each source is encoded with five base codecs and two codecs from an extended set, yielding 490 source–codec pairs in 17 coding configurations. For each pair, 20 distorted images are selected to approximate uniform 0.2-JND spacing over 0.2–4.0 JND, with JND values obtained from ColorVideoVDP scores through a power-law mapping fitted to the JPEG AIC-3 subjective dataset. The dataset comprises 9,618 distorted images and is publicly released. The paper further reports an extensive objective evaluation using 36 full-reference IQA metrics, documenting substantial inter-metric disagreement for fine-grained quality differences, particularly for learning-based codecs.

Significance. If the perceptual-spacing claim were validated, AIC2026 would be a uniquely large and diverse fine-grained benchmark covering both conventional and learning-based codecs, with substantial potential impact on IQA benchmarking and codec development. The manuscript has notable strengths: the dataset is publicly released with encoding recipes and an interactive visualization platform; the source-selection procedure is documented in detail; and the pairwise metric-correlation table is a useful reference. The paper is also honest about several limitations, including the extrapolated mapping and incomplete coverage for some codecs. However, the headline claim of "20 perceptually uniform distortion levels" rests on an unvalidated, extrapolated CVVDP-to-JND mapping; the current evidence supports a CVVDP-calibrated dataset rather than a perceptually calibrated one. With appropriate reframing or validation, this can still be a very valuable resource.

major comments (3)
  1. [Section IV-A, Eq. (2)] The central claim that the 20 levels are "approximately perceptually uniform" and span 0.2–4.0 JND rests entirely on the power-law mapping CVVDP_JND = 3.1889(10−CVVDP)^1.0129, fitted to AIC-3 subjective data. The paper itself notes that AIC-3 data cover only about 0–2.5 JND and that CVVDP does not fully capture supra-threshold contrast constancy [46]; levels above 2.5 JND are therefore extrapolated, and no AIC2026-specific subjective validation is provided. Some codecs (learning-based, WebP, palette PNG) do not even reach 0.2 JND. To support the headline claim, the authors should either add a validation experiment on a subset of AIC2026 using the AIC-3 PTC/BTC protocol, or consistently describe the levels as "CVVDP-calibrated" rather than "perceptually uniform JND levels" in the abstract, Section IV, and conclusion.
  2. [Section V, Eq. (3), Fig. 10] The inter-metric-disagreement analysis in Fig. 10 is partly self-referential. The 20 distortion levels are selected to be uniformly spaced in CVVDP-mapped JND units via Eq. (2); for any metric that correlates strongly with CVVDP, the observed decrease in disagreement with coarser spacing is partly imposed by the selection procedure rather than an independent property of human perception or of the metric set. Please report the analysis with CVVDP excluded from the metric set, or at least provide per-codec / per-metric-group breakdowns, and state this caveat near Eq. (3). Without this, the conclusion that "inter-metric disagreement increases as the spacing between distortion levels decreases" is overstated.
  3. [Section IV-A and Table V] The dataset coverage is uneven: the five base codecs cover all 70 sources, but each extended codec covers only 11–13 sources. Thus the 490 source–codec pairs are dominated by the base set, and cross-codec comparisons involving extended codecs are possible only on small subsets. The paper should include a precise table of achieved JND-range coverage per codec and per source–codec pair, including how many pairs reach the nominal 0.2 JND lower bound and 4.0 JND upper bound. The current Fig. 5 shows only aggregate medians, which can hide pairs where the target range is not achieved. This information is needed to calibrate all claims about "20 perceptually uniform distortion levels spanning 0.2–4.0 JND".
minor comments (5)
  1. [Section III-E / Table IV] The relationship between the three processing categories and the final dimensions could be clearer. For example, how many of the 70 sources are exactly 840×944 versus larger? A short table listing final resolution ranges and counts would help.
  2. [Throughout] The codec name "A VIF" appears with a space in several places (e.g., Table I, Section IV). This should be corrected to "AVIF" consistently. Also, consider standardizing the codec acronyms used in Table V and the text.
  3. [Section V, Eq. (3)] The notation uses m for both the number of subsets in Eq. (3) and the quality-score vectors m_s^i introduced in Eq. (1). Use a different symbol (e.g., M or R) for the number of subsets to avoid ambiguity.
  4. [Fig. 1] The x-axis mixing raw CVVDP scores and JND-mapped values is visually confusing, especially because the mapping in Eq. (2) is nonlinear. Consider separate panels or clear dual-axis labeling.
  5. [Table VI] The full 36×36 correlation table is useful but very dense. Consider making it available in the supplementary material and keeping in the main text a compact version with only representative metrics, to improve readability.

Circularity Check

0 steps flagged

No significant circularity: the JND calibration is an external fit to AIC-3 subjective data, and AIC2026's construction is not a derivation from its own outputs.

full rationale

The paper's derivation chain is not circular. The 20 distortion levels are produced by sweeping encoder parameters and then selecting decoded images whose CVVDP scores, mapped to JND by Eq. 2, match targets of 0.2–4.0 JND. Eq. 2 is fitted to subjective scores from the AIC-3 dataset, which is external human-subject data published by an overlapping group but not derived from AIC2026 or from the conclusions of this paper. The dataset is therefore built on an external calibration, not on its own evaluation. The paper explicitly discloses the limitation of this calibration: 'Because the AIC-3 subjective data span approximately within 0–2.5 JND, mapped values beyond this range are extrapolated. Also, Hammou et al. [46] showed that CVVDP does not fully capture contrast constancy for supra-threshold distortions, limiting its accuracy for higher distortion levels.' It also states that true subjective validation is deferred to future work ('a large-scale crowdsourced subjective study following the JPEG AIC-3 methodology will be conducted'). Fig. 10's x-axis is labeled 'CVVDP-estimated JND units,' so the granularity analysis is openly a function of the calibration scale rather than a hidden re-importation of the conclusion. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported, and no ansatz is smuggled in via citation. The central risks—extrapolation beyond the calibrated range and possible inadequacy of CVVDP for learned-codec artifacts—are validity threats, not circularity.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 0 invented entities

No new physical entities are introduced. The dataset is a collection of real compressed images, not a new theoretical object. The main load-bearing free parameters are the CVVDP-to-JND mapping coefficients, which determine the perceptual spacing of the distortion levels.

free parameters (1)
  • CVVDP-to-JND mapping scale and exponent = 3.1889, 1.0129
    Parameters in Eq. 2 fitted by nonlinear least squares to AIC-3 subjective scores; used to convert CVVDP quality scores to JND units and select the 20 distortion levels per source-codec pair.
axioms (4)
  • ad hoc to paper CVVDP scores can be mapped to perceptual JND units via a power law calibrated on AIC-3 subjective data
    Eq. 2 uses parameters fitted to AIC-3 data; this mapping is assumed to generalize to all 17 codecs and 70 sources. Section IV-A acknowledges that mapped values beyond 0-2.5 JND are extrapolated.
  • domain assumption The 17 codec configurations and encoding settings are representative of the relevant compression artifacts
    Section IV-A lists configurations chosen with expert consultation, but there is no subjective evidence that they cover all artifacts of interest, especially for learned codecs.
  • ad hoc to paper Inter-metric disagreement on a 4-codec, 7-quality subset predicts informativeness of source images for the full codec set
    Section III-C uses IMD on JPEG/JPEG2000/JPEGXL/AVIF and 7 metrics to select source images; this assumes the proxy correlates with disagreement on the final 17-codec dataset.
  • domain assumption Manual inspection and replacement of candidate images does not introduce bias
    Section III-D describes manual replacement of images with defects or sensitive content; the final 70-source set is partly subjectively curated.

pith-pipeline@v1.3.0-alltime-deepseek · 23424 in / 8524 out tokens · 81860 ms · 2026-08-01T05:30:25.690895+00:00 · methodology

0 comments
read the original abstract

Recent advances in conventional and learning-based image coding have increased the demand for benchmark datasets that support fine-grained assessment of compressed image quality, particularly for learning-based image compression methods. This paper introduces Assessment of Image Coding 2026 (AIC2026), a large-scale dataset for high-fidelity image compression containing 70 source images selected from 2,787 candidates using semantic clustering, inter-metric disagreement among objective image quality assessment (IQA) methods, and manual inspection and refinement. The dataset covers a wide range of compression artifacts produced by eight conventional and four learning-based codecs across 17 coding configurations. Each source image is encoded using seven codecs. For each source-codec pair, decoded images are provided at 20 perceptually spaced distortion levels, corresponding approximately to 0.2-4.0 just-noticeable difference (JND) units using the ColorVideoVDP (CVVDP) metric for distortion estimation, yielding 9,618 distorted images. This fine-grained sampling enables analysis of rate-distortion behavior and objective metric evaluation for subtle quality differences across a wide range of compression artifacts. We report an extensive objective analysis using 24 conventional and 12 learning-based IQA methods. The results show substantial disagreement among current IQA methods for fine-grained quality differences, particularly for artifacts introduced by learning-based codecs. The complete dataset is publicly available at https://doi.org/10.18419/DARUS-6156.

Figures

Figures reproduced from arXiv: 2607.22783 by Alexander Karabutov, Ant\'onio Pinheiro, Dietmar Saupe, Elena Alshina, Jo\~ao Ascenso, Jon Sneyers, Mohsen Jenadeleh, Osamu Watanabe, Panqi Jia, Thomas Richter, Touradj Ebrahimi.

Figure 1
Figure 1. Figure 1: Distributions of ColorVideoVDP (CVVDP) scores for distorted images in the TID2013, JPEG AIC-3, and proposed AIC2026 datasets. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Schematic overview of the scope of JPEG AIC-3 and AIC-4. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: 2D UMAP projection of deep visual embeddings for [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Thumbnails of the 70 source images. JND at intervals of approximately 0.2 JND. The upper limit of 4.0 JND was selected to fully cover the approximate 0–3 JND range targeted by the JPEG AIC-3 standard [29], while allowing additional headroom for possible CVVDP prediction errors, particularly for compression artifacts that differ from those represented in its calibration data. Because the AIC-3 subjective da… view at source ↗
Figure 5
Figure 5. Figure 5: Histogram of JND-mapped CVVDP scores of the distorted images in the AIC2026 dataset, colored by distortion level. [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Selected distorted images and distortion-rate results for [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Distortion-rate curves for source 10 obtained with Cool [PITH_FULL_IMAGE:figures/full_fig_p009_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Example distortion-rate curves for source 61, with metric disagreement 0.137. [PITH_FULL_IMAGE:figures/full_fig_p011_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Example distortion-rate curves for source 45, with metric disagreement 0.252. [PITH_FULL_IMAGE:figures/full_fig_p011_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: shows the effect of granularity, represented by the JND spacing δ between consecutive distortion levels, on inter-metric disagreement for n = 4. Finer distortion￾level spacing yields higher inter-metric disagreement, whereas IMD decreases as the spacing becomes coarser. This trend is consistent across all source images. The fine-grained distortion-level spacing of AIC2026 makes it particularly valuable fo… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

79 extracted references · 4 linked inside Pith

  1. [1]

    The JPEG still picture compression standard,

    G. K. Wallace, “The JPEG still picture compression standard,”Commu- nications of the ACM, vol. 34, no. 4, pp. 30–44, 1991

  2. [2]

    An overview of the JPEG 2000 still image compression standard,

    M. Rabbani and R. Joshi, “An overview of the JPEG 2000 still image compression standard,”Signal Process. Image Commun., vol. 17, no. 1, pp. 3–48, 2002

  3. [3]

    Overview of the high efficiency video coding (HEVC) standard,

    G. J. Sullivan, J.-R. Ohm, W.-J. Han, and T. Wiegand, “Overview of the high efficiency video coding (HEVC) standard,”IEEE Trans. Circuits Syst. Video Technol., vol. 22, no. 12, pp. 1649–1668, 2012

  4. [4]

    An evaluation of the next-generation image coding standard A VIF,

    N. Barman and M. G. Martini, “An evaluation of the next-generation image coding standard A VIF,” in12th Int. Conf. Quality Multimedia Experience, 2020, pp. 1–4

  5. [5]

    Versatile video coding standard: A review from coding tools to consumers deployment,

    W. Hamidouche, T. Biatek, M. Abdoliet al., “Versatile video coding standard: A review from coding tools to consumers deployment,”IEEE Consumer Electronics Magazine, vol. 11, no. 5, pp. 10–24, 2022

  6. [6]

    The JPEG XL image coding system: History, features, coding tools, design rationale, and future,

    J. Sneyers, J. Alakuijala, L. Versari, Z. Szabadka, S. Boukorttet al., “The JPEG XL image coding system: History, features, coding tools, design rationale, and future,”arXiv:2506.05987, 2025

  7. [7]

    Overview of variable rate coding in JPEG AI,

    P. Jia, F. Brand, D. Yuet al., “Overview of variable rate coding in JPEG AI,”IEEE Trans. Circuits Syst. Video Technol., vol. 35, no. 9, pp. 9460–9474, 2025

  8. [8]

    An overview of the JPEG AI learning-based image coding standard,

    S. Esenlik, Y . Wu, Z. Zhanget al., “An overview of the JPEG AI learning-based image coding standard,”IEEE Trans. Circuits Syst. Video Technol., vol. 36, no. 2, pp. 2520–2537, 2026

  9. [9]

    Frequency-aware transformer for learned image compression,

    H. Li, S. Li, W. Dai, C. Li, J. Zou, and H. Xiong, “Frequency-aware transformer for learned image compression,” inProc. Int. Conf. Learn. Represent., vol. 2024, 2024, pp. 30 447–30 465

  10. [10]

    Cool-chic: Coordinate-based low complexity hierarchical image codec,

    T. Ladune, P. Philippe, F. Henry, G. Clare, and T. Leguay, “Cool-chic: Coordinate-based low complexity hierarchical image codec,” inProc. IEEE/CVF Int. Conf. Comput. Vis., 2023, pp. 13 515–13 522

  11. [11]

    Cool-chic 5.0: Faster encoding and inter-feature entropy modeling for overfitted image compression,

    T. Ladune, P. Philippe, P. Jaffuer, T. Blard, S. Kervadec, F. Henry, and G. Clare, “Cool-chic 5.0: Faster encoding and inter-feature entropy modeling for overfitted image compression,”arXiv 2605.02726, 2026

  12. [12]

    Good, cheap, and fast: Overfitted image compression with Wasserstein distortion,

    J. Ball ´e, L. Versari, E. Dupont, H. Kim, and M. Bauer, “Good, cheap, and fast: Overfitted image compression with Wasserstein distortion,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2025, pp. 23 259–23 268

  13. [13]

    Generative adversarial networks for extreme learned image compres- sion,

    E. Agustsson, M. Tschannen, F. Mentzer, R. Timofte, and L. V . Gool, “Generative adversarial networks for extreme learned image compres- sion,” inProc. IEEE/CVF Int. Conf. Comput. Vis., 2019, pp. 221–231

  14. [14]

    Subjective visual quality assessment for high-fidelity learning-based image compression,

    M. Jenadeleh, J. Sneyers, P. Jiaet al., “Subjective visual quality assessment for high-fidelity learning-based image compression,” in17th Int. Conf. Quality Multimedia Experience, 2025, pp. 1–7

  15. [15]

    Image and video compression with neural networks: A review,

    S. Ma, X. Zhang, C. Jiaet al., “Image and video compression with neural networks: A review,”IEEE Trans. Circuits Syst. Video Technol., vol. 30, no. 6, pp. 1683–1698, 2019

  16. [16]

    Image quality assessment: from error visibility to structural similarity,

    Z. Wang, A. Bovik, H. Sheikh, and E. Simoncelli, “Image quality assessment: from error visibility to structural similarity,”IEEE Trans. Image Process., vol. 13, no. 4, pp. 600–612, 2004

  17. [17]

    Toward a practical perceptual video quality metric,

    Z. Li, A. Aaronet al., “Toward a practical perceptual video quality metric,” Netflix TechBlog, 2016, https://netflixtechblog.com/ toward-a-practical-perceptual-video-quality-metric-653f208b9652

  18. [18]

    HDR- VDP-2: A calibrated visual metric for visibility and quality predictions in all luminance conditions,

    R. K. Mantiuk, K. J. Kim, A. G. Rempel, and W. Heidrich, “HDR- VDP-2: A calibrated visual metric for visibility and quality predictions in all luminance conditions,”ACM Trans. Graphics, vol. 30, no. 4, pp. 40:1–40:14, 2011

  19. [19]

    ColorVideoVDP: A visual difference predictor for image, video and display distortions,

    R. K. Mantiuk, A. Chapiro, A. Kaplanyan, J. Kim, and T. O. Aydin, “ColorVideoVDP: A visual difference predictor for image, video and display distortions,”ACM Trans. Graphics, vol. 42, no. 4, 2023

  20. [20]

    A statistical evaluation of recent full reference image quality assessment algorithms,

    H. R. Sheikh, M. F. Sabir, and A. C. Bovik, “A statistical evaluation of recent full reference image quality assessment algorithms,”IEEE Trans. Image Process., vol. 15, no. 11, pp. 3440–3451, 2006

  21. [21]

    Image database TID2013: Peculiarities, results and perspectives,

    N. Ponomarenko, L. Jin, O. Ieremeievet al., “Image database TID2013: Peculiarities, results and perspectives,”Signal Processing: Image Com- munication, vol. 30, pp. 57–77, 2015

  22. [22]

    Fine-grained image quality assess- ment: A revisit and further thinking,

    X. Zhang, W. Lin, and Q. Huang, “Fine-grained image quality assess- ment: A revisit and further thinking,”IEEE Trans. Circuits Syst. Video Technol., vol. 32, no. 5, pp. 2746–2759, 2022

  23. [23]

    Fine-grained quality assessment for compressed images,

    X. Zhang, W. Lin, S. Wang, J. Liu, S. Ma, and W. Gao, “Fine-grained quality assessment for compressed images,”IEEE Trans. Image Process., vol. 28, no. 3, pp. 1163–1175, 2019

  24. [24]

    Percep- tual quality assessment for fine-grained compressed images,

    Z. Zhang, W. Sun, W. Wu, Y . Chen, X. Min, and G. Zhai, “Percep- tual quality assessment for fine-grained compressed images,”J. Visual Communication and Image Representation, vol. 90, p. 103696, 2023

  25. [25]

    2AFC prompting of large multimodal models for image quality assessment,

    H. Zhu, X. Sui, B. Chen, X. Liu, P. Chen, Y . Fang, and S. Wang, “2AFC prompting of large multimodal models for image quality assessment,” IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 12, pp. 12 873– 12 878, 2024

  26. [26]

    Fine-grained HDR image quality assessment from noticeably distorted to very high fidelity,

    M. Jenadeleh, J. Sneyers, D. Lazzarottoet al., “Fine-grained HDR image quality assessment from noticeably distorted to very high fidelity,” in 2025 17th Int. Conf. Quality Multimedia Experience, 2025, pp. 1–7

  27. [27]

    Subjective image quality assessment with boosted triplet comparisons,

    H. Men, H. Lin, M. Jenadeleh, and D. Saupe, “Subjective image quality assessment with boosted triplet comparisons,”IEEE Access, vol. 9, pp. 138 939–138 975, 2021

  28. [28]

    On the mDCT-PSNR image quality index,

    T. Richter, “On the mDCT-PSNR image quality index,” inProc. Int. Workshop Qual. Multimedia Exp., 2009, pp. 53–58

  29. [29]

    Information technology — JPEG AIC Assessment of image coding — Part 3: Subjective quality assessment of high-fidelity images,

    ISO/IEC 29170-3, “Information technology — JPEG AIC Assessment of image coding — Part 3: Subjective quality assessment of high-fidelity images,” 2026

  30. [30]

    Final call for proposals on objective image quality assessment (aic-4),

    ISO/IEC JTC 1/SC 29/WG 1, “Final call for proposals on objective image quality assessment (aic-4),” 2025, 107th Meeting, Brussels

  31. [31]

    Statistical study on perceived JPEG image quality via MCL-JCI dataset construction and analysis,

    L. Jin, J. Y . Lin, S. Hu, H. Wang, P. Wang, I. Katsavounidis, A. Aaron, and C.-C. J. Kuo, “Statistical study on perceived JPEG image quality via MCL-JCI dataset construction and analysis,”Electronic Imaging, vol. 2016, no. 13, pp. 1–9, 2016

  32. [32]

    JND-Pano: Database for just noticeable difference of JPEG compressed panoramic images,

    X. Liu, Z. Chen, X. Wang, J. Jiang, and S. Kowng, “JND-Pano: Database for just noticeable difference of JPEG compressed panoramic images,” inPacific Rim Conference on Multimedia. Springer, 2018, pp. 458–468

  33. [33]

    Picture-level just noticeable difference for symmetrically and asymmetrically compressed stereo- scopic images: Subjective quality assessment study and datasets,

    C. Fan, Y . Zhang, H. Zhanget al., “Picture-level just noticeable difference for symmetrically and asymmetrically compressed stereo- scopic images: Subjective quality assessment study and datasets,”J. Vis. Commun. Image Represent., vol. 62, pp. 140–151, 2019

  34. [34]

    A JND dataset based on VVC com- pressed images,

    X. Shen, Z. Ni, W. Yanget al., “A JND dataset based on VVC com- pressed images,” inProc. IEEE Int. Conf. Multimedia Expo Workshops, 2020, pp. 1–6

  35. [35]

    Large-scale crowdsourced subjective assessment of picturewise just noticeable difference,

    H. Lin, G. Chen, M. Jenadelehet al., “Large-scale crowdsourced subjective assessment of picturewise just noticeable difference,”IEEE Trans. Circuits Syst. Video Technol., vol. 32, no. 9, pp. 5859–5873, 2022

  36. [36]

    JPEG AIC-3 dataset: towards defining the high quality to nearly visually lossless quality range,

    M. Testolina, V . Hosu, M. Jenadelehet al., “JPEG AIC-3 dataset: towards defining the high quality to nearly visually lossless quality range,” inInt. Conf. Qual. Multimedia Exp., 2023, pp. 55–60

  37. [37]

    Crowdsourced estimation of collective just noticeable difference for compressed video with the flicker test and QUEST+,

    M. Jenadeleh, R. Hamzaoui, U.-D. Reips, and D. Saupe, “Crowdsourced estimation of collective just noticeable difference for compressed video with the flicker test and QUEST+,”IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 10, pp. 10 135–10 151, 2024

  38. [38]

    Evaluation of objec- tive image quality metrics for high-fidelity image compression,

    S. Mohammadi, M. Jenadeleh, J. Sneyerset al., “Evaluation of objec- tive image quality metrics for high-fidelity image compression,”IEEE Access, vol. 14, pp. 35 651–35 668, 2026

  39. [39]

    Information technology — Advanced image coding and evaluation — Part 2: Evaluation procedure for nearly lossless coding,

    ISO/IEC 29170-2, “Information technology — Advanced image coding and evaluation — Part 2: Evaluation procedure for nearly lossless coding,” 2015. 14

  40. [40]

    A new standard method of subjective assessment of barely visible image artifacts and a new public database,

    D. M. Hoffman and D. Stolitzka, “A new standard method of subjective assessment of barely visible image artifacts and a new public database,” J. Society for Information Display, vol. 22, no. 12, pp. 631–643, 2014

  41. [41]

    DINOv2: learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanniet al., “DINOv2: learning robust visual features without supervision,”Trans. Mach. Learn. Res., pp. 1–31, 2024

  42. [42]

    Toward a better quality metric for the video community,

    Z. Li, K. Swansonet al., “Toward a better quality metric for the video community,” Netflix TechBlog, 2020, https://netflixtechblog.com/ toward-a-better-quality-metric-for-the-video-community-7ed94e752a30

  43. [43]

    Butteraugli, a tool for measuring perceived differences be- tween images,

    J. Alakuijala, “Butteraugli, a tool for measuring perceived differences be- tween images,” 2016-2025, original version: https://github.com/google/ butteraugli; current version: https://github.com/libjxl/libjxl/blob/main/ lib/jxl/butteraugli/butteraugli.cc

  44. [44]

    SSIMULACRA 2.1: Structural similarity unveiling lo- cal and compression related artifacts,

    J. Sneyers, “SSIMULACRA 2.1: Structural similarity unveiling lo- cal and compression related artifacts,” https://github.com/cloudinary/ ssimulacra2, 2022

  45. [45]

    Common Test Conditions on Objective Image Quality Assessment v2.0,

    ISO/IEC JTC 1/SC29/WG1 N101246, “Common Test Conditions on Objective Image Quality Assessment v2.0,” 2025, https://jpeg.org/aic/ documentation.html

  46. [46]

    Evaluating quality metrics through the lenses of psychophysical measurements of low-level vision,

    D. Hammou, Y . Cai, P. Madhusudanaraoet al., “Evaluating quality metrics through the lenses of psychophysical measurements of low-level vision,”arXiv preprint arXiv:2503.16264, 2026

  47. [47]

    JPEG on STEROIDS: Common optimization techniques for JPEG image compression,

    T. Richter, “JPEG on STEROIDS: Common optimization techniques for JPEG image compression,” inProc. Int. Conf. Image Process., 2016, pp. 61–65

  48. [48]

    JPEG AI Common Training and Test Conditions,

    ISO/IEC JTC 1/SC 29/WG 1 N100600, “JPEG AI Common Training and Test Conditions,” Jul. 2023, Version 8.0

  49. [49]

    JPEG AI reference software, https://gitlab.com/wg1/jpeg-ai/ jpeg-ai-reference-software

  50. [50]

    Cool-chic, https://github.com/Orange-OpenSource/Cool-Chic

  51. [51]

    FTIC, https://github.com/qingshi9974/ICLR2024-FTIC

  52. [52]

    Users prefer Jpegli over same-sized libjpeg-turbo or MozJPEG ,

    M. Bruse, L. Versari, Z. Szabadka, and J. Alakuijala, “Users prefer Jpegli over same-sized libjpeg-turbo or MozJPEG ,”arXiv preprint arXiv:2403.18589, 2024

  53. [53]

    libwebp: WebP codec, version 1.2.4,

    Google, “libwebp: WebP codec, version 1.2.4,” https://chromium. googlesource.com/webm/libwebp/+/refs/tags/v1.2.4, 2022, includes the cwebpencoder

  54. [54]

    pngquant, version 2.14.1,

    K. Lesi ´nski, “pngquant, version 2.14.1,” https://github.com/kornelski/ pngquant/releases/tag/2.14.1

  55. [55]

    New full-reference quality metrics based on HVS,

    K. Egiazarian, J. Astola, N. Ponomarenkoet al., “New full-reference quality metrics based on HVS,” inProc. 2nd Int. Workshop Video Process. Qual. Metrics, vol. 4, 2006, pp. 1–6

  56. [56]

    RGBA structural similarity: DSSIM version 3.3.4,

    K. Lesi ´nski, “RGBA structural similarity: DSSIM version 3.3.4,” 2024, https://github.com/kornelski/dssim

  57. [57]

    Information content weighting for perceptual image quality assessment,

    Z. Wang and Q. Li, “Information content weighting for perceptual image quality assessment,”IEEE Trans. Image Process., vol. 20, no. 5, pp. 1185–1198, 2011

  58. [58]

    FSIM: A feature similarity index for image quality assessment,

    L. Zhang, L. Zhang, X. Mou, and D. Zhang, “FSIM: A feature similarity index for image quality assessment,”IEEE Trans. Image Process., vol. 20, no. 8, pp. 2378–2386, 2011

  59. [59]

    VMAF v1: Good is not good enough,

    C. G. Bampis, Z. Li, K. Swanson, N. F. Miret, and P. Madhusudanarao, “VMAF v1: Good is not good enough,” Netflix TechBlog, 2026, https:// netflixtechblog.com/vmaf-v1-good-is-not-good-enough-60d7e4244ea8

  60. [60]

    Image information and visual quality,

    H. R. Sheikh and A. C. Bovik, “Image information and visual quality,” IEEE Trans. Image Process., vol. 15, no. 2, pp. 430–444, 2006

  61. [61]

    The CIEDE2000 color-difference formula: Implementation notes, supplementary test data, and mathemat- ical observations,

    G. Sharma, W. Wu, and E. N. Dalal, “The CIEDE2000 color-difference formula: Implementation notes, supplementary test data, and mathemat- ical observations,”Color Res. Appl., vol. 30, no. 1, pp. 21–30, 2005

  62. [62]

    VSI: A visual saliency-induced index for perceptual image quality assessment,

    L. Zhang, Y . Shen, and H. Li, “VSI: A visual saliency-induced index for perceptual image quality assessment,”IEEE Trans. Image Process., vol. 23, no. 10, pp. 4270–4281, 2014

  63. [63]

    Gradient magnitude similarity deviation: A highly efficient perceptual image quality index,

    W. Xue, L. Zhang, X. Mou, and A. C. Bovik, “Gradient magnitude similarity deviation: A highly efficient perceptual image quality index,” IEEE Trans. Image Process., vol. 23, no. 2, pp. 684–695, 2014

  64. [64]

    Perceptual image quality assessment using a normalized Laplacian pyramid,

    V . Laparra, J. Ball ´e, A. Berardino, and E. P. Simoncelli, “Perceptual image quality assessment using a normalized Laplacian pyramid,” in Proc. IS&T Electronic Imaging, 2016, pp. 1–6

  65. [65]

    A Haar wavelet-based perceptual similarity index for image quality assessment,

    R. Reisenhofer, S. Bosse, G. Kutyniok, and T. Wiegand, “A Haar wavelet-based perceptual similarity index for image quality assessment,” Signal Process. Image Commun., vol. 61, pp. 33–43, 2018

  66. [66]

    FLIP: A difference evaluator for alternating images

    P. Andersson, J. Nilsson, T. Akenine-M ¨olleret al., “FLIP: A difference evaluator for alternating images.”Proc. ACM Comput. Graph. Interact. Tech., vol. 3, no. 2, pp. 1–23, 2020

  67. [67]

    HDR-VDP-3: A multi-metric for predicting image differences, quality and contrast distortions in high dynamic range and regular content,

    R. K. Mantiuk, D. Hammou, and P. Hanji, “HDR-VDP-3: A multi-metric for predicting image differences, quality and contrast distortions in high dynamic range and regular content,”arXiv:2304.13625, 2023

  68. [68]

    The unreasonable effectiveness of deep features as a perceptual metric,

    R. Zhang, P. Isola, A. A. Efroset al., “The unreasonable effectiveness of deep features as a perceptual metric,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2018, pp. 586–595

  69. [69]

    Shift-tolerant perceptual similarity metric,

    A. Ghildyal and F. Liu, “Shift-tolerant perceptual similarity metric,” in Proc. Eur. Conf. Comput. Vis., 2022, pp. 89–105

  70. [70]

    PieAPP: Perceptual image- error assessment through pairwise preference,

    E. Prashnani, H. Cai, Y . Mostofi, and P. Sen, “PieAPP: Perceptual image- error assessment through pairwise preference,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2018, pp. 1808–1817

  71. [71]

    Deep neural networks for no-reference and full-reference image quality assessment,

    S. Bosse, D. Maniry, K.-R. M ¨ulleret al., “Deep neural networks for no-reference and full-reference image quality assessment,”IEEE Trans. Image Process., vol. 27, no. 1, pp. 206–219, 2018

  72. [72]

    Image quality assessment: Unifying structure and texture similarity,

    K. Ding, K. Ma, S. Wang, and E. P. Simoncelli, “Image quality assessment: Unifying structure and texture similarity,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 5, pp. 2567–2581, 2022

  73. [73]

    Attentions help CNNs see better: Attention-based hybrid image quality assessment network,

    S. Lao, Y . Gong, S. Shi, S. Yang, T. Wu, X. Wang, Y . Xia, and J. Gu, “Attentions help CNNs see better: Attention-based hybrid image quality assessment network,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, 2022, pp. 1140–1149

  74. [74]

    DeepDC: Deep distance correlation as a perceptual image quality evaluator,

    H. Zhu, B. Chen, L. Zhu, S. Wang, and W. Lin, “DeepDC: Deep distance correlation as a perceptual image quality evaluator,”IEEE Trans. Image Process., vol. 34, pp. 7859–7873, 2025

  75. [75]

    DreamSim: Learning new dimensions of human visual similarity using synthetic data,

    S. Fu, N. Tamir, S. Sundaramet al., “DreamSim: Learning new dimensions of human visual similarity using synthetic data,”Adv. Neural Inf. Process. Syst., vol. 36, pp. 50 742–50 768, 2023

  76. [76]

    TOPIQ: A top-down approach from semantics to distortions for image quality assessment,

    C. Chen, J. Mo, J. Houet al., “TOPIQ: A top-down approach from semantics to distortions for image quality assessment,”IEEE Trans. Image Process., vol. 33, pp. 2404–2418, 2024

  77. [77]

    Wasserstein distortion: Unifying fidelity and realism,

    Y . Qiu, A. B. Wagner, J. Ball ´eet al., “Wasserstein distortion: Unifying fidelity and realism,” inProc. Conf. Inf. Sci. Syst., 2024, pp. 1–6

  78. [78]

    Toward generalized image quality assessment: Relaxing the perfect reference quality assumption,

    D. Chen, T. Wu, K. Maet al., “Toward generalized image quality assessment: Relaxing the perfect reference quality assumption,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2025, pp. 12 742– 12 752

  79. [79]

    Debiased mapping for full-reference image quality assessment,

    B. Chen, H. Zhu, L. Zhu, S. Wang, J. Pan, and S. Wang, “Debiased mapping for full-reference image quality assessment,”IEEE Trans. Multimedia, vol. 27, pp. 2638–2649, 2025