Pith. sign in

REVIEW 3 major objections 5 minor 49 references

This paper publishes the first open-access cone-beam CT dataset in which 1,764 image slices carry expert quality ratings, providing a shared benchmark for testing image-quality measures against human perception.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 10:43 UTC pith:7QTN62TJ

load-bearing objection A genuinely useful first CBCT IQA dataset, but the benchmark conclusions are anchored to a reference image that the paper itself suggests is noisy. the 3 major comments →

arxiv 2607.29253 v1 pith:7QTN62TJ submitted 2026-07-31 physics.med-ph cs.CV

CBCT-IQ: A Publicly Available Annotated Cone-Beam CT Dataset for Image Quality Assessment and Benchmarking

classification physics.med-ph cs.CV
keywords cone-beam CTimage quality assessmentexpert annotationsbenchmark datasetfull-reference IQAno-reference IQACBCTphantom study
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper aims to give the CBCT research community a much-needed public benchmark: a systematically varied set of CBCT images with expert clinical annotations of image quality. The authors scanned thorax and pelvis phantoms on a C-arm CBCT system, varying tube voltage, pulse width, and reconstruction type, and had three experts rate overall and region-of-interest quality on a four-level scale relative to a device-standard reference image. They computed 26 full-reference and no-reference IQA measures and compared them to the expert scores, finding that the strongest correlations are driven by reconstruction-type differences rather than by tube voltage or pulse width changes. They also introduce an exploratory IQA-based ranking that agrees across measures on subtle quality differences that human raters could not distinguish. If accepted, the dataset gives researchers a shared testbed for developing and validating quantitative image-quality measures, with a baseline for comparison.

Core claim

The central discovery is the dataset: 1,764 annotated CBCT slices from 36 volumes with systematic variations in kV, pulse width, and reconstruction type, each graded by three experts for overall and ROI quality on a 1–4 scale, anchored to a device-standard reference acquisition. Benchmarking 27 IQA implementations shows LPIPS, DISTS, FSIM, VSI, VIF, DSS, and Haar-based measures correlate most strongly with expert scores, but several respond mainly to background noise rather than structure. The high overall correlations are driven almost entirely by reconstruction-type variation; within a fixed reconstruction type, expert ratings barely change and IQA correlations drop sharply. To fill this g

What carries the argument

The central object is the paired reference/degraded image set: every degraded slice has a matching reference slice from the same phantom and anatomical position, acquired at the device-standard settings (90 kV, 8.0 ms, Normal reconstruction). This pairing enables both the expert relative-quality annotations (each test image is graded while displayed next to its reference) and all full-reference IQA computations. The paper's quantitative machinery is Spearman rank correlation between the three experts' z-scored ratings and each of 27 IQA measure implementations, applied to the whole image, a predefined ROI, and a background-noise patch; the ROI and noise-patch analyses are what reveal which m

Load-bearing premise

The load-bearing premise is that the single reference acquisition (90 kV, 8.0 ms, Normal reconstruction) is the highest-quality image for every phantom and slice, so that all expert relative-quality ratings and all full-reference IQA measures inherit that one anchor.

What would settle it

A concrete check: take a subset of slices and have the same three experts re-rate the degraded images paired against a different reference (for instance, 125 kV or the Smooth reconstruction). If the relative quality rankings of the other variants change substantially, or if full-reference IQA measures no longer correlate with expert scores, then the default-protocol anchor is not a stable ground truth and the dataset's benchmark conclusions weaken.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Any new CBCT IQA method can be tested on the dataset and its Spearman correlation against expert OIQ and ROI scores compared directly with the 27 measures benchmarked here.
  • The finding that LPIPS, DISTS, and RPIAXIS track background noise rather than structure warns researchers against using those measures alone for CBCT protocol evaluation.
  • Since expert ratings detect reconstruction-type changes but not kV/pulse-width changes, the dataset separates 'visible' quality differences from 'measureable but invisible' ones, providing two complementary test regimes.
  • The IQA-based consensus rating offers a way to rank acquisition settings for subtle quality differences when human perception is insensitive, potentially informing dose-optimization studies.
  • Because reference and degraded slices come from the same phantom and slice position, the dataset can also be used for training and validating deep learning models on image-quality prediction.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Extension: if the reference acquisition (90 kV, 8.0 ms, Normal) is not actually the clinically best image for a given task, then both the expert relative ratings and all full-reference IQA correlations shift; re-running the expert study with a different reference (e.g., 125 kV or Smooth reconstruction) would test this.
  • Extension: the dataset's single scanner and phantom geometry means its IQA rankings may not transfer to other CBCT systems with different detectors or reconstruction algorithms; cross-scanner validation is needed before treating these results as a universal CBCT IQA benchmark.
  • Extension: the consistent IQA consensus on subtle variants suggests a route toward automated acquisition-protocol optimization, where a set of measures could serve as a proxy objective in closed-loop dose optimization.
  • Extension: the noise-patch analysis shows some measures respond primarily to background noise, so the dataset could be used to develop structure-aware IQA measures that explicitly penalize noise-only responses—a testable direction the authors hint at but do not pursue.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces CBCT-IQ, a publicly available cone-beam CT dataset for image quality assessment. The authors acquired CBCT volumes of thorax and pelvis phantoms on a Siemens ARTIS pheno system, systematically varying tube voltage (60/90/125 kV), pulse width (3.2/5.0/6.4/8.0 ms), and reconstruction type (Normal/Smooth/Very Smooth), resulting in 36 acquisition configurations. From these, they selected 1,764 2D slices (71 per chest configuration, 27 per pelvis configuration) and had three expert raters assign four-level quality scores for both overall image quality and a predefined ROI, with the default acquisition (90 kV, 8.0 ms, Normal) used as the reference. The paper further benchmarks 26 full-reference and no-reference IQA measures (27 implementations) against the expert ratings and proposes an exploratory IQA-measure-based ranking intended to capture subtle quality differences not perceived by human raters. The dataset and annotations are deposited on Zenodo.

Significance. If the annotations and benchmark are trustworthy, CBCT-IQ would fill a clear gap: a standardized, openly available CBCT dataset with expert quality labels, enabling development and comparison of IQA measures in a modality where such resources are scarce. The systematic variation of acquisition and reconstruction parameters, the use of three clinical experts, the provision of OIQ and ROI scores, and the inclusion of computed IQA values and software versions are concrete strengths that support reproducibility and benchmarking. The paper also provides a useful baseline comparison of 26 IQA measures on CBCT data. However, the validity of the expert labels and of the full-reference IQA benchmark depends critically on the choice of the reference image; the manuscript contains an internal contradiction about the noise level of that reference, which raises serious concerns about the benchmark's interpretation.

major comments (3)
  1. [Image Acquisition Protocols / Performance of IQA measures in presence of background noise] The paper declares the default acquisition (kV=90, PW=8.0 ms, RT Normal) as the 'device-standard, high-quality' reference (Section 'Image Acquisition Protocols') and uses it as the anchor for all expert relative-quality ratings and for every full-reference IQA measure. Yet Section 'Performance of IQA measures in presence of background noise' states that 'noise in the reconstruction type Normal was more pronounced visually compared to the other two types (Smooth and Very Smooth).' These statements directly contradict each other: Smooth and Very Smooth images, which are labelled as 'degraded', are less noisy than the reference. Since raters were asked to grade each test image relative to this reference, a noisier reference biases the quality labels: a smoother image may be scored lower not because it is diagnostically worse, but because it differs from the reference. Likewise, full-referen
  2. [IQA measure–based rating, Table 7] The text claims that 'all IQA measures agreed and identified PW6.4-KV90 combination as the best quality' and that 'Results on the ROI also confirmed this agreement.' Table 7 itself contradicts this. In the VIF-python row, the values are: PW8-KV60=1.258, PW8-KV125=0.696, PW5-KV90=0.934, PW6.4-KV90=0.937, PW3.2-KV90=0.921. For VIF, higher values indicate better quality, so VIF-python selects PW8-KV60 as the best, not PW6.4-KV90. The same issue appears in Table 8 (ROI) where VIF-python gives PW8-KV60=1.475 versus PW6.4-KV90=0.918. Thus the 'strong agreement' claim is not supported by the reported data. The IQA-based rating must be derived from per-measure rankings with a clearly defined aggregation rule, and the text should be corrected to avoid overclaiming consensus.
  3. [Tables 3–6: Benchmark statistics] The benchmark results report Spearman rank correlation coefficients as point estimates without confidence intervals or significance tests. Many of the SRCC values in these tables are close to each other (e.g., the top performers in Table 3 differ by a few hundredths), and with 1,764 samples but heavily dependent ratings, these differences may not be meaningful. The paper makes practical recommendations about which IQA measures are 'best-performing' (e.g., in the noise-exclusion analysis), so the absence of uncertainty quantification is a substantive omission. The authors should provide bootstrap confidence intervals for the SRCC values and, where appropriate, tests for differences between dependent correlations, so that readers can judge whether the reported ordering of IQA measures is statistically reliable. Without this, the benchmark's conclusions are under-supported.
minor comments (5)
  1. [Image Acquisition Protocols] The text says '36 volumes consisting of 378 slices were generated,' but the dataset contains 1,764 slices (18 chest configurations × 71 slices + 18 pelvis configurations × 27 slices). Please clarify what the 378 refers to or correct the sentence; as written it is internally inconsistent.
  2. [Abstract and Methods] The abstract says '26 full reference- and no reference-based IQA measures,' while the Methods section mentions '26 commonly used as well as task-based IQA measures (27 IQA measure implementations).' Please standardize the count throughout (e.g., '26 measures, 27 implementations') to avoid confusion.
  3. [Data Records] In the Data Records section, the reference path is described as containing 'kernel EE' (e.g., './DATASETNAME/SLICE/EE/090/8.0/NORMAL.nii.gz'). The term 'kernel' has not been introduced; the paper discusses reconstruction types (Normal, Smooth, VSmooth). Please clarify what 'EE' denotes and whether it is a reconstruction kernel name.
  4. [Figures and Table 1] Several typos appear: 'Nomal' in Table 1 and Figure 2/3 captions; 'Norml' in Figure 2; 'boy anatomy' should likely be 'bony anatomy' in the Phantoms section; 'Slides' appears instead of 'slices' in the abstract/Data Records. Please proofread the manuscript.
  5. [Image annotation and reader study] The annotation software section says the main task category is 'Image-guided therapy' and subcategory 'Region of Interest (ROI),' but the raters assessed both OIQ and ROI. Clarify how the OIQ task was configured in Speedy IQA, since the described configuration appears to name only the ROI subcategory.

Circularity Check

0 steps flagged

No significant circularity: the CBCT dataset and expert ratings are independent inputs, while the IQA-measure ranking is explicitly exploratory rather than a prediction.

full rationale

This paper is a data descriptor, not a derivation chain. The central product is a publicly available CBCT dataset with expert quality annotations; no fitted parameter, normalization, or equation generates the expert ratings or the IQA benchmark. The 27 IQA measures are established external algorithms applied post hoc, not fit to (or derived from) the expert annotations, so their correlation with the annotations is an empirical benchmark rather than a circular validation. The IQA-measure-based ranking is computed from the same measures, but the authors explicitly label it 'exploratory' and state that 'it does not necessarily indicate a closer correspondence to diagnostic image quality than expert reader assessments' (Section 'IQA measure–based rating'), so it is not presented as an independent prediction derived from the dataset. Self-citations (e.g., refs. 3, 11, 12, 14, 21, 23, 25, 38, 49) are contextual and not load-bearing; no uniqueness theorem or ansatz is imported from them. The only notable problem is an internal-validity concern, not circularity: the default protocol (kV 90, PW 8.0 ms, RT Normal) is declared the 'reference' and 'device-standard, high-quality' image in 'Image Acquisition Protocols', yet the paper later observes that 'noise in the reconstruction type Normal was more pronounced visually compared to the other two types' in 'Performance of IQA measures in presence of background noise'. This could bias both expert relative ratings and full-reference IQA scores, but it does not reduce any derived result to its own inputs by construction. Therefore, no circular step is exhibited, and the circularity score is 0.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

No fitted model parameters or invented physical entities are introduced. The central claim rests on hand-chosen ROI/noise-patch/slice-sampling choices and domain assumptions about reference-image quality, expert ground truth, and transferability to clinical CBCT. These choices are stated in Methods but are not independently validated or supported by uncertainty analysis.

free parameters (3)
  • ROI crop coordinates = Chest 160-352x250-410; Pelvis 60-452x260-420
    Hand-selected 'heuristically to best fit the anatomical area'; all ROI-based IQA correlations depend on this choice.
  • Noise patch coordinates = 250-349x100-199
    Chosen to contain only background noise; used to classify which IQA measures respond primarily to noise rather than anatomy.
  • Slice sampling interval and ranges = Every third slice; chest 85-295, pelvis 193-271
    Hand-selected to reduce redundancy while capturing anatomical variation; affects the score distribution and correlation results.
axioms (4)
  • domain assumption The device-standard acquisition (kV=90, PW=8 ms, RT Normal) is the highest-quality reference for each phantom/slice.
    Used as the reference for expert ratings and full-reference IQA measures; if false, all relative quality scores and correlations shift (Section 'Image Acquisition Protocols').
  • domain assumption Expert visual ratings on a 4-point scale relative to the reference are a valid ground truth for IQA benchmarking.
    Three raters; inter-rater SRCC is high for OIQ (>0.80) but lower for ROI (0.63-0.82), so ROI ground truth is weaker (Table 2).
  • domain assumption Phantom CBCT images acquired on a single Siemens ARTIS pheno system transfer to clinical CBCT across systems and tasks.
    The paper positions the dataset as a general CBCT IQA benchmark, but data come from two phantoms and one scanner (Methods; Technical Validation).
  • domain assumption The 27 IQA measures were computed with correct implementations and parameters as implied by software_versions.txt.
    No code is shipped; benchmark results depend on third-party implementations whose exact configuration is only listed as version names.

pith-pipeline@v1.3.0-daily-deepseek · 13953 in / 15531 out tokens · 137372 ms · 2026-08-03T10:43:02.916158+00:00 · methodology

0 comments
read the original abstract

Medical image quality plays a critical role in diagnostic accuracy, especially in X-ray-based imaging modalities such as cone-beam computed tomography (CBCT), where image quality must be balanced against radiation dose. While expert visual evaluation remains the clinical standard for image quality evaluation, it is time-consuming, subjective and affected by inter-observer variability, emphasizing the need for reliable quantitative image quality assessment (IQA) methods. However, the development and validation of such IQA methods have been limited by the lack of publicly available CBCT datasets with expert image quality annotations. In this study, we provide the first open-access CBCT IQA dataset containing 1,764 annotated image slices acquired using systematic variations in image acquisition and reconstruction parameters. Three clinical experts graded the overall image quality and a predefined regions of interest (ROI) using a four-level scoring scheme. In addition, we benchmark 26 full reference- and no reference-based IQA measures against expert annotations and introduce an exploratory IQA measure-based ranking capable of distinguishing subtle image quality differences. This dataset introduced a standardized benchmark for future CBCT IQA research and provides a valuable resource for the development and validation of new IQA methods, enabling reproducible research and advancing CBCT IQA.

Figures

Figures reproduced from arXiv: 2607.29253 by Afshin Mohammadi, Alfred Pohl, Ali Abbasian Ardakani, Ander Biguri, Anna Breger, Birgit Pohn, Carola-Bibiane Sch\"onlieb, Clemens Karner, Gernot Kronreif, Laura Haddad, Martin Buschmann, Paul Apfaltrer, Poorya MohammadiNasab, Sepideh Hatamikia, Stephanie Nougaret, Tess Reynolds, Wolfgang Birkfellner.

Figure 1
Figure 1. Figure 1: The overall proposed study design including data acquisition, reader study, image quality assessment (IQA) measures and IQA measure-based rating. Methods Phantoms Two anatomical phantoms, a thorax and pelvis (Kyoto Kagaku), were used to generate a collection of CBCT images. The thorax phantom contains bony anatomy, including the spine, ribs, and sternum, soft tissue equivalent material, and simplified bron… view at source ↗
Figure 2
Figure 2. Figure 2: Visualization of some example used CBCT slices from thorax phantom when changing kV (60, 90, 125), pulse width (PW) (3.2, 5, 6.4, 8 ms), and reconstruction type (Norml, Smooth and very smooth (VSmooth) for overal image, Region of interest (ROI) and noise patch [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 4
Figure 4. Figure 4: A screenshot from open-source software Speedy IQA for overall image quality (OIQ) assessment [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

49 extracted references · 3 canonical work pages

  1. [1]

    S., & Paramesran, R

    Chow, L. S., & Paramesran, R. (2016). Review of medical image quality assessment. Biomedical signal processing and control, 27, 145-154

  2. [2]

    Catalano, C., Francone, M., Ascarelli, A., Mangia, M., Iacucci, I., & Passariello, R. (2007). Optimizing radiation dose and image quality. European Radiology Supplements, 17(Suppl 6), 26-32

  3. [3]

    RadRepro CBCT: An Open -Access CBCT Phantom Dataset for Improved Standardization and Reproducibility of Radiomics Research

    Hatamikia S, Steiner E, Muniya EJ, Elmirad S, Borji A, Kronreif G, Birkfellner W, Buschmann M. RadRepro CBCT: An Open -Access CBCT Phantom Dataset for Improved Standardization and Reproducibility of Radiomics Research. Sci Data. 2026 Feb 14;13(1):454

  4. [4]

    Herath, H. M. S. S., Herath, H. M. K. K. M. B., Madusanka, N., & Lee, B. I. (2025). A Systematic Review of Medical Image Quality Assessment. Journal of Imaging, 11(4), 100

  5. [5]

    Solomon, J., & Samei , E. (2016). Correlation between human detection accuracy and observer model -based image quality measuresin computed tomography. Journal of Medical Imaging, 3(3), 035506-035506

  6. [6]

    Source -detector trajectory optimization for FOV extension in dental CBCT imaging

    Islam SMRS, Biguri A, Landi C, Di Domenico G, Schneider B, et al. Source -detector trajectory optimization for FOV extension in dental CBCT imaging. Comput Struct Biotechnol J. 2024 Nov 8;24:679-689

  7. [7]

    Trajectory generation for ROI expansion in dental CBCT imaging: An out-of-FOV ROI annotation guided arc scan-based approach

    Islam SMRS, Biguri A, Landi C, Di Domenico G, Grün P, et al. Trajectory generation for ROI expansion in dental CBCT imaging: An out-of-FOV ROI annotation guided arc scan-based approach. Comput Struct Biotechnol J. 2025 Jun 9;28:211-216

  8. [8]

    Bailey, J., Solan, M., & Moore, E. (2022). Cone -beam computed tomography in orthopaedics. Orthopaedics and Trauma, 36(4), 194-201

  9. [9]

    H., Sharashidze, V ., Narayan, V ., Ali, A.,

    Raz, E., Nossek, E., Sahlein, D. H., Sharashidze, V ., Narayan, V ., Ali, A., ... & Shapiro, M. (2023). Principles, techniques and applications of high resolution cone beam CT angiography in the neuroangio suite. Journal of neurointerventional surgery, 15(6), 600-607

  10. [10]

    P., Otake, Y ., Azizian, M., Wagner, O

    Liu, W. P., Otake, Y ., Azizian, M., Wagner, O. J., Sorger, J. M., Armand, M., & Taylor, R. H. (2015). 2D–3D radiograph to cone-beam computed tomography (CBCT) registration for C- arm image-guided robotic surgery. International journal of computer assisted radiology and surgery, 10(8), 1239-1252

  11. [11]

    Toward on -the-fly trajectory optimization for C -arm CBCT under strong kinematic constraints

    Hatamikia S, Biguri A, Kronreif G, Figl M, Russ T, Kettenbach J, Buschmann M, Birkfellner W. Toward on -the-fly trajectory optimization for C -arm CBCT under strong kinematic constraints. PLoS One. 2021 Feb 9;16(2):e0245508

  12. [12]

    Optimization for customized trajectories in cone beam computed tomography

    Hatamikia S, Biguri A, Kronreif G, Kettenbach J, Russ T, Furtado H, Shiyam Sundar LK, Buschmann M, Unger E, Figl M, Georg D, Birkfellner W. Optimization for customized trajectories in cone beam computed tomography. Med Phys. 2020 Oct;47(10):4786-4799

  13. [13]

    IEEE Access 11, 14154–14168 (2023)

    Kastryulin, S., Zakirov, J., Pezzotti, N., Dylov, D.V .: Image quality assessment for magnetic resonance imaging. IEEE Access 11, 14154–14168 (2023)

  14. [14]

    Breger, A., Karner, C., Selby, I., Gröhl, J., Dittmer, S., Lilley, E., Babar, J., Beckford, J., Sadler, T.J., Shahipasand, S., Thavakumar, A., Roberts, M., Schönlieb, C. -B.. A study on the adequacy of common IQA measures for medical images. Proceedings of 2024 International Conference on Medical Imaging and Computer -Aided Diagnosis (MICAD), Springer Lec...

  15. [15]

    AAPM Truth -based CT (TrueCT) reconstruction grand challenge

    Abadi E, Segars WP, Felice N, Sotoudeh-Paima S, Hoffman EA, Wang X, Wang W, Clark D, Ye S, Jadick G, Fryling M, Frush DP, Samei E. AAPM Truth -based CT (TrueCT) reconstruction grand challenge. Med Phys. 2025 Apr;52(4):1978-1990

  16. [16]

    LoDoPaB-CT, a benchmark dataset for low- dose computed tomography reconstruction

    Leuschner J, Schmidt M, Baguer DO, Maass P. LoDoPaB-CT, a benchmark dataset for low- dose computed tomography reconstruction. Sci Data. 2021 Apr 16;8(1):109

  17. [17]

    Advancing the Frontiers of Deep Learning for Low-Dose 3D Cone-Beam CT Reconstruction,

    A. Biguri et al., "Advancing the Frontiers of Deep Learning for Low-Dose 3D Cone-Beam CT Reconstruction," in IEEE Open Journal of Signal Processing, vol. 6, pp. 942-949, 2025

  18. [18]

    https://www.aimsciences.org/article/doi/10.3934/ammc.2023010

  19. [19]

    M., Reynolds, M., Sabet, M., Kendrick, J., Rowshanfarzad , P., & Ebert, M

    Rusanov, B., Hassan, G. M., Reynolds, M., Sabet, M., Kendrick, J., Rowshanfarzad , P., & Ebert, M. (2022). Deep learning methods for enhancing cone‐beam CT image quality toward adaptive radiation therapy: a systematic review. Medical Physics, 49(9), 6019-6054

  20. [20]

    & Lambrichts, I

    Liang, X., Jacobs, R., Hassan, B., Li, L., Pauwels, R., Corpas, L., ... & Lambrichts, I. (2010). A comparative evaluation of cone beam computed tomography (CBCT) and multi -slice CT (MSCT): Part I. On subjective image quality. European journal of radiology, 75(2), 265-269

  21. [21]

    S., Selby, I., Amberg, N., Brunner, E.,

    Breger, A., Biguri, A., Landman, M. S., Selby, I., Amberg, N., Brunner, E., ... & Schönlieb, C. B. (2025). A study of why we need to reassess full reference image quality assessment with medical images. Journal of Imaging Informatics in Medicine, 38(6), 3444-3469

  22. [22]

    A., Boone, J

    Johnston, A., Mahesh, M., Uneri, A., Rypinski, T. A., Boone, J. M., & Siewerdsen, J. H. (2024). Objective image quality assurance in cone‐beam CT: test methods, analysis, and workflow in longitudinal studies. Medical physics, 51(4), 2424-2443

  23. [23]

    -S., Duchêne, M., Weir-McCall, J., Schönlieb, C.-B.: PhotIQA: A photoacoustic image data set with image quality ratings, accepted to appear in Scientific Data (Nature), 2026

    Breger, A., Gröhl, J., Karner, C., Else, T.R., Selby, I., Rix, T., Witt, L. -S., Duchêne, M., Weir-McCall, J., Schönlieb, C.-B.: PhotIQA: A photoacoustic image data set with image quality ratings, accepted to appear in Scientific Data (Nature), 2026

  24. [24]

    Low-dose computed tomography perceptual image quality as - sessment grand challenge dataset (miccai 2023),

    Lee, W., Wagner, F., “ Low-dose computed tomography perceptual image quality as - sessment grand challenge dataset (miccai 2023),” https://doi.org/10.5281/zenodo.7833096

  25. [25]

    https://github.com/selbs/speedy_iqa

  26. [26]

    A Statistical Evaluation of Recent Full Reference Image Quality Assessment Algorithms,

    H. R. Sheikh, M. F. Sabir and A. C. Bovik, "A Statistical Evaluation of Recent Full Reference Image Quality Assessment Algorithms," in IEEE Transactions on Image Processing, vol. 15, no. 11, pp. 3440-3451, Nov. 2006

  27. [27]

    C., Sheikh, H

    Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4), 600-612

  28. [28]

    P., & Bovik, A

    Wang, Z., Simoncelli, E. P., & Bovik, A. C. (2003). Multiscale structural similarity for image quality assessment. In The thrity-seventh asilomar conference on signals, systems & computers, 2003 (Vol. 2, pp. 1398-1402). Ieee

  29. [29]

    Wang, Z., & Li, Q. (2010). Information content weighting for perceptual image quality assessment. IEEE Transactions on image processing, 20(5), 1185-1198

  30. [30]

    P., Wang, Z., Gupta, S., Bovik, A

    Sampat, M. P., Wang, Z., Gupta, S., Bovik, A. C., & Markey, M. K. (2009). Complex wavelet structural similarity: A new image similarity index. IEEE transactions on image processing, 18(11), 2385-2401

  31. [31]

    Girod, B. (1991). Psychovisual aspects of image processing: What's wrong with mean squared error?. In Proceedings of the Seventh Workshop on Multidimensional Signal Processing (p. 3)

  32. [32]

    Willmott, C. J. (1982). Some comments on the evaluation of model performance. Bulletin of the American meteorological society, 63(11), 1309-1313

  33. [33]

    Ding, K., Ma, K., Wang, S., & Simoncelli, E. P. (2020). Image quality assessment: Unifying structure and texture similarity. IEEE transactions on pattern analysis and machine intelligence, 44(5), 2567-2581

  34. [34]

    Balanov, A., Schwartz, A., Moshe, Y., & Peleg, N. (2015). Image quality assessment based on DCT subband similarity. In 2015 IEEE international conference on image processing (ICIP) (pp. 2105-2109). IEEE

  35. [35]

    Zhang, L., Zhang, L., Mou, X., & Zhang, D. (2011). FSIM: A feature similarity index for image quality assessment. IEEE transactions on Image Processing, 20(8), 2378-2386

  36. [36]

    Xue, W., Zhang, L., Mou, X., & Bovik, A. C. (2013). Gradient magnitude similarity deviation: A highly efficient perceptual image quality index. IEEE transactions on image processing, 23(2), 684-695

  37. [37]

    Reisenhofer, R., Bosse, S., Kutyniok, G., & Wiegand, T. (2018). A Haar wavelet-based perceptual similarity index for image quality assessment. Signal Processing: Image Communication, 61, 33-43

  38. [38]

    Karner, C., Gröhl, J., Selby, I., Babar, J., Beckford, J., Else, T. R., ... & Breger, A. (2025, April). Parameter choices in HaarPSI for IQA with medical images. In 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI) (pp. 1-5). IEEE

  39. [39]

    A., Shechtman, E., & Wang, O

    Zhang, R., Isola, P., Efros, A. A., Shechtman, E., & Wang, O. (2018). The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 586-595)

  40. [40]

    Z., Shahkolaei, A., Hedjam, R., & Cheriet, M

    Nafchi, H. Z., Shahkolaei, A., Hedjam, R., & Cheriet, M. (2016). Mean deviation similarity index: Efficient and reliable full-reference image quality evaluator. Ieee Access, 4, 5579-5590

  41. [41]

    R., & Bovik, A

    Sheikh, H. R., & Bovik, A. C. (2005, January). A visual information fidelity approach to video quality assessment. In The first international workshop on video processing and quality measuresfor consumer electronics (Vol. 7, No. 2, pp. 2117-2128). sn

  42. [42]

    Zhang, L., Shen, Y., & Li, H. (2014). VSI: A visual saliency-induced index for perceptual image quality assessment. IEEE Transactions on Image processing, 23(10), 4270-4281

  43. [43]

    D., & Ramli, A

    Chen, S. D., & Ramli, A. R. (2003). Minimum mean brightness error bi-histogram equalization in contrast enhancement. IEEE transactions on Consumer Electronics, 49(4), 1310-1319

  44. [44]

    completely blind

    Mittal, A., Soundararajan, R., & Bovik, A. C. (2012). Making a “completely blind” image quality analyzer. IEEE Signal processing letters, 20(3), 209-212

  45. [45]

    T., Bernecky, W

    Bosworth, B. T., Bernecky, W. R., Nickila, J. D., Adal, B., & Carter, G. C. (2008). Estimating signal-to-noise ratio (SNR). IEEE Journal of Oceanic Engineering, 33(4), 414- 418

  46. [46]

    Ying, Z., Niu, H., Gupta, P., Mahajan, D., Ghadiyaram, D., & Bovik, A. (2020). From patches to pictures (PaQ-2-PiQ): Mapping the perceptual space of picture quality. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 3575-3585)

  47. [47]

    C., & Swarup, S

    Venkatanath, N., Praneeth, D., Sumohana, S. C., & Swarup, S. M. (2015, February). Blind image quality evaluation using perception based features. In 2015 twenty first national conference on communications (NCC) (pp. 1-6). IEEE

  48. [48]

    & Choi, J

    Lee, W., Wagner, F., Galdran, A., Shi, Y., Xia, W., Wang, G., ... & Choi, J. H. (2025). Low-dose computed tomography perceptual image quality assessment. Medical Image Analysis, 99, 103343

  49. [49]

    Hatamikia, S., Breger, A., Karner, C., Pohn, B., MohammadiNasab, P., Buschmann, M., Nougaret, S., Haddad, L., Abbasian Ardakani, A., Mohammadi, A., Apfaltrer, P., Birkfellner, W., Pohl, A., Biguri, A., Kronreif, G., Schönlieb, C.-B.& Reynolds, T. (2026). CBCT-IQ: A Publicly Available Annotated Cone-Beam CT Dataset for Image Quality Assessment and Benchmar...