REVIEW 3 major objections 5 minor 9 references
DODA: A Database of Datasets for Aesthetics Research
T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read DODA is a filterable online catalog of over 120 image datasets annotated for aesthetics, with standardized metadata and precomputed image statistics.
desk verdict Useful catalog for aesthetics dataset discovery, with a few internal inconsistencies that need fixing before it becomes the standard reference. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is DODA itself: a structured, filterable table of dataset profiles, each described by a fixed set of metadata columns (e.g., number of images, raters per image, annotation scale, image source, field of common use). The most mathematically distinctive piece is the heterogeneity score: for each dataset, the average Shannon entropy of several bounded quantitative image properties (RMS contrast, color channel means, symmetry, balance, homogeneity, DCM distance) is normalized to [0,1] and averaged, giving a semantics-agnostic measure of how varied the images are in low-level statistics. This score is meant to help users anticipate how dataset diversity could affect statistical
What would settle it
Check DODA's coverage against a systematic literature search of empirical and computational aesthetics papers published in a given window (e.g., 2015–2025): if a clearly relevant dataset is absent, or if a random sample of entries disagrees with the source papers on basic fields like image count or rater number, the catalog's reliability as a find-and-filter tool is called into question.
Extended reading notes
Core claim
The central claim is that an openly accessible online database, DODA, can serve as the field's common reference point for finding and selecting image datasets for aesthetics research. For each of the over 120 entries, DODA reports standardized metadata covering year, citation, number of images, number of raters, ratings per image, semantic classes, image style, content, annotation format and scale, source, resolution, and task, plus a link to the original dataset. Where images are available and the set is not prohibitively large, the database also provides precomputed quantitative image properties and a heterogeneity score, computed as the mean normalized Shannon entropy across selected boun
Load-bearing premise
The entire usefulness of DODA rests on the manual curation being both sufficiently complete and correct; if substantial numbers of aesthetics datasets are missing or their metadata are wrong, the tool cannot reliably guide researchers to the right dataset.
Editorial extensions
If this is right
- Researchers can filter across all major aesthetics-annotated image sets at once, cutting the time spent hunting through papers and shared folders for basic dataset facts.
- Precomputed image statistics and heterogeneity scores make it possible to assess dataset diversity before committing to download, which is especially useful for large collections.
- The standardized metadata schema reveals gaps and overlaps among existing datasets (e.g., which sets share the same source images), supporting conscious reuse.
- The invitation for community submissions makes DODA a living catalog that tracks new datasets as they appear.
- By lowering the cost of finding suitable datasets, DODA encourages researchers to reuse rather than create one-off stimulus sets, improving comparability across studies.
Reading between the lines
- If DODA becomes the standard portal for aesthetics datasets, future papers may cite the catalog entry rather than describing dataset characteristics from scratch, making reporting more uniform.
- The heterogeneity score could be repurposed as a general-purpose dataset descriptor in other fields that rely on image statistics, such as image quality assessment or material perception.
- A natural extension would be to let users upload their own metadata or verify existing entries, turning the catalog into a collaborative registry with version control.
- The authors' decision to cover both large computational sets and small empirical stimulus sets positions DODA as a bridge between two research cultures that often do not share data; if adopted, it could uncover regularities that only appear when results are compared across dataset sizes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces DODA (Database of Datasets for Aesthetics), a web application embedded in the Aesthetics Toolbox that catalogs image datasets annotated for aesthetic variables. It claims to let researchers browse 'all important datasets' in empirical and computational aesthetics, with metadata on dataset size, annotations, raters, image source, resolution, formats, and scales. For many datasets, DODA also provides precomputed quantitative image properties (QIPs) computed via the Aesthetics Toolbox. The paper further proposes a heterogeneity score based on the mean normalized Shannon entropy of selected bounded QIPs, and illustrates it on a subset of datasets. The manuscript includes a supplementary list of datasets and a list of open museum collections, and describes copyright considerations and dataset-reuse arguments.
Significance. If the database is complete and accurate, DODA fills a real gap: there is currently no centralized, filterable registry of aesthetics-annotated image sets spanning both the small controlled stimulus sets of empirical aesthetics and the large machine-learning corpora of computational aesthetics. The paper is pragmatic and useful, and the Shannon-entropy heterogeneity measure is a simple, parameter-light descriptor that could be useful for dataset comparison. The authors openly provide the web application, the dataset list, and precomputed QIPs, which is a concrete open-science contribution. However, the central claim of a reliable, comprehensive catalog rests on manual curation that is not described with enough rigor, and the manuscript contains internal inconsistencies about the number of datasets and the coverage of QIPs. These issues are fixable but must be addressed before the database can be relied upon as the authoritative resource the paper promises.
major comments (3)
- [Introduction / Introducing the Database (pp. 8–9)] The paper states 'DODA currently encompasses over 120 datasets' (p. 8) but later says 'Based on a thorough literature review we compiled a list of over 90 datasets' (p. 9). The acknowledgements add that Huckle's GitHub list allowed inclusion of 25 additional image sets; 90 + 25 = 115, not 120. This is not just a wording issue: the central claim is that DODA lets users browse 'all important datasets', so the inventory count and its provenance are load-bearing. Please give the exact number with a date, describe the systematic literature search, inclusion/exclusion criteria, and how completeness is assessed, and specify whether the '90' is an earlier version of the inventory.
- [Data Available (p. 15) vs. Resolution, Image Quality and Image Fidelity (p. 21)] The manuscript contradicts itself on QIP coverage. The 'Data Available' section says DODA provides QIPs 'whenever images were available and image set size is equal to or below 17,000 images', but the later section says 'we provide pre-calculated QIPs for all datasets found on DODA'. The abstract says 'for many of them'. Since users rely on DODA to know whether QIPs are present, please state unambiguously what fraction of datasets have QIPs, list the cutoff criterion, and make the mapping from datasets to QIP availability explicit in the web interface and in the paper.
- [Heterogeneity Score, Eq. (2)] The heterogeneity score is not reproducible as written. Equation (2) uses the empirical frequency p_q(v) of 'each observed value v' and defines K_q = |V_q| as the number of possible discrete values, but the selected QIPs (RMS contrast, mean RGB/L, symmetry, balance, homogeneity, DCM distance) are continuous. The paper does not specify the binning scheme, bin width, number of bins, or how K_q is determined for continuous variables. It also does not state how datasets with missing QIPs are treated in the mean over M. Without this information, Figure 7 cannot be reproduced and the stated 'parameter-free' nature of the score is unverifiable. Please add the full discretization protocol and the handling of missing data.
minor comments (5)
- [Supplementary Materials, Table 1] The text refers to 'Table A in Supplementary Materials' (p. 8), but the supplementary table is labeled 'Table 1'. Please align the cross-reference.
- [Figure 4] The caption says 'Fifty four exemplary DODA datasets' but does not explain why these 54 were selected or how the others are excluded. State the selection criterion.
- [Supplementary Materials, Table 1] Typo: 'ArtBrench-10' should be 'ArtBench-10', consistent with the main text and the cited reference.
- [Data Category, p. 17] The hierarchical classification rule is described but not formalized. Please specify the exact precedence order and how ties or mixed datasets are resolved, since 'Data Category' is a filterable field.
- [General] Because DODA is 'continuously growing', please provide a version number, snapshot date, or DOI for the database as described in the paper, so users can cite a stable version and know when the inventory was last updated.
Circularity Check
No significant circularity: DODA is a curated catalog, and its heterogeneity score is a standard entropy computation with no fitted parameters or prediction that reduces to its inputs.
full rationale
The paper’s only quantitative derivation is the heterogeneity score in Eqs. (1)-(3). This is a descriptive application of normalized Shannon entropy to precomputed quantitative image properties (QIPs); it involves no fitted constants, no learned parameters, and no quantity that is being predicted from a subset of the same data. The QIPs are supplied by the Aesthetics Toolbox, cited as Redies et al. (2025); this is a self-citation, but it is not load-bearing in a circular sense. DODA’s usefulness does not depend on the Toolbox being independently derived here, and the citation is to an externally published software/tool paper, not to a self-imposed uniqueness theorem or an ansatz that is smuggled in as an external fact. The central claim — that DODA is an open, filterable catalog of aesthetics datasets with metadata and precomputed image properties — is a curation claim, not a derivation. Dataset entries were compiled from literature review and existing overviews, and the metadata are not computed from the paper’s own outputs. The stated inconsistencies (e.g., 'over 90 datasets' vs. 'over 120 datasets'; 'pre-calculated QIPs for all datasets' vs. the 17,000-image cutoff) are completeness and internal-consistency concerns that an auditor should investigate, but they are not circularity. No equation or claim in the paper reduces by construction to its own input, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- QIP computation cutoff =
17,000 images
assumptions (3)
- domain assumption Shannon entropy of binned continuous QIP values quantifies dataset heterogeneity (Eq. 2).
- ad hoc to paper Selected bounded-range QIPs (RMS contrast, RGB/mean L, symmetry, balance, homogeneity, DCM distance) span complementary image-structure aspects; unbounded QIPs are excluded because they 'can lead to unstable binning'.
- domain assumption The metadata and coverage of DODA are accurate and complete enough for dataset discovery.
Cite this review
Pith. "Pith review of DODA: A Database of Datasets for Aesthetics Research." pith.science (2026). https://pith.science/paper/UNDGASWE
@misc{pith2026260800089,
author = {Pith},
title = {Pith review of: DODA: A Database of Datasets for Aesthetics Research},
year = {2026},
howpublished = {\url{https://pith.science/paper/UNDGASWE}},
note = {Machine review of arXiv:2608.00089}
}
read the original abstract
With rapid growth in the fields of empirical and computational aesthetics we have seen a vast increase in large image datasets annotated for aesthetics. As the image databases differ widely in many respects (e.g., different standards for annotation), it can be tedious to find the dataset that fits one's research needs best. The absence of a centralized open-science search system causes additional problems. Currently, researchers typically share dataset links in papers or on diverse platforms like OSF, GitHub or Dropbox. Manually searching for details like image quality and content often requires downloading all datasets. Therefore, we present the Database Of Datasets for Aesthetics (DODA), an intuitive Web application in which researchers can browse all important datasets for aesthetics research. DODA provides general information about these datasets (size, resolution, type of annotation, number of annotators, etc.) and for many of them also precomputed quantitative image properties. We discuss relevant criteria for selecting a suitable dataset with DODA and illustrate the benefits of reusing datasets. Our approach facilitates collaboration across the fields of empirical and computational aesthetics. Keywords: empirical aesthetics, computational aesthetics, machine learning, image annotation, quantitative image properties, Open Science
Reference graph
Works this paper leans on
-
[8]
https://doi.org/10.1037/aca0000398 Datta, R., Joshi, D., Li, J., & Wang, J. Z. (2006). Studying aesthetics in photographic images using a computational approach. In A. Leonardis, H. Bischof, & A. Pinz (Eds.), Computer Vision – ECCV 2006 (pp. 288–301). Springer. https://doi.org/10.1007/11744078_23 Datta, R., Li, J., & Wang, J. Z. (2008). Algorithmic infere...
arXiv 2006
-
[9]
https://doi.org/10.1027/1618-3169/a000432 Gartus, A., & Leder, H. (2013). The small step toward asymmetry: Aesthetic judgment of broken symmetries. I-Perception, 4(5), 361–364. https://doi.org/10.1068/i0588sas Ghosal, K., Rana, A., & Smolic, A. (2019). Aesthetic image captioning from weakly- labelled photographs. Proceedings - 2019 International Conferenc...
arXiv 2013
-
[67]
R., Kim, O., Meletaki, V., & Chatterjee, A
https://doi.org/10.1080/00223980.1972.9923789 Estrada Gonzalez, V., Bobrow, I., Cardillo, E. R., Kim, O., Meletaki, V., & Chatterjee, A. (2025). A normed art database that incorporates diverse cultures and genres: 33 The Penn Center for Neuroaesthetics artwork repository. Psychology of Aesthetics, Creativity, and the Arts. https://doi.org/10.1037/aca00007...
arXiv 1972
-
[125]
https://doi.org/10.1027/1618-3169/a000432 Gartus, A., & Leder, H. (2013). The small step toward asymmetry: Aesthetic judgment of broken symmetries. I-Perception, 4(5), 361–364. https://doi.org/10.1068/i0588sas Geiger, C., & Schönherr, F. (2017). Consumers’ Frequently Asked Questions (FAQS) on Copyright (pp. 1–96) [Summary Report]. The European Union Intel...
arXiv 2013
-
[201]
https://doi.org/10.1016/j.actpsy.2011.10.004 Bartho, R., Thoemmes, K., & Redies, C. (2023). Predicting beauty, liking, and aesthetic quality: A comparative analysis of image databases for visual aesthetics research. arXiv.https://arxiv.org/abs/2307.00984v1 Bertamini, M., Palumbo, L., Gheorghes, T. N., & Galatsidas, M. (2016). Do observers like curvature o...
arXiv 2011
-
[337]
https://doi.org/10.1037/aca0000398 De Winter, S., & Koßmann, L. (2026, September 14–18). A conservation approach to digital image datasets: Safeguarding cultural integrity in the digital age [Paper presentation]. 21st International Council of Museums Committee for Conservation (ICOM-CC) Triennial Conference, Oslo, Norway. Ding, K., Ma, K., Wang, S., & Sim...
-
[350]
https://doi.org/10.3389/fnhum.2016.00350 Srinivasa Desikan, B., Shimao, H., & Miton, H. (2022). WikiArtVectors: Style and color representations of artworks for cultural analysis via information theoretic measures. Entropy, 24(9), 1175. https://doi.org/10.3390/e24091175 Sun, M., & Ying, H. (2023). Color’s perceptual diversity and categorical harmony improv...
arXiv 2016
-
[508]
https://doi.org/10.1348/0007126042369811 Leder, H., & Nadal, M. (2014). Ten years of a model of aesthetic appreciation and aesthetic judgments: The aesthetic episode – Developments and challenges in empirical aesthetics. British Journal of Psychology, 105(4), 443–464. https://doi.org/10.1111/BJOP.12084 Liao, P., Li, X., Liu, X., & Keutzer, K. (2022). The ...
arXiv 2014
Show all 9 references
-
[680]
https://bmva- archive.org.uk/bmvc/2025/assets/papers/Paper_680/paper.pdf Chipman, S. F. (1977). Complexity and structure in visual patterns. Journal of Experimental Psychology: General, 106(3), 269–301. https://doi.org/10.1037/0096-3445.106.3.269 Chipman, S. F., & Mendelson, M...
2025 doi
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.