REVIEW 4 major objections 5 minor 2 cited by
APDDv2: Aesthetics of Paintings and Drawings Dataset with Artist Labeled Scores and Comments
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper introduces APDDv2, a 10,023-image dataset of paintings and drawings rated by experts on ten aesthetic attributes across 24 artistic categories, and reports that the ArtCLIP model trained on it outperforms prior art-aesthetic…
desk verdict A genuinely useful painting-aesthetics dataset with detailed curation, but the missing inter-annotator reliability means the benchmark-quality claim is not yet supported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two mechanisms carry the argument. The first is the annotation rubric: 24 artistic categories cross ten aesthetic attributes, and each category has a benchmark table that ties every score range to representative example images, so expert annotators can apply one continuous standard that also stays comparable to the earlier APDDv1 release. The second is the ArtCLIP training scheme: attribute-aware contrastive learning that pairs an image with comments sampled from the matching attribute category, a multimodal fusion module that merges image and text embeddings, and final regression fine-tuning on APDDv2. The benchmark tables make the averaged expert scores meaningful, while the contrastive pre-training is what imports aesthetic language from the photography domain into painting.
What would settle it
Compute per-attribute and per-category inter-annotator agreement from the released raw score records; if alpha or intraclass correlation falls below roughly 0.5 on the overall aesthetic score, the averaged ground truth is too noisy to support the reported model comparisons. Alternatively, re-run the Table 4 comparison using split-half annotator averages as the target: if ArtCLIP no longer beats AANSPS when the scores come from a different half of the annotators, the ranking is an artifact of the averaging procedure.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that a painting-aesthetics dataset can reach the scale and annotation depth needed to move the field: APDDv2 contains 10,023 images in 24 categories, with 85,191 expert score annotations averaged from at least six annotators per image and 6,249 language comments covering ten defined attributes. The paper further reports that ArtCLIP, pre-trained with attribute-aware contrastive learning on photographic captions and then fine-tuned on APDDv2, achieves lower mean squared error and higher Spearman correlation than AANSPS and SAAN across all reported score types, which it reads as evidence that the dataset carries genuine aesthetic signal.
Load-bearing premise
The dataset's value and ArtCLIP's reported edge both rest on one unstated premise: that averaging at least six experts' scores yields stable, unbiased ground truth, and the paper gives no agreement statistics to show that the experts actually concur.
Editorial extensions
If this is right
- APDDv2 becomes the first painting-domain benchmark that supports both score prediction and comment generation, so multimodal models can be trained and evaluated on art rather than only on photographs.
- ArtCLIP's reported margins over AANSPS and SAAN imply that fine-grained expert annotation, not just larger image counts, is what lifts art-aesthetic prediction quality.
- The per-category scoring standards give later builders a reusable protocol for adding new images to the dataset without drifting away from earlier versions.
- The applications the authors pursue, teacher-assisted art instruction, children's art education, and quality control for AI-generated art, become testable once the dataset and model are public.
Reading between the lines
- Because the paper reports no inter-annotator agreement statistics, the published model margins could partly reflect noise or bias in the averaged scores; a natural extension is to compute per-category agreement from the released raw annotation records and show that the gains survive within each annotator group.
- The paper had to pre-train on a photography caption dataset for lack of painting commentary data, so the transfer gap between photographic and painterly aesthetics is untested; a painting-specific caption corpus is the obvious next step and would likely push ArtCLIP's scores further.
- The authors state that the dataset captures expert taste only, with no public opinion and thin cultural coverage; a concrete follow-up is comparing expert scores against crowd ratings on a shared image subset, which would reveal where professional and popular aesthetic judgment diverge.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents APDDv2, a dataset of 10,023 paintings and drawings annotated by a team of 37 experts with scores on 10 aesthetic attributes across 24 artistic categories, plus 6,249 free-form language comments. The authors also describe ArtCLIP, a CLIP-based multimodal model pre-trained on a photography caption dataset and fine-tuned on APDDv2, and report that it outperforms the AANSPS and SAAN models on APDDv2. The dataset and model are publicly released under CC BY 4.0.
Significance. If the annotations are reliable, APDDv2 would be a valuable resource: it is larger than previous painting-specific aesthetic datasets with attribute scores, it is the first painting dataset in this scope to include free-form aesthetic comments, and it provides category-specific scoring standards. The authors have shipped a permissively licensed dataset and code, which is a concrete strength. However, the paper currently lacks inter-annotator agreement measures, a fully specified train/test protocol, and error bars, so the reliability of the ground-truth scores and the reported model improvements are not yet established to the standard expected for a benchmark contribution.
major comments (4)
- [Section 3.3, Table 3] The paper reports that each image was rated by at least six annotators and that 533,513 raw records were collapsed by averaging into 85,191 scores, but no inter-annotator agreement statistics (e.g., ICC, Krippendorff's alpha, Fleiss' kappa) are reported for any of the 10 attributes or 24 categories. Because these averaged scores are used as ground truth for every model comparison in Table 4, the dataset's validity as a benchmark depends on demonstrating that annotators apply the scoring standards consistently. Averaging can reduce independent noise but cannot remove systematic rater bias. Please report agreement per attribute and per category, and discuss any categories or attributes with low agreement.
- [Section 4, Table 4] The experimental protocol is under-specified: the paper does not state how APDDv2 is divided into training and test partitions (e.g., ratio, stratification by category, random seed), and the checklist confirms that no error bars are reported. Without a precise split description the results in Table 4 are not reproducible, and without error bars or significance tests the differences between ArtCLIP, AANSPS, and SAAN (e.g., TAS MSE 0.68 vs. 0.88) could be within run-to-run variability. Provide the split details and report mean ± std over at least three seeds or an equivalent significance analysis.
- [Section 4, Section 7] The claim that ArtCLIP 'surpasses state-of-the-art techniques' is based solely on experiments on the APDDv2 test set, and all comparison models are also trained on APDDv2. This does not demonstrate generalization beyond the APDDv2 distribution. To support the state-of-the-art claim, evaluate on held-out painting datasets such as BAID or JenAesthetics, or at minimum temper the conclusion to say that ArtCLIP performs best among models trained and tested on APDDv2.
- [Section 3.3, Section 1] Section 3.3 states that 'each image was rated by at least six annotators and commented on by at least one annotator,' but the paper later reports only 6,249 comment records for 10,023 images, implying that roughly 38% of images have no comment. The abstract and introduction also describe 'over 40 experts,' while Section 3.2 lists 37 annotators. These inconsistencies affect the precise description of the dataset's coverage and should be reconciled.
minor comments (5)
- [Section 3.2] The phrase 'an labeling team' should be 'a labeling team.'
- [Table 1, Section 2.1] The dataset named 'V APS' in the table and text should be 'VAPS' (the Vienna Art Picture System).
- [Table 4] The score type 'L$S' appears to be a formatting artifact; it should be 'L&S' (Light and Shadow) to match Table 2.
- [Figure 8] The caption for Figure 8 does not define the axis labels; please specify what the horizontal and vertical axes represent.
- [Section 3.3] The labeling system URL 'http://103.30.78.19/' is unlikely to be accessible to readers; consider omitting it or providing a description of the interface.
Circularity Check
No significant circularity: the paper's claims are dataset construction and supervised model evaluation, not a derivation that reduces to its own inputs.
full rationale
APDDv2 is a dataset paper; its central contributions are the collection and annotation of 10,023 images and the fine-tuning of ArtCLIP on that dataset. The scoring standards are defined from expert collaboration and from benchmarks derived from APDDv1 annotations, but this is a calibration/provenance step rather than a mathematical derivation that presupposes the target result. The self-citations to the authors' prior APDDv1 work for the 24 categories and 10 attributes supply the taxonomy of the dataset, yet the paper also grounds these attributes in artist cognition and painting evaluation standards, so the citation is descriptive rather than load-bearing in a circular sense. The model comparison in Table 4 trains AANSPS, SAAN, and ArtCLIP on APDDv2 and reports metrics on that dataset; this is a standard supervised benchmark within the released data, and the paper does not claim to predict values outside the annotated set. No passage equates a fitted parameter with a prediction, and no uniqueness theorem or ansatz is smuggled in via self-citation. The absence of inter-annotator agreement statistics and error bars is a data-quality and reporting limitation, but it is not an instance of circular reasoning. Consequently, the derivation chain is self-contained and no circular step can be exhibited from the text.
Assumptions & free parameters
assumptions (3)
- domain assumption Expert annotations averaged over at least six raters are valid ground truth for aesthetic quality.
- domain assumption The 10 aesthetic attributes are comprehensive and meaningful across all 24 artistic categories.
- domain assumption CLIP features and contrastive learning are suitable for aesthetic assessment of art.
Cite this review
Pith. "Pith review of APDDv2: Aesthetics of Paintings and Drawings Dataset with Artist Labeled Scores and Comments." pith.science (2026). https://pith.science/paper/BXR2R6AF
@misc{pith2026241108545,
author = {Pith},
title = {Pith review of: APDDv2: Aesthetics of Paintings and Drawings Dataset with Artist Labeled Scores and Comments},
year = {2026},
howpublished = {\url{https://pith.science/paper/BXR2R6AF}},
note = {Machine review of arXiv:2411.08545}
}
read the original abstract
Datasets play a pivotal role in training visual models, facilitating the development of abstract understandings of visual features through diverse image samples and multidimensional attributes. However, in the realm of aesthetic evaluation of artistic images, datasets remain relatively scarce. Existing painting datasets are often characterized by limited scoring dimensions and insufficient annotations, thereby constraining the advancement and application of automatic aesthetic evaluation methods in the domain of painting. To bridge this gap, we introduce the Aesthetics Paintings and Drawings Dataset (APDD), the first comprehensive collection of paintings encompassing 24 distinct artistic categories and 10 aesthetic attributes. Building upon the initial release of APDDv1, our ongoing research has identified opportunities for enhancement in data scale and annotation precision. Consequently, APDDv2 boasts an expanded image corpus and improved annotation quality, featuring detailed language comments to better cater to the needs of both researchers and practitioners seeking high-quality painting datasets. Furthermore, we present an updated version of the Art Assessment Network for Specific Painting Styles, denoted as ArtCLIP. Experimental validation demonstrates the superior performance of this revised model in the realm of aesthetic evaluation, surpassing its predecessor in accuracy and efficacy. The dataset and model are available at https://github.com/BestiVictory/APDDv2.git.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 2 Pith papers
-
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
ArtiMuse is an MLLM that jointly scores image aesthetics and writes expert-style 8-attribute critiques, trained on a new 10,000-image expert-annotated dataset with a token-based continuous scoring method.
-
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
Ming-Lite-Uni couples a frozen multimodal LLM with a learnable diffusion model via multi-scale learnable tokens to perform text-to-image generation and instruction-based image editing.
Reference graph
Works this paper leans on
-
[1]
Paintings and drawings aesthetics assessment with rich attributes for various artistic categories
Xin Jin, Qianqian Qiao, Yi Lu, Huaye Wang, Shan Gao, Heng Huang, and Guangdong Li. Paintings and drawings aesthetics assessment with rich attributes for various artistic categories. In Proceedings of the 2024 International Joint Conference on Artificial Intelligence (IJCAI), August 2024
work page 2024
-
[2]
Development trends of image aesthetic quality evaluation techniques
Xin Jin, Bin Zhou, Dongqing Zou, Xiaodong Li, Hongbo Sun, and Le Wu. Development trends of image aesthetic quality evaluation techniques. Science and Technology Review, 36 0 (9): 0 36--45, 2018
work page 2018
-
[3]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748--8763. PMLR, 2021
2021
-
[4]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \'a r, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740--755. Springer, 2014
2014
-
[5]
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE international conference on computer vision, pages 2641--2649, 2015
work page 2015
-
[6]
Aesclip: Multi-attribute contrastive learning for image aesthetics assessment
Xiangfei Sheng, Leida Li, Pengfei Chen, Jinjian Wu, Weisheng Dong, Yuzhe Yang, Liwu Xu, Yaqian Li, and Guangming Shi. Aesclip: Multi-attribute contrastive learning for image aesthetics assessment. In Proceedings of the 31st ACM International Conference on Multimedia, pages 1117--1126, 2023
work page 2023
-
[7]
Towards artistic image aesthetics assessment: a large-scale dataset and a new method
Ran Yi, Haoyuan Tian, Zhihao Gu, Yu-Kun Lai, and Paul L Rosin. Towards artistic image aesthetics assessment: a large-scale dataset and a new method. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22388--22397, 2023
work page 2023
-
[8]
Aacp: Aesthetics assessment of children’s paintings based on self-supervised learning
Shiqi Jiang, Ning Li, Chen Shi, Liping Guo, Changbo Wang, and Chenhui Li. Aacp: Aesthetics assessment of children’s paintings based on self-supervised learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 2534--2542, 2024
work page 2024
Show all 14 references
-
[9]
The vienna art picture system (vaps): A data set of 999 paintings and subjective ratings for art and aesthetics research
Anna Fekete, Matthew Pelowski, Eva Specker, David Brieber, Raphael Rosenberg, and Helmut Leder. The vienna art picture system (vaps): A data set of 999 paintings and subjective ratings for art and aesthetics research. Psychology of Aesthetics, Creativity, and the Arts, 2022
2022
-
[10]
Jenaesthetics subjective dataset: analyzing paintings by subjective scores
Seyed Ali Amirshahi, Gregor Uwe Hayn-Leichsenring, Joachim Denzler, and Christoph Redies. Jenaesthetics subjective dataset: analyzing paintings by subjective scores. In Computer Vision-ECCV 2014 Workshops: Zurich, Switzerland, September 6-7 and 12, 2014, Proceedings, Part I 13...
2014
-
[11]
Color: A crucial factor for aesthetic quality assessment in a subjective dataset of paintings
Seyed Ali Amirshahi, Gregor Uwe Hayn-Leichsenring, Joachim Denzler, and Christoph Redies. Color: A crucial factor for aesthetic quality assessment in a subjective dataset of paintings. arXiv preprint arXiv:1609.05583, 2016
2016 arXiv
-
[12]
In the eye of the beholder: employing statistical analysis and eye tracking for analyzing abstract paintings
Victoria Yanulevskaya, Jasper Uijlings, Elia Bruni, Andreza Sartori, Elisa Zamboni, Francesca Bacci, David Melcher, and Nicu Sebe. In the eye of the beholder: employing statistical analysis and eye tracking for analyzing abstract paintings. In Proceedings of the 20th ACM inter...
2012
-
[13]
Wiki art gallery, inc.: A case for critical thinking
Fred Phillips and Brandy Mackintosh. Wiki art gallery, inc.: A case for critical thinking. Issues in Accounting Education, 26 0 (3): 0 593--608, 2011
2011
-
[14]
Aesthetically relevant image captioning
Zhipeng Zhong, Fei Zhou, and Guoping Qiu. Aesthetically relevant image captioning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 3733--3741, 2023
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.