Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

APDDv2: Aesthetics of Paintings and Drawings Dataset with Artist Labeled Scores and Comments

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper introduces APDDv2, a 10,023-image dataset of paintings and drawings rated by experts on ten aesthetic attributes across 24 artistic categories, and reports that the ArtCLIP model trained on it outperforms prior art-aesthetic…

desk verdict A genuinely useful painting-aesthetics dataset with detailed curation, but the missing inter-annotator reliability means the benchmark-quality claim is not yet supported. read the letter →

arxiv 2411.08545 v1 pith:BXR2R6AF submitted 2024-11-13 cs.CV

classification cs.CV
keywords paintingaestheticsdatasetexpertannotationaestheticattributesimageassessmentCLIPfine-tuningcontrastivelearningcomments
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to establish APDDv2 as the first painting-domain benchmark that pairs fine-grained expert scores with free-form aesthetic commentary: 10,023 images spanning 24 artistic categories, each scored by at least six expert annotators on a per-category subset of ten aesthetic attributes, plus 6,249 written comments. Its companion claim is that ArtCLIP, a CLIP-based model fine-tuned on this data, surpasses prior art-aesthetic scorers, AANSPS and SAAN, in accuracy and correlation. This matters because automatic aesthetic evaluation of paintings lags far behind photography, and existing painting datasets are small, thinly annotated, and contain no language. If the claims hold, the field gains a reusable evaluation standard and a model that can score, describe, and potentially guide art creation, education, and AI-generated-image quality control.

What carries the argument

Two mechanisms carry the argument. The first is the annotation rubric: 24 artistic categories cross ten aesthetic attributes, and each category has a benchmark table that ties every score range to representative example images, so expert annotators can apply one continuous standard that also stays comparable to the earlier APDDv1 release. The second is the ArtCLIP training scheme: attribute-aware contrastive learning that pairs an image with comments sampled from the matching attribute category, a multimodal fusion module that merges image and text embeddings, and final regression fine-tuning on APDDv2. The benchmark tables make the averaged expert scores meaningful, while the contrastive pre-training is what imports aesthetic language from the photography domain into painting.

What would settle it

Compute per-attribute and per-category inter-annotator agreement from the released raw score records; if alpha or intraclass correlation falls below roughly 0.5 on the overall aesthetic score, the averaged ground truth is too noisy to support the reported model comparisons. Alternatively, re-run the Table 4 comparison using split-half annotator averages as the target: if ArtCLIP no longer beats AANSPS when the scores come from a different half of the annotators, the ranking is an artifact of the averaging procedure.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that a painting-aesthetics dataset can reach the scale and annotation depth needed to move the field: APDDv2 contains 10,023 images in 24 categories, with 85,191 expert score annotations averaged from at least six annotators per image and 6,249 language comments covering ten defined attributes. The paper further reports that ArtCLIP, pre-trained with attribute-aware contrastive learning on photographic captions and then fine-tuned on APDDv2, achieves lower mean squared error and higher Spearman correlation than AANSPS and SAAN across all reported score types, which it reads as evidence that the dataset carries genuine aesthetic signal.

Load-bearing premise

The dataset's value and ArtCLIP's reported edge both rest on one unstated premise: that averaging at least six experts' scores yields stable, unbiased ground truth, and the paper gives no agreement statistics to show that the experts actually concur.

Editorial extensions

If this is right

  • APDDv2 becomes the first painting-domain benchmark that supports both score prediction and comment generation, so multimodal models can be trained and evaluated on art rather than only on photographs.
  • ArtCLIP's reported margins over AANSPS and SAAN imply that fine-grained expert annotation, not just larger image counts, is what lifts art-aesthetic prediction quality.
  • The per-category scoring standards give later builders a reusable protocol for adding new images to the dataset without drifting away from earlier versions.
  • The applications the authors pursue, teacher-assisted art instruction, children's art education, and quality control for AI-generated art, become testable once the dataset and model are public.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the paper reports no inter-annotator agreement statistics, the published model margins could partly reflect noise or bias in the averaged scores; a natural extension is to compute per-category agreement from the released raw annotation records and show that the gains survive within each annotator group.
  • The paper had to pre-train on a photography caption dataset for lack of painting commentary data, so the transfer gap between photographic and painterly aesthetics is untested; a painting-specific caption corpus is the obvious next step and would likely push ArtCLIP's scores further.
  • The authors state that the dataset captures expert taste only, with no public opinion and thin cultural coverage; a concrete follow-up is comparing expert scores against crowd ratings on a shared image subset, which would reveal where professional and popular aesthetic judgment diverge.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents APDDv2, a dataset of 10,023 paintings and drawings annotated by a team of 37 experts with scores on 10 aesthetic attributes across 24 artistic categories, plus 6,249 free-form language comments. The authors also describe ArtCLIP, a CLIP-based multimodal model pre-trained on a photography caption dataset and fine-tuned on APDDv2, and report that it outperforms the AANSPS and SAAN models on APDDv2. The dataset and model are publicly released under CC BY 4.0.

Significance. If the annotations are reliable, APDDv2 would be a valuable resource: it is larger than previous painting-specific aesthetic datasets with attribute scores, it is the first painting dataset in this scope to include free-form aesthetic comments, and it provides category-specific scoring standards. The authors have shipped a permissively licensed dataset and code, which is a concrete strength. However, the paper currently lacks inter-annotator agreement measures, a fully specified train/test protocol, and error bars, so the reliability of the ground-truth scores and the reported model improvements are not yet established to the standard expected for a benchmark contribution.

major comments (4)
  1. [Section 3.3, Table 3] The paper reports that each image was rated by at least six annotators and that 533,513 raw records were collapsed by averaging into 85,191 scores, but no inter-annotator agreement statistics (e.g., ICC, Krippendorff's alpha, Fleiss' kappa) are reported for any of the 10 attributes or 24 categories. Because these averaged scores are used as ground truth for every model comparison in Table 4, the dataset's validity as a benchmark depends on demonstrating that annotators apply the scoring standards consistently. Averaging can reduce independent noise but cannot remove systematic rater bias. Please report agreement per attribute and per category, and discuss any categories or attributes with low agreement.
  2. [Section 4, Table 4] The experimental protocol is under-specified: the paper does not state how APDDv2 is divided into training and test partitions (e.g., ratio, stratification by category, random seed), and the checklist confirms that no error bars are reported. Without a precise split description the results in Table 4 are not reproducible, and without error bars or significance tests the differences between ArtCLIP, AANSPS, and SAAN (e.g., TAS MSE 0.68 vs. 0.88) could be within run-to-run variability. Provide the split details and report mean ± std over at least three seeds or an equivalent significance analysis.
  3. [Section 4, Section 7] The claim that ArtCLIP 'surpasses state-of-the-art techniques' is based solely on experiments on the APDDv2 test set, and all comparison models are also trained on APDDv2. This does not demonstrate generalization beyond the APDDv2 distribution. To support the state-of-the-art claim, evaluate on held-out painting datasets such as BAID or JenAesthetics, or at minimum temper the conclusion to say that ArtCLIP performs best among models trained and tested on APDDv2.
  4. [Section 3.3, Section 1] Section 3.3 states that 'each image was rated by at least six annotators and commented on by at least one annotator,' but the paper later reports only 6,249 comment records for 10,023 images, implying that roughly 38% of images have no comment. The abstract and introduction also describe 'over 40 experts,' while Section 3.2 lists 37 annotators. These inconsistencies affect the precise description of the dataset's coverage and should be reconciled.
minor comments (5)
  1. [Section 3.2] The phrase 'an labeling team' should be 'a labeling team.'
  2. [Table 1, Section 2.1] The dataset named 'V APS' in the table and text should be 'VAPS' (the Vienna Art Picture System).
  3. [Table 4] The score type 'L$S' appears to be a formatting artifact; it should be 'L&S' (Light and Shadow) to match Table 2.
  4. [Figure 8] The caption for Figure 8 does not define the axis labels; please specify what the horizontal and vertical axes represent.
  5. [Section 3.3] The labeling system URL 'http://103.30.78.19/' is unlikely to be accessible to readers; consider omitting it or providing a description of the interface.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims are dataset construction and supervised model evaluation, not a derivation that reduces to its own inputs.

full rationale

APDDv2 is a dataset paper; its central contributions are the collection and annotation of 10,023 images and the fine-tuning of ArtCLIP on that dataset. The scoring standards are defined from expert collaboration and from benchmarks derived from APDDv1 annotations, but this is a calibration/provenance step rather than a mathematical derivation that presupposes the target result. The self-citations to the authors' prior APDDv1 work for the 24 categories and 10 attributes supply the taxonomy of the dataset, yet the paper also grounds these attributes in artist cognition and painting evaluation standards, so the citation is descriptive rather than load-bearing in a circular sense. The model comparison in Table 4 trains AANSPS, SAAN, and ArtCLIP on APDDv2 and reports metrics on that dataset; this is a standard supervised benchmark within the released data, and the paper does not claim to predict values outside the annotated set. No passage equates a fitted parameter with a prediction, and no uniqueness theorem or ansatz is smuggled in via self-citation. The absence of inter-annotator agreement statistics and error bars is a data-quality and reporting limitation, but it is not an instance of circular reasoning. Consequently, the derivation chain is self-contained and no circular step can be exhibited from the text.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free numeric parameters in a derivation sense; the central claim rests on the validity of the annotation process, the chosen attribute schema, and the suitability of CLIP for art aesthetics. All of these are domain assumptions made without direct evidence such as inter-annotator agreement or cross-dataset validation.

assumptions (3)
  • domain assumption Expert annotations averaged over at least six raters are valid ground truth for aesthetic quality.
    Section 3.3 describes averaging scores from at least six annotators per image, but no inter-annotator agreement statistics are reported, so reliability is assumed.
  • domain assumption The 10 aesthetic attributes are comprehensive and meaningful across all 24 artistic categories.
    Section 3.1 defines the attributes and Figure 4 maps them to categories, but Section 6 acknowledges cultural and subjective variation may limit this schema.
  • domain assumption CLIP features and contrastive learning are suitable for aesthetic assessment of art.
    Section 4 uses CLIP as a backbone, a practical assumption borrowed from AesCLIP and not derived in this paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of APDDv2: Aesthetics of Paintings and Drawings Dataset with Artist Labeled Scores and Comments." pith.science (2026). https://pith.science/paper/BXR2R6AF

@misc{pith2026241108545,
  author       = {Pith},
  title        = {Pith review of: APDDv2: Aesthetics of Paintings and Drawings Dataset with Artist Labeled Scores and Comments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BXR2R6AF}},
  note         = {Machine review of arXiv:2411.08545}
}
read the original abstract

Datasets play a pivotal role in training visual models, facilitating the development of abstract understandings of visual features through diverse image samples and multidimensional attributes. However, in the realm of aesthetic evaluation of artistic images, datasets remain relatively scarce. Existing painting datasets are often characterized by limited scoring dimensions and insufficient annotations, thereby constraining the advancement and application of automatic aesthetic evaluation methods in the domain of painting. To bridge this gap, we introduce the Aesthetics Paintings and Drawings Dataset (APDD), the first comprehensive collection of paintings encompassing 24 distinct artistic categories and 10 aesthetic attributes. Building upon the initial release of APDDv1, our ongoing research has identified opportunities for enhancement in data scale and annotation precision. Consequently, APDDv2 boasts an expanded image corpus and improved annotation quality, featuring detailed language comments to better cater to the needs of both researchers and practitioners seeking high-quality painting datasets. Furthermore, we present an updated version of the Art Assessment Network for Specific Painting Styles, denoted as ArtCLIP. Experimental validation demonstrates the superior performance of this revised model in the realm of aesthetic evaluation, surpassing its predecessor in accuracy and efficacy. The dataset and model are available at https://github.com/BestiVictory/APDDv2.git.

Figures

Figures reproduced from arXiv: 2411.08545 by the authors.

Figure 1
Figure 1. Samples from the APDDv2 dataset. ∗Corresponding author: Heng Huang (hecate@mail.ustc.edu.cn) Submitted to the 38th Conference on Neural Information Processing Systems (NeurIPS 2024) Track on Datasets and Benchmarks. Do not distribute. arXiv:2411.08545v1 [cs.CV] 13 Nov 2024 [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The five-layered tasks of IAQA exhibit an overall inverted triangular distribution in terms [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. 24 Artistic Categories in the APDD Dataset. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Correspondence between artistic categories and aesthetic attributes. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Labeling Team Composition. As shown in [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Scoring benchmark table for "Oil Painting - Symbolism - Still Life" category. [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Benchmark table for language comments in "Oil Painting - Symbolism - Still Life" category. [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Distribution of Total Aesthetic Score. Most images received scores between 50 and 70, [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: ArtCLIP samples two comments from different aesthetic attribute perspectives and further [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: Test samples. Predicted represents the predicted score of the ArtCLIP output. GT represents the ground-truth score. 5 Prospects for APDDv2 and ArtCLIP applications The integration of APDDv2 and ArtCLIP provides a powerful multimodal platform for the aesthetic assessme…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ArtiMuse is an MLLM that jointly scores image aesthetics and writes expert-style 8-attribute critiques, trained on a new 10,000-image expert-annotated dataset with a token-based continuous scoring method.

  2. Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Ming-Lite-Uni couples a frozen multimodal LLM with a learnable diffusion model via multi-scale learnable tokens to perform text-to-image generation and instruction-based image editing.

Reference graph

Works this paper leans on

14 extracted references · 11 canonical work pages · cited by 2 Pith papers

  1. [1]

    Paintings and drawings aesthetics assessment with rich attributes for various artistic categories

    Xin Jin, Qianqian Qiao, Yi Lu, Huaye Wang, Shan Gao, Heng Huang, and Guangdong Li. Paintings and drawings aesthetics assessment with rich attributes for various artistic categories. In Proceedings of the 2024 International Joint Conference on Artificial Intelligence (IJCAI), August 2024

  2. [2]

    Development trends of image aesthetic quality evaluation techniques

    Xin Jin, Bin Zhou, Dongqing Zou, Xiaodong Li, Hongbo Sun, and Le Wu. Development trends of image aesthetic quality evaluation techniques. Science and Technology Review, 36 0 (9): 0 36--45, 2018

  3. [3]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748--8763. PMLR, 2021

  4. [4]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \'a r, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740--755. Springer, 2014

  5. [5]

    Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

    Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE international conference on computer vision, pages 2641--2649, 2015

  6. [6]

    Aesclip: Multi-attribute contrastive learning for image aesthetics assessment

    Xiangfei Sheng, Leida Li, Pengfei Chen, Jinjian Wu, Weisheng Dong, Yuzhe Yang, Liwu Xu, Yaqian Li, and Guangming Shi. Aesclip: Multi-attribute contrastive learning for image aesthetics assessment. In Proceedings of the 31st ACM International Conference on Multimedia, pages 1117--1126, 2023

  7. [7]

    Towards artistic image aesthetics assessment: a large-scale dataset and a new method

    Ran Yi, Haoyuan Tian, Zhihao Gu, Yu-Kun Lai, and Paul L Rosin. Towards artistic image aesthetics assessment: a large-scale dataset and a new method. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22388--22397, 2023

  8. [8]

    Aacp: Aesthetics assessment of children’s paintings based on self-supervised learning

    Shiqi Jiang, Ning Li, Chen Shi, Liping Guo, Changbo Wang, and Chenhui Li. Aacp: Aesthetics assessment of children’s paintings based on self-supervised learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 2534--2542, 2024

Show all 14 references
  1. [9]

    The vienna art picture system (vaps): A data set of 999 paintings and subjective ratings for art and aesthetics research

    Anna Fekete, Matthew Pelowski, Eva Specker, David Brieber, Raphael Rosenberg, and Helmut Leder. The vienna art picture system (vaps): A data set of 999 paintings and subjective ratings for art and aesthetics research. Psychology of Aesthetics, Creativity, and the Arts, 2022

  2. [10]

    Jenaesthetics subjective dataset: analyzing paintings by subjective scores

    Seyed Ali Amirshahi, Gregor Uwe Hayn-Leichsenring, Joachim Denzler, and Christoph Redies. Jenaesthetics subjective dataset: analyzing paintings by subjective scores. In Computer Vision-ECCV 2014 Workshops: Zurich, Switzerland, September 6-7 and 12, 2014, Proceedings, Part I 13...

  3. [11]

    Color: A crucial factor for aesthetic quality assessment in a subjective dataset of paintings

    Seyed Ali Amirshahi, Gregor Uwe Hayn-Leichsenring, Joachim Denzler, and Christoph Redies. Color: A crucial factor for aesthetic quality assessment in a subjective dataset of paintings. arXiv preprint arXiv:1609.05583, 2016

  4. [12]

    In the eye of the beholder: employing statistical analysis and eye tracking for analyzing abstract paintings

    Victoria Yanulevskaya, Jasper Uijlings, Elia Bruni, Andreza Sartori, Elisa Zamboni, Francesca Bacci, David Melcher, and Nicu Sebe. In the eye of the beholder: employing statistical analysis and eye tracking for analyzing abstract paintings. In Proceedings of the 20th ACM inter...

  5. [13]

    Wiki art gallery, inc.: A case for critical thinking

    Fred Phillips and Brandy Mackintosh. Wiki art gallery, inc.: A case for critical thinking. Issues in Accounting Education, 26 0 (3): 0 593--608, 2011

  6. [14]

    Aesthetically relevant image captioning

    Zhipeng Zhong, Fei Zhou, and Guoping Qiu. Aesthetically relevant image captioning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 3733--3741, 2023

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.