{"id":"482da0cd-2a03-4752-9a90-987ad614e80b","arxiv_id":"2411.08545","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"APDDv2 is a 10,023-image painting dataset with expert scores on 10 aesthetic attributes and language comments, plus ArtCLIP, a CLIP-based model that outperforms prior art-scoring models on it.","lead":"This paper releases APDDv2, a dataset of about 10,000 paintings and drawings scored by expert artists on 10 aesthetic attributes and annotated with language comments. It also presents ArtCLIP, a CLIP-based model for scoring art, and reports that it beats two previous models on this dataset.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No inter-annotator reliability is reported; the averaged expert scores that serve as ground truth are unvalidated, so the central claim of a high-quality benchmark is not yet established.","rationale":"The reader's weakest assumption—annotation reliability—is indeed the single most load-bearing concern. A dataset built from subjective aesthetic ratings must demonstrate that raters can agree; otherwise benchmark scores and model comparisons are comparisons of noise. The paper's careful process (expert teams, benchmark tables, at least six annotators per image) supports plausibility but not validity. The complaint is not that the scores are definitely wrong; it is that the necessary evidence is missing. A conditional verdict is appropriate because the missing IRR statistics, data-split details, and error bars are all empirically addressable. The central dataset and model claims should not be accepted unconditionally until reliability of the ground-truth scores is shown.","tokens_in":8916,"tokens_out":4540,"duration_ms":41178,"concrete_test":"Compute inter-annotator agreement on the pre-averaging annotations: for each of the 10 attributes and 24 categories, calculate ICC(2,k) for the continuous scores and Krippendorff's alpha for the comment annotations (or release the per-annotator scores so reviewers can compute these). Set a pre-specified threshold, e.g., ICC >= 0.7 for the Total Aesthetic Score and for the majority of attributes. If agreement falls below the threshold, the averaged scores are unstable and the dataset/model claims require revision; if agreement meets the threshold, this concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 describes averaging scores from at least six annotators per image (Table 3: 533,513 raw records collapsed to 85,191 averaged scores) but reports no inter-annotator agreement statistic (ICC, Krippendorff's alpha, Fleiss' kappa, etc.) for any of the 10 attributes or 24 categories. This is load-bearing because the dataset's entire value and every model ranking in Table 4 use these averaged scores as ground truth. Averaging only reduces independent noise; it cannot remove systematic rater bias, and if annotators disagree strongly, the 'ground truth' is essentially arbitrary. The benchmark tables in Figure 6 may establish shared standards, but they do not establish that annotators actually follow them consistently. The Limitations section concedes subjective and cultural variation, and the checklist explicitly admits that no error bars are reported. Consequently, the central claims—that APDDv2 is a valid, high-quality expert-annotated benchmark and that ArtCLIP 'surpasses state-of-the-art' on it—are not yet supported by the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents APDDv2, a dataset of 10,023 paintings and drawings annotated by a team of 37 experts with scores on 10 aesthetic attributes across 24 artistic categories, plus 6,249 free-form language comments. The authors also describe ArtCLIP, a CLIP-based multimodal model pre-trained on a photography caption dataset and fine-tuned on APDDv2, and report that it outperforms the AANSPS and SAAN models on APDDv2. The dataset and model are publicly released under CC BY 4.0.","tokens_in":9095,"tokens_out":5386,"duration_ms":43712,"significance":"If the annotations are reliable, APDDv2 would be a valuable resource: it is larger than previous painting-specific aesthetic datasets with attribute scores, it is the first painting dataset in this scope to include free-form aesthetic comments, and it provides category-specific scoring standards. The authors have shipped a permissively licensed dataset and code, which is a concrete strength. However, the paper currently lacks inter-annotator agreement measures, a fully specified train/test protocol, and error bars, so the reliability of the ground-truth scores and the reported model improvements are not yet established to the standard expected for a benchmark contribution.","major_comments":[{"comment":"The paper reports that each image was rated by at least six annotators and that 533,513 raw records were collapsed by averaging into 85,191 scores, but no inter-annotator agreement statistics (e.g., ICC, Krippendorff's alpha, Fleiss' kappa) are reported for any of the 10 attributes or 24 categories. Because these averaged scores are used as ground truth for every model comparison in Table 4, the dataset's validity as a benchmark depends on demonstrating that annotators apply the scoring standards consistently. Averaging can reduce independent noise but cannot remove systematic rater bias. Please report agreement per attribute and per category, and discuss any categories or attributes with low agreement.","section":"Section 3.3, Table 3"},{"comment":"The experimental protocol is under-specified: the paper does not state how APDDv2 is divided into training and test partitions (e.g., ratio, stratification by category, random seed), and the checklist confirms that no error bars are reported. Without a precise split description the results in Table 4 are not reproducible, and without error bars or significance tests the differences between ArtCLIP, AANSPS, and SAAN (e.g., TAS MSE 0.68 vs. 0.88) could be within run-to-run variability. Provide the split details and report mean ± std over at least three seeds or an equivalent significance analysis.","section":"Section 4, Table 4"},{"comment":"The claim that ArtCLIP 'surpasses state-of-the-art techniques' is based solely on experiments on the APDDv2 test set, and all comparison models are also trained on APDDv2. This does not demonstrate generalization beyond the APDDv2 distribution. To support the state-of-the-art claim, evaluate on held-out painting datasets such as BAID or JenAesthetics, or at minimum temper the conclusion to say that ArtCLIP performs best among models trained and tested on APDDv2.","section":"Section 4, Section 7"},{"comment":"Section 3.3 states that 'each image was rated by at least six annotators and commented on by at least one annotator,' but the paper later reports only 6,249 comment records for 10,023 images, implying that roughly 38% of images have no comment. The abstract and introduction also describe 'over 40 experts,' while Section 3.2 lists 37 annotators. These inconsistencies affect the precise description of the dataset's coverage and should be reconciled.","section":"Section 3.3, Section 1"}],"minor_comments":[{"comment":"The phrase 'an labeling team' should be 'a labeling team.'","section":"Section 3.2"},{"comment":"The dataset named 'V APS' in the table and text should be 'VAPS' (the Vienna Art Picture System).","section":"Table 1, Section 2.1"},{"comment":"The score type 'L$S' appears to be a formatting artifact; it should be 'L&S' (Light and Shadow) to match Table 2.","section":"Table 4"},{"comment":"The caption for Figure 8 does not define the axis labels; please specify what the horizontal and vertical axes represent.","section":"Figure 8"},{"comment":"The labeling system URL 'http://103.30.78.19/' is unlikely to be accessible to readers; consider omitting it or providing a description of the interface.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The dataset is potentially useful and the authors have been transparent about several limitations in the checklist. The main risk is the missing inter-annotator agreement analysis and under-specified experimental protocol; both are standard expectations for a datasets-and-benchmarks submission and should be addressable in a revision. I recommend major revision rather than rejection because the core dataset construction is described in detail and the identified issues are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"APDDv2 is a serious dataset contribution: it roughly doubles APDDv1 to 10k paintings, adds language comments, and ships a benchmark table system plus a contrastive model. The curation process is described in unusual detail — three annotator groups, 24 category-specific benchmark tables, at least six scores per image. That is real work, and the dataset fills a genuine gap: no other painting aesthetics dataset has both multi-attribute scores and free-text comments at this scale. ArtCLIP is not conceptually new — it is AesCLIP adapted to art — but the paper does not oversell that; it positions it as an updated model.\n\nThe soft spots are the ones the stress-test flags. No inter-annotator agreement statistic appears anywhere, and the ground truth is a mean of six raters. That is load-bearing: every model comparison in Table 4 uses these means as labels, and rater bias is not removed by averaging. The checklist admits no error bars, and the train/test split is described only as 'data splits' with no actual description in the paper. For a benchmark paper, that is a significant omission, not a nitpick. The 'superior performance' claim also rests on training and evaluating all models on the same APDDv2 distribution, so it is in-distribution superiority; cross-dataset validation (e.g., on BAID or JenAesthetics) would be needed to say anything general.\n\nI disagree lightly with the reader's circularity burden of 3.0: for a dataset paper, training and testing on the same distribution is standard if the claim is 'our model works on our benchmark.' The problem is not circularity per se, it is that without reliability statistics we cannot know whether the benchmark itself is stable. The self-citation pattern is fine — APDDv1 is the direct predecessor and is cited properly.\n\nBottom line: this deserves a serious referee. The dataset is likely to become a standard resource regardless of the missing statistics, and the authors can plausibly add ICC or Krippendorff's alpha, error bars, and a clear split description. As submitted, the benchmark-quality claim is not yet established. I would send it to review with a request for those analyses rather than desk-reject.","headline":"A genuinely useful painting-aesthetics dataset with detailed curation, but the missing inter-annotator reliability means the benchmark-quality claim is not yet supported.","tokens_in":9651,"tokens_out":1802,"would_cite":true,"duration_ms":16129,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces APDDv2, a 10,023-image dataset of paintings and drawings rated by experts on ten aesthetic attributes across 24 artistic categories, and reports that the ArtCLIP model trained on it outperforms prior art-aesthetic…","keywords":["painting aesthetics dataset","expert annotation","aesthetic attributes","image aesthetic assessment","CLIP fine-tuning","contrastive learning","aesthetic comments"],"falsifier":"Compute per-attribute and per-category inter-annotator agreement from the released raw score records; if alpha or intraclass correlation falls below roughly 0.5 on the overall aesthetic score, the averaged ground truth is too noisy to support the reported model comparisons. Alternatively, re-run the Table 4 comparison using split-half annotator averages as the target: if ArtCLIP no longer beats AANSPS when the scores come from a different half of the annotators, the ranking is an artifact of the averaging procedure.","tokens_in":8716,"feed_emoji":"🎨","tokens_out":11712,"duration_ms":87287,"temperature":0.7,"pith_summary":"The paper aims to establish APDDv2 as the first painting-domain benchmark that pairs fine-grained expert scores with free-form aesthetic commentary: 10,023 images spanning 24 artistic categories, each scored by at least six expert annotators on a per-category subset of ten aesthetic attributes, plus 6,249 written comments. Its companion claim is that ArtCLIP, a CLIP-based model fine-tuned on this data, surpasses prior art-aesthetic scorers, AANSPS and SAAN, in accuracy and correlation. This matters because automatic aesthetic evaluation of paintings lags far behind photography, and existing painting datasets are small, thinly annotated, and contain no language. If the claims hold, the field gains a reusable evaluation standard and a model that can score, describe, and potentially guide art creation, education, and AI-generated-image quality control.","feed_headline":"10,023 paintings get expert beauty scores and critiques","feed_subtitle":"First painting benchmark pairing ten attribute scores with written comments; its CLIP model tops prior art scorers.","key_machinery":"Two mechanisms carry the argument. The first is the annotation rubric: 24 artistic categories cross ten aesthetic attributes, and each category has a benchmark table that ties every score range to representative example images, so expert annotators can apply one continuous standard that also stays comparable to the earlier APDDv1 release. The second is the ArtCLIP training scheme: attribute-aware contrastive learning that pairs an image with comments sampled from the matching attribute category, a multimodal fusion module that merges image and text embeddings, and final regression fine-tuning on APDDv2. The benchmark tables make the averaged expert scores meaningful, while the contrastive pre-training is what imports aesthetic language from the photography domain into painting.","core_discovery":"On the paper's own terms, the discovery is that a painting-aesthetics dataset can reach the scale and annotation depth needed to move the field: APDDv2 contains 10,023 images in 24 categories, with 85,191 expert score annotations averaged from at least six annotators per image and 6,249 language comments covering ten defined attributes. The paper further reports that ArtCLIP, pre-trained with attribute-aware contrastive learning on photographic captions and then fine-tuned on APDDv2, achieves lower mean squared error and higher Spearman correlation than AANSPS and SAAN across all reported score types, which it reads as evidence that the dataset carries genuine aesthetic signal.","pith_inferences":["Because the paper reports no inter-annotator agreement statistics, the published model margins could partly reflect noise or bias in the averaged scores; a natural extension is to compute per-category agreement from the released raw annotation records and show that the gains survive within each annotator group.","The paper had to pre-train on a photography caption dataset for lack of painting commentary data, so the transfer gap between photographic and painterly aesthetics is untested; a painting-specific caption corpus is the obvious next step and would likely push ArtCLIP's scores further.","The authors state that the dataset captures expert taste only, with no public opinion and thin cultural coverage; a concrete follow-up is comparing expert scores against crowd ratings on a shared image subset, which would reveal where professional and popular aesthetic judgment diverge."],"forward_implications":["APDDv2 becomes the first painting-domain benchmark that supports both score prediction and comment generation, so multimodal models can be trained and evaluated on art rather than only on photographs.","ArtCLIP's reported margins over AANSPS and SAAN imply that fine-grained expert annotation, not just larger image counts, is what lifts art-aesthetic prediction quality.","The per-category scoring standards give later builders a reusable protocol for adding new images to the dataset without drifting away from earlier versions.","The applications the authors pursue, teacher-assisted art instruction, children's art education, and quality control for AI-generated art, become testable once the dataset and model are public."],"supporting_citations":[{"why":"APDDv1, the predecessor dataset that supplies the 24 categories, 10 attributes, and scoring standards APDDv2 extends, and the AANSPS baseline ArtCLIP is compared against.","marker":"[Jin et al., 2024]"},{"why":"CLIP, the image-text model whose encoder ArtCLIP fine-tunes for aesthetic scoring.","marker":"[Radford et al., 2021]"},{"why":"AesCLIP, the multi-attribute contrastive learning approach that ArtCLIP's pretraining design explicitly follows.","marker":"[Sheng et al., 2023]"},{"why":"DPC2022, the aesthetic caption dataset from photography used to pre-train ArtCLIP before fine-tuning on APDDv2.","marker":"[Zhong et al., 2023]"},{"why":"SAAN and the BAID dataset, the state-of-the-art art-aesthetic model and large art dataset that ArtCLIP claims to surpass.","marker":"[Yi et al., 2023]"}],"fun_headline_variants":["10k paintings, 24 styles, 10 aesthetics: new benchmark","ArtCLIP beats prior models with 10k-painting aesthetic dataset","85k expert beauty scores and 6k critiques for paintings","Dataset pairs 10 attribute scores with written artist comments","New painting dataset fuels better aesthetic scoring model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's value and ArtCLIP's reported edge both rest on one unstated premise: that averaging at least six experts' scores yields stable, unbiased ground truth, and the paper gives no agreement statistics to show that the experts actually concur.","fun_headline_variants_meta":{"raw":{"variants":["10k paintings, 24 styles, 10 aesthetics: new benchmark","ArtCLIP beats prior models with 10k-painting aesthetic dataset","85k expert beauty scores and 6k critiques for paintings","Dataset pairs 10 attribute scores with written artist comments","New painting dataset fuels better aesthetic scoring model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1346,"prompt_tokens":905,"completion_tokens":441,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":357}},"tokens_in":521,"tokens_out":441,"duration_ms":5304,"temperature":1.0,"reasoning_tokens":357,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:30:53.974754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute per-attribute and per-category inter-annotator agreement from the released raw score records; if alpha or intraclass correlation falls below roughly 0.5 on the overall aesthetic score, the averaged ground truth is too noisy to support the reported model comparisons. Alternatively, re-run the Table 4 comparison using split-half annotator averages as the target: if ArtCLIP no longer beats AANSPS when the scores come from a different half of the annotators, the ranking is an artifact of the averaging procedure.","supporting_citations":[{"cited_title":"Paintings and drawings aesthetics assessment with rich attributes for various artistic categories","cited_arxiv_id":null,"evidence_quote":"APDDv1, the predecessor dataset that supplies the 24 categories, 10 attributes, and scoring standards APDDv2 extends, and the AANSPS baseline ArtCLIP is compared against."},{"cited_title":"Aesclip: Multi-attribute contrastive learning for image aesthetics assessment","cited_arxiv_id":null,"evidence_quote":"AesCLIP, the multi-attribute contrastive learning approach that ArtCLIP's pretraining design explicitly follows."},{"cited_title":"Aesthetically relevant image captioning","cited_arxiv_id":null,"evidence_quote":"DPC2022, the aesthetic caption dataset from photography used to pre-train ArtCLIP before fine-tuning on APDDv2."}],"review_version":1}