{"id":"9bcd3e53-1aa3-4707-a02b-6481ccf95649","arxiv_id":"2411.15560","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Four LLMs show high agreement when scoring and ranking alternative uses and do not favor their own outputs, but the accuracy benchmark is derived from the generation prompts rather than human judgment.","lead":"Four large language models were asked to score and rank creative alternative uses for everyday objects, and they largely agreed with one another. The paper's stronger claim, that this agreement validates the reliability of LLMs as creativity judges, rests on a benchmark built from the same prompts that generated the answers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Oracle-based 'accuracy' is circular: the ground truth in §3.2.4 is derived from self-cited prior work with no independent human validation, so the reliability claim rests on an unverified assumption.","rationale":"The reader's weakest_assumption correctly identifies the oracle as the load-bearing element: the accuracy/reliability claim is only as strong as the ground truth, and that ground truth is constructed from a self-cited prior study with overlapping authors and no human validation. My stress-test confirms this is the single most important concern. The inter-model agreement result is a legitimate descriptive finding, but the abstract's stronger claim—'validating the reliability of LLMs in creativity assessment'—hinges on the oracle. Since the reader already assigned a CONDITIONAL verdict based on this concern, and since the issue is addressable with human ratings or an independent re-derivation of the oracle, the verdict should remain unchanged. I do not identify an additional concern that would warrant rejection: the paper is transparent about its oracle construction, and the descriptive agreement result is internally consistent. The proposed concrete test—independent human ratings—would settle whether the concern lands. If the human ordering aligns with the oracle, the paper's reliability claim gains real support; if not, the central claim must be weakened to 'LLMs agree with each other,' which is a much narrower contribution.","tokens_in":15065,"tokens_out":2371,"duration_ms":23453,"concrete_test":"Recruit independent human raters (at least 3, blind to prompt category) to score a random sample of the 300 generated AUs on a standard AUT rubric (originality and utility, 1–5). Compute (a) whether mean human scores satisfy common < creative < highly_creative for each object, and (b) Spearman correlations between each LLM's raw scores and mean human ratings, with bootstrapped confidence intervals. If the human category ordering does not reproduce the oracle ordering, or if LLM-human correlations are substantially lower than the reported LLM-oracle correlations, the oracle is not valid ground truth and the reliability claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—'validating the reliability of LLMs in creativity assessment'—depends on the oracle defined in §3.2.4 (Equation 1). This oracle orders the 60 AUs per object as [highly_creative_1..4, creative_1..4, common_1..4], justified solely by 'relying on the results of [2]'. Reference [2] is Góes et al. (ICCC 2023), which shares co-authors with this paper and whose 'results' are themselves LLM evaluations generated by forceful prompts, not human creativity judgments. Therefore, the reported oracle correlations (0.77–0.97) measure agreement with an assumed category ordering, not with actual creativity as judged by humans. If the three prompt categories do not correspond to true creativity differences (e.g., 'creative' is not actually more creative than 'common' in human judgment), then the 'accuracy' claim collapses to a statement about self-consistency among LLMs. The inter-model agreement finding (Figure 7) survives independently, but it is a descriptive consensus result, not evidence of reliability for automated creativity assessment. The absence of human validation is especially consequential because the aggregation step—averaging 5 AUs per model and category into 12 values per object before computing Spearman correlation—can inflate correlations and mask within-category disagreements. The paper acknowledges no human baseline and provides no uncertainty estimates, so the abstract's conclusion overreaches what the data can support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether four LLMs (Claude 3.5 Sonnet, ChatGPT-4, ChatGPT-4o, Gemini 1.5 Flash) agree in evaluating the creativity of Alternative Uses Test (AUT) responses. AUs were generated by prompting each model to produce common, creative, and highly creative uses for five objects, using a forceful-prompt technique from a prior paper [2]. Each model then evaluated all AUs under four conditions: scoring vs. ranking and comprehensive (60 AUs) vs. segmented (12 AUs). Spearman rank correlations were computed between each model's evaluations and an oracle defined by the generation categories, and between models. The paper reports average oracle correlations above 0.77 and inter-model correlations above 0.7, and concludes that LLMs are impartial and reliable for automated creativity assessment.","tokens_in":15356,"tokens_out":5663,"duration_ms":52871,"significance":"If the inter-model agreement result is robust, the paper provides a useful descriptive finding about consensus among current LLMs in ranking AUT responses. The benchmark construction methodology (category-based oracle from forceful prompts) is also of interest, although its validity as ground truth is not established. The main contribution is a descriptive consensus study, not a validation of LLM creativity assessment, because no human ground truth is involved. The work is potentially relevant to practitioners considering LLM-based evaluation, but the accuracy/reliability claims overreach the evidence.","major_comments":[{"comment":"The oracle is defined entirely by the generation-prompt categories, and its ordering is justified only by reference to [2], which shares authors with this manuscript. There is no independent human validation that the 'highly_creative', 'creative', and 'common' categories actually differ in creativity. Consequently, the reported oracle correlations (0.77–0.97) largely measure whether the evaluator reproduces the assumed category ordering, not accuracy in assessing creativity. Because the oracle assigns tied positions to all four highly_creative entries, all four creative entries, and all four common entries, a perfect Spearman correlation only requires the three category-level means to be ordered; within-category ordering is never tested. This is load-bearing for the abstract's claim that the results 'validate the reliability of LLMs in creativity assessment'. The paper should either supply human ratings of the AUs to validate the oracle or explicitly restrict its claims to inter-model agreement.","section":"§3.2.4, Eq. (1)"},{"comment":"All Spearman correlations are computed on n=12 aggregated values per object (the average score or rank of five AUs per model-category cell), and no confidence intervals, standard errors, or significance tests are reported. The abstract's claim of correlations 'averaging above 0.7' is a descriptive statistic with no uncertainty quantification. Given the small effective sample size, the aggregation can inflate correlation estimates, and the high values may be driven by one or two objects. The authors should report bootstrap confidence intervals, per-object correlation distributions, or a multilevel analysis to support the claimed magnitude and stability of the correlations.","section":"§3.2.3, Tables 2–4, Figure 7"},{"comment":"The statement that 'LLMs do not favor their own responses' is not supported by a statistical test. The evidence consists of informal comparisons of average scores in Table 5 and in the appendix tables. A model may rate its own AUs higher because its AUs are genuinely more creative, and a proper impartiality test would compare each model's self-evaluation with the average evaluation of the same AUs by the other models, with an error term accounting for per-AU variability. Without such a test, the impartiality claim in the abstract and Section 4.5 is not established.","section":"§4.5, Table 5"}],"minor_comments":[{"comment":"The manuscript header reads 'DO LLM S AGREE' instead of 'DO LLMs AGREE'; this should be corrected.","section":"Title"},{"comment":"The computation of the Spearman correlation on the 12 averaged values per object should be stated explicitly before the results are presented; currently the reader must infer this from Figure 6 and the described aggregation.","section":"§3.2.3"},{"comment":"The text refers to 'top left in the table' for the heatmap; ensure that all heatmap entries include numeric values and that a color scale legend is provided for accessibility.","section":"§4.1, Table 2"},{"comment":"The phrase 'the highest score observed in our evaluation experiments' is vague; specify the exact correlation value and the condition to which it refers.","section":"§4.1"},{"comment":"The sentence 'the standard deviation ranged between 0.10 and 1.03' does not specify whether this is across models, across AUs, or across categories; clarify the population over which the standard deviation is computed.","section":"§4.2"},{"comment":"The statement that LLM evaluation performances are 'far superior' to semantic-distance-based evaluations is strong; consider citing the relevant empirical comparisons or softening the wording.","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"The oracle constructed in §3.2.4 is a major concern for the paper's central claim: it is derived from a self-cited prior paper with overlapping authors, and no independent human validation is provided. If the authors can add a human-rating study or explicitly reframe the contribution as measuring inter-model consensus rather than accuracy, the paper would be much stronger. The inter-model agreement results are the most defensible contribution and should be foregrounded in a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The inter-model agreement finding is the solid core here: four LLMs, scoring and ranking, two set sizes, and Spearman correlations mostly above 0.8 across conditions. That is a legitimate descriptive result, and the apparent absence of self-favoring is worth a footnote. The paper also asks a sensible question that extends the LLM-as-judge literature beyond single-model setups.\n\nThe soft spot is exactly where the stress-test note lands. The oracle in §3.2.4 is defined as the ordering [highly_creative, creative, common] justified solely by citing [2], a paper with overlapping authors that itself used LLM evaluations to validate its forceful prompts. So the reported 0.77–0.97 oracle correlations largely measure whether the evaluator reproduces the expected category order. Without a human baseline, calling this “accuracy” or “reliability” for creativity assessment is circular. The paper does not flag this limitation; it presents the oracle as ground truth.\n\nTwo more issues, in proportion. First, the aggregation step—averaging five AUs per model and category into 12 values per object before computing Spearman—can inflate correlations and masks within-category spread. The paper gives no confidence intervals or significance tests on any of the correlations, which matters when the sample is effectively 12 per object. Second, the impartiality claim (“models do not favor their own responses”) is based on inspecting average ratings across evaluators, not on any statistical test. It might be true, but the paper doesn't demonstrate it beyond eyeballing low standard deviations.\n\nWhat is genuinely new is the multi-model comparison itself: the dataset, the agreement heatmaps, and the observation that consensus holds across evaluation formats. That part is a useful contribution to the computational-creativity evaluation literature. The overreach is in the abstract's conclusion about “validating the reliability of LLMs in creativity assessment.”\n\nMy bottom line: worth serious peer review, but not in this form. The fix is straightforward—either add human creativity judgments as a baseline, or explicitly reframe the oracle as a proxy for prompt-defined categories and drop the reliability language. Add uncertainty estimates and a proper self-preference test. The paper would then be a solid empirical note.","headline":"The cross-model agreement result is real; the oracle-based accuracy claim is circular and the paper overstates it.","tokens_in":15841,"tokens_out":2233,"would_cite":false,"duration_ms":22854,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that four large language models produce mutually consistent and oracle-matched creativity judgments for alternative uses of everyday objects, and that no model favors its own outputs.","keywords":["Alternative Uses Test","LLM-as-a-judge","creativity evaluation","Spearman correlation","oracle benchmark","inter-model agreement","self-preference","scoring vs ranking"],"falsifier":"Have human raters rank or score the same 300 alternative uses and compare their average ordering with the oracle. If many highly_creative responses fall below creative or common responses in human judgment, the claim that LLMs accurately evaluate creativity fails, even if inter-model agreement remains high.","tokens_in":14872,"feed_emoji":"🤖","tokens_out":7732,"duration_ms":67484,"temperature":0.7,"pith_summary":"The paper asks whether different large language models agree with one another when they judge the creativity of alternative uses in the Alternative Uses Test, and whether they judge fairly when the responses come from other models. It builds a dataset of 300 alternative uses for five objects, generated by four LLMs under three prompt conditions labeled common, creative, and highly creative, and defines an oracle ordering from those labels. Each model then scores and ranks all responses, alone and in smaller batches. The paper reports average inter-model rank correlations above 0.7, correlations with the oracle above 0.77, and no systematic self-preference, concluding that LLMs can be reliable and impartial automated creativity assessors.","feed_headline":"Four AI models agree on creativity rankings of alternative uses","feed_subtitle":"Four AI judges give near-identical creativity scores and rankings, and none favors its own answers.","key_machinery":"The evaluation oracle is the load-bearing object: a fixed ordering that places all highly_creative responses first, then all creative responses, then all common responses, with each of the four models' responses kept together inside each tier. The paper measures every LLM's output against this oracle and against every other model using Spearman's rank correlation, so 'accuracy' reduces to whether the model reconstructs the oracle tier order and 'agreement' reduces to pairwise rank similarity. The forceful-prompt generation procedure from [2] is the machinery that produces the three creativity tiers in the first place.","core_discovery":"The paper's central claim, stated on its own terms, is that four different LLMs converge on a shared ordering of alternative uses by creativity and that this ordering matches a benchmark oracle built from the prompt conditions used to generate the responses. In the comprehensive scoring condition, every model's Spearman correlation with the oracle was above 0.95 for the average across objects, and two objects produced a perfect correlation of 1.0; across all four experimental conditions, correlations with the oracle averaged above 0.77 and inter-model correlations averaged above 0.7. The paper also reports that no model assigned systematically higher scores or ranks to its own generated responses over those of other models, and it interprets this as impartiality. It further finds that scoring outperforms ranking in oracle alignment and that the comprehensive setting generally matches or exceeds the segmented setting.","pith_inferences":["An implicit consequence I draw is that the inter-model agreement may reflect shared training on similar human judgments, so it establishes consensus among LLMs but does not by itself establish that the consensus matches human creativity judgments.","A test the paper leaves open is human validation: scoring the same 300 alternative uses by human raters would show whether the oracle's tier order matches human judgment or only the forcefulness of the generation prompts.","The framework could be extended to other divergent-thinking tasks, such as poetry or joke evaluation; in that extension, the same oracle construction would need independent per-task creativity levels rather than inherited AUT prompts."],"forward_implications":["A single LLM evaluator can be used in place of a panel of different models for AUT scoring tasks, since the models' judgments are near-interchangeable.","Because no self-preference was observed, LLM-generated alternative uses can be judged by another LLM without applying a bias correction for authorship.","Scoring should be preferred over full ranking when designing automated creativity evaluation prompts, because scoring produced higher oracle correlations in this study.","Batch size had limited impact: scoring 60 responses at once performed at least as well as scoring groups of 12, which supports using longer evaluation prompts when needed.","The benchmark construction can be reused to compare future LLM judges on the same 300-response dataset or on new objects generated under the same prompt tiers."],"supporting_citations":[{"why":"Supplies the forceful-prompt generation technique and the common/creative/highly_creative tiers from which the dataset and oracle ordering are built.","marker":"[2]"},{"why":"Provides the AUT scoring approach with LLMs and the precedent that LLM scores beat semantic-distance baselines.","marker":"[27]"},{"why":"Provides the ranking evaluation method for creative content that the ranking prompts are modeled on.","marker":"[9]"},{"why":"Supplies the Spearman rank correlation procedure used to compare model evaluations with the oracle.","marker":"[32]"},{"why":"Establishes the precedent of measuring how closely LLM evaluations align with expected rankings, the same measure the paper applies.","marker":"[6]"},{"why":"Defines the Alternative Uses Test, the task whose responses are being evaluated.","marker":"[24]"},{"why":"Shows a single LLM can assess AUT flexibility with high correlation to human ratings, the single-model result this paper extends to multi-model agreement.","marker":"[3]"},{"why":"Provides semantic distance, the traditional AUT scoring baseline that prior work found LLM-based evaluation to outperform.","marker":"[25]"}],"fun_headline_variants":["AI judges agree on creativity, no model favors its own","Four LLMs reach near-identical creativity scores","Creativity ranking: AI models show high consensus","No self-preference: LLMs align on creativity ratings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the oracle's tier order—highly_creative above creative above common, based on the generation prompt's forcefulness—really reflects increasing creativity; if it does not, the reported accuracy correlations only show agreement with a label ordering, not with creative quality.","fun_headline_variants_meta":{"raw":{"variants":["AI judges agree on creativity, no model favors its own","Four LLMs reach near-identical creativity scores","Creativity ranking: AI models show high consensus","No self-preference: LLMs align on creativity ratings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1364,"prompt_tokens":929,"completion_tokens":435,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":372}},"tokens_in":545,"tokens_out":435,"duration_ms":4174,"temperature":1.0,"reasoning_tokens":372,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:09:12.405485+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human raters rank or score the same 300 alternative uses and compare their average ordering with the oracle. If many highly_creative responses fall below creative or common responses in human judgment, the claim that LLMs accurately evaluate creativity fails, even if inter-model agreement remains high.","supporting_citations":[{"cited_title":"Pushing gpt’s creativity to its limits: Alternative uses and torrance tests,","cited_arxiv_id":null,"evidence_quote":"Supplies the forceful-prompt generation technique and the common/creative/highly_creative tiers from which the dataset and oracle ordering are built."},{"cited_title":"Is GPT-4 good enough to evaluate jokes?","cited_arxiv_id":null,"evidence_quote":"Provides the ranking evaluation method for creative content that the ranking prompts are modeled on."},{"cited_title":"Leveraging llm-respondents for item evaluation: a psychometric analysis,","cited_arxiv_id":null,"evidence_quote":"Supplies the Spearman rank correlation procedure used to compare model evaluations with the oracle."},{"cited_title":"Using large language models to evaluate alternative uses task flexibility score,","cited_arxiv_id":null,"evidence_quote":"Shows a single LLM can assess AUT flexibility with high correlation to human ratings, the single-model result this paper extends to multi-model agreement."}],"review_version":1}