{"id":"f47668d0-1363-4a0d-84e4-0b6e3dd2be97","arxiv_id":"2504.15784","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A reference-based Likert scoring method with analyze-rate prompting improves LLM judges' agreement with human creativity rankings on the TTCW benchmark, though the headline result is partly fitted to the test set.","lead":"This paper tests an automated way to score creative writing from large language models by comparing each machine story against a human-written reference story on the Torrance Test of Creative Writing. The method lifts AI judge agreement with human rankings to 0.75 pairwise accuracy, but the reported gain depends on tuning a cutoff on the same test set.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline improvement is not protected from selection: the cutoff and Likert scale are tuned on the same test set, only three of ten evaluator models are reported, and no held-out validation or significance test is provided; the reported +15% gain may be a selection artifact.","rationale":"The paper proposes a clear, simple method and provides ablation and a second dataset as positive signals; I do not question the authors' intent. The load-bearing weakness is methodological: every reported number in Table 1 is a maximum over choices made using the same 12 human-ranked plots that define the target. Section 5.1 selects the cutoff by evaluating four candidates on the full test set; Section 5.2 selects the 5-point scale on one test model using the same data; and Section 4.3 reports 'Ours' for only three of the ten evaluator models, with the headline 0.75 being the best of those three. With 12 plots, the effective sample is tiny, so the gap between the best and the true expected performance can be large. The reader's concern about noisy human ground truth is valid, but even noise-free labels would not cure test-set selection; therefore I focus on the selection protocol. If a held-out protocol preserves the improvement, the method is valuable; if not, the central claim fails. Given the paper's contributions, the appropriate verdict is CONDITIONAL rather than REJECT, because the issue is testable and the method may generalize.","tokens_in":7479,"tokens_out":19069,"duration_ms":185397,"concrete_test":"Run a leave-one-plot-out cross-validation over the 12 stories. Within each training fold, choose the binary cutoff (and, if kept as a tuned component, the Likert granularity) using only the held-in plots; then score the held-out plot with the chosen settings. Apply this protocol to all ten evaluator models in Table 1, not just Claude 3.5, GPT-4o, and Qwen-2-72B-Chat, and report mean held-out Spearman, Kendall, and pairwise accuracy plus a paired significance test (e.g., Wilcoxon signed-rank) against the same model's baseline. If the held-out gain is not positive across models, the reported +15% is a tuning artifact rather than a property of the method.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim rests on a small, tuned evaluation. All headline numbers come from 12 New Yorker plots. The authors select the binary-conversion cutoff by searching Cutoff = -3, -2, -1, 0 on the same human rankings they later report (Section 5.1 and Table 4), and they select the 5-point Likert granularity using one of the three test models on the same data (Section 5.2). They then report 'Ours' for only three of the ten evaluator models listed in Table 1, and the headline 0.75 is the best of these three. With 12 plots and only 3 or 4 items ranked per plot, the standard error of Spearman and pairwise accuracy is large; searching over 4 cutoffs, 3 scales, and 3 models makes the maximum likely to exceed the true expected value substantially. For example, Claude 3.5's Spearman varies from 0.20 to 0.49 across cutoffs. No confidence intervals, significance tests, or held-out tuning are reported. Therefore the reported +15% / +0.11 gain could be a selection artifact rather than a property of the reference-based method. This is load-bearing because if the gain disappears under proper validation, the paper's central claim fails regardless of how reliable the human ground truth is.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an automated creativity evaluation method for LLMs based on the Torrance Test of Creative Writing (TTCW). For each of 14 binary TTCW tests, an LLM judge compares a candidate story against a high-quality reference story using a Likert scale, with the order of candidate and reference alternated across two trials. A test is counted as passed if the score difference meets a cutoff, and the total number of passed tests ranks the candidate models. The authors compare these rankings with human expert rankings using Spearman's correlation, Kendall's tau, and pairwise accuracy. With an analyze-rate prompting strategy, they report a pairwise accuracy of 0.75 (+15%) for Claude 3.5, with smaller gains for GPT-4o and Qwen-2-72B-Chat, and they include ablations and a generalization experiment on a second dataset.","tokens_in":7710,"tokens_out":4599,"duration_ms":44429,"significance":"If the reported gains hold under proper validation, the reference-based Likert-style TTCW evaluation would be a practical, scalable alternative to manual creativity annotation, and the paper would make a useful contribution to automated evaluation of creative writing. The method is simple, grounded in established psychometric tests, and the inclusion of ablations and a second dataset are strengths. However, the central quantitative claim is currently not protected from selection effects: the binary-conversion cutoff is chosen on the same test set used to report the headline results, the Likert scale granularity is chosen on the same data, and only three of ten evaluator models are reported for the proposed method. As a result, the reported +15% pairwise-accuracy improvement may reflect test-set fitting rather than a genuine property of the method. The paper would be significant after the empirical claims are made robust through held-out validation or equivalent safeguards.","major_comments":[{"comment":"The score cutoff is selected by searching Cutoff = -3, -2, -1, 0 on the same test set that is later used to report the main results. For example, Claude 3.5's Spearman correlation under the proposed method ranges from 0.20 to 0.49 across cutoffs, and the reported 0.49 is the maximum. With only 12 stories and 3-4 candidate texts ranked per story, the variance of Spearman and pairwise accuracy is large, and selecting the cutoff, Likert scale, and evaluator model on the same data makes the reported maximum likely to exceed the true expected value. Please provide a held-out validation procedure, cross-validation, or at minimum report all cutoff results with confidence intervals or bootstrap estimates so the reader can assess the selection effect.","section":"Section 5.1, Table 4, Eq. (1)"},{"comment":"The 5-point Likert scale is selected using qwen2-72b-chat on the same TTCW test data, with 3-point and 7-point scales yielding lower Spearman correlations. This is an additional hyperparameter selected on the test set, and the reported results for all three models use that selected scale. The claim that the 5-point scale is the 'more effective choice' is therefore not validated independently. Please report the sensitivity of all reported metrics to the Likert granularity, or validate the choice on a separate development set.","section":"Section 5.2"},{"comment":"The 'Ours' rows are reported for only three of the ten evaluator models listed in the baseline, and these three models appear to be the ones with the largest improvements. The abstract and conclusion claim that 'our method significantly improves the alignment between LLM evaluations and human assessments' in general, but no results are shown for the other seven models. Reporting only the favorable subset, without a multiple-comparison correction, materially weakens the generalization claim. Please report results for all ten models, or clearly frame the claim as applying to the three models tested.","section":"Section 4.3, Table 1"},{"comment":"The human expert rankings from Chakrabarty et al. (2024) are used as ground truth without reporting inter-annotator agreement or any measure of their reliability. If those rankings are noisy or biased, all reported Spearman, Kendall, and pairwise accuracy values inherit that noise, and the 0.75 pairwise accuracy is not a meaningful measure of evaluation quality in isolation. No confidence intervals or significance tests are provided for any of the headline numbers, which is particularly concerning given the small sample of 12 stories. Please include reliability statistics for the human ground truth and uncertainty estimates for the reported alignment metrics.","section":"Section 4.1 and Section 4.3"}],"minor_comments":[{"comment":"The notation is inconsistent: the subscripts in L^{k,+}_{i,j} are later written as L^{k,+}_{i,k}, and the prose says a test passes if the average score across two assessments is higher than the cutoff, while Eq. (1) uses the difference L^{k,+} - L^{k,-}. Please clarify whether the two trials are averaged or differenced, and align the notation and the equation.","section":"Section 3.2, Eq. (1)"},{"comment":"The prompt template shown in Appendix A.1 contains the placeholder 'Stories and Question...' instead of the actual instruction block. Please provide the complete prompt used in the experiments so the method is reproducible.","section":"Appendix A.1"},{"comment":"The per-story results are only shown as sorted averages in the figures; it would be more informative to display the individual story-level values or add error bars so that variability across the 12 stories is visible.","section":"Appendix A.4, Figures 2-4"},{"comment":"In the reference for Summers-Stay et al. (2023), the author name is rendered as 'Clare R. V oss' with an extra space; please fix the typesetting.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper has a sensible core idea and the reference-based comparison is worth exploring, but the empirical evidence as presented does not support the headline claim because of test-set-based selection of the cutoff and Likert scale, selective reporting of three of ten models, and the absence of uncertainty quantification. I am not recommending rejection because these issues could be addressed within the scope of a revision, for example by adding a held-out validation split, reporting all models, and providing confidence intervals. If the authors are unable to add such validation, the claims should be substantially softened in the abstract and conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this paper combines known pieces into a sensible reference-based Likert evaluation for TTCW, with position-flip averaging and analyze-rate prompting. That combination is new, and the paper is clearly written and honest about its limitations. But the headline +15% is not protected from selection: the cutoff and Likert scale are chosen on the test set, and 'Ours' is reported for only three of ten evaluator models. The stress-test note is on target.\n\nWhat is genuinely good: the method is simple and intuitive, the ablations show both the reference anchor and the analyze-rate prompt help, and the second dataset (Gómez-Rodríguez & Williams) gives some evidence that the approach transfers beyond TTCW. The paper also includes full tables in the appendix, so the reader can see exactly how the numbers move. The citation pattern is fine—TTCW, BERTScore/BARTScore, and analyze-rate are all credited appropriately.\n\nThe soft spots. First, the cutoff search in Section 5.1 and Table 4: the authors scan Cutoff = -3, -2, -1, 0 on the very same human rankings they later report as the result. Claude 3.5's Spearman goes from 0.20 to 0.49 depending on cutoff. Picking the best cutoff is post-hoc fitting, and with 12 stories the noise is high. Second, the 5-point Likert scale is selected using qwen2-72b-chat on the same data (Section 5.2). Third, 'Ours' appears for only three models; the abstract's +15% is the best of those three, and the pairwise accuracy gain is +0.11 absolute—the +15% is a relative gain, which is a bit misleading. Fourth, no confidence intervals, significance tests, or held-out tuning. The human rankings being noisy is a secondary concern; even if they are perfect, the reported gain could be a selection artifact.\n\nI don't think this is fatal. The method may generalize, and the underlying idea is reasonable. But the central claim needs better support: a held-out cutoff selection or a tight prior on the threshold, results for all ten models, and uncertainty measures. The paper also has a useful limitation section and a risks section, which suggests the authors are thinking about scope.\n\nWho should read this: people working on automated LLM evaluation, especially for creative writing or subjective qualities. It's a small but real step, and the methodological lesson about tuning on the test set is worth discussing.\n\nMy recommendation: send it to peer review. It's not desk-reject material—the method is novel and the flaws are fixable with a revision. I'd tell the authors to redo the analysis with proper validation and report all models.","headline":"A plausible reference-based creativity evaluator whose headline gain is likely inflated by tuning on the test set; worth a revision, not a reject.","tokens_in":8282,"tokens_out":1889,"would_cite":false,"duration_ms":19380,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Comparing generated stories against human-written references, test by test, makes LLM creativity judgments align with human experts more closely than rating stories alone.","keywords":["creativity evaluation","large language models","Torrance Test of Creative Writing","reference-based evaluation","Likert scale","LLM-as-judge","story generation","automated evaluation"],"falsifier":"Run the proposed reference-based TTCW protocol on a new set of 50 storylines where human rankings come from multiple independent expert judges with measured inter-annotator agreement; if the method's pairwise accuracy against one noisy judge set drops below 0.60, the claimed 0.75 alignment is not generalizable.","tokens_in":1244,"feed_emoji":"✍️","tokens_out":1417,"duration_ms":55557,"temperature":0.7,"pith_summary":"This paper argues that LLM judges can evaluate creative writing almost as reliably as human experts when the judgment is reframed as a relative comparison rather than an absolute rating. Each machine-written story is scored against a high-quality human-written reference on the 14 binary tests of the Torrance Test of Creative Writing, covering Fluency, Flexibility, Originality, and Elaboration. The judge model analyzes both stories against the criterion before choosing a five-level Likert verdict, the presentation order is alternated to cancel position bias, and a test counts as passed when the average verdict clears a cutoff of -2. On 12 storylines with human expert rankings as ground truth, the method lifts average pairwise accuracy to 0.75, a 15 percent improvement over direct prompting, and raises the best judge's average Spearman correlation from 0.15 to 0.49. If correct, this offers a scalable automated benchmark for creativity that does not require manual annotation.","feed_headline":"Reference-based scoring lifts LLM creativity judgments to 75% human agreement","feed_subtitle":"A comparative Likert protocol over the TTCW tests outperforms direct rating, offering a scalable creativity benchmark.","key_machinery":"The machinery is a reference-based Likert comparison over the 14 binary TTCW tests. The tests are organized into four dimensions: Fluency, Flexibility, Originality, and Elaboration, each with a guiding question about the story. The judge sees the candidate and reference together, evaluates one test at a time, and outputs one of five verdicts; two orderings are averaged; the test is passed if the mean score is at least -2; and the sum over the 14 tests becomes the creativity score. This converts subjective judgment into a finite countable score, which is what makes automated and scalable evaluation possible.","core_discovery":"The central claim is that creativity in generated stories can be measured automatically as a product-level property using the TTCW's fourteen binary questions, provided each question is answered comparatively rather than absolutely. For each question, an LLM judge reads the candidate and a human-authored reference, states which is better on a five-level scale, and repeats the comparison with the order swapped; the question counts as passed only when the average verdict meets a cutoff of -2. The total creativity score is the number of tests passed. With this protocol plus an analyze-then-rate prompt, the paper reports that three LLM judges align better with human expert rankings than the original direct-prompt TTCW, with the best judge reaching 0.75 pairwise accuracy, 0.49 average Spearman, and 0.44 average Kendall's tau.","pith_inferences":["Because every score is computed relative to a single high-quality reference per storyline, the method probably measures proximity to the reference's aesthetic as much as creativity in the abstract; using multiple or deliberately diverse reference anchors could reveal how much of the ranking is style-dependent.","The 0.75 pairwise accuracy is estimated from only 12 storylines, so its uncertainty is unknown; re-running the protocol on 50 or more storylines with bootstrap confidence intervals would give a more stable effect size and would test whether the 15-point gain over baseline persists.","The paper itself flags, in its limitation section, that reference dependence restricts open-ended use and that the method collapses when all candidates are far worse than the reference; a practical extension would be adaptive reference selection or multiple reference anchors to recover ranking signal in that regime.","The same reference-based comparison could be repurposed as a distillation monitor, scoring a student model's creative outputs against a teacher model's outputs; the paper mentions this direction only in passing."],"forward_implications":["LLM creativity can be tracked across model versions cheaply and without human annotation, as long as high-quality reference texts exist for the relevant domain.","Model rankings become more comparable across research groups because scores are anchored to the same published reference set rather than to each judge's private taste.","The procedure can be ported to other creative writing forms, such as poetry, dialogue, or advertising copy, by adapting TTCW-style test batteries and reference collections.","The choice of Likert granularity matters: the paper reports that the 5-point scale outperforms 3-point and 7-point scales, so scale design is a practical parameter for future evaluation suites.","Automated creativity scoring could support reference-based quality control in generation systems, flagging outputs that fall below the quality bar of a chosen reference style."],"supporting_citations":[{"why":"Introduces the TTCW framework, the 12 New Yorker reference stories, the human expert rankings used as ground truth, and the direct-prompt baseline that this paper improves on.","marker":"Chakrabarty et al., 2024"},{"why":"Supplies the analyze-rate prompting strategy that the paper adopts and shows improves evaluation accuracy with LLM judges.","marker":"Chiang and yi Lee, 2023"},{"why":"Provides the second, 10-point-scale creative writing dataset used to test generalization beyond TTCW's binary format.","marker":"Gómez-Rodríguez and Williams, 2023"},{"why":"Defines the rank correlation metric used to measure alignment between automated scores and human rankings.","marker":"Spearman, 1904"},{"why":"Defines Kendall's tau, the second rank-correlation metric reported in the main results.","marker":"Kendall, 1938"}],"fun_headline_variants":["Reference-based TTCW scoring raises LLM creativity alignment to 75%","Automated creativity evaluation: reference-based Likert outperforms direct rating","LLM creativity grading: comparative scoring hits 75% pairwise agreement with humans","New TTCW protocol benchmarks LLM creativity by reference comparison"],"cache_read_input_tokens":10368,"weakest_assumption_plain":"The human expert rankings from Chakrabarty et al. are treated as reliable and noise-free ground truth, but the paper does not report inter-annotator agreement for those rankings, so all reported correlations inherit whatever noise or bias those rankings contain.","fun_headline_variants_meta":{"raw":{"variants":["Reference-based TTCW scoring raises LLM creativity alignment to 75%","Automated creativity evaluation: reference-based Likert outperforms direct rating","LLM creativity grading: comparative scoring hits 75% pairwise agreement with humans","New TTCW protocol benchmarks LLM creativity by reference comparison"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000504,"raw_usage":{"total_tokens":2400,"prompt_tokens":826,"completion_tokens":1574,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":442,"completion_tokens_details":{"reasoning_tokens":1496}},"tokens_in":442,"tokens_out":1574,"duration_ms":11346,"temperature":1.0,"reasoning_tokens":1496,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:17:11.879789+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the proposed reference-based TTCW protocol on a new set of 50 storylines where human rankings come from multiple independent expert judges with measured inter-annotator agreement; if the method's pairwise accuracy against one noisy judge set drops below 0.60, the claimed 0.75 alignment is not generalizable.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the TTCW framework, the 12 New Yorker reference stories, the human expert rankings used as ground truth, and the direct-prompt baseline that this paper improves on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the rank correlation metric used to measure alignment between automated scores and human rankings."}],"review_version":1}