{"id":"1e56b0da-724e-4f37-9141-c515bc0d8d9d","arxiv_id":"2506.08172","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper proposes and tests GrAImes, a literary-theory-based 15-item questionnaire for evaluating Spanish microfictions, finding modest inter-rater consistency across expert and enthusiast groups.","lead":"This paper presents GrAImes, a 15-question evaluation protocol for judging the literary quality of Spanish-language microfictions, whether written by humans or generated by AI. It reports two small validation experiments with literary experts and reading enthusiasts, and argues that such a protocol can ground AI fiction assessment in literary theory.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reliability evidence for GrAImes is confounded: expert/enthusiast groups and human/AI text sets are never crossed, so the protocol's general reliability across its stated scope is unsupported.","rationale":"The reader's conditional verdict is appropriate; the concern I identify strengthens the condition but does not change the verdict. The reader's weakest assumption focuses on the questionnaire's construct validity and small sample sizes. My more specific objection is structural: the two experiments do not estimate the protocol's reliability for AI-generated texts with experts or for human-written texts with enthusiasts, because evaluator group and text source are perfectly confounded. The paper reports 'good to acceptable internal consistency' separately for the Expert group on human texts and the Enthusiast group on AI texts, but this only shows that those particular combinations were somewhat consistent. It cannot support the central claim that GrAImes is a reliable framework across both settings. Additionally, the per-microfiction Cronbach's alpha with n=5 experts is statistically highly unstable; the paper acknowledges this sensitivity but still reports the point estimates as confirmatory evidence. The concrete check I propose—a crossed design with the same raters and text set, plus bootstrap confidence intervals—would directly test whether the observed reliability is an artifact of the confound. Since the paper's own conclusion is hedged ('could become a reliable framework'), the conditional verdict stands, but the acceptance conditions should explicitly require crossed validation before GrAImes is adopted or compared against other systems.","tokens_in":23080,"tokens_out":5909,"duration_ms":72334,"concrete_test":"Run a crossed validation study using the same 5 experts and 16 enthusiasts: each rater evaluates the same balanced set of 12 Spanish microfictions—6 human-written and 6 AI-generated—using the unchanged GrAImes questionnaire. Compute ICC(2,k), Cronbach's alpha, and Kendall's W separately for each of the four cells (Expert×Human, Expert×AI, Enthusiast×Human, Enthusiast×AI), with bootstrap 95% confidence intervals. Pre-register that the protocol is considered reliable only if the lower bound of ICC and alpha exceeds 0.70 in all four cells. If reliability is acceptable only in the original two cells, the original 'good to acceptable' conclusion is an artifact of the confound; if it fails in the crossed cells, GrAImes needs revision before being claimed as a general framework.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that GrAImes 'could become a reliable framework for assessing the literary quality of human written and AI-generated microfictions'—is supported only by internal-consistency statistics from two experiments in which evaluator population and text source are perfectly confounded. Section 3.2.1 defines two distinct rater groups (5 PhD literary experts; 16 enthusiasts), and Section 3.2.2 assigns the experts exclusively the 6 human-written microfictions and the enthusiasts exclusively the 6 AI-generated microfictions. Therefore no measurement cell exists for Expert×AI or Enthusiast×Human. The conclusion that GrAImes has 'good to acceptable internal consistency' for both human and AI texts cannot be separated from the fact that expert reliability was measured on human texts only and enthusiast reliability on AI texts only. If the protocol's reliability depends on rater expertise or on text type, the current design cannot detect it. A second, related problem is that Cronbach's alpha is computed separately for each microfiction (Tables 7 and 13), i.e., with 5 experts or 16 enthusiasts as the sample per text. With 5 raters, alpha's 95% confidence interval spans a range that includes both 'good' and 'unacceptable', so the reported 0.80/0.79 values are not stable evidence of acceptable reliability. The authors acknowledge sample-size sensitivity in Section 3.3 but still treat the point estimates as confirmatory. Finally, the 15-item instrument mixes aesthetic questions (Q3, Q7–Q9) with commercial/editorial ones (Q11–Q15), so high consistency would not by itself establish that 'literary value' is being measured.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GrAImes, a 15-item evaluation protocol for Spanish microfictions, grounded in literary theory and editorial practice, and reports two validation experiments. In the first, five PhD-holding literary experts rated six human-authored microfictions; in the second, sixteen literature enthusiasts rated six AI-generated microfictions (three from ChatGPT-3.5 and three from Monterroso, a GPT-2 baseline fine-tuned on Spanish microfiction). The authors report internal-consistency statistics (ICC, Cronbach's alpha, Kendall's W), a Sentence-BERT comparison of open-answer responses, and expert feedback on the protocol. They conclude that GrAImes 'could become a reliable framework' for assessing literary quality and that ChatGPT-3.5 texts were slightly favored over Monterroso texts. The paper includes a GitHub repository for reproducibility.","tokens_in":23368,"tokens_out":4796,"duration_ms":52337,"significance":"The paper addresses a genuine gap: automated metrics such as BLEU, ROUGE, and perplexity are not designed to capture literary qualities, and the protocol's grounding in reception theory and editorial criteria is a welcome contribution. The use of real PhD-level literary experts, the inclusion of a fine-tuned Spanish microfiction baseline, and the public GitHub repository for replication are strengths. If the validation were properly designed, GrAImes could be a useful instrument for the computational-creativity community. However, the current evidence does not support the broad reliability claim because the experimental design confounds rater expertise with text source, and the sample sizes are too small for the reported reliability statistics to be stable.","major_comments":[{"comment":"The validation design is confounded between rater group and text source: the five experts evaluated only the six human-written microfictions, and the sixteen enthusiasts evaluated only the six AI-generated microfictions. The paper's central claim that GrAImes has 'good to acceptable internal consistency' for assessing both human and AI-authored texts (Abstract; Section 4) is therefore not established, because reliability could depend on rater expertise, text type, or their interaction, and the design contains no Expert×AI or Enthusiast×Human cell. The authors should either cross the design or explicitly restrict the reliability claim to the measured cells.","section":"Sections 3.2.1-3.2.2 and 4"},{"comment":"The reliability statistics are computed on very small samples: Cronbach's alpha is estimated per microfiction with only five expert raters (Table 7) and sixteen enthusiast raters (Table 13), and the reported point values (e.g., 0.80, 0.79) have wide confidence intervals that include unacceptable levels of consistency. The paper acknowledges sample-size sensitivity in Section 3.3 but still treats these point estimates as confirmatory; the authors should report confidence intervals or bootstrap estimates and temper the conclusions accordingly.","section":"Section 3.3, Tables 6-7 and 13"},{"comment":"The labeling of the human-authored microfictions is internally inconsistent: Table 3 assigns MF3 and MF6 to the 'Medium' experience author, but the text in Section 4.1 describes MF3 and MF5 as written by 'low expertise' authors, describes MF6 as by an 'emerging author', and repeats a contradictory sentence about MF4 and MF6 versus MF3 and MF6. These contradictions undermine the reported correlation between author expertise and expert evaluations and make the results non-reproducible; the authors must correct the mismatches between the table and the prose.","section":"Section 4.1, Tables 3-7"},{"comment":"Tables 14-16 are headed 'Literary experts' responses' to Monterroso and ChatGPT-3.5 microfictions, which contradicts Section 3.2.2 where the AI-generated texts are assigned to the Enthusiast group. The surrounding text alternates between 'Enthusiast group leaders' and 'literary experts', making it unclear which rater population produced these data. In addition, the comparison between ChatGPT-3.5 and Monterroso is based on descriptive averages only, with no significance test; the claim that ChatGPT texts were 'slightly favored' is not statistically supported. The authors should clarify the rater groups and provide appropriate inferential statistics.","section":"Section 4.2, Tables 14-16"}],"minor_comments":[{"comment":"The section title and text contain several typos: 'GrAlmes' instead of 'GrAImes' and 'Cronbanch' instead of 'Cronbach' in Table 7.","section":"Section 4.1"},{"comment":"The captions contain typos: 'nthusiast' in Figure 10 and 'evlauation' in Figure 14 should be corrected.","section":"Figure 10 and Figure 14"},{"comment":"The sentence 'a critical stand is needeed' contains a typo and should read 'needed'.","section":"Section 2.2"},{"comment":"The entry for MF3, Question 6 reads '4.3 1.' with an incomplete decimal; this should be corrected to a complete value.","section":"Table 8"},{"comment":"The phrase 'MF 3, generated by Program A' should read 'generated by Monterroso' for consistency with the rest of the section.","section":"Section 4.2"},{"comment":"The description of the Enthusiast group mentions '16 literary enthousiasts, plus the group leader and booktuber', but the results in Section 4.2 sometimes refer to 'group leaders' without specifying whether the leader is included in the reported statistics; this ambiguity should be resolved.","section":"Section 3.2.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript shows signs of insufficient proofreading: the acknowledgments list 'Jorge Luis Borges' as a participant in the studies, which is impossible given his death in 1986, and footnote 3 reads like an editorial insertion rather than a scholarly note. I recommend that the editor request a careful revision of the text and a re-analysis of the reliability statistics with proper uncertainty quantification before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the GrAImes questionnaire: fifteen Spanish-language items, grounded in reception theory and editorial practice, that ask about interpretive depth, technical execution, and commercial viability. That is a real departure from BLEU, perplexity, and the usual \"does it make sense\" Likert scales. The authors also ship the materials on GitHub, and they are unusually candid in their limitations section about sample-size sensitivity. This is an honest, useful artifact, and the paper deserves a serious referee.\n\nThe problem is that the central reliability claim—\"good to acceptable internal consistency\" for both human and AI microfiction—is built on a design where the two evaluator groups and the two text sources are perfectly crossed in the wrong way: the five PhD experts rated only the six human-written microfictions, and the sixteen enthusiasts rated only the six AI-generated ones. So the reported alphas and ICCs can't separate \"experts agree on human texts\" from \"enthusiasts agree on AI texts.\" The stress-test note is right; this confound is load-bearing for the validation. The authors acknowledge small samples but still treat the point estimates as if they settle the matter.\n\nSecond, Cronbach's alpha computed per microfiction with five raters is an unstable statistic. The paper says so in Section 3.3, but then uses 0.80 and 0.79 as confirmatory values. A confidence interval would likely span \"unacceptable\" to \"good.\"\n\nThird, there is a genuine internal-consistency problem in the manuscript itself: Section 4.2 is titled as the Enthusiast group, yet several tables in that section are labeled \"Literary experts responses to Microfictions from Monterroso and ChatGPT-3.5.\" That makes it genuinely unclear which data come from which raters. The reader's conditional verdict is fair, and the labeling needs to be fixed before the numbers can be trusted.\n\nOn the content side, mixing aesthetic and commercial items is a defensible design choice, but a high alpha on the whole instrument does not by itself show that \"literary value\" is being measured. That is a narrower interpretive caveat rather than a fatal flaw. The critique of Porter and Machery is interesting but somewhat tangential.\n\nWho is this for: computational creativity researchers and literary scholars who want an alternative to surface metrics. It is not ready to be used as a benchmark, but the protocol is worth engaging with. I would send it to peer review with a request for major revision: cross the design or at least report reliability separately by rater group and text type, add uncertainty around the reliability statistics, and clean up the table labels.","headline":"A promising literary-theory-based evaluation protocol for Spanish microfiction, but the reliability validation is confounded by a two-arm design where experts see only human texts and enthusiasts only AI texts.","tokens_in":23922,"tokens_out":3635,"would_cite":false,"duration_ms":44030,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a fifteen-question protocol grounded in editorial practice can reliably assess the literary value of human-written and AI-generated Spanish microfictions, with expert ratings that track author experience.","keywords":["artificial intelligence","creative writing","evaluation protocol","microfiction","GrAImes","literary reception","large language models","inter-rater reliability"],"falsifier":"Run the same protocol with a larger panel (on the order of thirty experts and a matched set of texts at each experience level); if Cronbach's alpha for expert-authored texts no longer separates from emerging-writer texts, or the ICC values for the Likert items fall below acceptable thresholds, the claimed reliability collapses. A second decisive check is test-retest: have the same raters score the same microfictions twice, weeks apart; if individual ratings drift substantially, the instrument measures transient preference rather than stable literary judgment. A third check is a control panel of readers with no literary training rating the human-written texts; if their scores track the experts', the instrument is registering generic fluency rather than literary quality.","tokens_in":22922,"feed_emoji":"📚","tokens_out":10363,"duration_ms":108144,"temperature":0.7,"pith_summary":"This paper claims that the literary quality of microfiction can be scored with a reusable instrument rather than by intuition or surface text metrics. The instrument, GrAImes, is a fifteen-item questionnaire — ten Likert-scale items and five open answers — spanning literary interpretation, technical craft, and editorial or commercial appeal, modeled on how publishers decide whether to accept a manuscript. The authors validate it in two small experiments: five literature experts rated six human-written microfictions, and sixteen reading enthusiasts rated six AI-generated microfictions, three from ChatGPT-3.5 and three from a GPT-2 model fine-tuned on Spanish microfiction. They report good to acceptable internal consistency across raters, expert ratings that tracked the authors' experience level, and enthusiast ratings that slightly favored ChatGPT-3.5 on commercial appeal. If the protocol holds, it would give computational creativity research a literary-theory-grounded way to compare human and machine writing, and a counterweight to crowd-based claims that AI poetry already outranks the canon.","feed_headline":"15-question rubric claims to grade AI microfiction as literature","feed_subtitle":"Five experts and 16 readers tested the rubric on Spanish microfiction; expert scores tracked author skill.","key_machinery":"The load-bearing object is the GrAImes questionnaire: fifteen items — ten answered on a 1-to-5 Likert scale and five as open answers — split into three dimensions: story overview and text complexity (thematic coherence, clarity, interpretive depth), technical assessment (credibility, reader cooperation, originality of reality, genre, and language), and editorial or commercial quality (intertextual familiarity, desire for more, recommendation, gift-worthiness, publisher fit). Its design mirrors the editorial report a publisher commissions on unsolicited manuscripts, converting the editor's parameters — content clarity, technical value, and relevance — into scored items. The argument is carried by three reliability statistics applied to the raters' responses: the intraclass correlation coefficient for agreement, Cronbach's alpha for the internal consistency of each text's scores, and Kendall's W as a concordance measure chosen because it is less affected by small sample sizes. The protocol's claim to objectivity lives in the inter-rater agreement these statistics report, and its claim to literary validity lives in the questions themselves, which are grounded in reception theory and editorial practice rather than corpus statistics.","core_discovery":"GrAImes is presented as a reliable framework for assessing the literary value of human-written and AI-generated microfictions, something the authors argue current natural-language metrics cannot do because BLEU, ROUGE, and perplexity measure surface similarity rather than metaphor, symbolism, or stylistic originality. The protocol turns the publishing industry's editorial report into fifteen questions organized in three dimensions — story overview and textual complexity, technical assessment, and editorial or commercial quality — so that literary value is operationalized as thematic coherence, interpretive depth, technical execution, and market viability. Validation rests on three inter-rater statistics: the intraclass correlation coefficient, Cronbach's alpha, and Kendall's W, applied to two experiments with five expert and sixteen enthusiast raters. The authors report good to acceptable internal consistency, a correlation between author expertise and scores from the expert panel, and a slight enthusiast preference for ChatGPT-3.5 microfictions over the fine-tuned baseline in editorial and commercial appeal, while the baseline scored slightly higher on technical quality. They position GrAImes as a challenge to crowd-sourced findings that non-experts prefer AI-generated poetry to canonical human poetry, arguing that evaluations by readers without literary training measure immediate readability rather than interpretive depth.","pith_inferences":["Inference: A test-retest study — the same evaluators re-rating the same microfictions weeks apart — would separate stable literary judgments from familiarity or mood effects; the paper reports inter-rater agreement but no intra-rater stability, so its reliability claim covers only one axis of reliability.","Inference: The instrument's diagnostic profile (high technical scores, low innovation scores across both panels) suggests it could steer generation: fine-tuning or prompting could target the weakest dimensions, such as proposing a new vision of the genre, which scored lowest almost everywhere.","Inference: The negative ICC values for one item in each panel — Question 13 on gift-worthiness at -0.72 among experts and Question 8 on genre innovation at -0.44 among enthusiasts — indicate that some items actively depress the summary reliability figures; a revision that rewrites or drops these items would likely change the headline 'good to acceptable' verdict.","Inference: Applying GrAImes beyond Spanish would require re-norming rather than translation alone, since several items (publisher fit, gift-worthiness, intertextual recognition) presuppose a specific literary market and readership; otherwise cross-linguistic comparisons would confound literary quality with cultural familiarity."],"forward_implications":["GrAImes gives researchers a single fifteen-item instrument for comparing human-written, AI-generated, and AI-assisted microfictions on the same literary criteria, replacing surface metrics like BLEU and perplexity for this genre.","Because expert ratings tracked author experience (Cronbach's alpha of 0.80 and 0.79 for expert-authored texts versus 0.34 and 0.13 for emerging writers), the protocol can separate more accomplished writing from less accomplished writing.","On the enthusiast panel, ChatGPT-3.5 microfictions scored higher on editorial and commercial appeal while the fine-tuned baseline scored slightly higher on technical quality, implying that general audiences reward fluency and marketability more than structural craft.","Evaluator composition changes the result: experts emphasized originality and technical execution, enthusiasts emphasized accessibility, so any comparison of human versus AI literary output must report who did the judging.","The protocol is positioned as a check on crowd-sourced claims that AI poetry outranks canonical human poetry, by measuring interpretive depth rather than immediate preference."],"supporting_citations":[{"why":"Supplies the editorial-process model (clarity, technical value, relevance) that the three GrAImes dimensions mirror.","marker":"[14]"},{"why":"Provides the 300-word definition of microfiction that fixes the genre parameters for the evaluation.","marker":"[21]"},{"why":"The crowd-based AI-poetry-preference study whose method GrAImes is designed to challenge.","marker":"[4]"},{"why":"Reception theory, the literary basis for weighting readers' expertise in the protocol.","marker":"[10]"},{"why":"GPT-2, the base architecture behind the fine-tuned baseline generator used in experiment 2.","marker":"[38]"},{"why":"The Spanish GPT-2 model used to build the fine-tuned baseline's language capacity.","marker":"[57]"},{"why":"ChatGPT-3.5, the state-of-the-art generator whose microfictions the enthusiasts compared against the baseline.","marker":"[58]"},{"why":"Sentence-BERT embeddings used to measure semantic agreement among the open-answer responses.","marker":"[60]"}],"fun_headline_variants":["GrAImes rubric scores AI microfiction on literary value","Fifteen literary questions put AI microfiction to the test","AI microfiction gets a literary grade from expert panel","Borges benchmark: 15-question rubric for AI microfiction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The validation rests on two linked premises: that fifteen questions about interpretation, technique, and marketability capture what makes a microfiction literary, and that five experts and sixteen enthusiasts rating six texts each is a large enough sample for the reliability statistics to mean anything.","fun_headline_variants_meta":{"raw":{"variants":["GrAImes rubric scores AI microfiction on literary value","Fifteen literary questions put AI microfiction to the test","AI microfiction gets a literary grade from expert panel","Borges benchmark: 15-question rubric for AI microfiction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000575,"raw_usage":{"total_tokens":2714,"prompt_tokens":942,"completion_tokens":1772,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":1703}},"tokens_in":558,"tokens_out":1772,"duration_ms":15449,"temperature":1.0,"reasoning_tokens":1703,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:16:14.177545+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same protocol with a larger panel (on the order of thirty experts and a matched set of texts at each experience level); if Cronbach's alpha for expert-authored texts no longer separates from emerging-writer texts, or the ICC values for the Likert items fall below acceptable thresholds, the claimed reliability collapses. A second decisive check is test-retest: have the same raters score the same microfictions twice, weeks apart; if individual ratings drift substantially, the instrument measures transient preference rather than stable literary judgment. A third check is a control panel of readers with no literary training rating the human-written texts; if their scores track the experts', the instrument is registering generic fluency rather than literary quality.","supporting_citations":[{"cited_title":"What editors do: the art, craft, and business of book editing; University of Chicago Press, 2017","cited_arxiv_id":null,"evidence_quote":"Supplies the editorial-process model (clarity, technical value, relevance) that the three GrAImes dimensions mirror."},{"cited_title":"Cómo escribir un microrrelato; Siglo XXI Editores, 2023","cited_arxiv_id":null,"evidence_quote":"Provides the 300-word definition of microfiction that fixes the genre parameters for the evaluation."},{"cited_title":"AI-generated poetry is indistinguishable from human-written poetry and is rated more favorably","cited_arxiv_id":null,"evidence_quote":"The crowd-based AI-poetry-preference study whose method GrAImes is designed to challenge."},{"cited_title":"The act of reading: A theory of aesthetic response","cited_arxiv_id":null,"evidence_quote":"Reception theory, the literary basis for weighting readers' expertise in the protocol."},{"cited_title":"Language models are unsupervised multitask learners","cited_arxiv_id":null,"evidence_quote":"GPT-2, the base architecture behind the fine-tuned baseline generator used in experiment 2."},{"cited_title":"GPT2-spanish","cited_arxiv_id":null,"evidence_quote":"The Spanish GPT-2 model used to build the fine-tuned baseline's language capacity."}],"review_version":1}